ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation
ChemDIRT is a comprehensive benchmark designed to evaluate the robustness of LLMs in chemistry reasoning, addressing limitations of existing narrow benchmarks The benchmark systematically measures model performance across variations in instructions and molecular representations across eight chemistry task categories Evaluation includes both accuracy and consistency metrics under controlled perturbations, providing more reliable assessments than single-format benchmarks Results reveal substantial
Analysis
TL;DR
- ChemDIRT is a comprehensive benchmark designed to evaluate the robustness of LLMs in chemistry reasoning, addressing limitations of existing narrow benchmarks
- The benchmark systematically measures model performance across variations in instructions and molecular representations across eight chemistry task categories
- Evaluation includes both accuracy and consistency metrics under controlled perturbations, providing more reliable assessments than single-format benchmarks
- Results reveal substantial prompt sensitivity, representation dependence, and uneven performance across different task families in both open- and closed-source LLMs
- The framework aims to prevent overestimation of model capabilities by testing consistency across realistic variations in problem formulation
Why It Matters
This benchmark addresses a critical gap in AI evaluation for scientific domains, where current metrics may overestimate model capabilities due to narrow testing formats. For AI practitioners working on chemistry applications, ChemDIRT provides a more realistic assessment framework that accounts for real-world variability in how problems are presented and represented. The findings have implications for anyone deploying LLMs in scientific research or education, where robustness across different formulations is essential.
Technical Details
- ChemDIRT spans eight categories of chemistry tasks, systematically varying both instruction formats and molecular representations to test model robustness
- The evaluation framework measures both accuracy and consistency under controlled perturbations, going beyond simple correctness metrics
- A diverse set of open-source and closed-source LLMs were benchmarked, revealing significant performance variations
- The benchmark specifically tests prompt sensitivity and representation dependence, factors often overlooked in conventional single-format chemistry benchmarks
- Submitted to arXiv on August 21, 2026, under cs.LG (Machine Learning) category
Industry Insight
- AI developers should prioritize robustness testing across multiple input formats and representations when deploying LLMs for scientific applications, as single-format benchmarks may produce misleading performance estimates
- The substantial prompt sensitivity revealed by ChemDIRT suggests that prompt engineering and format standardization will remain critical for reliable chemistry AI applications
- Researchers building chemistry-focused LLMs should consider consistency metrics alongside accuracy, as uneven performance across task families indicates current models lack generalizable chemical reasoning capabilities
Disclaimer: The above content is generated by AI and is for reference only.