Research Papers 论文研究 5h ago Updated 45m ago 更新于 45分钟前 45

ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation ChemDIRT:面向鲁棒化学大语言模型评估的多样化指令、表示与任务基准

ChemDIRT is a comprehensive benchmark designed to evaluate the robustness of LLMs in chemistry reasoning, addressing limitations of existing narrow benchmarks The benchmark systematically measures model performance across variations in instructions and molecular representations across eight chemistry task categories Evaluation includes both accuracy and consistency metrics under controlled perturbations, providing more reliable assessments than single-format benchmarks Results reveal substantial 提出ChemDIRT基准测试,系统评估LLM在化学领域的鲁棒推理能力,弥补现有基准测试的不足 现有化学基准测试仅关注有限任务集,忽视模型对问题表述和化学表示变化的鲁棒性,可能高估模型真实能力 ChemDIRT涵盖8类化学任务,系统性测量模型在指令和分子表示变化下的准确性与一致性 基准测试多种开源和闭源LLM,发现模型存在显著的提示敏感性、表示依赖性和任务家族间的不均衡表现

58
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • ChemDIRT is a comprehensive benchmark designed to evaluate the robustness of LLMs in chemistry reasoning, addressing limitations of existing narrow benchmarks
  • The benchmark systematically measures model performance across variations in instructions and molecular representations across eight chemistry task categories
  • Evaluation includes both accuracy and consistency metrics under controlled perturbations, providing more reliable assessments than single-format benchmarks
  • Results reveal substantial prompt sensitivity, representation dependence, and uneven performance across different task families in both open- and closed-source LLMs
  • The framework aims to prevent overestimation of model capabilities by testing consistency across realistic variations in problem formulation

Why It Matters

This benchmark addresses a critical gap in AI evaluation for scientific domains, where current metrics may overestimate model capabilities due to narrow testing formats. For AI practitioners working on chemistry applications, ChemDIRT provides a more realistic assessment framework that accounts for real-world variability in how problems are presented and represented. The findings have implications for anyone deploying LLMs in scientific research or education, where robustness across different formulations is essential.

Technical Details

  • ChemDIRT spans eight categories of chemistry tasks, systematically varying both instruction formats and molecular representations to test model robustness
  • The evaluation framework measures both accuracy and consistency under controlled perturbations, going beyond simple correctness metrics
  • A diverse set of open-source and closed-source LLMs were benchmarked, revealing significant performance variations
  • The benchmark specifically tests prompt sensitivity and representation dependence, factors often overlooked in conventional single-format chemistry benchmarks
  • Submitted to arXiv on August 21, 2026, under cs.LG (Machine Learning) category

Industry Insight

  • AI developers should prioritize robustness testing across multiple input formats and representations when deploying LLMs for scientific applications, as single-format benchmarks may produce misleading performance estimates
  • The substantial prompt sensitivity revealed by ChemDIRT suggests that prompt engineering and format standardization will remain critical for reliable chemistry AI applications
  • Researchers building chemistry-focused LLMs should consider consistency metrics alongside accuracy, as uneven performance across task families indicates current models lack generalizable chemical reasoning capabilities

TL;DR

  • 提出ChemDIRT基准测试,系统评估LLM在化学领域的鲁棒推理能力,弥补现有基准测试的不足
  • 现有化学基准测试仅关注有限任务集,忽视模型对问题表述和化学表示变化的鲁棒性,可能高估模型真实能力
  • ChemDIRT涵盖8类化学任务,系统性测量模型在指令和分子表示变化下的准确性与一致性
  • 基准测试多种开源和闭源LLM,发现模型存在显著的提示敏感性、表示依赖性和任务家族间的不均衡表现

为什么值得看

本文揭示了当前科学领域LLM评估的重要盲点,为构建更可靠的化学AI评估体系提供了系统性框架。对AI从业者和化学研究者而言,ChemDIRT提供了更贴近真实应用场景的模型能力评估方法。

技术解析

  • ChemDIRT(Diversified Instruction, Representation, and Task Benchmark)是一个综合评估框架,通过控制变量法系统测量模型在指令变化和分子表示变化下的性能表现
  • 基准测试涵盖8类化学任务,包括分子性质预测、反应预测、合成规划等核心化学推理任务
  • 评估维度不仅包括准确率,还强调一致性(consistency),即在相同问题不同表述下的稳定输出能力
  • 实验覆盖了多种开源和闭源LLM,揭示了模型在化学领域的系统性弱点

行业启示

  • 科学AI评估需要从单一准确率指标转向多维度鲁棒性评估,特别是在输入格式多样化的真实场景中
  • 化学领域的LLM应用应重视模型对分子表示形式(如SMILES、InChI等)的适应能力,这直接影响实际部署效果
  • 未来基准测试设计应更加关注模型的一致性和泛化能力,而非仅追求特定格式下的最优性能

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Benchmark 基准测试 Evaluation 评测 Dataset 数据集 Research 科学研究