Research Papers 论文研究 5h ago Updated 1h ago 更新于 1小时前 45

SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction SciReC:面向多模态多轮关系推理的自适应交互诊断评估

SciReC is a model-adaptive multimodal academic dialog benchmark designed to evaluate relational reasoning in multimodal large language models (MLLMs) across analogical, structural, and cause-effect categories DMRA (Deficit-based Multimodal Relational Analysis) is a diagnostic framework that quantifies the contribution of visual understanding, domain knowledge, and memory recall to identify root causes of model failures Claude 4.6 leads with 73% overall relational score, followed by GPT 5.4 at 68 提出SciReC基准测试,用于评估多模态大语言模型(MLLM)在关系推理任务上的表现,涵盖类比、结构和因果关系等推理类别 提出DMRA诊断框架,量化视觉理解、知识表现和记忆回忆等组件的贡献,定位模型失败的主要原因 Claude 4.6以73%的整体关系推理得分领先,GPT 5.4以68%紧随其后 开源模型在空间关系上得分最低,专有模型在层次和序列关系上表现较差;跨领域中天文学表现最差、心理学最佳 DMRA分析显示关系推理是所有模型的主要错误来源,其次是记忆限制

58
Hot 热度
72
Quality 质量
63
Impact 影响力

Analysis 深度分析

TL;DR

  • SciReC is a model-adaptive multimodal academic dialog benchmark designed to evaluate relational reasoning in multimodal large language models (MLLMs) across analogical, structural, and cause-effect categories
  • DMRA (Deficit-based Multimodal Relational Analysis) is a diagnostic framework that quantifies the contribution of visual understanding, domain knowledge, and memory recall to identify root causes of model failures
  • Claude 4.6 leads with 73% overall relational score, followed by GPT 5.4 at 68%, revealing a notable performance gap between top proprietary models and the ceiling on relational reasoning
  • Open-source models perform worst on spatial relations, while proprietary models struggle more with hierarchical and sequential relations; Astronomy is the weakest domain and Psychology the strongest across all models
  • Relational reasoning is the primary source of error across all evaluated models, with memory limitations ranking as the second most significant failure factor

Why It Matters

This work addresses a critical gap in MLLM evaluation by moving beyond single-turn, static benchmarks to assess multi-turn, adaptive relational reasoning—a capability essential for real-world scientific and academic applications. The DMRA diagnostic framework provides practitioners with a actionable methodology for decomposing model failures into interpretable components, enabling targeted improvements rather than opaque performance tracking.

Technical Details

  • SciReC Benchmark: A model-adaptive multimodal academic dialog benchmark that dynamically adjusts interaction difficulty based on model performance, covering relational reasoning categories including analogical, structural, and cause-effect reasoning across multiple academic domains
  • DMRA Framework: A deficit-based diagnostic methodology that isolates and quantifies the contribution of three key components—visual understanding, domain knowledge/exhibiting knowledge, and memory recall—to pinpoint the primary cause of unsuccessful reasoning cases
  • Evaluation Scope: Tests span multiple domains (Astronomy through Psychology), with performance measured on overall relational scores, spatial/hierarchical/sequential relation categories, and multi-turn adaptive dialog interactions
  • Model Coverage: Evaluates both proprietary models (Claude 4.6, GPT 5.4) and open-source MLLMs, revealing divergent failure patterns between model categories

Industry Insight

  • The persistent difficulty across all models with relational reasoning—especially spatial, hierarchical, and sequential relations—suggests that current MLLM architectures still lack robust mechanisms for multi-hop, structured inference; investment in reasoning-specific training objectives and architectural modifications should be prioritized
  • The domain-specific performance gap (Astronomy weakest, Psychology strongest) indicates that models benefit from domains with more structured, rule-based knowledge representations, guiding dataset curation and fine-tuning strategies for specialized applications
  • DMRA's diagnostic decomposition approach offers a practical blueprint for teams to move beyond aggregate benchmark scores and systematically identify whether failures stem from perception, knowledge retrieval, or reasoning—enabling more efficient resource allocation in model development

TL;DR

  • 提出SciReC基准测试,用于评估多模态大语言模型(MLLM)在关系推理任务上的表现,涵盖类比、结构和因果关系等推理类别
  • 提出DMRA诊断框架,量化视觉理解、知识表现和记忆回忆等组件的贡献,定位模型失败的主要原因
  • Claude 4.6以73%的整体关系推理得分领先,GPT 5.4以68%紧随其后
  • 开源模型在空间关系上得分最低,专有模型在层次和序列关系上表现较差;跨领域中天文学表现最差、心理学最佳
  • DMRA分析显示关系推理是所有模型的主要错误来源,其次是记忆限制

为什么值得看

该研究首次系统性地评估了多模态大模型在复杂关系推理任务上的能力瓶颈,为模型优化提供了明确的诊断方向。DMRA框架的提出有助于AI从业者理解模型失败的根本原因,而非仅关注表面准确率。

技术解析

  • SciReC是一个模型自适应的多模态学术对话基准,涵盖多个学科领域(如天文学、心理学等),通过多轮对话形式评估模型的关系推理能力
  • DMRA(Deficit-based Diagnostic Framework)是一种基于缺陷的诊断框架,能够量化视觉理解、知识表现和记忆回忆三个组件对最终推理结果的贡献度
  • 关系推理被细分为多个类别:类比推理、结构推理和因果关系推理,每个类别捕捉高阶理解的不同方面
  • 评测结果显示Claude 4.6和GPT 5.4在整体关系推理任务上表现最佳,但各模型在不同推理类型和领域上存在显著差异

行业启示

  • 多模态大模型在复杂关系推理方面仍存在明显瓶颈,尤其是空间关系和层次关系,这提示未来模型优化应重点关注推理能力的提升而非仅依赖规模扩张
  • 开源模型与专有模型在推理类型上存在差异化短板,开发者应根据具体应用场景选择合适的模型类型
  • DMRA诊断框架为模型迭代提供了可操作的评估工具,建议AI团队在开发多模态系统时引入类似的缺陷分析机制

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Multimodal 多模态 Evaluation 评测 Benchmark 基准测试 Research 科学研究 LLM 大模型