Research Papers 论文研究 6h ago Updated 1h ago 更新于 1小时前 45

ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge ORQA:面向LLM专业知识的职业现实问答框架

ORQA is a novel benchmark framework that evaluates LLMs on occupation-level professional knowledge by linking O*NET occupations to trusted, occupation-specific websites (regulatory agencies, licensing bodies, professional organizations) to generate source-traceable QA pairs The benchmark covers 116 occupations across all 21 major SOC groups, with 480 questions sourced from 187 different websites, combining an automated pipeline with human review Top frontier models (Claude Opus 4.6, GPT-5.4, Cla ORQA是一种评估LLM职业级别知识的新方法,通过将O*NET职业与可信的职业特定网站(监管机构、许可机构、专业组织、政府出版物)连接,生成可追溯来源的问答对 数据集覆盖116个职业(21个SOC主要组),包含480个问题,来源自187个不同网站 测试15个前沿和开源模型,Claude Opus 4.6、GPT-5.4和Claude Sonnet 4.6表现最佳(58-62%),较小开源模型约33-41% 职业间性能差异显著:医疗相关职业最高(78%),办公室行政支持约40%,钣金工和鱼类野生动物管理员等职业几乎为零 开放式问题和按工资总额加权对模型排名无显著影响

60
Hot 热度
70
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • ORQA is a novel benchmark framework that evaluates LLMs on occupation-level professional knowledge by linking O*NET occupations to trusted, occupation-specific websites (regulatory agencies, licensing bodies, professional organizations) to generate source-traceable QA pairs
  • The benchmark covers 116 occupations across all 21 major SOC groups, with 480 questions sourced from 187 different websites, combining an automated pipeline with human review
  • Top frontier models (Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6) achieve ~58-62% accuracy, while smaller open-weight models score ~33-41%, revealing a significant performance gap
  • Performance varies dramatically across occupations: healthcare-related roles achieve 78% accuracy while Office/Administrative Support reaches only ~40%, and some occupations (e.g., Sheet Metal Workers, Fish and Game Wardens) score near zero
  • Open-ended questions and wage-bill weighting do not significantly affect model rankings, suggesting the benchmark is robust to these variations

Why It Matters

This benchmark addresses a critical gap in AI evaluation by moving beyond abstract skill assessments to measure real-world, occupation-specific professional knowledge grounded in authoritative sources. For AI practitioners and researchers, ORQA provides a scalable methodology for evaluating whether LLMs can perform credibly in professional domains, which is essential as these models are increasingly deployed in workplace settings. The findings also highlight which occupational domains LLMs are and are not ready to support, informing both deployment decisions and future research directions.

Technical Details

  • ORQA connects O*NET occupation classifications to trusted occupation-specific websites—including regulatory agencies, licensing bodies, professional organizations, and government publications—and converts this content into source-traceable question-answer pairs through an automated pipeline supplemented by human review
  • The dataset spans 116 occupations across all 21 major Standard Occupational Classification (SOC) groups, with 480 questions drawn from 187 distinct websites, each designed to probe real-world skill questions relevant to the target occupation
  • Evaluation was conducted on 15 state-of-the-art frontier and open-weight models, with performance measured as accuracy on the occupation-specific QA pairs
  • The study found that varying question format (open-ended vs. closed) and weighting questions by occupation wage bill did not significantly alter model rankings, indicating benchmark stability across these design choices

Industry Insight

  • The near-zero performance on certain occupations (e.g., Sheet Metal Workers, Fish and Game Wardens) suggests that LLMs have severe knowledge gaps in specialized or niche professional domains, which should caution against deploying these models in such roles without significant domain-specific fine-tuning or retrieval augmentation
  • The strong performance on healthcare-related occupations (78%) indicates that LLMs are better equipped for knowledge-intensive, text-heavy professional domains with abundant online authoritative sources, pointing to where current AI capabilities are most viable for professional assistance
  • The scalable methodology of leveraging existing trusted websites rather than relying on expensive expert annotation offers a practical blueprint for building future occupation-specific benchmarks, enabling continuous evaluation as both LLMs and occupational knowledge evolve

TL;DR

  • ORQA是一种评估LLM职业级别知识的新方法,通过将O*NET职业与可信的职业特定网站(监管机构、许可机构、专业组织、政府出版物)连接,生成可追溯来源的问答对
  • 数据集覆盖116个职业(21个SOC主要组),包含480个问题,来源自187个不同网站
  • 测试15个前沿和开源模型,Claude Opus 4.6、GPT-5.4和Claude Sonnet 4.6表现最佳(58-62%),较小开源模型约33-41%
  • 职业间性能差异显著:医疗相关职业最高(78%),办公室行政支持约40%,钣金工和鱼类野生动物管理员等职业几乎为零
  • 开放式问题和按工资总额加权对模型排名无显著影响

为什么值得看

ORQA提供了一种可扩展的职业级AI评估方法,解决了传统方法依赖抽象任务定义或昂贵专家知识的局限。该框架为衡量LLM在真实职业场景中的专业能力提供了可追溯、可量化的基准,对AI行业评估模型专业应用价值具有重要参考意义。

技术解析

  • 数据来源与构建:ORQA连接O*NET职业分类系统与可信职业特定网站(监管机构、许可机构、专业组织、政府出版物),通过自动化管道结合人工审核生成高质量问答对,确保问题具有真实职业场景相关性
  • 数据集规模:覆盖116个职业,来自21个SOC主要组,480个问题源自187个不同网站,每个问题设计用于测试与特定职业相关的真实技能问题
  • 模型评估:测试了15个前沿和开源模型,Claude Opus 4.6、GPT-5.4和Claude Sonnet 4.6表现最佳(58-62%),较小开源模型约33-41%
  • 职业差异分析:医疗相关职业表现最高(78%),办公室行政支持约40%,钣金工和鱼类野生动物管理员等职业性能几乎为零,显示模型在专业领域知识存在显著不均衡
  • 实验验证:研究发现开放式问题和按工资总额加权对模型排名无显著影响,验证了该基准的稳定性

行业启示

  • 职业级AI评估成为新趋势:ORQA方法为衡量LLM在专业领域的应用能力提供了可扩展的评估框架,未来职业级AI性能评估将成为行业关注重点
  • 模型专业能力存在显著不均衡:医疗领域表现优异而其他职业几乎为零,提示AI开发者需关注模型在特定职业领域的知识覆盖和训练数据偏差问题
  • 可追溯来源的评估方法更具可信度:通过连接可信职业网站生成问答对的方法,为AI评估提供了可验证、可追溯的基准构建范式,值得行业借鉴推广

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Benchmark 基准测试 Dataset 数据集 Research 科学研究