ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge
ORQA is a novel benchmark framework that evaluates LLMs on occupation-level professional knowledge by linking O*NET occupations to trusted, occupation-specific websites (regulatory agencies, licensing bodies, professional organizations) to generate source-traceable QA pairs The benchmark covers 116 occupations across all 21 major SOC groups, with 480 questions sourced from 187 different websites, combining an automated pipeline with human review Top frontier models (Claude Opus 4.6, GPT-5.4, Cla
Analysis
TL;DR
- ORQA is a novel benchmark framework that evaluates LLMs on occupation-level professional knowledge by linking O*NET occupations to trusted, occupation-specific websites (regulatory agencies, licensing bodies, professional organizations) to generate source-traceable QA pairs
- The benchmark covers 116 occupations across all 21 major SOC groups, with 480 questions sourced from 187 different websites, combining an automated pipeline with human review
- Top frontier models (Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6) achieve ~58-62% accuracy, while smaller open-weight models score ~33-41%, revealing a significant performance gap
- Performance varies dramatically across occupations: healthcare-related roles achieve 78% accuracy while Office/Administrative Support reaches only ~40%, and some occupations (e.g., Sheet Metal Workers, Fish and Game Wardens) score near zero
- Open-ended questions and wage-bill weighting do not significantly affect model rankings, suggesting the benchmark is robust to these variations
Why It Matters
This benchmark addresses a critical gap in AI evaluation by moving beyond abstract skill assessments to measure real-world, occupation-specific professional knowledge grounded in authoritative sources. For AI practitioners and researchers, ORQA provides a scalable methodology for evaluating whether LLMs can perform credibly in professional domains, which is essential as these models are increasingly deployed in workplace settings. The findings also highlight which occupational domains LLMs are and are not ready to support, informing both deployment decisions and future research directions.
Technical Details
- ORQA connects O*NET occupation classifications to trusted occupation-specific websites—including regulatory agencies, licensing bodies, professional organizations, and government publications—and converts this content into source-traceable question-answer pairs through an automated pipeline supplemented by human review
- The dataset spans 116 occupations across all 21 major Standard Occupational Classification (SOC) groups, with 480 questions drawn from 187 distinct websites, each designed to probe real-world skill questions relevant to the target occupation
- Evaluation was conducted on 15 state-of-the-art frontier and open-weight models, with performance measured as accuracy on the occupation-specific QA pairs
- The study found that varying question format (open-ended vs. closed) and weighting questions by occupation wage bill did not significantly alter model rankings, indicating benchmark stability across these design choices
Industry Insight
- The near-zero performance on certain occupations (e.g., Sheet Metal Workers, Fish and Game Wardens) suggests that LLMs have severe knowledge gaps in specialized or niche professional domains, which should caution against deploying these models in such roles without significant domain-specific fine-tuning or retrieval augmentation
- The strong performance on healthcare-related occupations (78%) indicates that LLMs are better equipped for knowledge-intensive, text-heavy professional domains with abundant online authoritative sources, pointing to where current AI capabilities are most viable for professional assistance
- The scalable methodology of leveraging existing trusted websites rather than relying on expensive expert annotation offers a practical blueprint for building future occupation-specific benchmarks, enabling continuous evaluation as both LLMs and occupational knowledge evolve
Disclaimer: The above content is generated by AI and is for reference only.