ORQA paper tests LLM knowledge across 116 occupations; top models score just ~60%
soumitrashukla9 · x · 2026-09-15
- New working paper ORQA (Occupation-Realistic QA) presents a scalable method to test LLM occupation-level knowledge: linking ONET occupations to trusted professional sites (regulators, licensing bodies) and auto-generating source-traceable QA pairs with human review. It covers 116 occupations across all 21 SOC groups: 480 questions from 187 websites.
- Testing 15 frontier and open-weight models: Claude Opus 4.6, GPT-5.4 and Claude Sonnet 4.6 lead at 58-62%, while smaller open-weight models reach only 33-41%.
- Performance varies sharply by occupation: healthcare-related roles score highest (78%), Office and Administrative Support lowest (40%). Paper and dashboard are public.
More from AGI Musings
- Microsoft's "it is not conscious" stance is logically incoherent, critic argues — rgblong · 2026-09-15
- "Twitter doesn't offer this filter": transformer-gatekeeping take stirs debate — generativist · 2026-09-15
- Can't sketch a transformer? Then stop weighing in on AI, researcher says — generativist · 2026-09-15
- Thom Wolf clashes with researcher over whether AI labs are only profit-driven — mervenoyann · 2026-09-15
- Researcher joins Stanford DigEconLab to simulate the economy with AI agents — soumitrashukla9 · 2026-09-15
- A 5-minute video of AI extinction without rogue AI: humans hand over control one step at a time — ronbodkin · 2026-09-15