New scalable method aims to measure LLM capabilities across any occupation, beyond GDPVal-style expert evals
danielrock · x · 2026-09-15
Daniel Rock (co-author of "GPTs are GPTs") reposts abhishekn's thread: LLM capabilities are jagged — great in some areas, poor in others. Two existing ways to measure fitness for real work: mapping abstract capability benchmarks onto jobs (the GPTs are GPTs approach), or expensive expert-led evals for select occupations like OpenAI's GDPVal. The team asks whether model capabilities can be examined for any occupation at random in a scalable way — and answers: turns out we can. This post teases the method, with details to follow.
More from Research
- ECCV 2026: DF3DV-1K dataset and DI²FIX benchmark for distractor-free novel view synthesis — rsasaki0109 · 2026-09-15
- Dev implements Gaussian splatting as a single from-scratch Rust megakernel via cuda-oxide — bilawalsidhu · 2026-09-15
- Zhejiang University's Mechanist: AI studies its own mechanisms like a neuroscientist — jiqizhixin · 2026-09-15
- GlossoGen: new framework studies when LLM agents develop languages humans can't understand — lucy3_li · 2026-09-15
- Researcher open-sources theseus, a human-language architecture research framework — pratyusha_PS · 2026-09-15
- Article argues intelligence has a speed limit, curbing recursive self-improvement hype — docmilanfar · 2026-09-15