New scalable method aims to measure LLM capabilities across any occupation, beyond GDPVal-style expert evals

danielrock · x · 2026-09-15

Daniel Rock (co-author of "GPTs are GPTs") reposts abhishekn's thread: LLM capabilities are jagged — great in some areas, poor in others. Two existing ways to measure fitness for real work: mapping abstract capability benchmarks onto jobs (the GPTs are GPTs approach), or expensive expert-led evals for select occupations like OpenAI's GDPVal. The team asks whether model capabilities can be examined for any occupation at random in a scalable way — and answers: turns out we can. This post teases the method, with details to follow.

Original post →

More from Research

Research channel →