Meta’s GAMUT benchmark says top models still miss half the needed facts
dair_ai · x · 2026-07-22
Meta’s GAMUT benchmark finds models still miss factual completeness
Meta researchers introduced GAMUT, a benchmark focused not on whether answers are correct, but on whether they are complete.
What the paper proposes
- A two-level meta-rubric that captures structure, importance, open-ended sets, ordered procedures, and fact relationships.
- The rubric is mechanically compiled into a flat checklist of binary items that LLM judges can score more reliably.
- The approach is designed for long-form open-ended generation, where simple fact-checking misses what users actually care about.
Benchmark details
- 1,813 questions grounded in real wearable imagery.
- Coverage across 10 domains.
- Each question has an expert-verified rubric.
- A text-only variant is also released.
Main result
Across 14 frontier and open models, the best score is only 58.7%.
The authors say the benchmark remains challenging and that results are stable even when different judge models are used.
Why it matters
The paper argues that factuality is not just precision; coverage is the part many systems fail at. The rubric-compilation trick may also transfer to other tasks with hierarchical requirements.
Related event: Meta Introduces GAMUT Benchmark for Evaluating Long-Form Text Completeness(2 posts)→
More from Research
- Helmholtz Munich and TUM open fully funded PhD spots in trustworthy AI for science — zeynepakata · 2026-07-22
- Paper frames AI alignment as a moving sociotechnical target — weballergy · 2026-07-22
- AI coding tools are making fast social-science foresight experiments much cheaper — weballergy · 2026-07-22
- Researchers warn static alignment could cause value lock-in and societal stagnation — weballergy · 2026-07-22
- A population model suggests AI alignment can slow social progress under strong lock-in — weballergy · 2026-07-22
- Paper argues AI alignment breaks when human values keep evolving — weballergy · 2026-07-22