Meta’s GAMUT benchmark says top models still miss half the needed facts
dair_ai · x · 2026-07-22
Meta’s GAMUT benchmark finds models still miss factual completeness
Meta researchers introduced GAMUT, a benchmark focused not on whether answers are correct, but on whether they are complete.
What the paper proposes
- A two-level meta-rubric that captures structure, importance, open-ended sets, ordered procedures, and fact relationships.
- The rubric is mechanically compiled into a flat checklist of binary items that LLM judges can score more reliably.
- The approach is designed for long-form open-ended generation, where simple fact-checking misses what users actually care about.
Benchmark details
- 1,813 questions grounded in real wearable imagery.
- Coverage across 10 domains.
- Each question has an expert-verified rubric.
- A text-only variant is also released.
Main result
Across 14 frontier and open models, the best score is only 58.7%.
The authors say the benchmark remains challenging and that results are stable even when different judge models are used.
Why it matters
The paper argues that factuality is not just precision; coverage is the part many systems fail at. The rubric-compilation trick may also transfer to other tasks with hierarchical requirements.
Related event: Meta Introduces GAMUT Benchmark for Evaluating Long-Form Text Completeness(2 posts)→
More from Research
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Five tells that still make AI video read as AI, from physics glitches to missing operators — NewPhoneWhotiz · 2026-09-11