Meta’s GAMUT benchmark scores long answers on missing facts, and the best model gets 58.7%
rohanpaul_ai · x · 2026-07-24
Meta proposes GAMUT to measure what factual answers leave out
Meta’s new paper argues that factuality is not just about avoiding wrong claims — it also means including the important facts an answer should contain.
- GAMUT introduces a two-level rubric for evaluating open-ended generation.
- The framework first builds a structured meta-rubric covering required facts, acceptable choices, ordered steps, relationships, and importance.
- That structure is then compiled into machine-checkable pass/fail rubrics that an LLM judge can score more consistently.
- The benchmark includes 1,813 questions grounded in wearable imagery across 10 domains.
- Across 14 strong models, the best score was only 58.7%; missing information caused more failures than false statements.
- Meta says the setup is modality-agnostic and also released a text-only variant.
More from Research
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11