Meta’s GAMUT benchmark scores long answers on missing facts, and the best model gets 58.7%
rohanpaul_ai · x · 2026-07-24
Meta proposes GAMUT to measure what factual answers leave out
Meta’s new paper argues that factuality is not just about avoiding wrong claims — it also means including the important facts an answer should contain.
- GAMUT introduces a two-level rubric for evaluating open-ended generation.
- The framework first builds a structured meta-rubric covering required facts, acceptable choices, ordered steps, relationships, and importance.
- That structure is then compiled into machine-checkable pass/fail rubrics that an LLM judge can score more consistently.
- The benchmark includes 1,813 questions grounded in wearable imagery across 10 domains.
- Across 14 strong models, the best score was only 58.7%; missing information caused more failures than false statements.
- Meta says the setup is modality-agnostic and also released a text-only variant.
More from Research
- Dive into LLMs tutorial repo jumps to 44,837 GitHub stars — Lordog · 2026-07-24
- Hugging Face releases The Stack v3, a 114 TB open code corpus with train and full buckets — Nunki08 · 2026-07-24
- Bifrost says its real2sim pipeline can rebuild a site video into a simulation-ready 3D world in 30 minutes — jnack · 2026-07-24
- Masked Visual Actions turns 15 hours of robot video into a zero-shot world model — jbhuang0604 · 2026-07-24
- New subnet design lets miners compete on training data instead of weights — const_reborn · 2026-07-24
- Cohere Labs opens a 48-hour model challenge on language learning and reasoning — Cohere_Labs · 2026-07-24