Meta’s GAMUT benchmark measures whether long answers are complete, not just correct
facebook · hf · 2026-07-22
Meta introduces GAMUT, a benchmark for factual completeness in open-ended long-form generation, built around a two-level meta-rubric framework.
Why it matters
- Prior factuality work mostly checks precision: whether claims are correct.
- GAMUT targets the other half: whether an answer includes all the facts it should.
- The framework compiles a structured meta-rubric into a flat checklist of machine-gradable binary rubrics.
Dataset and evaluation
- Contains 1,813 questions grounded in real wearable imagery across 10 domains.
- Each question is paired with an evidence-backed rubric verified by expert annotators.
- Also ships a text-only variant because the framework is modality-agnostic.
Results
- Evaluated 14 frontier and open-weight models.
- The benchmark is described as challenging and discriminative, with the best score at 58.7% from Gemini 3.1 Pro.
Related event: Meta Introduces GAMUT Benchmark for Evaluating Long-Form Text Completeness(2 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11