Meta’s GAMUT benchmark measures whether long answers are complete, not just correct
facebook · hf · 2026-07-22
Meta introduces GAMUT, a benchmark for factual completeness in open-ended long-form generation, built around a two-level meta-rubric framework.
Why it matters
- Prior factuality work mostly checks precision: whether claims are correct.
- GAMUT targets the other half: whether an answer includes all the facts it should.
- The framework compiles a structured meta-rubric into a flat checklist of machine-gradable binary rubrics.
Dataset and evaluation
- Contains 1,813 questions grounded in real wearable imagery across 10 domains.
- Each question is paired with an evidence-backed rubric verified by expert annotators.
- Also ships a text-only variant because the framework is modality-agnostic.
Results
- Evaluated 14 frontier and open-weight models.
- The benchmark is described as challenging and discriminative, with the best score at 58.7% from Gemini 3.1 Pro.
Related event: Meta Introduces GAMUT Benchmark for Evaluating Long-Form Text Completeness(2 posts)→
More from Research
- A 2019 paper on five principles for AI in society resurfaces — ArtificialOther · 2026-07-23
- Fingerprint analysis says Kimi K3 and Fable 5 write more alike than sibling models — alex_verem · 2026-07-23
- A deep dive into how MCP tool calling works under the hood — jeffiql · 2026-07-23
- Open-source agent memory separates third-party claims from user facts — Deep-Thinker-01 · 2026-07-23
- CryptanalysisBench tests LLMs on 191 real cryptographic schemes — thegautamkamath · 2026-07-23
- Bittensor’s Swarm is training 30 drone-flight challenge combinations in parallel — bittingthembits · 2026-07-23