Meta’s GAMUT benchmark says top models still miss half the needed facts

dair_ai · x · 2026-07-22

Meta’s GAMUT benchmark finds models still miss factual completeness

Meta researchers introduced GAMUT, a benchmark focused not on whether answers are correct, but on whether they are complete.

What the paper proposes

Benchmark details

Main result

Across 14 frontier and open models, the best score is only 58.7%.

The authors say the benchmark remains challenging and that results are stable even when different judge models are used.

Why it matters

The paper argues that factuality is not just precision; coverage is the part many systems fail at. The rubric-compilation trick may also transfer to other tasks with hierarchical requirements.

Related event: Meta Introduces GAMUT Benchmark for Evaluating Long-Form Text Completeness(2 posts)→

Original post →

More from Research

Research channel →