Meta’s GAMUT benchmark measures whether long answers are complete, not just correct

facebook · hf · 2026-07-22

Meta introduces GAMUT, a benchmark for factual completeness in open-ended long-form generation, built around a two-level meta-rubric framework.

Why it matters

Dataset and evaluation

Results

Related event: Meta Introduces GAMUT Benchmark for Evaluating Long-Form Text Completeness(2 posts)→

Original post →

More from Research

Research channel →