Meta’s GAMUT benchmark scores long answers on missing facts, and the best model gets 58.7%

rohanpaul_ai · x · 2026-07-24

Meta proposes GAMUT to measure what factual answers leave out

Meta’s new paper argues that factuality is not just about avoiding wrong claims — it also means including the important facts an answer should contain.

Original post →

More from Research

Research channel →