Paper: Gavel reveals long-context failures in frontier LLMs

cocoweixu · x · 2026-08-28

Accepted to EMNLP 2026, the Gavel paper studies frontier LLMs on legal case summarization with inputs up to 500K tokens. Models struggle beyond 256K tokens and are more likely to omit key info than hallucinate. Even the best model, Gemini 2.5 Pro, scores only 50. The authors also built Gavel-Agent, an agentic, reference-free evaluator that reduces token usage by 36% using Qwen3.

Original post →

More from Research

Research channel →