Paper: Gavel reveals long-context failures in frontier LLMs
cocoweixu · x · 2026-08-28
Accepted to EMNLP 2026, the Gavel paper studies frontier LLMs on legal case summarization with inputs up to 500K tokens. Models struggle beyond 256K tokens and are more likely to omit key info than hallucinate. Even the best model, Gemini 2.5 Pro, scores only 50. The authors also built Gavel-Agent, an agentic, reference-free evaluator that reduces token usage by 36% using Qwen3.
More from Research
- Qwen3.8-Next paper: matches 397B predecessor with 1/9 the training FLOPs — NielsRogge · 2026-08-28
- How 'Making It Worse' Made $20M: The Tech Behind Bodycam — aakashgupta · 2026-08-28
- NLP sarcasm detection challenge: Who is building SarcasmBench? — paul_cal · 2026-08-28
- PILOT Enables Live Self-Improvement for Long-Horizon Agents via Skill Distillation — PolyUHK · 2026-08-28
- Skild vs GEN: Both Call It In-Context Learning, Very Different Paths — gan_chuang · 2026-08-28
- Paper analyzes thermal tuning overhead in optical interconnects for MoE training — jwt0625 · 2026-08-28