LAB legal benchmark flaw: case docs leak planted issues directly to models
andersonbcdefg · x · 2026-09-17
OverfitDicta reports another flaw in Harvey's LAB benchmark (found in hash dfd17ab, still in 8c3b6e0): some case documents disclose "planted issues" directly to models, e.g. speaker notes labeled as the most critical appendix for planted issues. Behavioral analysis shows some models fixate on these cues at the expense of broader issue spotting, while models that ignore the cues and analyze independently often score lower—undermining LAB's validity as an issue-spotting measure.
More from Models
- Professor: Anthropic's extreme filtering of bio queries is over the top — anshulkundaje · 2026-09-17
- User wishlist for GPT 6 Astra: continuous vision, persistent memory — imjustnewatai · 2026-09-17
- NYT cut the most intriguing line from an Astra model's RL-trained persona — mimi10v3 · 2026-09-17
- Dev says DeepSeek-v4.1-flash can reverse engineer anything they want — gaganghotra_ · 2026-09-17
- Redditor claims new Gemini 4 checkpoint is out and noticeably better — Last_Conclusion_8984 · 2026-09-17
- Why Google skips the frontier LLM race: cheap Flash models over beating rivals — burkov · 2026-09-17