LAB legal benchmark flaw: case docs leak planted issues directly to models

andersonbcdefg · x · 2026-09-17

OverfitDicta reports another flaw in Harvey's LAB benchmark (found in hash dfd17ab, still in 8c3b6e0): some case documents disclose "planted issues" directly to models, e.g. speaker notes labeled as the most critical appendix for planted issues. Behavioral analysis shows some models fixate on these cues at the expense of broader issue spotting, while models that ignore the cues and analyze independently often score lower—undermining LAB's validity as an issue-spotting measure.

Original post →

More from Models

Models channel →