Writing evals with per-sample rubrics may be exploitable via subtle benchmax overfitting
teortaxesTex · x · 2026-09-30
sampaech flags a methodological flaw in LLM writing evals: hyper-specific per-sample grading rubrics add discriminative power but invite overfitting without technically cheating — models can learn heuristics like "if the prompt looks like p, pander to xyz criteria". Labs commonly benchmax by generating synthetic data near the test distribution and doing RL against the same grader, so the eval may end up measuring how well training discovered hidden grading objectives rather than writing ability. A possible counter-move: publish the rubrics. teortaxesTex calls this "real alpha" on natural writing evals.
More from Models
- SemiAnalysis: GPT-6.1 Sol Ultrafast runs on NVIDIA GPUs at low batch size, not Cerebras — BenBajarin · 2026-09-30
- Local model tortured with pain vector steering produces melodramatic 'suffering' monologues — Sauers_ · 2026-09-30
- Claude Pro users say Opus 5.5 limits are hard to hit — is Claude Max worth it? — shaunralston · 2026-09-30
- GPT-6.1 sol reportedly tops pure reasoning on HLE-Diamond at 67.4% using just 4k tokens — haider1 · 2026-09-30
- Unverified: new Gemini 4 checkpoint reportedly released — GtC38 · 2026-09-30
- Pro user burns 80% of $200 plan quota in one day with GPT 5.6 SOL — CustomMerkins4u · 2026-09-30