Writing evals with per-sample rubrics may be exploitable via subtle benchmax overfitting

teortaxesTex · x · 2026-09-30

sampaech flags a methodological flaw in LLM writing evals: hyper-specific per-sample grading rubrics add discriminative power but invite overfitting without technically cheating — models can learn heuristics like "if the prompt looks like p, pander to xyz criteria". Labs commonly benchmax by generating synthetic data near the test distribution and doing RL against the same grader, so the eval may end up measuring how well training discovered hidden grading objectives rather than writing ability. A possible counter-move: publish the rubrics. teortaxesTex calls this "real alpha" on natural writing evals.

Original post →

More from Models

Models channel →