Models 'eat the harness' because they trained on harness-guided outputs — it's a self-fulfilling loop
arjunrajlab · x · 2026-10-05
Responding to the popular claim that models 'eat the harness', the author offers a subtler read: models likely bypass or exploit harnesses because their training data itself contains outputs guided by those harnesses.
It's a self-fulfilling loop — we collect harness-guided data, train on it, and the model then internalizes the harness's patterns. He half-jokingly wonders whether harness-based evals are still worth doing. A sharp observation on the entanglement between eval validity and training-data contamination.
More from Models
- Anthropic investigating elevated errors on Claude Mythos 5.1 and Fable 5.1 — ClaudeAI-mod-bot · 2026-10-05
- Qwen 3.8 keeps hallucinating it's out of context at 25% usage on local setup — TastesLikeOwlbear · 2026-10-05
- Unverified claim: GPT-6 Astra uses ~12k tokens per task vs Opus 5.5's ~66k — VraserX · 2026-10-05
- One Kimi hides among four Claudes: behavioral fingerprinting and hash commitments expose the mole — MarketingNetMind · 2026-10-05
- OpenRouter hits 146 trillion tokens a week, up 30x in a year — IndyDevDan · 2026-10-05
- Model release gap fell from 70 days to 11 — but Gary Marcus's profit-skepticism call scores correct — GaryMarcus · 2026-10-05