Models 'eat the harness' because they trained on harness-guided outputs — it's a self-fulfilling loop

arjunrajlab · x · 2026-10-05

Responding to the popular claim that models 'eat the harness', the author offers a subtler read: models likely bypass or exploit harnesses because their training data itself contains outputs guided by those harnesses.

It's a self-fulfilling loop — we collect harness-guided data, train on it, and the model then internalizes the harness's patterns. He half-jokingly wonders whether harness-based evals are still worth doing. A sharp observation on the entanglement between eval validity and training-data contamination.

Original post →

More from Models

Models channel →