No model is good enough to supervise itself, and multi-model harnesses don't fix it

alexcovo_eth · x · 2026-09-20

The author argues no model — GPT, Claude, Gemini or whatever comes next — can supervise itself, and stuffing multiple models into one harness (Planner-Implementer-Reviewer, even with different models) doesn't fix it.

In practice, the planner says "we probably need a recovery layer," the implementer builds it, the reviewer checks it was built right — everyone did their job, but nobody asked why the thing was being built. A task like "finish the mail integration" drifts within 40 minutes into "build a resilient recovery and orchestration framework." No skill or prompt fixes this.

Original post →

More from coding & agent

coding & agent channel →