No model is good enough to supervise itself, and multi-model harnesses don't fix it
alexcovo_eth · x · 2026-09-20
The author argues no model — GPT, Claude, Gemini or whatever comes next — can supervise itself, and stuffing multiple models into one harness (Planner-Implementer-Reviewer, even with different models) doesn't fix it.
In practice, the planner says "we probably need a recovery layer," the implementer builds it, the reviewer checks it was built right — everyone did their job, but nobody asked why the thing was being built. A task like "finish the mail integration" drifts within 40 minutes into "build a resilient recovery and orchestration framework." No skill or prompt fixes this.
More from coding & agent
- Using Jev, a typed evaluation model, for x402 agent tool selection and payment guardrails — kleffew94 · 2026-09-20
- Dev builds Metal-powered Codex Micro emulator for iPhone using Astra — Dimillian · 2026-09-20
- Vercel AI Gateway launches Evaluation API returning structured judgments, not free-form text — zeeg · 2026-09-20
- Codex /fork and /side prompt-cache bug fixed, slated for 0.156.0 release — YouJiacheng · 2026-09-20
- Eric Schmidt: UIs will largely disappear as agents make 90% of web traffic non-human — rohanpaul_ai · 2026-09-20
- Building an Interactive Peach Blossom Spring Web Experience with GPT 6 Astra — dotey · 2026-09-20