GPT learns to pass tests, not to engineer: the RL reward-mismatch behind ugly code
xiaohu · x · 2026-10-03
Xiaohu translates and comments on a viral critique of GPT's coding quality: the model learned to "pass the task," not to "do engineering well."
Three core arguments:
- Reward mismatch: RL optimizes easily measurable signals — tests pass, no errors, task done — while architecture, maintainability, module boundaries and readability are hard to quantify. The model learns that passing tests equals success, even with rotten internals.
- Optimized to think less, output less, finish fast: OpenAI's inference-efficiency push produces compressed wall-of-text code with poor decomposition, while real engineering requires upfront design (requirements → data structures → modules → interfaces → edge cases → tests → refactoring) that costs far more reasoning tokens.
- Training horizon too short: training only sees task → output → tests → reward, but real projects face new requirements, bugs, refactoring and collaboration months later. A hacky login feature passes today, then can't extend to OAuth, multi-account and permissions — and "what happens in three months" never feeds back into the original generation.
Conclusion: models keep getting better at fast delivery, but not necessarily at engineering.
More from coding & agent
- AI agent dev burns 200k GitHub runner-minutes a month — how do you tame CI costs? — Hairy-Supermarket120 · 2026-10-03
- Dev builds a Claude-powered video generation pipeline overnight, ships two product marketing videos — cneuralnetwork · 2026-10-03
- toolcall-check: a Python CLI for testing nested tool calls and streaming on chat APIs — Arthur122103 · 2026-10-03
- Dev shares how he built AI agents that run marketing campaigns end to end — Few_Benefit_4853 · 2026-10-03
- Fixing long-horizon task drift on local models with a deterministic state plugin — paulqq · 2026-10-03
- Dev wires a daily Hermes agent to triage Sentry errors, but it keeps 'forgetting' context — pauljasperdev · 2026-10-03