Dev's verdict: Claude forgets tasks, Codex is too literal and adds unrequested extras
rickasaurus · x · 2026-09-12
Developer rickasaurus summed up his experience with two leading coding models: Claude tends to forget to do things or misinterpret instructions, while Codex is super literal and sometimes adds extra stuff you never asked for.
More from Models
- Perplexity trusts GPT-6 Astra with end-to-end production systems — OpenAI News · 2026-09-12
- Specific Labs Launches Real-SWE, a Benchmark on Private Enterprise Codebases — zainhas · 2026-09-12
- GLM-5.3 Scores 28.8% on New RealSWE Benchmark, Closing In on GPT-6 Astra — zainhas · 2026-09-12
- Sentry CEO on Meta's Muse Spark: fast and pleasant but skimps on reasoning, needs heavy hand-holding — zeeg · 2026-09-12
- Open models flop on Terminal-Bench Science: best scores just 4/70 — teortaxesTex · 2026-09-12
- Anthropic Publishes Most Detailed Threat Intel Report Yet, Says It Disrupted Every Claude Misuse Case — whurley · 2026-09-12