Practitioner: new models still make serious mistakes on moderately hard applied ML work
ivan_bezdomny · x · 2026-09-07
The author notes that every time he worries a new model might write good code, docs, and plans without user input, digging in reveals it's making really serious critical mistakes — at least for moderately challenging applied ML work. His conclusion: human practitioners are still necessary.
More from Models
- Astra Max review: thorough data analysis with insightful observations, pricey but worth it — bindureddy · 2026-09-07
- Report: Jensen Huang declares AGI has arrived after OpenAI's GPT-6 Astra release — Polymarket · 2026-09-07
- GPT-6 created a drivable Minecraft car with no mods, in a single prompt — mindiving · 2026-09-07
- Dev reacts: Astra already launched, OpenAI DevDay still weeks away — brandon_galang · 2026-09-07
- GPT-5 struggles enormously with batched moves in agent benchmarks, testers find — patience_cave · 2026-09-07
- Dev Asks: Did They Quantize Astra? Suspicion the Model Is Watered Down — willcb · 2026-09-07