antirez: judge new models by whether they fix real blocking bugs, not three.js demos
antirez · x · 2026-09-04
Redis creator antirez argues that demos of what new models can generate with three.js don't matter. What matters: when you hit a blocking problem in real software that old models couldn't fix over many rounds, does the newly released model actually improve the situation? He advocates evaluating models on real engineering blockers rather than flashy demos.
More from coding & agent
- GPT-6 Astra debuts at No.1 on Terminal-Bench, 1.9% ahead of Claude Fable 5.1 — sandersted · 2026-09-04
- Perplexity API lands in Stripe Projects: one CLI command provisions key and credits — jeff_weinstein · 2026-09-04
- Stripe's Link agent wallet gets official docs: one-time credentials let agents pay online — jeff_weinstein · 2026-09-04
- Reasoning effort switching without breaking cache is live in Codex and Claude — altryne · 2026-09-04
- Fighting semantic rot in agent memory: supersedes metadata plus weekly dedup sweeps — PennyLawrence946 · 2026-09-04
- GPT-6 early access review: multi-session collaboration, pruning AGENTS.md made it 100x better — patricksrail · 2026-09-04