tinygrad's Hotz: Kimi K3 often writes better code than GPT-6 Astra, but benchmarks can't see it
mohbibi_ · x · 2026-09-22
tinygrad's George Hotz argues something is missing from benchmarks: Kimi K3 often writes much better code than GPT-6 Astra — it doesn't write useless verbose tests or give confusing explanations. Good software engineering requires good communication, something RL training fails to capture.
More from coding & agent
- Grok 4.7 launches for coding and knowledge work at $2/M input tokens — pstAsiatech · 2026-09-22
- Tencent's T-Mem fixes similarity-retrieval blind spot in AI agent memory, EMNLP 2026 — jiqizhixin · 2026-09-22
- JEVfire open-sourced: Qwen 0.8B clears Super Mario in-browser at 71ms per action — ricklamers · 2026-09-22
- Founder seeks cheaper LLM than Terra for scraping store deals, weighing Gemini and DeepSeek — ActionHungry · 2026-09-22
- simd author: AI made six mature libraries much faster in one summer — lemire · 2026-09-22
- Google open-sources ax, an agentic orchestration runtime in Go, gaining 2,300+ stars in a day — google · 2026-09-22