tinygrad's Hotz: Kimi K3 often writes better code than GPT-6 Astra, but benchmarks can't see it

mohbibi_ · x · 2026-09-22

tinygrad's George Hotz argues something is missing from benchmarks: Kimi K3 often writes much better code than GPT-6 Astra — it doesn't write useless verbose tests or give confusing explanations. Good software engineering requires good communication, something RL training fails to capture.

Original post →

More from coding & agent

coding & agent channel →