Simon Willison's 2026 LLM keynote: coding agents crossed the reliability line
dl_weekly · x · 2026-10-07
Simon Willison published annotated slides from his closing keynote at WeAreDevelopers World Congress NA, tracing 2026's LLM trends. Key turning point: Claude Opus 4.5 and GPT-5.1 (Nov 2025) pushed coding agents Claude Code and Codex from 'often make mistakes' to 'reliable enough for daily use'. He also tracked the pelican-on-a-bicycle SVG benchmark — as of November Claude still couldn't draw a proper bicycle — and the first commit to an obscure repo called Warelay.
More from AGI Musings
- Yudkowsky: no general intelligence until a mind proves Riemann, ABC, Collatz and Goldbach without chain-of-thought — aran_nayebi · 2026-10-07
- Hadfield-Menell: New 'research taste' benchmark measures metric hill-climbing, not taste — dhadfieldmenell · 2026-10-07
- 61% of 18 reviewed manuscripts had undisclosed deviations, technical errors, or heavy AI writing — paulnovosad · 2026-10-07
- Fields Medalist Daniel Litt warns LLM proofs could cause the thermal death of mathematics — littmath · 2026-10-07
- Former OpenAI and Anthropic researchers tell NYC Council humanity likely to lose control of advanced AI — fortune · 2026-10-07
- ~90% of frontier lab compute now goes to post-training and inference as pretraining corpus exhausts — rbhar90 · 2026-10-07