Armin Ronacher: reasoning levels live in the system prompt, and switching them busts your KV cache
mitsuhiko · x · 2026-09-06
Armin Ronacher points out a non-obvious mechanism: reasoning effort is baked into the system prompt for most models, so switching levels mid-session invalidates KV caches. The main exception is Fable 5.1, which allows mid-conversation reasoning switches. His linked post 'What Is Reasoning' explains that reasoning traces are just ordinary text emitted into a scratchpad (visible via GPT-OSS's Harmony analysis/final channels), that early token-budget APIs made effort look like a sampling property, and that his curiosity was sparked by a paper on extracting reasoning traces from closed-weight models.
Related event: Switching reasoning levels invalidates KV cache, developers find(3 posts)→
More from coding & agent
- ETH Zurich tested 100 developers: written communication predicts vibe coding success — _akpiper · 2026-09-06
- Image + video + agents combined: new step in agentic video creation opens for feedback — Lianhuiq · 2026-09-06
- Ghost agent watches your PC 24/7, proposes fixes, and deploys after your voice approval — BLUECOW009 · 2026-09-06
- Marketing firm with 2,000 engineers caps AI coding plans at $200/month — Crafty-Sugar2507 · 2026-09-06
- 8 software books that matter more now that AI writes the code — bibryam · 2026-09-06
- Benchmarking 8 public remote MCP servers: latency, schema token bloat, silent 500s — FreeMembership5107 · 2026-09-06