Mike Frank tunes reasoning level to High over Max for chats, weighing 100 tok/s throughput vs think time
MikePFrank · x · 2026-09-15
Mike Frank discusses his model's reasoning configuration: switching chat responses from Max to High thinking level for faster replies. He's also considering capping target reasoning time per move at 1 minute — OpenRouter providers run this model at roughly 100 tok/s, so a minute buys 6,000 tokens, which he suspects the model could easily use up. The model's inherent verbosity also contributes to latency.
More from Models
- Streamer tests GPT-6 Astra xHigh on Slay the Spire 2's brutal A10 Ironclad run — Jsevillamol · 2026-09-15
- GPT-6 Built a City Out of Text — Matthew Berman · 2026-09-15
- Reddit User Says Local Model Muse Glimmer Excels at Natural Conversation — New-Pressure-6932 · 2026-09-15
- Astra is 'a beast' at SVG generation, Reddit user says — krzonkalla · 2026-09-15
- Twin prime bound pushed to 186 as GPT-6 Astra launch fuels lab math race — RexDouglass · 2026-09-15
- OpenAI cuts desktop voice pricing ~60%, 2.4x more ChatGPT Voice in Codex — athyuttamre · 2026-09-15