Mike Frank tunes reasoning level to High over Max for chats, weighing 100 tok/s throughput vs think time

MikePFrank · x · 2026-09-15

Mike Frank discusses his model's reasoning configuration: switching chat responses from Max to High thinking level for faster replies. He's also considering capping target reasoning time per move at 1 minute — OpenRouter providers run this model at roughly 100 tok/s, so a minute buys 6,000 tokens, which he suspects the model could easily use up. The model's inherent verbosity also contributes to latency.

Related event: Mike Frank Tunes Model Reasoning Time: 1.5 Minutes per Chess Move, High for Chat(2 posts)→

Original post →

More from Models

Models channel →