Letting the Model Pick Its Own Thinking Effort: An Adaptive Harness Experiment
HeDo88TH · reddit · 2026-09-06
Prompted by the Qwen 3.8 27B thinking-levels debate, a Reddit developer suggests a harness modification that lets the LLM itself decide when to raise or lower reasoning effort. The system dynamically switches between low and xhigh based on success/failure streaks and task phase — dropping to low on stable success streaks or mechanical steps, escalating on meaningful failures or when entering DEBUG/RECOVER phases, with structured JSON logs tracking each transition. Early results show much faster task completion, though no reliable quality evaluation has been collected yet; he muses about before/after GPQA Diamond testing and asks whether anyone has tried something similar.
More from coding & agent
- Cognition rumored to launch a new model soon, per AI insider hunch — realsohamparekh · 2026-09-06
- Redditor builds free AskSary creative studio with game engine GPT can play and patch live — Beneficial-Cow-7408 · 2026-09-06
- Do the fun terminal work yourself, let Claude handle the boring chores — 4310sy · 2026-09-06
- Heavy AI user's cost ledger: 100+ agents a day, $70 of DeepSeek in two days — simmon_charlie · 2026-09-06
- First untrained agent run on local Qwen 3.8 Flash, no skills configured — jasonkneen · 2026-09-06
- Is a schema-aware memory graph 'overfitting'? Dev asks for the cleanest leakage test — chaachans · 2026-09-06