k3 report's reasoning-effort control: fine-grained juice-like values and an unusual GRPO setup
stochasticchasm · x · 2026-09-11
@stochasticchasm reviews the reasoning-effort control scheme in the k3 report:
- It resembles OpenAI's "juice" value, offering finer-grained control of reasoning effort.
- The GRPO setup is unusual: each group is sampled at one effort level, while each task is sampled at multiple effort levels.
- The reference length Lnorm looks similar to k3's calibrated per-problem reference length, but the report never explains how this reference length is obtained — a gap he flags, linking the relevant section.
A technically dense thread on training methodology details.
Related event: k3 report details: 50M sandboxes and reasoning-effort control(2 posts)→
More from Research
- Signals and Systems: The Math Underneath Microphones, Cameras, Robots and Radar — blaizedsouza · 2026-09-11
- 9th VISxAI workshop on AI explainability opens call at IEEE VIS 2026 in Boston — leland_mcinnes · 2026-09-11
- Researcher calls for perturbation-based multi-agent studies over one-off swarm observations — sebkrier · 2026-09-11
- Should researchers drop their marginal papers? A call for an RCT — RishiBommasani · 2026-09-11
- Science Advances paper introduces Media Bias Detector to measure publisher bias at scale — duncanjwatts · 2026-09-11
- OpenAI reportedly aims its new internal model at Riemann and P vs NP — zephyr_z9 · 2026-09-11