Novel Reasoning Effort Control Scheme Analyzed: Graded GRPO Training

stochasticchasm · x · 2026-09-11

stochasticchasm analyzes a novel reasoning effort control scheme, likening it to OpenAI's "juice" value for fine-grained control: each GRPO group samples at one effort level while each task samples at multiple levels, with an "Lnorm" reference length similar to k3's calibrated per-problem lengths — and suspects batch-invariant inference explains v4's throughput gains.

Related event: DeepSeek V4.1 Tech Report Deep Dive: KV Compression and Numeric Reasoning Effort Steal the Show(9 posts)→

Original post →

More from Models

Models channel →