V4.1's Reasoning Curve Isn't Linear; Agent Teams Explicitly RL-Trained to Collaborate
nrehiew_ · x · 2026-09-11
Continuing his V4.1 breakdown, nrehiew notes the reasoning-performance plot is not fully linear (unlike FrontierCode's code-quality penalties, these three benches seem to lack similar penalties), and that the Agent team mode is explicitly trained with an RL reward combining task performance, a collaboration bonus for delegation and inter-agent communication, and a latency penalty for efficient coordination.
More from Models
- Forcing models to always max effort is like humans evolving on Adderall, researcher argues — voooooogel · 2026-09-11
- Pushing models to always show 'maximum effort' drags along its corollaries, dev argues — voooooogel · 2026-09-11
- Claude is the distillation target of choice because agentic RL seed data is scarce — teortaxesTex · 2026-09-11
- Why Chinese labs distill from Anthropic: Claude's agent data is the scarce training signal — teortaxesTex · 2026-09-11
- ApprenticeBench: closed model scores 72% vs open Kimi K3 at 18% on real jobs — ysu_nlp · 2026-09-11
- CursorBench 4.0 launches; Muse Spark 1.3 matches Sol at under 40% the cost — jyangballin · 2026-09-11