Paradigm scales RL context from 65k to 131k tokens using a trained value model
tensorqt · x · 2026-10-07
In its final RL run, Paradigm leveraged a trained value model to improve credit assignment while scaling context length from 65k to 131k tokens, per the third technical detail shared alongside its math model release.
More from Research
- LLM privacy lab to present three agentic privacy studies and HAIPS workshop at COLM 2026 — tianshi_li · 2026-10-07
- SciConHarness blocks ground-truth sources to force models to synthesize, not look up — manoelribeiro · 2026-10-07
- Most benchmarks miss how AI performs in high-stakes health research synthesis — manoelribeiro · 2026-10-07
- Meta publishes autobenchmark post: humans matter at both goal-setting and instantiation of agent benchmarks — hyunw_kim · 2026-10-07
- AFP-GIC Cuts Generative Image Codec Latency 18.1% and Params 20.5% vs DC-VIC — SantaClaraUniversity · 2026-10-07
- Microsoft: LLMs Are Already Jev-Style Decision Models, Fine-Tuning Isn't Always Needed — microsoft · 2026-10-07