Paradigm scales RL context from 65k to 131k tokens using a trained value model

tensorqt · x · 2026-10-07

In its final RL run, Paradigm leveraged a trained value model to improve credit assignment while scaling context length from 65k to 131k tokens, per the third technical detail shared alongside its math model release.

Original post →

More from Research

Research channel →