Paradigma's 1B model scores 74.25% on BeyondAIME, trained on under 300B tokens

tensorqt · x · 2026-09-22

Paradigma released Limite 1B (Violetto), a 1B-parameter dense autoregressive transformer trained from scratch on under 300B curated tokens with 131k context. Using synthetic data, curated SFT and RL post-training — plus an architecture inspired by pre-training speedrun advances — it averages 74.25% on BeyondAIME, beating the 30B MUSE-Glimmer at 70%. Deliberately lightly instruction-tuned for single-turn math use, it ships with model weights, a training value model, and a custom vLLM inference plugin; a tech report is coming.

Original post →

More from Models

Models channel →