1B model scores 94% on AIME 2026: Paradigma open-sources Limite 1B Violetto
tensorqt · x · 2026-10-07
Paradigma released the technical report and open-sourced (Apache 2.0) its first model, Limite 1B - Violetto, built to maximize mathematical reasoning per parameter and per joule.
- Scale & data: 1B parameters trained from scratch on a curated math-centric corpus of 237B tokens, with up to 131k context; 6 weeks from first experiments to release.
- Recipe: post-training combines distillation, targeted SFT and RLVR on a curriculum; pre-training interleaves SFT with DPPO RL, with architecture borrowed from nanogpt speedrun (NorMuon, MUDD variant, XSA).
- Results: 94.01% on AIME 2026 and 74.25% on BeyondAIME, matching models many times its size on their eval suite.
- Positioning: deliberately lightly instruction-tuned, meant as a single-turn math solver and a fast parallel component in larger systems rather than a general assistant.
- Model, base and value checkpoints, and an optimized serving implementation are released.
Related event: Paradigm Releases Violetto, Its First 1B Math Model(5 posts)→
More from Models
- Paradigm releases tech report for Violetto, a 1B model trained from scratch for math — tensorqt · 2026-10-07
- Mistral's Le Chonk tops blind human code review among open models, second only to Opus 5 — qtnx_ · 2026-10-07
- DeepSeek V4.1 Flash hits 72.9% on ARC-AGI-2 at $0.13/task, costing 250% more — teortaxesTex · 2026-10-07
- Mistral claims Large 4 is one of the world's strongest AI models for cybersecurity — scaling01 · 2026-10-07
- Mistral Large 4 solves 18 of 19 CTF challenges in official speedrun with tool calls — MistralAI · 2026-10-07
- Community poll tiers AI labs: Anthropic and OpenAI frontier, Mistral and Amazon judged 3 generations behind — NathanpmYoung · 2026-10-07