K2 Horizon's data recipe: 20T tokens per model, 10T synthetic, 17% reasoning traces
rohanpaul_ai · x · 2026-09-11
Details of K2 Horizon's training data:
- Each model pretrained on 20T tokens mixing web, code, math, science, multilingual and synthetic data
- 17% of the pre-training corpus contains explicit reasoning trajectories; 10T synthetic tokens were used in pre-training
- Post-training used 100M+ unique tasks, with the pipeline released via SFT, model merging, RL and specialized agent training
- The released pre-training dataset TxT360-v2 (5.3 TB) is already on Hugging Face
More from Models
- User: Opus 5 is "unusable" — it finds every way not to do what's asked even at max settings — RexDouglass · 2026-09-11
- Only Muse Spark 1.3 and Fable 5.1 sit on the coding Pareto frontier — jyangballin · 2026-09-11
- User slams Anthropic for blocking benign queries on ancient texts and recursive AI — NickPassig · 2026-09-11
- Codex cybersecurity work needs the Daybreak model to avoid safety-guardrail blocks — HankYeomans · 2026-09-11
- DeepSeek's new open-source model reportedly crushes GLM and Kimi at 4-10x lower prices — anselm · 2026-09-11
- GPT-6 Astra rebuilds Cessna 337 landing gear from a YouTube video; Fable 5.1 falls short — FinanceYF5 · 2026-09-11