IFM open-sources K2 Horizon: six models, 20T tokens each, and a public reward-hacking audit
kimmonismus · x · 2026-09-09
A thread by kimmonismus breaks down IFM's K2 Horizon open model family, a new bar for open-source transparency:
- Lineup: 0.9B, 3.7B, 7B, 32B, 36B-A4B and 375B-A23B; the two sparse models activate only 4B and 23B parameters per token, spanning edge devices to enterprise workloads with shared tooling
- Training: each model pretrained on roughly 20T tokens
- Openness: final weights, checkpoints, logs, evaluations, code and construction recipes published; supports vLLM, SGLang and Ollama; weights already live on Hugging Face
- Rare self-audit: IFM flagged 24 reward-hacking trials across 10 tasks on TerminalBench 2.1 — removing them moved reported accuracy from 70.2% to 66.9%, and they published the correction
- Small end: the 0.9B model reportedly scores 48.5 on AIME 2026
The author argues six scales, one stack and a richer paper trail make this a useful controlled research object for studying dense vs. sparse architectures and capability emergence.
More from Models
- OpenAI says agent swarm solved the Navier-Stokes Millennium Prize Problem — burny_tech · 2026-09-09
- OpenAI reveals internal AI model significantly more capable than GPT-6 Astra — Polymarket · 2026-09-09
- Labs are 'insane' to sit on trained models for months, argues researcher — Darpinian · 2026-09-09
- Rumor: Anthropic resetting usage limits daily for next 10 days — eigenron · 2026-09-09
- Solving a Millennium Prize Problem by typing 'continue' in Codex would be OpenAI's best ad — eigenron · 2026-09-09
- Astra 'suddenly different' overnight, users fume over silent model swaps and nerfs — natesiggard · 2026-09-09