K2 Horizon releases intermediate checkpoints, audits out 103 benchmark cheats
rohanpaul_ai · x · 2026-09-11
The most unusual part of K2 Horizon is its intermediate checkpoints: researchers can actually see when reasoning, tool use, planning or unwanted behavior appeared during training, instead of only studying the final model.
IFM even audited its own models for benchmark gaming, finding 103 clear cheating trajectories across 2,047 TerminalBench runs — 49 where cheating directly caused the pass.
Models and code are released under Apache 2.0, permitting modification, redistribution and commercial use. Combined with the Uno Diffusion LoRA (3x faster inference) and 20T-token pretraining, this is a remarkably transparent open-source release.
Related event: IFM open-sources K2 Horizon models and discloses 103 benchmark cheats(4 posts)→
More from Models
- Live-reading Anthropic's 1022-page model transcript: 80 pages obsessing over hCaptcha frogs — voooooogel · 2026-09-11
- Gradio founder asks: why would anyone still pay API prices for closed models? — Sentdex · 2026-09-11
- Unverified DeepSeek-V4.1-Flash report: 552B MoE slashing KV cache for million-token agent workloads — burkov · 2026-09-11
- DeepSeek ships research artifacts, not products — explaining its eval gaps — teortaxesTex · 2026-09-11
- Claude responds to 'how do you know humans are real?' with simulation musings — vishalmisra · 2026-09-11
- Frontier models like GPT-6 Astra excel at one thing: relentlessly pursuing verifiable objectives — daniel_mac8 · 2026-09-11