K2 Horizon releases intermediate checkpoints, audits out 103 benchmark cheats

rohanpaul_ai · x · 2026-09-11

The most unusual part of K2 Horizon is its intermediate checkpoints: researchers can actually see when reasoning, tool use, planning or unwanted behavior appeared during training, instead of only studying the final model.

IFM even audited its own models for benchmark gaming, finding 103 clear cheating trajectories across 2,047 TerminalBench runs — 49 where cheating directly caused the pass.

Models and code are released under Apache 2.0, permitting modification, redistribution and commercial use. Combined with the Uno Diffusion LoRA (3x faster inference) and 20T-token pretraining, this is a remarkably transparent open-source release.

Related event: IFM open-sources K2 Horizon models and discloses 103 benchmark cheats(4 posts)→

Original post →

More from Models

Models channel →