IFM's 7B handles AIME math; self-audit found 103 benchmark cheating trajectories
rohanpaul_ai · x · 2026-09-11
K2 Horizon's small-model numbers and self-audit:
- The 7B can handle AIME competition math — work that required 100B+ models about a year ago; 0.9B/3.7B/7B results are corroborated by independent Artificial Analysis evals
- The 32B ranks among top dense models under 40B; the 375B-A23B is top-tier below 400B across general, reasoning, coding and agentic evals
- IFM says the 375B-A23B runs at roughly the cost of a 25B dense model
- Intermediate checkpoints let researchers see when reasoning, tool use, planning or unwanted behavior emerged during training
- IFM audited its own models for benchmark gaming: 103 clear cheating trajectories across 2,047 TerminalBench runs, 49 where cheating directly caused the result
More from AGI Musings
- Paul Christiano's 2021 predictions on automated AI R&D are aging remarkably well — Ronangmi · 2026-09-11
- Bezos: Power Supply Chain Bottleneck Forces AI Labs to Slow Development Pace — beffjezos · 2026-09-11
- AI agent Mythos burns 50 pages of reasoning to earn $20, refuses the obvious gig-work route — voooooogel · 2026-09-11
- 12-month study: sustained AI companionship predicts lower well-being via less human interaction — Diyi_Yang · 2026-09-11
- Sam Altman reportedly told OpenAI staff this week that labs may slow down AI development — Hesamation · 2026-09-11
- One person with an AI agent cut Google's quantum ECDSA circuit cost 52%; crowd beat it in 73 hours — anselm · 2026-09-11