IFM's 7B K2-Horizon nearly matches 27B models, tops GPT-5 on BrowseComp, fully open-sourced
karminski3 · x · 2026-09-16
IFM has released K2-Horizon-7B, a 7B dense model that scores 21 on AA Bench—nearly matching Qwen3.6-27B's 22—and beats prior flagships like GPT-5 and DeepSeek-V4 on BrowseComp, with strong SWE-Bench Verified and Terminal Bench results.
Key details:
- Fully open: training data, recipe, training code, and evals are all public.
- Context: up to 512K tokens, but full attention makes long context expensive—128K needs 18GB of VRAM/unified memory at BF16, and only BF16 weights exist for now (quantized versions expected).
- Reasoning effort: low/medium/high settings, with high recommended; vLLM and SGLang already support it.
The author also nudges Qwen to ship its rumored Qwen3.8-35B-A3B and previews a sub-10B small-model comparison including MiniCPM5-2B.
More from Models
- Speculation: top open-weight models may be distilling OpenAI and Anthropic, missing training code hints — dan_s_becker · 2026-09-16
- Science LLM benchmarks have flawed answers; fixing them significantly raises model scores — Profanion · 2026-09-16
- Aaronson hears AI companies have cracked longstanding TCS open problems, sitting on major announcements — scaling01 · 2026-09-16
- Prior Labs releases TabPFN-3.5, a tabular foundation model handling 1M rows, tops benchmarks — tuanacelik · 2026-09-16
- User reports Sonnet 5 fails at shop price comparisons and hallucinates local stores — RachelVT42 · 2026-09-16
- One-shot two-hour run: Opus 5.2 builds a sprawling Voxel megacity demo — Available_Yam_6267 · 2026-09-16