FULL STORY

MBZUAI's K2 Horizon: Open-Weighs Model Family Launches, First Benchmarks Land

MBZUAI released the fully open-source K2 Horizon model family on Sept 3, spanning 0.9B and larger models. Artificial Analysis benchmarks the same day showed the 375B model with only 26% hallucination and leading agent performance.

2026-09-03 ~ 2026-09-03 · 2 episodes · 10 posts

Episode 1 · MBZUAI's IFM Open-Sources K2 Horizon Family from 0.9B to 375B (2026-09-03, 6 posts)

On September 3, MBZUAI's Institute for Foundation Models (IFM) in the UAE released the K2 Horizon model family under the slogan "frontier performance, radically open source," comprising six models ranging from 0.9B to 375B parameters. The flagship K2-Horizon-375B (MoE, A23B) scored 47 on the Artificial Analysis Intelligence Index—roughly 30 points higher than its predecessor K2 Think V2, according to Artificial Analysis's evaluation. The release quickly drew attention for its exceptional degree of openness.

Confirmed

  • The K2 Horizon family consists of six models spanning 0.9B to 375B parameters, with the MoE-architecture K2-Horizon-375B A23B as the flagship, now available on Hugging Face.
  • The open-source release goes beyond weights: training code, data recipes, intermediate checkpoints, training logs, and evaluation results are all public. @testingcatalog noted the transparency is unusually high.
  • The flagship scored 47 on the Artificial Analysis Intelligence Index, an improvement of about 30 points over K2 Think V2.
  • @kimmonismus relayed that the series shows strengths in coding and intelligence tasks (details in the post are limited).

Why it matters

Multiple posters called this "the largest fully open-source release in AI history"—releasing not just model weights but the entire pipeline from training code to data recipes. If the community successfully reproduces or continues training from these materials, it could significantly lower the barrier to frontier model research and set a new transparency benchmark for the open-source ecosystem.

Episode 2 · K2 Horizon Evaluation Details Show 26% Hallucination Rate, Agent Lead (2026-09-03, 4 posts)

Artificial Analysis details show K2 Horizon 375B A23B answers only 40% of knowledge questions, refusing rather than guessing, cutting hallucinations to 26%, while leading agent tasks with a GDPval-AA Elo of 1430.