K2 Horizon eval breakdown: declines rather than guesses, 26% hallucination rate
ArtificialAnlys · x · 2026-09-03
Artificial Analysis published the full per-eval breakdown of K2 Horizon 375B A23B. Key figure: the model attempts just 40% of AA-Omniscience questions, declining the rest rather than guessing — yielding a 26% hallucination rate, among the lowest measured, with accuracy of 18% essentially unchanged from K2 Think V2 (71% hallucination). The AA-Omniscience Index improved from -40 to -3.
Related event: K2 Horizon Evaluation Details Show 26% Hallucination Rate, Agent Lead(4 posts)→
More from Models
- GPT-6 Astra and Astra Aeon spotted in Codex; Aeon tipped as long-horizon agent model — VraserX · 2026-09-03
- Ant Group Open-Sources Finance-Enhanced Model Ling-3.0-flash-Fin with 124B Parameters — niacolhealth · 2026-09-03
- Should 'hard scientific problems solved' be the new LLM benchmark? — Dr_Singularity · 2026-09-03
- Ben Thompson on Gemini: 'OK, it's now or never' — kieranklaassen · 2026-09-03
- User Reports ChatGPT Randomly Outputs 'GOD IS COMING' and Won't Stop — jaypro1005 · 2026-09-03
- Ling-3.0-flash-Fin Weights Released: 124B Params, 5.1B Active, 256K Context — Bestlife73 · 2026-09-03