K2 Horizon declines rather than guesses, cutting hallucination to 26%
ArtificialAnlys · x · 2026-09-03
Artificial Analysis notes K2 Horizon 375B A23B attempts just 40% of AA-Omniscience questions and declines the rest rather than guessing, producing a 26% hallucination rate — among the lowest measured — with accuracy of 18% essentially unchanged. This is a huge improvement over K2 Think V2's 71%, moving the AA-Omniscience Index from -40 to -3.
Related event: K2 Horizon Evaluation Details Show 26% Hallucination Rate, Agent Lead(4 posts)→
More from Models
- Google Launches WeatherNext 3, Its Most Accurate AI Weather Model — Scobleizer · 2026-09-04
- Don't Fear AI Subscription Price Hikes: It's the Most Competitive Market on Earth — brandon_galang · 2026-09-04
- Google Gemini suffers a widespread outage, users report LLMs down — moonsandhues · 2026-09-04
- GPT-6 Astra and Astra Aeon spotted in Codex; Aeon tipped as long-horizon agent model — VraserX · 2026-09-03
- Ant Group Open-Sources Finance-Enhanced Model Ling-3.0-flash-Fin with 124B Parameters — niacolhealth · 2026-09-03
- Should 'hard scientific problems solved' be the new LLM benchmark? — Dr_Singularity · 2026-09-03