AA Intelligence Index v4.2 lands as 300-person Moonshot's Kimi closes in on Google
Hesamation · x · 2026-09-05
Artificial Analysis shipped Intelligence Index v4.2 as an interim update ahead of v5: new agentic knowledge-work eval AA-Briefcase with a private test set, GDP.pdf long-context reasoning across 4,592 PDF pages, removal of saturated GPQA Diamond, plus heavier weighting on held-out tests to prevent gaming.
The accompanying chart sparked buzz: ThePrimeagen marveled that Moonshot's Kimi, from a 300-person company, is beating Google, a company spending $200B on AI this year.
More from Models
- GPT-6 Astra flunks complex PCB routing after 2h20m and 15% of weekly limits in biggest public test — yacineMTB · 2026-09-05
- Fable 5.1 medium effort matches Fable 5 high, no longer breaks prompt cache — lydiahallie · 2026-09-05
- GPT-6 Astra's first task uncovers two bugs from GPT5.6 Sol fix — op7418 · 2026-09-05
- OpenAI's 'yapping penalty' cuts Astra's HealthBench lead nearly in half after length adjustment — imjustnewatai · 2026-09-05
- GPT-6 Astra demoed redesigning board components and traces — jasonkneen · 2026-09-05
- Investor: GPT-6 Astra shows OpenAI and Anthropic are far ahead, Gemini 'benchmaxxed' — firstadopter · 2026-09-05