Kimi K3 may have been under-elicited in a 100M-token UK AISI eval
morqon · x · 2026-07-27
The post argues that Kimi K3 may have been under-elicited by the UK AISI eval setup.
- The UK AISI budget was 100M tokens (including cache hits), but the writer says scores did not clearly plateau before that point.
- They describe Kimi K3 as token-inefficient and suggest that with 5× more budget it might have performed much better.
- In the writer’s own pentesting eval, Kimi K3 reportedly landed between Opus 4.8 and GPT 5.6 Sol.
- Their conclusion: it is not a drop-in replacement, but it looks like a strong base for more post-training.
Related event: Kimi K3 Cyber Capabilities Underestimated Due to Token Efficiency(2 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23