Kimi K3 may be strong on cyber, but token efficiency keeps it off UK AISIS
teortaxesTex · x · 2026-07-27
- The post highlights a discussion where Kimi K3 is described as potentially "mythos tier" for cyber.
- A key caveat is that it may be too reasoning-inefficient to show up on the UK AISIS eval, which is capped at 100M total tokens.
- The thread also notes other cyber-eval claims, including that K3 reportedly sits between Opus 4.8 and 5.6 Sonnet on Noah Lebovic’s eval.
- The broader takeaway: current models, open or closed, already appear capable of a lot in cyber, but eval design and token-budget constraints can hide that capability.
Related event: Kimi K3 Cyber Capabilities Underestimated Due to Token Efficiency(2 posts)→
More from Models
- French prize-winning novel suspected of AI: $1,000 challenge over detector results — Afinetheorem · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23