Kimi K3 may be strong on cyber, but token efficiency keeps it off UK AISIS
teortaxesTex · x · 2026-07-27
- The post highlights a discussion where Kimi K3 is described as potentially "mythos tier" for cyber.
- A key caveat is that it may be too reasoning-inefficient to show up on the UK AISIS eval, which is capped at 100M total tokens.
- The thread also notes other cyber-eval claims, including that K3 reportedly sits between Opus 4.8 and 5.6 Sonnet on Noah Lebovic’s eval.
- The broader takeaway: current models, open or closed, already appear capable of a lot in cyber, but eval design and token-budget constraints can hide that capability.
More from Models
- Poster says Kimi’s near-term outlook depends on a K3 base model release — _xjdr · 2026-07-27
- European ChatGPT Plus users are now seeing an “Extra High” quality option — PressPlayPlease7 · 2026-07-27
- Opus 5 reportedly aces a car-racing game test on the first try — soumitrashukla9 · 2026-07-27
- Claude Opus 5 arrives at half the price and tops Frontier-Bench claims — GregCook2011 · 2026-07-27
- Open models may beat closed ones for cyber defense, researchers argue as Kimi K3 impresses — eliebakouch · 2026-07-27
- Opus 5 notices when its own generated game looks bad — Angaisb_ · 2026-07-27