More compute can materially improve frontier models’ cyber benchmark performance
peterwildeford · x · 2026-07-21
- A thread argues that **more compute can mean better hacking** in frontier models. - The cited example is UK AISI’s **“The Last Ones”** cyber benchmark, where models behave very differently when given **10M tokens vs. 100M tokens** of budget. - The point is that cyber capability can change materially with extra inference budget, not just with model weights or prompt wording.
More from Research
- OCT-Bench sets 10,076 questions to test whether multimodal models really understand retinal scans — Baochen Fu · 2026-07-21
- LTX-2.3 face-and-voice LoRA training can work on 12GB VRAM with heavy tradeoffs — __alpha_____ · 2026-07-21
- Follow-up paper argues digital twins could make clinical trials more adaptive — techhalla · 2026-07-21
- Nature npj Digital Medicine paper maps causal inference and digital twins for trials — techhalla · 2026-07-21
- Nature NPJ Digital Medicine Explores Causal Inference and Digital Twins in Clinical Trials — MihaelaVDS · 2026-07-21
- AI performance is increasingly limited by materials science, not just compute — nordicinst · 2026-07-21