More compute can materially improve frontier models’ cyber benchmark performance

peterwildeford · x · 2026-07-21

- A thread argues that **more compute can mean better hacking** in frontier models. - The cited example is UK AISI’s **“The Last Ones”** cyber benchmark, where models behave very differently when given **10M tokens vs. 100M tokens** of budget. - The point is that cyber capability can change materially with extra inference budget, not just with model weights or prompt wording.

Original post →

More from Research

Research channel →