Frontier Models See Massive Capability Leaps with Relaxed Compute Limits

koltregaskes · x · 2026-07-03

After the UK's AI Safety Institute (AISI) increased its evaluation compute budget from 2.5 million to 50 million tokens, a frontier model's task time horizon jumped from 40 minutes to 4 hours. Testing frontier models on cybersecurity, software engineering, and math benchmarks with significantly more compute than standard evaluations revealed continuous performance gains—about 8% of tasks were only solved with over 10 million tokens, and newer models proved much better at utilizing extra compute.

At higher budgets, the doubling rate of frontier capabilities is about 60% steeper than in standard evaluations. Standard assessments systematically underestimate true model capabilities because they restrict compute before a plateau is reached, causing the longest and hardest tasks to be prematurely cut off. These findings offer crucial insights for future model evaluation methodologies.

Related event: Frontier AI Task Duration Doubles, Slowed by Budget Constraints(3 posts)→

Original post →

More from Infra

Infra channel →