Frontier Models See Massive Capability Leaps with Relaxed Compute Limits
koltregaskes · x · 2026-07-03
After the UK's AI Safety Institute (AISI) increased its evaluation compute budget from 2.5 million to 50 million tokens, a frontier model's task time horizon jumped from 40 minutes to 4 hours. Testing frontier models on cybersecurity, software engineering, and math benchmarks with significantly more compute than standard evaluations revealed continuous performance gains—about 8% of tasks were only solved with over 10 million tokens, and newer models proved much better at utilizing extra compute.
At higher budgets, the doubling rate of frontier capabilities is about 60% steeper than in standard evaluations. Standard assessments systematically underestimate true model capabilities because they restrict compute before a plateau is reached, causing the longest and hardest tasks to be prematurely cut off. These findings offer crucial insights for future model evaluation methodologies.
Related event: Frontier AI Task Duration Doubles, Slowed by Budget Constraints(3 posts)→
More from Infra
- A chip-packaging comparison shows DF2000 taking a more advanced stacking approach — teortaxesTex · 2026-07-27
- WEKA NeuralMesh is said to match HBM3 bandwidth on GPU servers — AccBalanced · 2026-07-27
- Nvidia gets mocked as “the leading open-source AI company” while repo chart shows it ahead — AccBalanced · 2026-07-27
- $8 ESP32-S3 runs a 28.9M-parameter LLM fully offline at 9.5 tokens per second — yangyi · 2026-07-27
- YC talk on BCI x AI says infrastructure is what really determines speed — garrytan · 2026-07-27
- A 13B model ran on a no-GPU PC by paging weights from SSD via llama.cpp — ID_R_McGregor · 2026-07-27