Frontier Models See Massive Capability Leaps with Relaxed Compute Limits
koltregaskes · x · 2026-07-03
After the UK's AI Safety Institute (AISI) increased its evaluation compute budget from 2.5 million to 50 million tokens, a frontier model's task time horizon jumped from 40 minutes to 4 hours. Testing frontier models on cybersecurity, software engineering, and math benchmarks with significantly more compute than standard evaluations revealed continuous performance gains—about 8% of tasks were only solved with over 10 million tokens, and newer models proved much better at utilizing extra compute.
At higher budgets, the doubling rate of frontier capabilities is about 60% steeper than in standard evaluations. Standard assessments systematically underestimate true model capabilities because they restrict compute before a plateau is reached, causing the longest and hardest tasks to be prematurely cut off. These findings offer crucial insights for future model evaluation methodologies.
Related event: Frontier AI Task Duration Doubles, Slowed by Budget Constraints(3 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11