ICML Winner FutureSim: GPT-5.6 Executes Over 15K Tool Calls in One Run

scaling01 · x · 2026-08-04

Recent updates from FutureSim, which won Best Paper at the ICML Forecasting Workshop, reveal that GPT-5.6-Sol leads in long-horizon tasks. The model can perform over 15,000 tool calls in a single run, taking more than a day to complete.

Furthermore, compared to Fable-5 and Claude models, GPT-5.6 demonstrates strong test-time adaptation, showing significant improvement during execution despite starting with similar initial accuracy.

Original post →

More from Models

Models channel →