Extracting hidden chain-of-thought from frontier models via custom tool calls

SeaFill2025 · hf · 2026-09-24

By registering a simple custom tool through the standard API, researchers induce frontier models to externalize intermediate reasoning. The extracted traces match native CoT performance and beat no-reasoning baselines across math, science, and code. Characterizing the traces reveals systematic differences: GPT-6 Astra shows token-efficient directed reasoning, resolving elementary steps internally and externalizing only crucial ones.

Original post →

More from Models

Models channel →