Extracting hidden chain-of-thought from frontier models via custom tool calls
SeaFill2025 · hf · 2026-09-24
By registering a simple custom tool through the standard API, researchers induce frontier models to externalize intermediate reasoning. The extracted traces match native CoT performance and beat no-reasoning baselines across math, science, and code. Characterizing the traces reveals systematic differences: GPT-6 Astra shows token-efficient directed reasoning, resolving elementary steps internally and externalizing only crucial ones.
More from Models
- NaceAI launches Drex, a sub-6B decision model that tops the public Decision Index at 51.73 — ordax · 2026-09-25
- Model audit showdown: Astra dominates, Opus and Fable close, Grok 4.7 and GPT-6 Sol lag far behind — ivan_bezdomny · 2026-09-25
- Uncensored local model Bonzai 2 27B tops benchmarks, runs on 12GB VRAM — alexcovo_eth · 2026-09-25
- Same prompt, Opus 5.5 one-shot video generation put to a public retest with different tools — drrickio · 2026-09-25
- Users blast GPT-5.2 for rampant false crisis flags and needless helpline redirects — ryunuck · 2026-09-25
- Agent Arena Ranks 43 Models on 2M+ Real-World Agentic Tasks; Claude Fable 5.1 Tops Board — arena · 2026-09-25