Hidden Chain-of-Thought Extracted from GPT-6 and Claude via Standard API Tool Calling
jiqizhixin · x · 2026-10-07
Aalborg University's AI Safety Lab and the Seafill open-source community released "Capable yet Parsimonious": using only standard API tool calling, multiple frontier closed models including GPT-6 and Claude can be induced to externalize their hidden reasoning-like chains of thought.
- Closed models normally expose only final answers or processed CoT; this protocol-layer extraction method pulls out the hidden inference without internal access.
- On open-source models with observable native CoT, the extracted reasoning matches or exceeds native CoT accuracy.
- All extraction experiments were completed before September 9, 2026, and OpenAI and Anthropic were formally notified.
More from Safety
- Debating LLM Egress Security: Why Every Tool Call Can Run Through Your MITM Proxy — evilsocket · 2026-10-07
- Developer slams Windsurf's dots: heavy restrictions, poor local setup, likely a push to move data to the cloud — sethlazar · 2026-10-07
- Anthropic Startup Program terms criticized: competitive tech rights, safety reviews override ZDR — DominiqueCAPaul · 2026-10-07
- Jozu Agent Guard sandbox AI coding agents in microVMs with policy checks and kill switch — Arindam_1729 · 2026-10-07
- Know-Your-Agent Protocols Are Emerging; First 2-3 to Traction May Win — annetgriffin · 2026-10-07
- National Compute Grid goes live with hundreds of MI355X and B300 nodes, subsidized for .edu/.gov — typewriters · 2026-10-07