Berkeley & DeepMind paper: Abstract Token Curriculum lets models build internal scratchpads without CoT
rbhar90 · x · 2026-09-24
A paper from Berkeley and DeepMind challenges our reliance on Chain of Thought: we burn enormous compute and human effort hand-crafting step-by-step reasoning tokens just to nudge transformers toward the right answer.
Instead, the paper introduces an Abstract Token Curriculum (ATC) — the model is never forced to emit human-readable tokens, but is fed problems through steadily harder distributions, forcing it to invent its own continuous internal scratchpads to bridge the gap.
On parity learning they prove that single-layer softmax attention naturally gravitates toward intermediate representations, showing this latent reasoning mechanism emerges on its own.
More from Research
- NYU launches Mathematics in the Age of AI seminar, Buckmaster's inaugural talk packed — thegautamkamath · 2026-09-24
- Tencent ARC releases GAE: geometry-native latents halve camera error, cut FVD up to 23% — CSProfKGD · 2026-09-24
- OpenRSI founders on why RSI is the missing piece of ASI, open to everyone — ChengleiSi · 2026-09-24
- OpenRSI-Index calls for domain leads and compute partners to scale open RSI — ChengleiSi · 2026-09-24
- Two Critical Steps Toward RSI: Better Research Answers vs. Improving the Research System Itself — ChengleiSi · 2026-09-24
- Team Launches Tool to Turn Research Projects into AI-Usable Environments — ChengleiSi · 2026-09-24