Scaffold CoT: a 4M-example structured reasoning dataset built for small models' free-form CoT failures
Saraozte01 · reddit · 2026-09-03
The author released Scaffold CoT, a 4M-example (3B token) dataset designed to fix free-form CoT failures in models under 5B parameters. Key points:
- Fixed scaffold: every example uses the same three-section think block — Inventory → Interaction → Execution — so models spend capacity on reasoning, and output validity can be regex-checked
- Depth tiers: four tiers with 5.6x spread so models learn when to think longer vs. shorter
- 18 domains, 798 fully-labelled subdomains: code (670k), general (644k), antihal anti-hallucination/calibration (502k), science, logic, strategy, tooluse with 87k really-executed tool chains, selfcheck, mathqual, formatstrict, and more
- Max 2048 tokens per example for consumer-hardware fine-tuning
Experiments are ongoing; the dataset is released early for the community.
More from Research
- 'Transformers are samplers, not runtimes' — why residual streams can't host structured computation — gerardsans · 2026-09-03
- Information Bottleneck podcast to livestream Jonas Geiping on recurrent-depth models — ziv_ravid · 2026-09-03
- ICLR 2027's deanonymization policy wins fans; researcher urges slashing submission cap to 5 — peter_richtarik · 2026-09-03
- motion-bricks.cpp hits GitHub: tiny model generates animations in realtime on a desktop CPU — zhengyiluo · 2026-09-03
- Mostik bridges frontier and small models in latent space, tops ARC-AGI 3 at 1/20th the cost — SimplyAnnisa · 2026-09-03
- Mechanism Design Could Shape AI Behavior Without Understanding Neural Nets, Argues Economist — morqon · 2026-09-03