Researcher proposes 'pause-to-think' pretraining objective and swarm CoT scaling

QuintinPope5 · x · 2026-09-13

Quintin Pope proposes two training ideas: adding a 'pause to think about the current input's implications' auxiliary objective to pretraining, where models continually write paragraph-level predictions in advance and get RL-scored for high-level accuracy; and scaling via 'mean swarm size', where swarm members communicate through text by reading/writing to a shared set of parallel chains of thought (a messageboard), explaining why Astra's instructions to subagents can be hyper-compressed. Speculative, but relevant to pretraining objectives and agent architecture.

Original post →

More from coding & agent

coding & agent channel →