Autoresearch for LLM pretraining needs humans in the loop; stacked changes are the killer

menhguin · x · 2026-10-12

Researcher menhguin shares hands-on experience running autoresearch for LLM pretraining, responding to iaindunning's observation that 'make this faster' works well while 'improve this frontier run' is deeply unimpressive. Key points: autoresearch strongly benefits from human-in-the-loop supervision and steering toward constraints; the killer is stacking combinatorial changes and hyperparameter sweeps — one tweak (e.g. sigmoid instead of softmax) may lower loss but its impact on logits, long context and post-training becomes unpredictable when combined. Designing a frontier-run improvement task means specifying desired state, mechanism and metric (with Exa search to anchor), but then it's not fully autonomous.

Related event: Researchers: AI auto-research speeds experiments but lags on frontier advances(3 posts)→

Original post →

More from coding & agent

coding & agent channel →