Guidance-TTT trains an 8B model at test time, beats best results on 4 discovery tasks
pliang279 · x · 2026-10-07
MIT-led researchers introduce Guidance-TTT, which separates idea generation from execution: a compact 8B guidance model is trained at test time via an adaptive group-relative RL objective to propose high-level strategic changes, while a frozen larger model turns them into executable solutions that are verified to reward the small model.
Key points:
- Avoids the cost of test-time training a large model (gradients, optimizer states, long structured outputs) and concentrates learning on a short strategy space
- Uses only open-weight models with no web access, and outperforms the best reported results on 4 scientific discovery tasks
- Paper (arXiv:2610.06269) and code are public
More from Research
- MEMOIR-VLM hits 91.3% balanced accuracy on dementia classification with missing brain-scan inputs — PTenigma · 2026-10-07
- Claude paper's prompter was an Anthropic employee, authors were just digesters, says user — analisereal · 2026-10-07
- Claude math result row: prompter was an Anthropic employee, authors say they were the "digesters" — analisereal · 2026-10-07
- ChipForge NPU Challenge Goes Live: Crowdsource AI Accelerator Designs Headed to Real Silicon — bittingthembits · 2026-10-07
- SIGGRAPH Asia 2026 Lineup: Disney's Markus Gross, Huawei's Ran Huang, Oxford's Andrea Vedaldi — AlexTensor · 2026-10-07
- CoreWeave RL Rollouts Hot-Loads Policy Weights Into Live Deployments, ~15x Faster Than Redeploys — _ScottCondron · 2026-10-07