Guidance-TTT trains an 8B model at test time, beats best results on 4 discovery tasks

pliang279 · x · 2026-10-07

MIT-led researchers introduce Guidance-TTT, which separates idea generation from execution: a compact 8B guidance model is trained at test time via an adaptive group-relative RL objective to propose high-level strategic changes, while a frozen larger model turns them into executable solutions that are verified to reward the small model.

Key points:

Original post →

More from Research

Research channel →