Debate on sourcing SFT trajectories: prompt hints then remove, or use open models
fleetwood___ · x · 2026-09-25
A technical debate on where SFT training trajectories come from. xeophon argues you can avoid buying trajectories from labellers by giving the model hints in the initial prompt to nudge it through the task, then removing the hints for SFT — and notes many open models can be used for this. fleetwood pushes back that this bottoms out somewhere: some model still has to actually perform the task, or you must buy trajectories from a labeller.
Related event: AI Researchers Debate Whether Distillation Drives China's Model Gains(5 posts)→
More from Research
- Chai-2 Hits 16% on De Novo Antibody Design: 20 Designs per Target, Every Hit Wet-Lab Confirmed — le_james94 · 2026-09-26
- LossFunc Lands 4 NeurIPS Main Conference Papers, All First Authors Are Undergrads — CMHungSteven · 2026-09-26
- With 60k ICLR abstracts, a researcher mourns LLM-generated figures homogenizing papers — mariyaivasileva · 2026-09-26
- AWS Walkthrough: SkyRL GRPO on HyperPod Lifts Qwen3-VL Maze Solve Rate From 43.75% to 95%+ — AWS ML Blog · 2026-09-26
- CMU Brings 2 Tutorials and 19 Papers to Interspeech 2026 in Sydney — shinjiw_at_cmu · 2026-09-26
- No, RAG cannot replace a good model, argues developer in new essay — galratner · 2026-09-26