SpatialOPSD self-distills coding-agent traces so MLLMs reason spatially without tools
ATH-MaaS · hf · 2026-10-09
- Spatial coding agents boost MLLM spatial reasoning via tool-generated verified traces, but suffer heavy inference overhead and tool dependencies. The paper's key observation: prompting an MLLM with summarized agent traces unlocks its internal spatial chain-of-thought.
- SpatialOPSD is an on-policy self-distillation framework that internalizes these verified traces as privileged information into a standalone, tool-free MLLM; Repetition-Aware Distillation (repetition masking + unlikelihood regularization) mitigates privileged-information leakage.
- Across benchmarks, self-distillation beats SFT and GRPO on spatial and OOD data with superior generalization.
More from coding & agent
- IBM's RIT-RAG induces document sub-trees from retrieved chunks, lifting RAG accuracy by up to 11.4 points — _reachsumit · 2026-10-09
- Reddit thread rounds up open-source coding agents like MonkeyCode and OpenHands for local use — Good_Full_Tmes · 2026-10-09
- OpenAI's 10,000-agent, 130B-token run pushed slime v0.4.0 to rethink RL infrastructure scale — teortaxesTex · 2026-10-09
- Voice agent stack breakdown: Twilio, Deepgram, Cresta and ElevenLabs with a 300ms latency rule — schwentker · 2026-10-09
- Refund tools as verbs, not pens: how 'Veronica' stops LLM agents from inventing payees — schwentker · 2026-10-09
- Hackathon build 'Veronica': a voice agent that handles 'you charged me 10x' calls — schwentker · 2026-10-09