A concrete path past the data wall: train agents on interaction trajectories, not just text
TuXinming · x · 2026-09-09
Responding to complaints that scaling talk is fortune-telling, TuXinming proposes a concrete loop: pretraining compresses world→observation→text, but tools offer act→feedback→train-on-trajectories. Labs already train on tool-use/agent traces — e.g., agents querying AlphaGenome to refine hypotheses; swap tools for real experiments and feedback comes from the world itself.
More from Research
- Stu-tU: RANSAC scoring that survives 128x scale miscalibration — ducha_aiki · 2026-09-09
- Test-time compute could be 10,000 model instances coordinating on a message board — mertdumenci · 2026-09-09
- Do transformers lose object-like 'seeing' of definitions through training? — MoonL88537 · 2026-09-09
- Marigold V2 retools diffusion transformers for sharper monocular depth estimation — huawei-bayerlab · 2026-09-09
- kornia-rs: Rust Low-Level 3D Vision Library Outperforms OpenCV and NVIDIA VPI — edgarriba · 2026-09-09
- Researcher cites more blogposts and HF links than papers, blames academia — antoine_chaffin · 2026-09-09