A concrete path past the data wall: train agents on interaction trajectories, not just text

TuXinming · x · 2026-09-09

Responding to complaints that scaling talk is fortune-telling, TuXinming proposes a concrete loop: pretraining compresses world→observation→text, but tools offer act→feedback→train-on-trajectories. Labs already train on tool-use/agent traces — e.g., agents querying AlphaGenome to refine hypotheses; swap tools for real experiments and feedback comes from the world itself.

Related event: Stanford professors clash over AI limits: data bottlenecks vs a recent convert's 'it just works'(10 posts)→

Original post →

More from Research

Research channel →