Why Prime Agent Chose RLM Trajectories Over Prompt Tuning

CShorten30 · x · 2026-09-02

On the Weaviate Podcast, Alex Zhang argues that assuming a model can use tools in code just because it uses them well in a ReAct loop is anthropomorphizing. Models excel at what they are trained on, and years of agents were built around ReAct. This shaped the work behind Prime Agent, leading the team to train on RLM (Reinforcement Learning from Model feedback) trajectories rather than tuning the prompt.

Original post →

More from coding & agent

coding & agent channel →