Open tutorial: training a search agent with GRPO, every rollout browsable and code open-sourced
dejavucoder · x · 2026-09-17
Jasper Lu published his first research blog post, "Training Search Agents with GRPO" — a hands-on RL-for-LLMs tutorial with browsable rollouts and open-source code.
Highlights:
- Built on the Harness-1 paper (Jiang et al.): the RL split has 3,453 queries over 37,113 SEC filings split into 2,115,106 chunks; ablations start with 256 queries, scaling runs use 1,024 queries over 237,533 chunks
- Training runs on Thinking Machines' Tinker API
- Core claim: search is a great RL learning domain where every lever visibly changes model behavior, and a well-designed reward function influences search more than system prompts or harness engineering
- Walks through the full thought process from learning rate sweeps to reward shaping
More from coding & agent
- LangChain Rebuilds LangSmith Trace Filtering for Faster Agent Observability — LangChain · 2026-09-18
- LangChain rebuilds LangSmith trace filtering for faster, more precise agent debugging — LangChain · 2026-09-18
- Dev: the 'smartest' model isn't the best at writing simple, mergeable code — timigod · 2026-09-18
- Dev: The Smartest Model Isn't the Best at Writing Simple, Mergeable Code — timigod · 2026-09-18
- LinkedIn to present a PyTorch-native GPU retrieval engine powering feed and search — PyTorch · 2026-09-18
- Jev + Kimi K3 cascade classifies 100 fraud emails at 96% accuracy for $0.07 — nutlope · 2026-09-17