Alternating prompt and model upgrades lift science agent from 42% to 73%
rohanpaul_ai · x · 2026-10-02
A new arXiv paper, ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents, argues you shouldn't pick between better prompts and a better model — improving both in turns is what works. Researchers constantly correct AI agents in chat, but those fixes usually die with the conversation instead of making the agent better next time.
ScienceBuddy, an AI assistant for scientists, turns user feedback into scored test tasks, then loops: rewrite the agent's prompts and skills, retrain the model, and repeat. With a small 4B model on biology tasks, prompt and skill changes alone raised accuracy from 31.1% to 51.1%; retraining alone raised the share of problems solved within 4 tries from 48.3% to 67.8%; combined, accuracy went from 42.2% to 73.3%.
Takeaway for agent builders: save every user correction as a test, and keep upgrading prompts and model in turns.
More from coding & agent
- Agent tool-routing encoder trained for just $6.60 of A100 compute released on Hugging Face — MaziyarPanahi · 2026-10-02
- Sam Altman says his AI agent Dot now triages his mornings and gives him deep-work time back — every · 2026-10-02
- Harness engineering explained: building the scaffolding around coding agents — tekbog · 2026-10-02
- Matt Pocock shares a retro prompt that makes coding agents improve repo navigability — mattpocockuk · 2026-10-02
- T3 Code, theo's AI coding tool, passes 400,000 users — 0xkarasy · 2026-10-02
- AgSpec Speeds Up Coding Agents Up to 4.76x with Retrieval-Based Speculative Decoding — Sumin Lee · 2026-10-02