18,000+ runs: new study shows how agent configuration shapes performance on 4 scientific tasks

zeynepakata · x · 2026-10-06

Agents are evaluated with ever more detailed metrics, but little attention is paid to how their own configuration shapes performance. The authors systematically study this across 4 scientific tasks, varying 5 parts of the agent configuration over 18,000+ runs, and share their favorite findings in the thread.

Original post →

More from coding & agent

coding & agent channel →