18,000+ runs: new study shows how agent configuration shapes performance on 4 scientific tasks
zeynepakata · x · 2026-10-06
Agents are evaluated with ever more detailed metrics, but little attention is paid to how their own configuration shapes performance. The authors systematically study this across 4 scientific tasks, varying 5 parts of the agent configuration over 18,000+ runs, and share their favorite findings in the thread.
More from coding & agent
- Using Claude Design to prototype complex interactions beats static mocks — austin_malerba · 2026-10-06
- The handoff test: approve, revoke, transfer to a fresh agent—does human authority survive? — tallmetommy · 2026-10-06
- Beam agent one-shots a full end-to-end Unsloth training pipeline in OpenCode with a single prompt — bhutanisanyam1 · 2026-10-06
- Open-source 'Clay killer' launched: 85% cheaper, top people-search accuracy, 25x faster — Scobleizer · 2026-10-06
- New Obsidian plugin qiaomu-ui-learn trains your vibe-coding UI taste with copyable prompts — vista8 · 2026-10-06
- Developer open-sources Semantica, a Palantir-style enterprise BI alternative (13.7k stars) — mdancho84 · 2026-10-06