Reproducible agent evals: harbor makes configs, trajectories and logs shareable
seanwbren · x · 2026-09-23
A discussion on reproducibility in agent benchmarking: the gold standard is sharing auditable experimental data — configs, trajectories, rewards and logs for every trial. ryanmarten's team built the open-source tool harbor to make this easy. Unreal's analysis (related to Melissa Pan's work comparing costs across harnesses with the model fixed) sets a good example, uploading raw harbor jobs so every result reproduces with a single command. seanwbren notes that cross-harness eval comparisons have been missing.
More from coding & agent
- Developer Praises OpenAI's New Models for Finishing Tasks Fast Without Excessive Tool Calls — mgostIH · 2026-09-23
- GPT-6 Sol and Luna land in Devin: 61% cheaper at parity, under $0.10 per task — _sholtodouglas · 2026-09-23
- Spring AI Ships TypeSafe Model Router: 4 Classes, One Dependency, ~300ms per Prompt — therealdanvega · 2026-09-23
- LangChain on turning scarce clinical expert review into durable agent evals with LangSmith — LangChain · 2026-09-23
- One-shotting a cyberpunk open-world 3D game with Opus 5.5 in 4 hours — mnm9678 · 2026-09-23
- MCP server that composes through ambiguous prompts instead of disclaiming them — dirtyjesus__ · 2026-09-23