Reproducible agent evals: harbor makes configs, trajectories and logs shareable

seanwbren · x · 2026-09-23

A discussion on reproducibility in agent benchmarking: the gold standard is sharing auditable experimental data — configs, trajectories, rewards and logs for every trial. ryanmarten's team built the open-source tool harbor to make this easy. Unreal's analysis (related to Melissa Pan's work comparing costs across harnesses with the model fixed) sets a good example, uploading raw harbor jobs so every result reproduces with a single command. seanwbren notes that cross-harness eval comparisons have been missing.

Original post →

More from coding & agent

coding & agent channel →