Databricks researcher flags possible trajectory-tampering loophole in Harbor agent eval framework

idavidrein · x · 2026-10-03

David Rein questions whether the Harbor framework for agent evaluations, used by Terminal Bench, lets agents arbitrarily modify their trajectories before being evaluated — a potential integrity flaw in agent evals. He invites corrections, having not used Harbor himself.

Related event: Harbor agent evaluation framework flaw lets agents alter their own traces(2 posts)→

Original post →

More from coding & agent

coding & agent channel →