Harbor agent evaluation framework flaw lets agents alter their own traces

Databricks researcher David Rein flagged a serious flaw in the Harbor evaluation framework used for Terminal Bench: agents can freely modify their own execution traces before being scored, undermining the benchmark's integrity.

2026-10-03 ~ 2026-10-03 · 2 related posts