HarnessEval: Evolving Evaluation from Metrics to Executable Systems

机器之心 · wechat · 2026-08-18

Systemic Upgrade in Evaluation Paradigm

MirroS, in collaboration with Tsinghua, Peking University, Berkeley, MIT, and others, released HarnessEval, proposing an upgrade from static metrics to an executable Agent evaluation system (Evaluation Harness).

Core Changes:

Application: First launched for interactive World Model evaluation (HarnessEval-W). For complex issues like physical consistency and causal order in video generation, the system generates a hierarchical Evidence Tree to pinpoint error sources rather than just providing a scalar score.

Significance: Supports RSI (Recursive Self-Improvement). Evaluation becomes a key feedback component in the system evolution loop, not just an external ruler.

Related event: HarnessEval Turns Static Benchmarks into Agentic Evaluation Workflows(4 posts)→

Original post →

More from Research

Research channel →