HarnessEval: Evolving Evaluation from Metrics to Executable Systems
机器之心 · wechat · 2026-08-18
Systemic Upgrade in Evaluation Paradigm
MirroS, in collaboration with Tsinghua, Peking University, Berkeley, MIT, and others, released HarnessEval, proposing an upgrade from static metrics to an executable Agent evaluation system (Evaluation Harness).
Core Changes:
- From Rules to Workflows: Instead of running pre-set scripts, an Evaluation Agent understands the case, dynamically plans the path, and selects/combines skills.
- Four Stages:
- Plan: Understand case context and intent.
- Route: Select applicable skills from a library.
- Decompose: Break down abstract questions into verifiable sub-problems.
- Verify: Audit evidence and output scores.
Application: First launched for interactive World Model evaluation (HarnessEval-W). For complex issues like physical consistency and causal order in video generation, the system generates a hierarchical Evidence Tree to pinpoint error sources rather than just providing a scalar score.
Significance: Supports RSI (Recursive Self-Improvement). Evaluation becomes a key feedback component in the system evolution loop, not just an external ruler.
Related event: HarnessEval Turns Static Benchmarks into Agentic Evaluation Workflows(4 posts)→
More from Research
- Peking U & Kling Team Release RefCaptioner, Tackling Video Understanding Hallucinations — jiqizhixin · 2026-08-19
- 'Scaling laws are not laws of nature': better data and architectures can still bend the curves — bookwormengr · 2026-08-19
- PixRestore: Unified Image Restoration via Pixel Diffusion Transformer — Lingchen Sun · 2026-08-19
- Cross-Model Memory Transfer via Target-Side Reader Adaptation — OLAResearchX · 2026-08-19
- DeepSeek Web Chat Wins Industrial Track in TAAC × KDD Cup — jiqizhixin · 2026-08-19
- Next-Gen Agent Frameworks: Auto-Integrated Coded Extensions — _philschmid · 2026-08-19