HarnessEval-W: An Agentic Benchmark for World Models Evaluation

青稞AI · wechat · 2026-08-21

MirroS team open-sourced HarnessEval-W, an agentic benchmark for evaluating world models. It introduces the Harness framework to automate human evaluation workflows, dynamically dispatching specialized sub-agents and tools via a main agent. The system evaluates across three dimensions: Observation Quality, State Transition Correctness, and World Persistence. Instead of a single score, it produces a traceable evidence tree to diagnose model errors. The team also detailed their agentic pipeline for dataset generation and discussed future directions like test-time scaling, skill library expansion, and recursive self-improvement.

Original post →

More from Research

Research channel →