HarnessEval-W: An Agentic Benchmark for World Models Evaluation
青稞AI · wechat · 2026-08-21
MirroS team open-sourced HarnessEval-W, an agentic benchmark for evaluating world models. It introduces the Harness framework to automate human evaluation workflows, dynamically dispatching specialized sub-agents and tools via a main agent. The system evaluates across three dimensions: Observation Quality, State Transition Correctness, and World Persistence. Instead of a single score, it produces a traceable evidence tree to diagnose model errors. The team also detailed their agentic pipeline for dataset generation and discussed future directions like test-time scaling, skill library expansion, and recursive self-improvement.
More from Research
- Defining Multi-Agent Environments: Orchestration vs. Swarms vs. Simulations — sebkrier · 2026-08-21
- Cornell Nested Architecture Cuts Training Compute by 36% — burkov · 2026-08-21
- SkillEvo: Self-renewing evolution gradients for sustained agent skill improvement — zju · 2026-08-21
- ShikharMurty: Pass@k gains may just sharpen distribution — ShikharMurty · 2026-08-21
- SenseNova U1.5-Lite release: Expert training with OPD distillation — SandyL925 · 2026-08-21
- FasterFASTA tool achieves 1.69 GB/s with multithreaded BGZ decoding — viglovikov · 2026-08-21