HarnessEval Turns Static Benchmarks into Agentic Evaluation Workflows

MirroS, together with Tsinghua, PKU, Berkeley and MIT, released HarnessEval and HarnessEval-W, open-source frameworks that convert static benchmarks into agentic workflows, using hierarchical sub-agents to produce transparent, auditable reasoning chains for evaluating visual world models.

2026-08-18 ~ 2026-08-19 · 4 related posts