HarnessEval-W: Sub-Agents Make World Model Evaluation Auditable

HarnessEval-W brings the harness paradigm to visual world model evaluation, using hierarchical sub-agents to decompose tasks into verifiable reasoning chains. This makes every score transparent and auditable.

2026-08-18 ~ 2026-08-18 · 2 related posts