Tencent Hunyuan's WebCraftBench tests web apps like software, matching human preference 85.3%

TencentHunyuan · x · 2026-09-22

Tencent Hunyuan released WebCraftBench, an interactive benchmark that evaluates LLM-generated web applications from a software testing perspective.

Why: static benchmarks credit functionality in source code that's unreachable at runtime; interactive ones miss implemented features due to incomplete exploration and conflate app defects with agent execution failures.

How:

Scale & results: 369 real-world user requirements and 5,088 acceptance criteria; on 197 human-validated pairs it matches human preference 85.3% of the time, and the team benchmarked 17 frontier LLMs.

Original post →

More from coding & agent

coding & agent channel →