Tencent Hunyuan releases ExplorationBench to measure how AI systems explore

TencentHunyuan · x · 2026-09-30

Tencent Hunyuan, with Fudan and Tsinghua, released ExplorationBench to measure genuine exploration ability — framing hypotheses, designing experiments, learning from results. They built verifiable "Alien Worlds" with executable rules (exact grading) that conflict with familiar knowledge (defeating recall). Two sandboxes: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched inference rules, 70 theorems), evaluated via a flawed manual, four rounds of self-designed probes, and closed-book tests after each round.

Original post →

More from Research

Research channel →