Tencent Hunyuan releases ExplorationBench to measure how AI systems explore
TencentHunyuan · x · 2026-09-30
Tencent Hunyuan, with Fudan and Tsinghua, released ExplorationBench to measure genuine exploration ability — framing hypotheses, designing experiments, learning from results. They built verifiable "Alien Worlds" with executable rules (exact grading) that conflict with familiar knowledge (defeating recall). Two sandboxes: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched inference rules, 70 theorems), evaluated via a flawed manual, four rounds of self-designed probes, and closed-book tests after each round.
More from Research
- AVSD accepted at NeurIPS 2026: adaptive multi-view self-distillation improves LLM reasoning RL — mohitban47 · 2026-09-30
- Open-source BabyCoder gives small local LLM agents a dreaming sleep loop that generates new questions — Illustrious_Matter_8 · 2026-09-30
- CoRL 2026 to stage scientific debate: do robots need world models? — vincesitzmann · 2026-09-30
- Anthropic AI solves percolation theory 'holy grail' days after Fields medalist predicted it — aran_nayebi · 2026-09-30
- Rust + Vulkan training backend now parity-verifies 14 architectures and a full PEFT/LoRA workflow, on a handheld's iGPU — PhysicsDisastrous462 · 2026-09-30
- Ben Recht Critiques NFL Win-Probability Models Claiming Three-Digit Precision — beenwrekt · 2026-09-30