ASI-Bench: First Benchmark for Generalist Scientific Research Capabilities

Junwei Zhou · hf · 2026-08-19

Researchers released ASI-Bench, the first benchmark to jointly evaluate AI capabilities in innovative exploration and autonomous scientific execution. Built by 40+ experts with 31,000+ hours of effort, it covers 60 project-level tasks across 11 scientific domains. Tests show a sharp score decline in SOTA models as methodological guidance decreases (from 50.91 to 26.62), revealing heavy reliance on human guidance.

Original post →

More from AGI Musings

AGI Musings channel →