ASI-Bench: First Benchmark for Generalist Scientific Research Capabilities
Junwei Zhou · hf · 2026-08-19
Researchers released ASI-Bench, the first benchmark to jointly evaluate AI capabilities in innovative exploration and autonomous scientific execution. Built by 40+ experts with 31,000+ hours of effort, it covers 60 project-level tasks across 11 scientific domains. Tests show a sharp score decline in SOTA models as methodological guidance decreases (from 50.91 to 26.62), revealing heavy reliance on human guidance.
More from AGI Musings
- Opinion: AI Is the Only Frontier, Pre-AI Investments Like Polishing Stone Tools — flowersslop · 2026-08-19
- AI-generated slop flooding repos makes credentials matter even more — miniapeur · 2026-08-19
- New Perspective Paper: Benchmark Leaderboard Race Is a Distraction for Science — ShenRaphael · 2026-08-19
- Schmidhuber Attacks Hinton's Nobel, Accuses of Plagiarism — SchmidhuberAI · 2026-08-19
- Hands-on GPT-5.6: From Assistant to Digital Employee Prototype — tengyanAI · 2026-08-19
- Borretti on AI alignment as a thought-terminating cliche — zetalyrae · 2026-08-19