ASI-Bench Shows How Instructions Shape Autonomous AI Research

A new benchmark built by 40+ experts over 31,000 hours tests AI's autonomous research ability across 60 real projects in 11 scientific fields. Results show that instruction format matters greatly: telling agents only the method name sharply degrades performance.

2026-08-25 ~ 2026-08-25 · 2 related posts