ValsAI launches MysteryMechanism, a benchmark testing whether AI agents can rediscover sealed math mechanisms
burny_tech · x · 2026-09-17
AI evaluation firm ValsAI released MysteryMechanism, a new benchmark targeting a core problem in AI for Science: novel scientific results are hard to verify and therefore hard to measure. The benchmark asks agents to rediscover sealed mathematical mechanisms through bounded experiments, making scientific discovery measurable in a verifiable way.
Results cut sharply across the frontier: Astra scores roughly 20 percentage points above Sol, exposing significant gaps between leading models.
More from AGI Musings
- A decade ago EAs dismissed AI risk as delusional, recalls RokoMijic — burny_tech · 2026-09-17
- Medicine's AI misalignment problem, through the lens of the Navier-Stokes debacle — davidjhwu · 2026-09-17
- Dario Amodei's 'We Must Pace the Frontier' essay draws fire as Anthropic opens models to third-party evaluators — alex_verem · 2026-09-17
- The internet killed information scarcity; AI is now killing scarcity of interpretation — signulll · 2026-09-17
- Manning: frontier AI is built inside 'three monasteries' hidden from society — chrmanning · 2026-09-17
- Manning: AI absorbed 2,000 years of knowledge but still learns far worse than humans — chrmanning · 2026-09-17