Shanghai AI Lab's Atria-Dawn-Preview agent ranks top on agent benchmarks while AI ran 96.5% of its own training
新智元 · wechat · 2026-09-14
Shanghai AI Lab, with Fudan, CAS and other universities, released Atria-Dawn-Preview, an open-weight agent model. It ranked first on AutomationBench and CyberGym, second on SkillsBench and Workspace-Bench-Lite. Demos include autonomously designing and training a 400M+ parameter weather model that predicts global weather a week ahead in 1 minute (partly beating NVIDIA's FourCastNet), building a MiniOS from a natural-language spec in 20 minutes, and full vulnerability discovery-exploit-fix cycles in a sandbox.
Notably, the team analyzed 769 task records from 56 researchers: AI was used in 96.5% of tasks, 33.2% were deemed impossible without AI, AI proposed 64.6% of technical solutions — yet humans made 85.5% of final decisions, and 76% of stuck tasks resumed with human help. The team frames this as a data point on the road to recursive self-improvement.
More from coding & agent
- AI agent finds $3,500/year car insurance savings in 5 minutes, buys it and cancels old policy — garrytan · 2026-09-15
- Pokee Isaac MCP now runs inside OpenCode, generating complete artifacts in one pass — Kyrannio · 2026-09-15
- DB veteran Craig Mullins: recovery beats backup, and test envs must match production scale — craigmullins · 2026-09-15
- Craig Mullins: the database remembers every bad design decision, for decades — craigmullins · 2026-09-15
- Craig Mullins: indexes aren't free, and the optimizer is only as smart as your statistics — craigmullins · 2026-09-15
- Craig Mullins: database performance is an ongoing discipline, not a one-time project — craigmullins · 2026-09-15