RSIArena Stage 1: GPT-6 Astra wins blind vote as 10 AI agents self-train on Nemotron
my_cat_can_code · x · 2026-10-12
RSIArena Stage 1 concluded: 10 models acting as "AI researchers" each got the same NVIDIA Nemotron base and 1,000 GPU-hours to pick data, write training code, run experiments and submit models. Blind voting drew 1,685 votes; OpenAI's GPT-6 Astra placed first, with Grok 4.7 and Zhipu's GLM-5.3 rounding out the top three. Grok 4.7 was cut off 41 hours early by an API credit limit with 70% of its GPU budget unused — yet still reached top two.
Stage 2 runs October 10–13: each model resumes from its checkpoint with human feedback from arena battles, 500 more GPU-hours and $150 in API credit, with agents deciding how to use the feedback. Training is livestreamed; checkpoints and data will be open-sourced on Hugging Face. Collaborators include Hugging Face, Scale AI, Notre Dame, UW and Stanford.
Related event: RSIArena Stage 1 Ends: Grok 4.7 Reaches Top Two Despite Early Cutoff(3 posts)→
More from AGI Musings
- Billion AI scientists will tackle aging, predicts AI commentator — Dr_Singularity · 2026-10-12
- Steve Yegge's favorite essay in 3 years: adaptive complexity as the supervalue — Steve_Yegge · 2026-10-12
- A chatbot's overly familiar greeting reveals why humans shouldn't default to trusting LLMs — relic_radiation · 2026-10-12
- AI's Real Disruption Is the Invisible World of Custom Software, Not Photoshop — bennash · 2026-10-12
- AlphaGo insider: Don't be fooled—today's LLMs still can't truly reason — luislamb · 2026-10-12
- New Yorker Cartoon: 'We Need to Rethink Our Strategy of Hoping AI Will Just Go Away' — c_valenzuelab · 2026-10-12