RSIArena Stage 1: GPT-6 Astra wins blind vote as 10 AI agents self-train on Nemotron

my_cat_can_code · x · 2026-10-12

RSIArena Stage 1 concluded: 10 models acting as "AI researchers" each got the same NVIDIA Nemotron base and 1,000 GPU-hours to pick data, write training code, run experiments and submit models. Blind voting drew 1,685 votes; OpenAI's GPT-6 Astra placed first, with Grok 4.7 and Zhipu's GLM-5.3 rounding out the top three. Grok 4.7 was cut off 41 hours early by an API credit limit with 70% of its GPU budget unused — yet still reached top two.

Stage 2 runs October 10–13: each model resumes from its checkpoint with human feedback from arena battles, 500 more GPU-hours and $150 in API credit, with agents deciding how to use the feedback. Training is livestreamed; checkpoints and data will be open-sourced on Hugging Face. Collaborators include Hugging Face, Scale AI, Notre Dame, UW and Stanford.

Related event: RSIArena Stage 1 Ends: Grok 4.7 Reaches Top Two Despite Early Cutoff(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →