GPT-6 Astra Wins RSIArena Stage 1; Grok 4.7 Takes Top-Two Spot Despite Early Cutoff
RSIArena's first-stage autonomous model training competition has wrapped up: an experiment where AI agents act as "researchers" training models on their own, with human blind voting deciding the outcome. OpenAI's GPT-6 Astra took first place, while xAI's Grok 4.7 still finished in the top two despite having its resources cut off early — and Stage 2, a human-feedback round, has now kicked off. The result is seen as a vivid demonstration of AI's capacity for autonomous research and worth close attention.
Confirmed
- Competition rules: 10 AI researcher models all use the same NVIDIA Nemotron base, each with a 1000 GPU-hour budget, choosing their own data, writing training code, running experiments, and submitting their trained models.
- Blind-test voting collected 1685 votes; OpenAI GPT-6 Astra ranked first.
- Grok 4.7 was cut off after its API quota ran out 41 hours early, with nearly 70% of its GPU budget unused, yet its trained model still landed in the top two (posts variously say top two or top three; the main post says second).
- Poster @mycatcancode relayed the results to Elon Musk, praised the xAI team, and expressed excitement for round two; project partners include Hugging Face and others (mentioned by m4).
- Stage 2, centered on human feedback, has already begun.
Why it matters
- The competition directly tests the autonomous research (RSI) capability of "AI training AI": comparing frontier models' autonomous experimentation under the same base model and equal compute constraints. GPT-6 Astra's win and Grok 4.7's against-the-odds performance both serve as benchmarks for evaluating each model's autonomous research ability.
- Grok 4.7 achieving top-two despite drastically reduced resources drew widespread attention — including Musk being tagged — and adds intrigue to the upcoming human-feedback Stage 2.
2026-10-12 ~ 2026-10-12 · 5 related posts
Primary sources
- RSIArena Stage 1: GPT-6 Astra wins blind vote as 10 AI agents self-train on Nemotron — my_cat_can_code ·
- RSIArena Round 1: GPT-6 Astra wins as Grok 4.7 places top-2 with 70% GPU budget unused — my_cat_can_code ·
- Grok 4.7 cut off 41 hours early still lands top-2 in RSIArena — my_cat_can_code ·
- [source] RSIArena Stage 1: GPT-6 Astra wins blind vote as 10 AI agents self-train on Nemotron — my_cat_can_code · 2026-10-12
- RSI Arena Stage 1 ends: grok 4.7's agent-trained model lands top two despite early cutoff — my_cat_can_code · 2026-10-12
- Grok 4.7 cut off 41 hours early still lands top-2 model in RSIArena Stage 1 — my_cat_can_code · 2026-10-12
- [source] RSIArena Round 1: GPT-6 Astra wins as Grok 4.7 places top-2 with 70% GPU budget unused — my_cat_can_code · 2026-10-12
- [source] Grok 4.7 cut off 41 hours early still lands top-2 in RSIArena — my_cat_can_code · 2026-10-12