RSIArena Round 1: GPT-6 Astra wins as Grok 4.7 places top-2 with 70% GPU budget unused
my_cat_can_code · x · 2026-10-12
- RSIArena Stage 1 wrapped: 10 AI-researcher agents started from the same NVIDIA Nemotron base, each with 1,000 GPU-hours to pick data, write training code, run experiments, and submit their own model.
- 1,685 blind votes: OpenAI's GPT-6 Astra took first; models trained by Grok 4.7 and Zhipu's GLM-5.3 rounded out the top three. Notably, Grok 4.7 ran out of API credit 41 hours early with nearly 70% of its GPU budget unused.
- Stage 2 is live (Oct 10–13): each model resumes from its own checkpoint with human feedback from arena battles, 500 more GPU-hours, and $150 API credit—totals of 1,500 GPU-hours and $450 per model across both rounds. The agents decide how to act on the feedback.
- Organized by Bake AI with Hugging Face, Scale AI, Notre Dame, UW, and Stanford; checkpoints and data will be open-sourced on Hugging Face, with training streamed live.
More from Models
- Researchers hide a credential-stealing backdoor in an open LLM for under $50 — emax · 2026-10-12
- Coinbase fine-tuned Qwen3.5-9B on fraud data, beating Opus 4.5 at half the latency — J0se · 2026-10-12
- Users complain Claude keeps opening replies with 'One important distinction' — Ordinary_Bridge9703 · 2026-10-12
- Fireworks: open models plus fine-tuning match closed ones — Cursor gets 13x faster inference — AI Engineer · 2026-10-12
- Decision 3.0 ported to Core ML: 1,000 PRs triaged on-device at 0.12s each — emax · 2026-10-12
- One C# radio-player prompt benchmarks Qwen 27B, Next Flash vs GPT 6.1 Sol one-shot — Perfect-Campaign9551 · 2026-10-12