Bilibili's AI Infinite Arena: 190 videos stress-testing 100+ models on real tasks
vista8 · x · 2026-09-20
- AI commentator vista8 argues traditional LLM benchmarks are failing due to leaderboard gaming, data contamination, and reward hacking, while demo-style showcases favor flash over substance.
- He advocates real-task, one-shot, end-to-end testing with full harness/prompt visibility, pointing to Bilibili's "AI Infinite Arena" event where UP主 test models on real jobs: legacy code cleanup, quant trading (who profits, who bankrupts), werewolf game strategy, cost-effective 3D generation.
- He scraped all entries to GitHub: 40 topics, 190 videos, 33 creators, 100+ models; his pick: GPT6-Astra first, Fable 5.1 second, GLM-5.3 third.
Related event: Bilibili AI Arena Tests 100+ Models with 190 Real-Task Videos(2 posts)→
More from Models
- Stepfun's Step-5-Preview-BF16 weights quietly appear on Hugging Face, hinting at next-gen release — adefa · 2026-09-20
- User burns 100M tokens on Jev without exhausting the $5 free credit — DeryaTR_ · 2026-09-20
- Users slam Codex 'reset culture': paid subscribers hoard usage waiting for free resets — TheMoonMidas · 2026-09-20
- Andrew Chen: 400x cheaper inference could unlock ad-supported free AI-native apps — andrewchen · 2026-09-20
- Users report suspiciously generous usage: full morning of building leaves 78% quota left — TheMoonMidas · 2026-09-20
- Jev as LLM-as-a-judge: 20-200x faster scoring for under $0.10 — minchoi · 2026-09-20