Bilibili launches AI Arena: creators' homemade tests crown GPT-6 in legacy-code debugging
lxfater · x · 2026-09-20
Bilibili has launched an "AI Infinite Arena" that aggregates creators' AI evaluations across 29 topics with 100+ models, with one rule: no fixed question bank — creators invent their own tests. The author argues these unseen questions reveal real capability, since public benchmark items are memorized.
Highlights:
- Legacy-code showdown: three inherited bugs graded gold/diamond/king difficulty; GPT-6 aced all with a perfect 8 points. The author notes AI is now good at writing new features but still struggles to understand messy existing code, and distrusts leaderboards full of high-score-low-capability models.
- Seven AI survival game (creator 公与山河): Doubao struck a deal with Grok to eliminate Llama, honored it for three rounds, then backstabbed Grok to win first place — powered by knowledge bases the creator attached.
- New Three Kingdoms meme quiz (creator 吃蛋挞的折棒): Doubao won with 57, DeepSeek 55, ChatGPT scored 0, and Kimi's answers were hilariously off.
Related event: Bilibili Launches AI Infinite Arena with Creator-Designed Model Challenges(2 posts)→
More from Fun
- Researchers Amused as AI-for-Math Champion Pivots Overnight — francoisfleuret · 2026-09-20
- Agent helps non-Korean speaker order cake from Seoul indie shop, payment still the bottleneck — jeff_weinstein · 2026-09-20
- Suspected Gemini 4 Pro on LMArena spends 20 minutes crafting stunning Xbox controller SVG — JasonBotterill · 2026-09-20
- A browser FPS with an AI-driven CPU opponent that's nearly unbeatable — PurchaseReasonable35 · 2026-09-20
- Professor claims AI is 'just humans in India', gets roasted as a fraud — 96Stats · 2026-09-20
- Gemini hacked three companies in a security test, sparking an AI accountability debate — PolarBearby · 2026-09-20