Bilibili creators pit top AI models against legacy spaghetti code; GPT-6 scores full marks
lxfater · x · 2026-09-20
A Bilibili creator ran a 'spaghetti code showdown' feeding legacy codebases to top AI models to fix three graded legacy bugs (gold/diamond/king tiers). GPT-6 scored a full eight points and won in one run. The takeaway: writing new features with AI is easy now — the hard part is making models understand messy code left by others, which also fuels doubts about leaderboard 'high score, low ability' models. Bilibili has grown a culture of creators stress-testing AI across content categories, and the platform recently launched an 'AI Infinite Arena' to collect these ad-hoc evaluations under one roof, with no fixed question bank or topic limits.
Related event: Bilibili Launches AI Infinite Arena with Creator-Designed Model Challenges(2 posts)→
More from Models
- Mozilla report: open-weight AI models now just four months behind the closed frontier — mark_k · 2026-09-20
- GPT Pro user finds model browses completely irrelevant sources, questions real retrieval — CompetentRaindeer · 2026-09-20
- Open-source model exec calls 'open-source overtaking frontier' traffic data nonsense — zephyr_z9 · 2026-09-20
- Founder: Fable 5.1 and Astra both misread data today — AI is nowhere near 'superior intelligence' — bindureddy · 2026-09-20
- Suspected Gemini 4 Pro on LMArena spends 20 minutes crafting stunning Xbox controller SVG — JasonBotterill · 2026-09-20
- Bilibili launches AI Arena: creators' homemade tests crown GPT-6 in legacy-code debugging — lxfater · 2026-09-20