Bilibili creators pit top AI models against legacy spaghetti code; GPT-6 scores full marks

lxfater · x · 2026-09-20

A Bilibili creator ran a 'spaghetti code showdown' feeding legacy codebases to top AI models to fix three graded legacy bugs (gold/diamond/king tiers). GPT-6 scored a full eight points and won in one run. The takeaway: writing new features with AI is easy now — the hard part is making models understand messy code left by others, which also fuels doubts about leaderboard 'high score, low ability' models. Bilibili has grown a culture of creators stress-testing AI across content categories, and the platform recently launched an 'AI Infinite Arena' to collect these ad-hoc evaluations under one roof, with no fixed question bank or topic limits.

Related event: Bilibili Launches AI Infinite Arena with Creator-Designed Model Challenges(2 posts)→

Original post →

More from Models

Models channel →