Researcher questions whether 'release first, improve later' metrics are worth hillclimbing

suchenzang · x · 2026-10-11

AI researcher Su Chenzang mocked the industry's 'release first, improve second' mindset, questioning whether the metric being hillclimbed is even a meaningful target. The remark highlights a broader problem in current model evaluation culture: leaderboard numbers treated as marketing tools rather than genuine capability measures.

Original post →

More from Models

Models channel →