LM Arena's Auto-Eval Tool Predicts New DeepSeek Model at Rank 41

Unusual_Guidance2095 · reddit · 2026-08-13

LM Arena used its new auto-evaluation tool, which emulates human preferences, to predict that the upcoming DeepSeek model will rank 41st on the leaderboard.

This sparked community discussion: does DeepSeek consistently perform poorly in chat conversations, or is there an underlying flaw and bias within the auto-evaluation tool itself?

Original post →

More from Models

Models channel →