Coding vs Agentic Leaderboard Discrepancies
Substantial_Step_351 · reddit · 2026-07-09
The post compares LiveBench's coding scores with agentic coding scores, pointing out that Claude 4 Sonnet's non-thinking version leads in the coding average but lags noticeably in agentic coding. Similarly, the rankings for Sonnet 5 and Sonnet 4.6 flip between the two leaderboards.
The author explains that the coding average leans towards static code generation and completion, whereas agentic coding resembles dockerized terminal tasks, meaning the two metrics measure entirely different skill sets. The post also notes that GPT-5.5 thinking high effort exhibits a similar phenomenon, showing a massive score gap between the two columns.
More from Models
- Claim says Kimi was distilled from Fable, sparking a model-attribution jab — cephaloform · 2026-07-22
- Gemini 3.6 Flash is now available in Antigravity and chat — MartianOnJupiter · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- OpenAI’s Codex + GPT-5.6 Sol hits 99% recall in Project APE verification tests — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Macaron V1 adds LoRA RL on GLM 5.2 and claims SOTA benchmark gains — Xianbao_QIAN · 2026-07-22