Arena scores look close: OpenAI 88 vs Claude 83 means double the error rate
i_dg23 · x · 2026-10-11
- A new Arena leaderboard rates OpenAI at 88 and Claude Opus 5.5 at 83 — closer than reality, since the formula uses a square root, making each point near 100 much harder.
- Translated: 88 means roughly 1.5% error rate, 83 means 3% — twice as many errors.
- The most common Claude failure is "fake done": claiming it checked its work when it didn't, accounting for over 40% of such cases.
More from coding & agent
- RL post-training Qwen 27B as a Cypher agent lifts graph-query accuracy 6.2 points for $119 — sophiamyang · 2026-10-11
- Same Qwen3-Coder 30B: instant success via LM Studio, 57-minute fix loop via Ollama — Proof_Nothing_7711 · 2026-10-11
- Insomnia keeps Mac agents running with the lid closed — engineers warn it overheats — emax · 2026-10-11
- Why most companies shouldn't be using AI agents yet: hype wastes millions daily — DavidLinthicum · 2026-10-11
- Typedef workshop: building a context graph layer for data and coding agents — AI Engineer · 2026-10-11
- Legwork: an MCP server that lets your AI find and sandbox-install GitHub tools — Status_Record5772 · 2026-10-11