Arena Freshness Factuality Leaderboard Updates
arena · x · 2026-07-16
New model comparison results have emerged for Arena's factuality metric:
- GPT-5.6 currently leads, with Grok-4.5 close behind.
- Opus-4.8 scores higher in factuality than Fable-5.
- Muse-Spark's score dropped significantly after the factuality weight was increased.
The post also jokingly suggests that some RL compute resources seem to have value that is "only realized over the long term."
Related event: Arena Adds Factuality to Model Rankings(12 posts)→
More from Models
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22