Artificial Analysis Accused of Tweaking Weights to Suppress Open-Source Models
Infinite-Local5435 · reddit · 2026-08-07
A user raised concerns about the impartiality of Artificial Analysis (AA) leaderboards. Previously, the open-source Qwen 3.8 max topped the platform's Agentic Index.
The user noted that AA subsequently released version "v4.1.1" of the index, adjusting the weights of benchmarks like gdpval and t3 banking to lower the open-source model's score below Anthropic's Opus. The user suspects foul play, implying the weight adjustments were made to favor proprietary models for commercial reasons.
More from Models
- Benchmarking 2-bit Quantization: Half the VRAM, Double the Speed with No Performance Loss — WigglyScrotum · 2026-08-07
- July's LLM Frenzy: Open-Source Hits Top Tier as Focus Shifts to Code and Agents — 创业邦 · 2026-08-07
- DeepSeek Sees Traffic Surge with Superior API Cache Hit Ratio — teortaxesTex · 2026-08-07
- Meta's Models Strike Gold in Five STEM Olympiad Competitions — ilkamoi · 2026-08-07
- DeepSeek-V4 Flash Delivers 80% of GPT-5.6 Luna Performance at 1/6 Cost — zainhas · 2026-08-07
- Report: Ilya's SSI Has Started Benchmarking Its First Model — zephyr_z9 · 2026-08-07