Qwen3.8-Max scores 19 on OpenRouter vs 26 on Alibaba — effort level bug suspected
PawelHuryn · x · 2026-09-12
Pawel Huryn spotted a suspicious score gap for Qwen3.8-Max: 19 via OpenRouter vs 26 via Alibaba directly. His prime suspect is the effort level not being served properly through OpenRouter, an issue he's seen before. He's running a probe to confirm and will update the thread in 50 minutes.
More from Models
- Devs split on GPT-6 Astra: fast and token-cheap but code feels 'alien'; costly stack proposed — beffjezos · 2026-09-12
- DeepSeek V4.1-Flash Runs 502GB Model on a Single RTX 5090 at 5-21 tok/s — AccBalanced · 2026-09-12
- Codex reset rolling out now, and GPT-Image 2.5 ships a new sketch feature — koltregaskes · 2026-09-12
- Orca releases uncensored MLX weights for DeepSeek V4.1 Flash, cutting refusals by 87-96% — AccBalanced · 2026-09-12
- Open 33B multimodal Agnes-3.0-Flash ships hybrid delta-rule attention with 262K context — Skyline34rGt · 2026-09-12
- DeepSeek V4.1 Flash reportedly served 3.1T paid tokens on day one at 230-300 TPS — AccBalanced · 2026-09-12