Qwen3.8-27B Released, Tops Hugging Face Trends
On August 13–14, Alibaba's Qwen team open-sourced Qwen3.8-27B, a natively multimodal dense model with just 27B parameters. The team claims it comprehensively outperforms Qwen3.7-Plus, natively supports 262K context (extendable to 1M via YaRN), and ships under Apache 2.0. The model quickly topped the Hugging Face trending chart, with FP8 and GGUF quantized versions launching alongside it, sparking community debate over whether it reaches 'local Opus 4.6' level.
Confirmed
- The model quietly appeared on ModelScope on August 13 (spotted by @Ok-Shower7286); Qwen released the FP8 quantized version on Hugging Face the same day, with the open-source announcement made official on the 14th
- Official introduction (@AlibabaQwen): a 27B dense hybrid-attention model with linear attention on 48/64 layers, natively multimodal (image-text to text), with a built-in MTP draft head; available day-one on vLLM 0.17.0, supporting 1M context on a single GPU
- The team says it excels in real-world coding and office workflows, comprehensively surpassing Qwen3.7-Plus
- Topped the Hugging Face trending chart on launch day, compatible with the transformers and safetensors ecosystems
- Unsloth released GGUF quantized weights enabling local deployment with llama.cpp and other tools; on August 15, @Tccybo compiled Hugging Face benchmark scores into multiple comparison charts
Unconfirmed
- The 'local Opus 4.6' label comes from community users' performance charts and reviews: @HugeConsideration211's charts show it beating Glimmer with major gains over its predecessor; @kimmonismus claims it surpasses Claude Opus 4.6 Max on benchmarks like agentic coding and computer use — all third-party claims, as the team has not directly benchmarked against closed-source flagships
- @erdaltoprak called it 'the best local dense model right now' — a personal opinion
Why It Matters
- A compact 27B parameter count, fully open under Apache 2.0, plus FP8/GGUF quantization and day-one vLLM support, dramatically lowers the barrier to running long-context multimodal models locally
- The hybrid linear attention architecture enables million-token context on a single GPU, giving the open-source community a new architectural reference
2026-08-13 ~ 2026-08-15 · 20 related posts
- Episode 1: Qwen3.8-27B Model Confirmed for September 3 Release(2026-08-11, 2 posts)
- Episode 2: Alibaba Teases Qwen3.8-27B with Unmatched Intelligence Density(2026-08-12, 11 posts)
- Episode 3: Qwen3.8-27B Released, Tops Hugging Face Trends(2026-08-13, 20 posts)
Primary sources
- Qwen3.8-27B Model Surfaces on ModelScope Ahead of Hype — Ok-Shower7286 · 2026-08-13
- [source] Qwen releases Qwen3.8-27B-FP8 quantized model with image-text-to-text support — Qwen · 2026-08-13
- Rumor: Alibaba's Qwen 3.8 27B Model Set for Friday Launch — Brilliant-Hall1387 · 2026-08-14
- Qwen3.8-27B FP8 Weights Released on Hugging Face — Certain-Cod-1404 · 2026-08-14
- Qwen 3.8 27B released: best local dense model yet — erdaltoprak · 2026-08-14
- Qwen3.8-27B Released, Claimed to Match Opus 4.6 — de4dee · 2026-08-14
- Qwen3.8-27B: 27B params rival frontier, beats Claude Opus 4.6 Max on agent benchmarks? — kimmonismus · 2026-08-14
- Alibaba Releases Qwen3.8-27B: 27B-Parameter Multimodal Model Beats Claude Opus 4.6 Max on Several Benchmarks — kimmonismus · 2026-08-14
- [source] Alibaba open-sources Qwen3.8-27B: 27B multimodal model beats Qwen3.7-Plus, Apache 2.0 — Alibaba_Qwen · 2026-08-14
- Unsloth Releases Qwen3.8-27B GGUF Weights for Easy Local Deployment — kevin_1994 · 2026-08-14
- Qwen3.8-27B Open-Sourced: 27B Params Outperform Qwen3.7-Plus, Native Multimodal — NVIDIAAI · 2026-08-14
- Unsloth releases Qwen3.8-27B quantized: NVFP4 1.5x faster, retains 92-97% accuracy — danielhanchen · 2026-08-14
- Qwen3.8-27B tops Hugging Face trending, Apache 2.0 open weights — Qwen · 2026-08-14
- Qwen 3.8 27B GGUF Quantized Version Trending on Hugging Face — unsloth · 2026-08-14
- Qwen3.8-27B Released, Hugging Face Collection Available — song91 · 2026-08-14
- Qwen3.8-27B open-sourced with SGLang Day-0 support, 206 tok/s on RTX 5090 — ying11231 · 2026-08-14
- [source] Qwen3.8-27B launches with 1M context on a single GPU, vLLM ready — Alibaba_Qwen · 2026-08-14
- Qwen3.8 benchmark charts visualize performance across tests — Tccybo · 2026-08-15
- Qwen3.8-27B launches with day-zero AMD support for local AI development — Alibaba_Qwen · 2026-08-15
1 near-duplicate retellings: Alibaba_Qwen