NVIDIA releases NVFP4-quantized Qwen3.8-Flash-Next: 125B MoE, 63% smaller
_akhaliq · x · 2026-09-05
NVIDIA released the NVFP4-quantized Qwen3.8-Flash-Next on Hugging Face: a 125B MoE with hybrid attention that is now 63% smaller with minimal accuracy loss.
More from Models
- Eric Horvitz: Astra's model card shows CoT-based abuse monitoring is getting harder — erichorvitz · 2026-09-05
- GPT-6 Astra lands day-zero on Databricks, touting SOTA agentic reasoning and document processing — matei_zaharia · 2026-09-05
- Unverified: 'Astra' model explodes Runescape bench scores, records on 10/16 skills — scaling01 · 2026-09-05
- The model hyped as AGI two months ago vs. an average GPT-6 Astra output — aidan_mclau · 2026-09-05
- Unverified GPT-6 Astra crushes cybersecurity benchmark: 29/32 CVEs, $4k for three runs — charliermarsh · 2026-09-05
- Altman clarifies GPT-6 Astra isn't the model paused over cybersecurity concerns — rohanpaul_ai · 2026-09-05