Ask: Is Qwen 3.8 Flash-Next Worth It Over 3.6 35B-A3B on 64GB RAM + 16GB VRAM?
msalsas · reddit · 2026-10-08
With an RTX 4080 16GB and 64GB RAM, the poster happily runs Qwen 3.6 35B-A3B (Unsloth quant, fully in VRAM) and asks whether anyone has tried 3.8 Flash-Next at Q2/Q3 on similar hardware — worth a 30GB+ download to switch, or better to stay put?
More from Infra
- Nebius up 160% vs CoreWeave's 10%: the AI infrastructure stock divergence explained — Beth_Kindig · 2026-10-08
- Microsoft's $5,999 Surface RTX Spark Dev Box preorders open, ships November — tomwarren · 2026-10-08
- Together scales open-source inference with IBM and NVIDIA on B300 cluster — togethercompute · 2026-10-08
- How a Self-Appending Summarizer Quadrupled Token Costs in a 1.4M-Conversation Agent — Good_Education4713 · 2026-10-08
- National Compute gifts $100M in compute credits to White House Genesis Mission — typewriters · 2026-10-08
- vLLM cuts TTFT nearly 70% at ~100K throughput with DeepSeek NVFP4 kernels and fused ops — vllm_project · 2026-10-08