Running Qwen3.8-Flash-Next GGUF locally on an M4 Pro 48GB Mac, dense 27B still faster
JLeonsarmiento · reddit · 2026-09-28
A user reports running Qwen3.8-Flash-Next on an M4 Pro 48GB Mac via the ISTA-DASLab GGUF quantization on Hugging Face. Their take: the dense 3.8-27B is actually faster — and possibly better due to lighter quantization.
More from Infra
- Is Agentic scores how AI-agent-ready your website is, via a single npx command — seanwbren · 2026-09-28
- TQ: calibration-free 4-bit quantization open-sourced, hits 92.4% top-1 on Qwen 27B — textclf · 2026-09-28
- Modal's free GPU glossary mini-book surfaces via Gergely Orosz — charles_irl · 2026-09-28
- Estimating 100M DAU infra for Meta Muse: 1-4GW of power, tiny $3B sandbox layer — SuB8u · 2026-09-28
- $350 Dell from 2007 beats $1500 RTX 5070 rig on agentic LLM tasks — Truth-Does-Not-Exist · 2026-09-28
- Rural town promises every household $10k if a data center gets built — JumpCrisscross · 2026-09-28