antirez Ships qwen3.8-flash-next GGUF as Redditor Builds Weekend MLX Server
challis88ocarina · reddit · 2026-09-14
A Redditor shares spending a weekend building an MLX inference server and links antirez's qwen3.8-flash-next GGUF release on Hugging Face.
The post's substance is the linked model artifact: a GGUF quantization of the 3.8B Qwen-based model from antirez, suited for local and on-device inference—a useful pointer for the local-LLM community.
More from Infra
- Sandbox tip: bake dependencies into the image instead of pip-installing at runtime — xeophon · 2026-09-14
- US's No.2 law firm Latham & Watkins builds in-house AI stack with Nvidia servers — ayushtweetshere · 2026-09-14
- Oracle Cuts Double-Digit % of Some Teams While Hiring Aggressively for Data Centers and AI — mkheck · 2026-09-14
- Running Two Models Across Strix Halo + r9700 Hits OOM: Full Config Shared — El_90 · 2026-09-14
- Musk: AI will be 99% of SpaceX's value within four to five years — XFreeze · 2026-09-14
- Weaker enterprise HBM demand could finally normalize DRAM and NAND pricing, argues analyst — eyishazyer · 2026-09-14