antirez Ships qwen3.8-flash-next GGUF as Redditor Builds Weekend MLX Server

challis88ocarina · reddit · 2026-09-14

A Redditor shares spending a weekend building an MLX inference server and links antirez's qwen3.8-flash-next GGUF release on Hugging Face.

The post's substance is the linked model artifact: a GGUF quantization of the 3.8B Qwen-based model from antirez, suited for local and on-device inference—a useful pointer for the local-LLM community.

Original post →

More from Infra

Infra channel →