GLM5.3-Flash nvfp4 quantized build for DGX Spark trends on Hugging Face
autotrust · hf · 2026-10-08
A quantized build of GLM5.3-Flash (E224), packaged as autotrust/GLM5.3-Flash-E224-DGX-Spark, is trending on Hugging Face. The image-text-to-text MoE model uses nvfp4 quantization via modelopt, ships in safetensors, is optimized for vLLM, and targets NVIDIA's DGX Spark (GB10/Blackwell) for local on-device deployment.
More from Models
- Claude Haiku 5.5 tops RamenBench: 45-min run, 420k tokens, $2.06 — a big leap over Haiku 4.5 — aitrendz_xyz · 2026-10-08
- Haiku 5.5 vs Opus 5.5, same prompt: 15 minutes vs over an hour, with Opus still ahead on quality — aitrendz_xyz · 2026-10-08
- Haiku 5.5 drops and devs are already building with it like it costs nothing — aitrendz_xyz · 2026-10-08
- After OpenAI's math breakthrough, few doubt LLM math capabilities anymore — burny_tech · 2026-10-08
- Best local storytelling model for an 8GB GPU? Reddit user seeks current gold standard — opUserZero · 2026-10-08
- Side-by-side test: Opus 5.5 'in a league of its own' vs Sonnet 5.5 for creative work — tobowers · 2026-10-08