Quantized GLM MoE model (W4A16 AWQ) trends on Hugging Face
AikidoSec · hf · 2026-09-22
AikidoSec's altar-1 is trending on Hugging Face as a text-generation model built on the GLM MoE architecture (tags include glmmoedsa and glm-5.3), quantized to AWQ INT4 (W4A16) in compressed-tensors format.
The quantization lets large MoE models run with much lower memory footprints, useful for local or self-hosted inference setups.
More from Infra
- bitsandbytes2 slightly delayed: zero-config lazy compression plus dynamic expert swaps for near-infinite KV cache — Tim_Dettmers · 2026-09-22
- Iterative sensitivity probing: how bitsandbytes2 finds each layer's compression limit — Tim_Dettmers · 2026-09-22
- Tim Dettmers releases runtime dynamic compression framework, hits 1.5-2.0 bit at high quality — Tim_Dettmers · 2026-09-22
- PSA: Non-US Users Should Consider Local AI in Case Governments Ban LLMs — TheMoonMidas · 2026-09-22
- Engram: A Local Encrypted Memory Vault Unifying Agent Memory Across AI Tools — Acceptable_Leg3950 · 2026-09-22
- DigitalOcean Managed Agents enters public preview with idle-pause billing — damianplayer · 2026-09-22