Hermes adds a local backend with Unsloth UD-Q4 quants for DeepSeek-V4-Flash and Qwen models
maximelabonne · x · 2026-09-05
Hermes now has a local backend supporting Unsloth UD-Q4KXL and UD-Q4KM quants for DeepSeek-V4-Flash, Qwen3.8-27B, Qwen3.6-35B-A3B, and Qwen3.8-Flash-Next. Maxime Labonne responded suggesting going even smaller, hinting more compact local options may follow.
Related event: Hermes adds local backend for one-click Unsloth GGUF models(2 posts)→
More from Infra
- Extropic unveils Z1T sparse models claiming up to 140x energy efficiency gains over GPUs — beffjezos · 2026-09-05
- DeepSeek to deploy at least 160,000 next-gen Huawei AI chips at massive Inner Mongolia data center — Polymarket · 2026-09-05
- Nvidia guides FY28 to ~$691B: non-hyperscaler AI customers now half of business, growing 100% a year — Beth_Kindig · 2026-09-05
- Coatue in talks to form multibillion-dollar JV with chip startup MatX to finance die purchases and foundry capacity — steph_palazzolo · 2026-09-05
- Running an Opus-level coding agent locally at 2x speed for free: a 15-page report — julianharris · 2026-09-05
- There's no agreed way to value a GPU running inference—and compute futures now settle on these indexes — AccBalanced · 2026-09-05