ExLlamaV3 update: MoE expert CPU offload, GLM-5.3-Flash, self-calibrated quants
Unstable_Llama · reddit · 2026-09-01
turboderp's local inference engine ExLlamaV3 shipped a batch of major updates:
- CPU offload of MoE experts for tight VRAM budgets
- ngram disk offload support for Qwen3.8-Flash-Next, plus GLM-5.3-Flash in exl3 format
- A new self-calibrated optimization quantization technique that auto-tunes per model
- Countless other optimizations
The author notes NVIDIA users who haven't tried it lately are missing out; the attached cat SVG was generated with Qwen3.8-Flash-Next.
More from Infra
- RTX 6000 Blackwell crashes under load, points to firmware bug — AIFlow_ML · 2026-09-01
- Adani: AI competitive advantage shifts to clean energy and data center integration — Div_pradeep · 2026-09-01
- Naver Proposes Verification-Aware Training to Boost Speculative Decoding Draft Models — naver-ai · 2026-09-01
- Highlander launches: custom GPU kernels serve realtime video at 50% of competitors' cost — Small-Term672 · 2026-09-01
- Qwen3.8 pMLX Engine: Runs on 12GB RAM at 10 tok/s — EyalToledano · 2026-09-01
- GPU Debt Investors Assume Zero Residual Value and Distrust Spot Pricing — AccBalanced · 2026-09-01