ExLlamaV3 update: MoE expert CPU offload, GLM-5.3-Flash, self-calibrated quants

Unstable_Llama · reddit · 2026-09-01

turboderp's local inference engine ExLlamaV3 shipped a batch of major updates:

The author notes NVIDIA users who haven't tried it lately are missing out; the attached cat SVG was generated with Qwen3.8-Flash-Next.

Original post →

More from Infra

Infra channel →