Exllamav3 benchmarks: 700tk/s on 8x3090 setup
Leflakk · reddit · 2026-09-02
User benchmarks Exllamav3 with GLM 5.3 Flash (Q4) on an 8x3090 setup, achieving 700 tk/s prefill and 42 tk/s decoding, outperforming lcp and vllm with stable quality over 30m tokens.
More from Infra
- MLPerf Storage v3.0 lands with 144 results, adding KV cache and vector DB tests — TheKanter · 2026-09-02
- Cloudflare Agents emit OpenTelemetry traces, route directly to Braintrust for evals — ritakozlov · 2026-09-02
- Rabbi's take on DC moratorium: Using bans as leverage for environmental and labor concessions — joshua_saxe · 2026-09-02
- Intel exec: AI era security requires silicon-level design, not afterthoughts — BenBajarin · 2026-09-02
- PyTorch 2.14 released with 2,995 commits from 487 contributors — PyTorch · 2026-09-02
- GLM-5.3 Model Gets GGUF Quantization Release for Edge Deployment — unsloth · 2026-09-02