TensorSharp hits 2x llama.cpp decode throughput on GLM-5.3-Flash
fuzhongkai · reddit · 2026-08-29
Benchmarks of GLM-5.3-Flash (unsloth GGUF UD-Q2KXL, 101GiB, dual GPUs) comparing TensorSharp vs llama.cpp (PR #27754, CUDA 12.8):
- Prefill is a wash: pp2048 2070 vs 2014 t/s; pp32768 1483 vs 1446 t/s.
- Decode doubles: tg64 73.5 vs 36.6 t/s (2.01×); 40.9 t/s after a 17.7K-token prompt, 28.2 t/s after 36K.
- Weights load in 17s from warm cache; CPU-MoE offload (first 10 layers' experts host-resident) decodes at 35–40 t/s.
The 2× comes from a shape-keyed LRU graph cache that replays CUDA-captured graphs instead of rebuilding per token — GLM's 34 KDA layers plus hyper-connections make per-token graphs deep in small ops, exactly where rebuild overhead hurts. Prefill is GEMM-bound, so both engines land within a few percent.
More from Infra
- Google Cloud launches Fault Injection Testing to automate cloud resilience checks — rseroter · 2026-08-29
- Opinion: Local Models Enable a New Class of Software with Embedded Intelligence — carsonfarmer · 2026-08-29
- Designing AI Event Routing: How System Architecture Mirrors Org Charts — zakelfassi · 2026-08-29
- OpenAI's 'Jalapeño' Chip Reportedly Beats Nvidia Blackwell — dylan522p · 2026-08-29
- Cerebras CEO explains why wafer-scale architecture is 2,500X faster than GPU for inference — rohanpaul_ai · 2026-08-29
- Formal verification aids AI safety but may accelerate hardware iteration race — geoffreyirving · 2026-08-29