TensorSharp outperforms llama.cpp for GLM-5.2 on long prompts
AdhesivenessWeird770 · reddit · 2026-08-20
Benchmarking GLM-5.2-UD-IQ2XXS on 3x RTX PRO 6000 Blackwell shows TensorSharp surpasses llama.cpp by up to 50% on long contexts (2048+ tokens) by optimizing MoE routing with larger micro-batches, while llama.cpp retains an edge on short prompts due to lower fixed overhead.
More from Infra
- Investor Applies Munger's Three-Basket Rule: Data Centers Are a Clear Yes — RachelVT42 · 2026-08-20
- Benchmark: FP8 Models Run 5x Faster Than GGUF on Low-End Hardware — ROBOTTTTT13 · 2026-08-20
- MiniMax H3 + Qwen Image Edit recreates Sherlock shots, 10s per gen on a 3090 — nikhilprasanth · 2026-08-20
- Ops Horror Stories: Redesigning Storage Layout on a Live Ceph Cluster Invites Chernobyl Jokes — l4rz · 2026-08-20
- MiniMax H3 runs locally on a 16GB Mac via VPIPE — TgoAI · 2026-08-20
- 35B MoE in 7GB: Mach-1 ships GGUFs and llama.cpp fork for edge devices — pmttyji · 2026-08-20