Idle models might be smarter: a musing on GEMV vs GEMM under low concurrency
karminski3 · x · 2026-09-17
An interesting speculation: when a model is completely idle (concurrency of 1, a single request), extremely quantized inference may take the GEMV path instead of GEMM, incurring lower accumulated precision loss — so the model might effectively be smarter. An unverified technical musing on the relationship between serving load, quantization, and precision.
More from Infra
- Local Dual-RTX Pro Inference Rig: 150 tok/s Decode, 10K tok/s Prefill, Full Build Notes — No_Run8812 · 2026-09-17
- Open-Source Rust+Vulkan Training Backend Supports 143 Modern Transformer Architectures Without CUDA — PhysicsDisastrous462 · 2026-09-17
- GPU prices keep climbing as AI compute demand overwhelming supply — firstadopter · 2026-09-17
- Nearly all top-10 PFAS makers plan production hikes to serve AI chips and data center cooling — jathansadowski · 2026-09-17
- University of Memphis study finds no major air quality deterioration around xAI's Colossus 1 — TinfoilTricorn · 2026-09-17
- Beam moves ~1TB every 30 minutes six months after launching on Bittensor — markjeffrey · 2026-09-17