Running 27B Model with 3bit On-the-fly Dequant at 24 TPS on A6000

cephaloform · x · 2026-07-29

A developer managed to achieve 24 TPS (tokens per second) when running a 27B model with 3-bit on-the-fly dequantization on an A6000 GPU. The setup was constrained to 100W due to poor cooling, but proved usable for reinforcement learning (RL) tasks.

Related event: 27B LLM Runs at 24 TPS on Single A6000 via 3-bit Dequantization(2 posts)→

Original post →

More from Infra

Infra channel →