A 27B model reaches 24 TPS with on-the-fly 3-bit dequantization on an A6000

cephaloform · x · 2026-07-29

Online 3-bit dequantization pushes a 27B model to 24 TPS on an A6000

The post says the author is getting good results with 4B and 9B setups and wants to scale further. In the reply context, they explain they are trying to make a 27B model usable for RL by doing on-the-fly 3-bit dequantization.

Related event: 27B LLM Runs at 24 TPS on Single A6000 via 3-bit Dequantization(2 posts)→

Original post →

More from Infra

Infra channel →