A 27B model reaches 24 TPS with on-the-fly 3-bit dequantization on an A6000
cephaloform · x · 2026-07-29
Online 3-bit dequantization pushes a 27B model to 24 TPS on an A6000
The post says the author is getting good results with 4B and 9B setups and wants to scale further. In the reply context, they explain they are trying to make a 27B model usable for RL by doing on-the-fly 3-bit dequantization.
- The current setup reaches 24 TPS.
- It runs on an A6000 capped at 100W because of poor cooling.
- The goal is to make a larger model practical for RL workloads through aggressive quantization/dequantization tricks.
Related event: 27B LLM Runs at 24 TPS on Single A6000 via 3-bit Dequantization(2 posts)→
More from Infra
- Cradle Codec: GPU-Native KV Cache Compression Explained — knowrohit07 · 2026-07-29
- Cheap local intelligence could shift AI workloads away from the cloud — PeterDiamandis · 2026-07-29
- Bull case says AMD profit could 10x as AI spend and inference demand scale — AccBalanced · 2026-07-29
- Cradle Codec compresses KV cache for Ethernet transport between GPU nodes — knowrohit07 · 2026-07-29
- Bittensor raises q from 0.61 to 0.75, easing its emission gate for mid-ranked subnets — markjeffrey · 2026-07-29
- OpenRouter’s moat comes from routing data and tooling it can refine multiple times a day — mmurph · 2026-07-29