Qwen3.8-27B Hits 140 tok/s on a Single RTX 3090 with a CUDA Megakernel, KL Divergence 0.0009

Adorable_Weakness_39 · reddit · 2026-10-11

An update on the CUDA megakernel project: Qwen3.8-27B (Q4KM) runs 1.4-1.9x faster than llama.cpp on a single RTX 3090, now with accuracy validation.

Original post →

More from Infra

Infra channel →