Qwen3.8-27B Hits 140 tok/s on a Single RTX 3090 with a CUDA Megakernel, KL Divergence 0.0009
Adorable_Weakness_39 · reddit · 2026-10-11
An update on the CUDA megakernel project: Qwen3.8-27B (Q4KM) runs 1.4-1.9x faster than llama.cpp on a single RTX 3090, now with accuracy validation.
- Accuracy: KL divergence 0.0007-0.0009 nats vs llama.cpp batched, >99% top-token match, perplexity within noise of reference.
- Speed: 140 vs 73 tok/s on code, 78 vs 51 in reasoning mode, 1600 vs 1100 tok/s prefill. Gains come from speculative decoding checking 4-5 draft tokens per pass (llama.cpp uses 2).
- Limits: single 24GB+ RTX 30/40/50 card, tuned only on 3090; 147K context supported.
- Quantization: unsloth Q4KM native; others fall back to llama.cpp. Open source in the open-jet repo, PRs welcome.
More from Infra
- Dev claims local models can now handle 90% of your work — TheZachMueller · 2026-10-11
- Jevons paradox in AI: falling token prices keep GPU demand tight — AccBalanced · 2026-10-11
- Local AI coding hits only 20-35% of Sonnet's speed in weeks-long app-build tests — julianharris · 2026-10-11
- A 37ms Gap Left GPUs Idle ~40% of the Time; Fixing It Removed the Tiny Idles — HankYeomans · 2026-10-11
- A tiny logger to find which feature ate the OpenAI budget — and a $113k bill horror story — Pangji1003 · 2026-10-11
- Musk courts chip engineers as Terafab targets 10 chip designs a year — elonmusk · 2026-10-11