Custom CUDA Megakernel Hits 140 tok/s on Qwen3.8-27B With a Single RTX 3090
Adorable_Weakness_39 · reddit · 2026-10-10
Using Claude Opus 5.5 as a coding assistant, the author wrote a CUDA megakernel that runs each whole speculative-decoding cycle in one kernel launch, making checking 4-5 drafted tokens cost about the same as one. Versus llama.cpp with MTP on the same 3090: 140 vs 73 tok/s writing code (1.9x), 95 vs 56 with a 4K-token file in context (1.7x), and 1,600 vs 1,100 tok/s prefill (1.4x). It ships an OpenAI-compatible server as a drop-in llama server replacement. Caveats: built for one model (unsloth Q4KM) on 3090s only, fp16 KVCache to 80k context then q8, outputs match llama.cpp barring rare rounding near-ties. Code and benchmarks are open-sourced at open-jet/megakernel on GitHub.
More from Infra
- Texas data center power queue hits 474GW, 90% from data centers — FinanceYF5 · 2026-10-10
- DDR5 hits $7,200 for 256GB as engineers treat RAM as an appreciating AI asset — CtrlAltDwayne · 2026-10-10
- Lemire reruns 2026 WebSocket benchmarks: his old Bun-vs-Node.js result was wrong, Anthropic bought Bun and Cloudflare bought Deno — lemire · 2026-10-10
- Cornell/IBM paper: shared KV cache cuts looped-transformer memory 76-79% while improving quality — yoavartzi · 2026-10-10
- 4GB VRAM local LLM users: is there anything faster than llama.cpp? — your_real_Fathe_ · 2026-10-10
- Cascade GPU Topology Can Be Slower: The PCIe Hop Trap in Multi-GPU P2P — TheZachMueller · 2026-10-10