Custom CUDA Megakernel Hits 140 tok/s on Qwen3.8-27B With a Single RTX 3090

Adorable_Weakness_39 · reddit · 2026-10-10

Using Claude Opus 5.5 as a coding assistant, the author wrote a CUDA megakernel that runs each whole speculative-decoding cycle in one kernel launch, making checking 4-5 drafted tokens cost about the same as one. Versus llama.cpp with MTP on the same 3090: 140 vs 73 tok/s writing code (1.9x), 95 vs 56 with a 4K-token file in context (1.7x), and 1,600 vs 1,100 tok/s prefill (1.4x). It ships an OpenAI-compatible server as a drop-in llama server replacement. Caveats: built for one model (unsloth Q4KM) on 3090s only, fp16 KVCache to 80k context then q8, outputs match llama.cpp barring rare rounding near-ties. Code and benchmarks are open-sourced at open-jet/megakernel on GitHub.

Original post →

More from Infra

Infra channel →