One RTX 3090 ran Qwen 27B autonomously for 3 weeks — it shipped working CUDA kernels

skeole · reddit · 2026-09-21

Reddit user skeole ran quantized Qwen 27B on a single RTX 3090 with a local agent loop for 21 days, tasking it to build a CUDA inference engine for its own GPU. Highlights:

Takeaway: a quantized 27B can hold a coherent engineering goal for weeks on one consumer card — protocol design matters more than model strength. 15GB dump + rulebook open-sourced on HuggingFace.

Original post →

More from coding & agent

coding & agent channel →