A pragmatic guide to local agentic LLMs: compile buun's llama.cpp free on GitHub runners, Qwen3.8 27B quants span 30x

apollo_mg · reddit · 2026-09-06

A detailed Reddit guide walks from 'saw a tweet' to 'useful local agentic LLM': buun's llama.cpp fork ships a workflowdispatch CUDA build on GitHub's free windows-2022 runners but its upload step is commented out—add a few lines of upload-artifact, then use gh CLI to fork, build, and download. The post also maps Qwen3.8 27B's quantization space (14 weight quants × 8 KV codecs = 112 combos; 56 fit on 16GB), where context length ranges 12,322–373,316 tokens—a 30× spread—and cites its 46.8 Artificial Analysis Agentic Index score.

Original post →

More from coding & agent

coding & agent channel →