A pragmatic guide to local agentic LLMs: compile buun's llama.cpp free on GitHub runners, Qwen3.8 27B quants span 30x
apollo_mg · reddit · 2026-09-06
A detailed Reddit guide walks from 'saw a tweet' to 'useful local agentic LLM': buun's llama.cpp fork ships a workflowdispatch CUDA build on GitHub's free windows-2022 runners but its upload step is commented out—add a few lines of upload-artifact, then use gh CLI to fork, build, and download. The post also maps Qwen3.8 27B's quantization space (14 weight quants × 8 KV codecs = 112 combos; 56 fit on 16GB), where context length ranges 12,322–373,316 tokens—a 30× spread—and cites its 46.8 Artificial Analysis Agentic Index score.
More from coding & agent
- OpenClaw v2026.9.2 ships with GPT-6 Astra support across 1,245 PRs from 232 contributors — steipete · 2026-09-06
- GPT-6 Astra mixes Computer Use and direct file writing to build LEGO models — dkundel · 2026-09-06
- Merge Agent Handler adds six callable connectors including TinyFish and Microsoft Graph Security — shensi · 2026-09-06
- Spec-to-APK entirely on an Android phone via Termux CLI and AI — Digital_Otorongo · 2026-09-06
- Astra review: executes hour-long agentic tasks while iterating on the plan without losing the thread — ivan_bezdomny · 2026-09-06
- Stop picking coding models on vibes: run your own A/B/C refactor tests — robleclerc · 2026-09-06