llama.cpp advanced usage: run Qwen3.8 27B on 32GB VRAM with speculative decoding
ggerganov · x · 2026-08-15
ggerganov shared detailed commands for running Qwen3.8 27B on 32GB VRAM (e.g., RTX 5090) using llama.cpp, including GGUF quantization, speculative decoding (draft-mtp), large context window (196k), and agent mode.
Related event: Community Shares Qwen3.8-27B Deployment on 32GB VRAM(4 posts)→
More from coding & agent
- App built and shipped via TestFlight in under 50 hours using Replit — amasad · 2026-08-15
- Training models on agent harnesses leverages general capabilities for domain-specific intuition — rosstaylor90 · 2026-08-15
- Closed-Loop Prompt Optimization Framework for Production AI Systems — blaizedsouza · 2026-08-15
- Agentic Engineering is just software engineering best practices — rseroter · 2026-08-15
- Hands-on with xAI Grok Bot: Autonomous Context & Collaboration — aakashgupta · 2026-08-15
- A practical guide to CLAUDE.md in Claude Code: hierarchy, loading, best practices — 4310sy · 2026-08-15