Qwen3.8-27B on 16GB VRAM: 50 tok/s with 85k Context, Config Included
brainExploded99 · reddit · 2026-08-19
A Reddit user shares running Qwen3.8-27B on an Nvidia 5070Ti 16GB, achieving 85k context and 50 tok/s by removing the MTP layer and tuning config. Includes full config and tips like q8 cache and ubatch-size ablation.
More from coding & agent
- Obsidian Plugin Hyo Adds Task Mode: Turn Agent Chats Into Manageable Tasks — evielync · 2026-08-19
- LangChain updates onboarding for managed deep agents — hwchase17 · 2026-08-19
- Perplexity's MCP server ships as an Agent Plugin for Codex, Cursor, and VS Code — inductionheads · 2026-08-19
- AI Agent Scaffolding Guide: Context, Tools, and Orchestration Patterns — tom_doerr · 2026-08-19
- Rewriting AI Agent in Rust: Memory-Safe with Secure Plugin System — doodlestein · 2026-08-19
- Transparent Headroom MITM config: no client changes needed — m0ntanoid · 2026-08-19