Squeezing Qwen 3.8 27B MTP Q8_0 + Vision into 48GB VRAM
Creative-Type9411 · reddit · 2026-08-17
A user shared specific parameters and configurations for running the Qwen 3.8 27B MTP Q80 + Vision model within 48GB VRAM (3x Tesla T4). By fine-tuning settings (e.g., disabling mmap, using mlock, draft-mtp), they achieved 19.2k context at 35t/s without offloading, asking others for optimization tips.
More from Infra
- antirez's DwarfStar: adaptive VRAM/RAM expert placement runs DeepSeek V4 at 45 t/s — antirez · 2026-08-17
- DeepSeek v4 PRO Q2 hits 45 t/s with VRAM/RAM split on DGX Station — antirez · 2026-08-17
- Model routing should be optimized at the harness layer, not the gateway — agihouse_org · 2026-08-17
- Qwen3.8-27B Benchmarks: RPC inference tests across AMD and Nvidia GPUs — tabletuser_blogspot · 2026-08-17
- a16z: How Neoclouds Transformed from Crypto Mining to AI Compute Powerhouses — rohanpaul_ai · 2026-08-17
- AI Spending Gap 625x: Top 1% Spend $7,500/Employee/Month vs Median $12 — rohanpaul_ai · 2026-08-17