Intel GPU Experience: Running Qwen3.8 27B with 116k Context on a Budget
Accomplished_Yard636 · reddit · 2026-08-20
The author shared their experience running Qwen3.8-27B BF16 locally on a $3K 64GB Intel GPU. It achieves 16 tps with 116k FP8 token context, which is sufficient for coding needs. The author praised the out-of-the-box suspend functionality, noting it is more convenient than Nvidia's solutions.
More from Infra
- Llama-Mobile: 2.7-Bit Quantization Shrinks Llama 3.2 Vision 11B to 3.7GB for Phones — Luka Ribar · 2026-08-24
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24