Running Qwen 3.8 27B on Dual RTX 3090s: 524k Context at 60-88 tok/s
elsung · reddit · 2026-09-05
Reddit user elsung shares a working local setup for Qwen 3.8 27B (HuiHui abliterated config) on dual RTX 3090s: 524k token context at c=1, roughly 60-88 tok/s, with claimed decent accuracy and Chain-of-Draft to curb overthinking. After fixing errors from an earlier post (confusion with Qwen 3.8 Flash Next), the author open-sourced the benchmark configs on GitHub for reproduction.
More from Infra
- Neural network runs on FPGA with no CPU, OS, or software — pure Verilog logic — blaizedsouza · 2026-09-05
- Mighty Heaton takes on the reasons you hate data centers — csuwildcat · 2026-09-05
- Extropic's Z1T models claim up to 140x energy efficiency over GPUs on probabilistic chips — beffjezos · 2026-09-05
- NVIDIA Nsight Compute Now Profiles CUDA Tile Kernels — Two Changes Cut Kernel Time 81% — NVIDIA Developer · 2026-09-05
- Beff Jezos says Alcatraz would make a fantastic spot for an AI datacenter — beffjezos · 2026-09-05
- Free tokens are fueling open-source and local AI, Jason argues citing Jensen — AccBalanced · 2026-09-05