Running quantized Qwen3.8 coder on a 16GB GPU with Strata
evilsocket · x · 2026-10-01
Security researcher evilsocket demos running an IQ1M-quantized qwen3.8-flash-next-coder on a 16GB NVIDIA GPU using Strata — aggressive quantization putting a large coding model on consumer hardware, a practical local-deployment trick.
More from coding & agent
- Context engineering's next step: knowledge system engineering — ShanRizvi · 2026-10-01
- Dev adds Apple's Genie effect to AI agent plugin with Opus, says he's basically built an agentOS — RileyRalmuto · 2026-10-01
- $5,400 eBay 8x V100 Server Hits 200+ tok/s on a 27B Model with FlashAttention — MzCWzL · 2026-10-01
- As OpenAI et al. ship Jev-style APIs, one dev trains a 149M router model for $6.60 — MaziyarPanahi · 2026-10-01
- I poked OpenAI's Dots all day: it has its own cloud computer and runs three agents in parallel — dry_towelette99 · 2026-10-01
- Claude Code gamified software dev: long-dead GitHub accounts suddenly flood with AI commits — andimarafioti · 2026-10-01