Run DeepSeek-V3 on $47 Hardware with Pruning and Quantization
StefanoGogioso · x · 2026-08-23
This repository demonstrates running DeepSeek-V3-Flash (mistakenly referred to as v4) on a budget of around $47, achieving 47 tok/s and a 400k context window.
Key techniques include:
- exl3 3-bit quantization.
- 18.5% model pruning.
- Stacking pruning on top of quantization is highlighted as crucial for pushing the hardware floor even lower.
Related event: DeepSeek-V3-Flash Runs at 47 tok/s on a Single $47 Setup(2 posts)→
More from Infra
- Meta spends hundreds of millions on Azure tokens — Beth_Kindig · 2026-08-23
- Over 50% of GitHub's Capacity Issues Caused by Inefficient CI Tasks — cramforce · 2026-08-23
- Call for Top Architects to Design Beautiful Datacenter Exteriors — EddyVGG · 2026-08-23
- 25% chance orbital data centers launch by end of next year — Polymarket · 2026-08-23
- VC proposes network of giant data centers along US-Mexico border — Polymarket · 2026-08-23
- Experiment proposed: Local Qwen model on Mac vs $10k cloud security scan — natesiggard · 2026-08-23