What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use
riceinmybelly · reddit · 2026-09-11
A Reddit user asks whether an 8GB VRAM card like a 2050 is still useful for local inference: loading embeddings, rerankers, and chat models sequentially for office work. He prioritizes tool use and multilingual support over world knowledge, with vision as a nice-to-have.
He feels small models have been neglected lately and wants recommendations for recent good small models.
More from Infra
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11
- Stanford and Together AI paper: hybrid local-cloud routing cuts AI cost and energy by 60-80% — rohanpaul_ai · 2026-09-11
- Rented GPU Bills: Host CPU and Script Defaults Made Costs 31x Higher — Worldly_North_7213 · 2026-09-11