Reddit survey: What GPUs do you use for local inference? 3090/4090 still dominate
ocean_protocol · reddit · 2026-07-30
A Reddit user surveys what GPUs people actually use for local inference. Most replies still use 3090/4090 for models like Llama 3 8B, Mistral 7B; a few use A100/H100 as used prices drop. For larger models like Llama 3 70B, some use multi-GPU or quantized GGUF on a single card.
More from Infra
- NVIDIA Scales Matrix Factorization to 1M×1M, Doubling Single-GPU Capacity — marc_stampfli · 2026-07-30
- Building Local AI with Dual AMD R9700s: Is ROCm a Viable Alternative to CUDA? — Syosse-CH · 2026-07-30
- Run Gemma 4 Locally with 16GB RAM: A Zero-Cost Fully Offline Setup Guide — FinanceYF5 · 2026-07-30
- 4090+5060 Ti Hybrid Inference: Runs 122B Model at 37 t/s — Dry_Long3157 · 2026-07-30
- OpenAI to Consume 40% of Global DRAM: The Rise of the Metered Intelligence Complex — Shimano-No-Kyoken · 2026-07-30
- AI Giants Accused of Hiding $1.65 Trillion in Off-Balance-Sheet Debt, Echoing Enron — marigo · 2026-07-30