Pairing a Blackwell RTX 4000 24GB with an Old RTX 3060 12GB for Local LLMs
Otherwise-Tangelo-52 · reddit · 2026-10-06
A user picked up a Blackwell RTX 4000 24GB slightly under MSRP and asks whether it can be combined with an old RTX 3060 12GB to load the largest possible local model for coding tasks with a safe VRAM buffer. The thread touches on tensor split across heterogeneous GPUs, memory bandwidth tradeoffs, and driver support.
More from Infra
- 462GB DeepSeek model runs on two desk-side DGX Sparks with experts squeezed to 2.77 bits — Teknium · 2026-10-06
- Wilderness Society calls for immediate moratorium on AI data centers on US public lands — Polymarket · 2026-10-06
- KCoral: Shared GPU Benchmark Environment Speeds Agentic Kernel Evaluation 2.58x on B200 — BeidiChen · 2026-10-06
- Enterprises were promised an AI infrastructure future, but the power grid wasn't ready — DavidLinthicum · 2026-10-06
- Dual MI50 64GB HBM2 local inference rig built around a PLX switch — Savantskie1 · 2026-10-06
- Clef Flash 9B Q8_0 on RTX 5060 Ti: llama.cpp ~30% faster prompt eval than Ollama (3,000 vs 2,100 t/s) — ngxson · 2026-10-06