3x RTX 3060 mining rig runs FlashNext at 38-40 t/s vs 13.2 t/s on llama.cpp
MD_Reptile · reddit · 2026-10-04
A Reddit user built an open-air rig from three used RTX 3060 12GB mining cards and reports FlashNext with strata hits 38-40 tokens/s at IQ3 quantization, versus just 13.2 t/s on llama.cpp — nearly 3x faster.
Hardware details: Kingwin 8-GPU mining frame stacked on an unraid server, Asus Prime Z370-P, 8th-gen i7, 64GB DDR4, 1000W PSU. The author argues old mining cards remain great value for local LLM inference and is soliciting similar setups.
More from Infra
- The rise of overfit inference engines: narrow runtimes ditching generality for raw speed — carteakey · 2026-10-04
- Federated Learning: The Underrated Enterprise AI Pattern That Keeps Data Local — DavidLinthicum · 2026-10-04
- Nvidia Shield TV Pro, a 7-year-old streamer, gets $100 price hike due to AI memory demand — nordicinst · 2026-10-04
- Nvidia's 7-year-old Shield TV just got a $100 price hike as AI squeezes memory supply — Wired AI · 2026-10-04
- RTX 4090 power connector melts after one month of AI generation loads — Roman53275 · 2026-10-04
- Fitting 120,000 NVIDIA B300 GPUs in Tour Montparnasse: a 304MW thought experiment — IgorCarron · 2026-10-04