3x RTX 3060 mining rig runs FlashNext at 38-40 t/s vs 13.2 t/s on llama.cpp

MD_Reptile · reddit · 2026-10-04

A Reddit user built an open-air rig from three used RTX 3060 12GB mining cards and reports FlashNext with strata hits 38-40 tokens/s at IQ3 quantization, versus just 13.2 t/s on llama.cpp — nearly 3x faster.

Hardware details: Kingwin 8-GPU mining frame stacked on an unraid server, Asus Prime Z370-P, 8th-gen i7, 64GB DDR4, 1000W PSU. The author argues old mining cards remain great value for local LLM inference and is soliciting similar setups.

Original post →

More from Infra

Infra channel →