Real-world test: 4-drive array achieves 2tok/s for LLM

carrigmat · x · 2026-08-27

Addressing skepticism about practical software support, the author shared real-world performance data running on llama-server:

The author admits that current llama.cpp software doesn't fully utilize the NVMe array's potential, acting as a bottleneck, but the hardware setup validates the feasibility of the approach.

Related event: Developer shows how to run full trillion-param LLMs locally on CPU for ~$6,000(14 posts)→

Original post →

More from Infra

Infra channel →