Real-world test: 4-drive array achieves 2tok/s for LLM
carrigmat · x · 2026-08-27
Addressing skepticism about practical software support, the author shared real-world performance data running on llama-server:
- Current Setup: A 4-drive NVMe array + 768GB RAM achieves a speed of 2 tokens/second.
- Projected Performance: Upgrading to a 16-drive array + 2x EPYC 9175F CPUs is projected to reach 4-6 tokens/second.
The author admits that current llama.cpp software doesn't fully utilize the NVMe array's potential, acting as a bottleneck, but the hardware setup validates the feasibility of the approach.
More from Infra
- Prefix Sliding enables efficient test-time scaling by cutting memory costs — Niklas Muennighoff · 2026-08-27
- LightningAI offers instant H100 access on its self-owned AI cloud — LightningAI · 2026-08-27
- Cohere releases Parse 5, claiming best price-performance for enterprise document parsing — irombie · 2026-08-27
- M7 Ultra may feature native FP8, potentially boosting GLM 5.3-flash performance — Brilliant-Hall1387 · 2026-08-27
- ChronoScale announces 50MW NVIDIA GB300 deployment with Microsoft for AI inference — r_jegaa · 2026-08-27
- Engineering Win: mxfp8 x mxfp4 Matmul Outperforms Standard mxfp8 — zephyr_z9 · 2026-08-27