Qwen NVFP4 Stress Tested on Four 5060Ti GPUs
joorklee · reddit · 2026-07-11
This is an inference stress test log for unsloth/Qwen3.6-27B-NVFP4, conducted on 4x 5060 Ti GPUs using vLLM and pipeline parallel size=4.
Test Configuration
- Concurrency: 1, 4, 8, 12, 16
- Input length: 6144 tokens
- Output length: 512 tokens
- Max batch tokens: 16384
- Enabled options like prefix caching, chunked prefill, and async scheduling
Purpose
The author primarily wants to observe how prefill / TTFT degrades as concurrency increases in a 4x 5060 Ti setup, using benchmark results to evaluate the actual throughput and latency of the new NVFP4 quantization.
Related event: Unsloth Brings Faster NVFP4 Quantization to Qwen3.6(3 posts)→
More from Infra
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22