llama.cpp Vulkan benchmarks: mining card P102-100 tops 76-GPU price-performance ranking
tabletuser_blogspot · reddit · 2026-10-09
A Reddit user compiled llama.cpp Vulkan benchmark data (76 GPUs, Llama 2 7B Q40) against 2026 secondhand prices into a local-LLM price-performance ranking. Nvidia's $40-50 P102-100 mining card tops the list (1.3 tok/s per dollar decode) thanks to its 320-bit bus; AMD Instinct MI50 and Radeon VII win on decode cost per token via 1TB/s HBM2. Wide-bus older flagships like GTX 1080 Ti and RTX 2080 Ti often out-decode modern cards under $300, while modern mid-range cards are bottlenecked by narrow 128/192-bit buses.
More from Infra
- Investor: Market Overestimates How Much AI Capex Depends on OpenAI and Anthropic Revenue — matt_slotnick · 2026-10-09
- tinygrad unveils new configurable tinybox starting at $7,000 — HankYeomans · 2026-10-09
- Nvidia claims Vera CPU delivers 2X throughput and memory bandwidth vs x86 for 100MW data centers — Beth_Kindig · 2026-10-09
- Running 12 GPUs off 4 PCIe x16 slots while bypassing the CPU — on paper — TheZachMueller · 2026-10-09
- Atomic Agent Desktop goes open-source: local Qwen/Gemma agents, cloud planning, 69.8% on GAIA L1 — testingcatalog · 2026-10-09
- YC hosts inference-focused Paper Club: naive vs tuned inference can differ 100x in cost — ycombinator · 2026-10-09