A builder runs 50 tokens/sec on a four-P100 local inference rig, with six GPUs planned
Odd_Caterpillar_2994 · reddit · 2026-07-26
Built a local inference system with four P100 GPUs and plans to expand it to six in a standard case.
- The builder shared photos of the custom setup and said it was tested in advance before scaling up.
- Current performance is about 50 tokens/sec and PP around 530–550 in real use.
- The next step is to add two more P100s and see how much throughput improves.
More from Infra
- PinchTab ships a 12MB Go browser-control binary for AI agents, with HTTP API and token-saving diff mode — Shruti_0810 · 2026-07-26
- AI Bubble Risk Could Be Worse Than the Dot-Com Bust Because So Much More of It Is Debt-Funded — rohanpaul_ai · 2026-07-26
- Jensen Huang says science was never too hard — it was too slow — r0ck3t23 · 2026-07-26
- Paper argues every microsecond matters for GPU collective latency — TheZachMueller · 2026-07-26
- Stage CTO Vinay plans to open-source in-house systems that saved crores in bills — jackedAJ · 2026-07-26
- SmolVM’s disposable macOS sandbox runtime tops r/macOSVMs as interest in ephemeral Mac environments grows — aniketmaurya · 2026-07-26