Architecting pipeline-parallel LLM inference across friends' PCs over the internet
BuildWithEren · reddit · 2026-09-01
A personal project proposal explores running a large model (e.g., 10B) by sharding it across multiple friends' PCs over the internet, as it doesn't fit on a single machine. The focus is on pipeline parallelism architecture under WAN constraints. Key questions include model partitioning on heterogeneous GPUs, latency/bandwidth impact, transport protocols, fault tolerance, and suitable base frameworks like llama.cpp or vLLM.
More from Infra
- Nvidia Earnings: Avoiding Consolidation and Dollars per Gigawatt — Stratechery · 2026-09-01
- TEAS benchmark: measures inference at natural lengths, reporting cost, accuracy, and energy — PontiEdoardo · 2026-09-01
- Nebius GTM on AI Bottlenecks and Inference Demand Explosion — demian_ai · 2026-09-01
- Beyond Model Speed: 19 Distributed Patterns to Optimize AI Latency — bibryam · 2026-09-01
- MongoDB CTO on Database Architecture Evolution and the Unsolved Problem of Agent Memory — The Cognitive Revolution · 2026-09-01
- MongoDB's Pete Johnson on How Retrieval Drives Agent Performance — The Cognitive Revolution · 2026-09-01