Startup routes inference across desktop devices and shards models across two PCs
punkyrockypocky · reddit · 2026-08-04
A local AI startup says it has built a “multiplayer model serving” network that routes requests across desktop devices, handles failures and recovery, manages device memory, and places models dynamically. It can even shard larger models across two devices, and the team has opened a public endpoint for consumer-facing app traffic. They are now looking for AI builders to test a private inference network before launch, with plans to let people contribute idle compute later.
More from Infra
- Thread questions which model providers can actually sustain 150+ tokens/sec — DanielLockyer · 2026-08-04
- SGLang creator’s exit from xAI leads to RadixArk and a broader AI infra push — hsu_byron · 2026-08-04
- A token-cost estimate says 15T daily tokens could fit in under two SuperPods — teortaxesTex · 2026-08-04
- Multi-agent systems may overwhelm laptop RAM and CPU, poster warns — zephyr_z9 · 2026-08-04
- Doota Mail self-hosting is now fork, secrets, and CI rerun — samgoodwin89 · 2026-08-04
- FL2VA 20B is the default local pick for video generation on a normal PC — cocktailpeanut · 2026-08-04