Mesh LLM: Distributed AI for pooling local compute
alex_verem · x · 2026-08-25
Mesh LLM is an open-source project that pools GPUs and memory across multiple machines (desktops, laptops, old rigs) to run a single AI model. Models too large for one machine are split across stages in the mesh. It exposes an OpenAI-compatible API, supports CUDA/AMD/Vulkan/Apple Silicon, and works with major open families like Qwen, Llama, DeepSeek, and GLM.
More from Infra
- Nvidia Vera CPU exec calls agentic AI the most complex computing workload in history — firstadopter · 2026-08-25
- Hyperscale Data Centers Use 1.5GWh/Day vs 50GWh for Steel — davidpattersonx · 2026-08-25
- NVIDIA blog: Gemma 4 hits 10,996 OTSU with Vera Rubin optimizations — ricklamers · 2026-08-25
- sPTC speeds up agents via speculative tool calling — a1zhang · 2026-08-25
- Speculative Programmatic Tool Calling Overlaps Code Gen and LLM Inference — a1zhang · 2026-08-25
- Vinci Hits 100k Physics Sims in 24 Hours on Single H200 Node — AnneliesGamble · 2026-08-25