flashtensors: Run hundreds of LLMs on one GPU with sub-2s cold starts

tom_doerr · x · 2026-08-03

flashtensors is a GitHub project that enables running hundreds of large models on a single GPU by loading them from SSD to VRAM up to 10x faster than alternative loaders, achieving cold starts under 2 seconds. Built on ServerlessLLM, it is licensed under Apache 2.0.

Original post →

More from Infra

Infra channel →