flashtensors: Run hundreds of LLMs on one GPU with sub-2s cold starts
tom_doerr · x · 2026-08-03
flashtensors is a GitHub project that enables running hundreds of large models on a single GPU by loading them from SSD to VRAM up to 10x faster than alternative loaders, achieving cold starts under 2 seconds. Built on ServerlessLLM, it is licensed under Apache 2.0.
More from Infra
- Chip Stocks Show Strong YTD Performance Amid AI Boom — firstadopter · 2026-08-03
- Minimax H3 Pruning Optimization: Massive VRAM Reduction with Zero Quality Loss — Valuable_Issue_ · 2026-08-03
- llama.cpp Adds MTP Support for Qwen3-Next, Enabling Full-Speed Inference — jacek2023 · 2026-08-03
- Developer Regrets Skipping 512GB Mac Studio, Hoards Cash for Local AI Compute — MannyKayy · 2026-08-03
- VRAM Bottleneck? Developer Stuck Running MiniMax Video Model Locally on RTX 4080 — witcherknight · 2026-08-03
- Open-Source System Serves VLA Models to 10+ Robots on a Single GPU — danfei_xu · 2026-08-03