SlimServe Open-Sourced: Runs DeepSeek at 1k tok/s on 4x A100s
QuixiAI · x · 2026-08-11
QuixiAI has open-sourced SlimServe, an inference serving framework, demonstrating impressive concurrency capabilities.
When running the DeepSeek v4 Flash 0731 model on 4x A100 GPUs, it achieved the following benchmarks:
- Single-request speed: 175 tokens/s
- Throughput for 64 concurrent requests: 1,000 tokens/s
The framework aims to provide a highly efficient solution for LLM deployment.
More from Infra
- Demand for AI Gateways and Model Routing Sees Explosive Growth — shensi · 2026-08-11
- OpenAI's Letter to Texas Governor on Responsible AI Infrastructure — borowcy · 2026-08-11
- Testing DGX Spark with H3: First Text-to-Video Run via ComfyUI — allaboutai-kris · 2026-08-11
- China's First Domestic Front-end Lithography Tool Enters Production, Competing with 90s ASML — pstAsiatech · 2026-08-11
- Under 10% of Enterprises Scale AI; Compute Shortage to Persist — BenBajarin · 2026-08-11
- Amazon Backs Texas Gas Plant That May Become Top US Climate Polluter for AI — Ars Technica AI · 2026-08-11