LithosAI ships Day-0 API inference for DeepSeek-V4.1-Flash at 250+ tokens/s per user
JiaZhihao · x · 2026-09-11
Inference provider LithosAI announced Day-0 public API support for DeepSeek-V4.1-Flash, live within 12 hours of the model's release. Standard tier delivers 250+ tokens/s per user with faster tiers coming, targeting ultra low-latency agent workloads.
More from Infra
- Scoop: Microsoft plans massive data center expansion to 38 gigawatts — dinabass · 2026-09-11
- Together launches preemptible compute for GPU clusters at 50% of on-demand price — togethercompute · 2026-09-11
- New Model Hits Opus-Level Benchmarks at Wild Efficiency, RL Infra Details Emerge — nrehiew_ · 2026-09-11
- 2,500 Concurrent Sandboxes Per Node, Cross-Scaffold Checkpoint Merging Yield Free Gains — nrehiew_ · 2026-09-11
- Only 15 Fused Kernels in Prefill: V4.1's Inference Stack and Minute-Scale SWA Cache — nrehiew_ · 2026-09-11
- DeepSeek V4.1 Flash infra: Siglip-style vision encoder and shadow indexer workers — nrehiew_ · 2026-09-11