YC-backed Prism launches agent-optimized inference cloud, serving DeepSeek V4.1 Flash at 547 tok/s
ycombinator · x · 2026-09-25
YC-backed startup Prism launched an inference cloud for open-source LLMs that uses agents to optimize deployments for cost, latency and throughput.
- Serves DeepSeek-V4.1-Flash (552B MoE, 1M context), GLM-5.3, Kimi K3, Qwen3.8 (2.4T MoE, 95B active) and more via OpenAI- and Anthropic-compatible APIs
- Claims up to 547 tok/s on DeepSeek V4.1 Flash via its Prism Engine, up to 5.8x faster than major providers
- Positioned as serverless inference aimed at agent workloads
More from Infra
- Google to launch TPUs into space next week on Falcon 9 to test orbital AI data centers — McDonaghMatthew · 2026-09-25
- NVIDIA now tops the list of America's biggest businesses after a decade-long climb — lemire · 2026-09-25
- LithosAI uses GPU virtualization to push the Pareto frontier of agentic inference — JiaZhihao · 2026-09-25
- Analog chip runs LLM attention 100x faster than H100 using 70,000x less power, Nature paper claims — anselm · 2026-09-25
- Healthcare AI's GPU dilemma: balancing latency-sensitive clinical inference against batch research workloads — Arindam_1729 · 2026-09-25
- LexiPanel: Open-Source Control Panel Runs Local LLM, Image and Audio AI on Your Own GPU — W61k3r · 2026-09-25