YC-backed Isoquant launches GLM-5.3-Flash inference at $0.07/M with 452ms TTFT
ycombinator · x · 2026-09-25
Isoquant, a YC-backed startup, launched its Inference Cloud serving what it claims is the fastest and cheapest GLM-5.3-Flash: $0.07/M input, $0.20/M output and $0.014/M cached tokens, with a 452ms P50 time-to-first-token and 158.9 tok/s throughput — far ahead of Together, CoreWeave, Fireworks and Baseten in its own benchmarks.
The startup says it optimized the full stack — GPU kernels, mixed precision, speculative decoding, KV cache and workload-aware load balancing. It also serves Qwen3.6-27B, offers OpenAI/Anthropic SDK-compatible APIs with streaming, tool calling, structured outputs, image inputs, automatic prompt caching and zero data retention.
More from Infra
- A 32-GPU motherboard for local AI shows up, and people want one in the closet — TheZachMueller · 2026-09-25
- Modern Microprocessors: A 90-Minute Guide Still the Best Crash Course for Systems Engineers — blaizedsouza · 2026-09-25
- GPU returns hit +76% a year as H100 rents jump 49% and B300 rates climb 66% — KyeGomezB · 2026-09-25
- openjev-sglang: SGLang Radix Cache lets Qwen3.6-35B-A3B run 64 Jev decisions in under 1s — multiply_matrix · 2026-09-25
- Ramp benchmarks Jev to replace LLM reranking: 10x lower tail latency at 300ms, 3x cheaper — multiply_matrix · 2026-09-25
- Qualcomm pitches the phone as the AI hub at Snapdragon Summit, aiming for Apple-like cross-device experience — BenBajarin · 2026-09-25