YC-backed Isoquant launches GLM-5.3-Flash inference at $0.07/M with 452ms TTFT

ycombinator · x · 2026-09-25

Isoquant, a YC-backed startup, launched its Inference Cloud serving what it claims is the fastest and cheapest GLM-5.3-Flash: $0.07/M input, $0.20/M output and $0.014/M cached tokens, with a 452ms P50 time-to-first-token and 158.9 tok/s throughput — far ahead of Together, CoreWeave, Fireworks and Baseten in its own benchmarks.

The startup says it optimized the full stack — GPU kernels, mixed precision, speculative decoding, KV cache and workload-aware load balancing. It also serves Qwen3.6-27B, offers OpenAI/Anthropic SDK-compatible APIs with streaming, tool calling, structured outputs, image inputs, automatic prompt caching and zero data retention.

Original post →

More from Infra

Infra channel →