YC's AI Office Hours: Wafer hits 379ms vs Cerebras 674ms, 44% lower latency
ycombinator · x · 2026-09-27
For YC's AI Office Hours — AI versions of its partners giving startup advice at conversational speed — Wafer ran GLM-5.2 at an average 379ms versus 674ms for Gemma 4 31B on Cerebras, 44% lower latency with a much larger model. After testing lightweight Gemma and OpenAI models, YC moved to a dedicated Wafer endpoint; Wafer agents tuned serving for request rate, cache usage, and prompt/response lengths. Users spent 2.5 minutes longer talking to the AI partners on Wafer.
More from Infra
- Databricks targets 100x cheaper warehouse AI functions for all DB vendors — sh_reya · 2026-09-27
- Ollaya runs decision models locally: 5 answers in 8–10 ms on an RTX 4090, no token generation — petrusenko_max · 2026-09-27
- DHH rewrites ttfx in asm with AI, hitting 8x mean speedup on an Intel 135U — olcan · 2026-09-27
- ASML Makes Zero Revenue From Europe as No Chip Factories Are Built There — pstAsiatech · 2026-09-27
- Agents Make GPUs Need More CPUs: 3 Data Points From the Last 10 Days — tengyanAI · 2026-09-27
- Cloudflare Workers adds zero-dependency OpenTelemetry tracing APIs — irvinebroque · 2026-09-27