YC built AI versions of its partners on GLM-5.2, cutting latency 31% vs OpenAI

ycombinator · x · 2026-09-17

YC built AI versions of its partners for its Office Hour Simulator to help more founders work through startup ideas. After testing lightweight Gemini and OpenAI models, the team moved to GLM-5.2 on a dedicated Wafer endpoint.

YC compared that deployment against GPT-4.1 mini on OpenAI and Gemma 4 31B on Cerebras, evaluating answer quality, latency, and conversation duration: the Wafer setup delivered 31% lower average LLM latency than OpenAI and 44% lower than Cerebras, and users talked to YC's AI partners 2.5 minutes longer on average.

Original post →

More from coding & agent

coding & agent channel →