Wally inference stack debuts: GLM-5.3 Max at 790 tok/s for open models

ycombinator · x · 2026-09-22

RunAnywhereAI launched Wally, an inference stack built to be the fastest place to run open frontier models. Performance snapshot: GLM-5.3 Flash at 380 tok/s, GLM-5.3 Max at 790 tok/s, Qwen3.8-27B at 485 tok/s, and DeepSeek-V4.1 Flash at 615 tok/s.

Original post →

More from Infra

Infra channel →