Wally inference stack debuts: GLM-5.3 Max at 790 tok/s for open models
ycombinator · x · 2026-09-22
RunAnywhereAI launched Wally, an inference stack built to be the fastest place to run open frontier models. Performance snapshot: GLM-5.3 Flash at 380 tok/s, GLM-5.3 Max at 790 tok/s, Qwen3.8-27B at 485 tok/s, and DeepSeek-V4.1 Flash at 615 tok/s.
More from Infra
- RL infra detail: re-prefill over PipelineRL's KV cache reuse, batch size for GPU utilization — stochasticchasm · 2026-09-22
- Raspberry Pi locks devices to original RAM size, blocking aftermarket memory upgrades — ngxson · 2026-09-22
- fal's H3 Max generates 5 seconds of frontier-quality video in just 3 seconds — gorkem · 2026-09-22
- NVIDIA's EPD Disaggregation Cuts Multimodal TTFT Up to 5x, E2E Latency 7x — dl_weekly · 2026-09-22
- Running MiniMax H3 locally on a 16GB Mac: 8-10s clips in 15-20 minutes — coberholzer · 2026-09-22
- A Wild Async RL Config: 30 Steps x 25K Rollouts Per Step at Parallelism 4 — willcb · 2026-09-22