Qwen 3.8 27B lands on Cerebras at 1,500 tokens per second

gibbonwalker · reddit · 2026-09-04

A Reddit user flagged a Hacker News post: Qwen 3.8 27B is now available on Cerebras, generating at 1,500 tokens per second — a speed the author calls insane. Notable for anyone needing ultra-low-latency inference on a mid-size open-weights model.

Original post →

More from Infra

Infra channel →