Qwen 3.8 27B lands on Cerebras at 1,500 tokens per second
gibbonwalker · reddit · 2026-09-04
A Reddit user flagged a Hacker News post: Qwen 3.8 27B is now available on Cerebras, generating at 1,500 tokens per second — a speed the author calls insane. Notable for anyone needing ultra-low-latency inference on a mid-size open-weights model.
More from Infra
- Nutanix CEO: Firms Will Ditch the 'Token Tax' by Running Open-Weight Models on Neoclouds — TiernanRayTech · 2026-09-04
- Sandbox readiness is the new bottleneck in agentic apps; git as state primitive — w_hgm · 2026-09-04
- NVIDIA's open-source PAIR beta routes local AI inference across PCs on your network — Codeblix_Ltd · 2026-09-04
- Computer Imports Hit ~$700B Annualized, Up 126% in a Year on AI Buildout — rjurney · 2026-09-04
- Grok outage traced to Memphis compute center failure, xAI says all systems restored — SpaceXAI · 2026-09-04
- Quartermaster: open-source local AI platform that auto-configures models by free VRAM — OneMoreName1 · 2026-09-04