Cerebras runs Qwen 3.8 27B at ~1,500 tokens/sec, making an AI assistant 19x faster at dinner reservations

Sethwinterroth · x · 2026-09-30

Cerebras demoed an AI personal assistant running Qwen 3.8 27B at 1,500 tokens/sec on its hardware, completing a dinner reservation task 19x faster than a suite of rivals including Grok Bot, Meta Muse, and Claude Cowork. Scott Belsky argues inference speed will increasingly differentiate consumer agent products.

Original post →

More from coding & agent

coding & agent channel →