Cerebras runs Qwen 3.8 27B at ~1,500 tokens/sec, making an AI assistant 19x faster at dinner reservations
Sethwinterroth · x · 2026-09-30
Cerebras demoed an AI personal assistant running Qwen 3.8 27B at 1,500 tokens/sec on its hardware, completing a dinner reservation task 19x faster than a suite of rivals including Grok Bot, Meta Muse, and Claude Cowork. Scott Belsky argues inference speed will increasingly differentiate consumer agent products.
More from coding & agent
- Codex desktop Linux hang bug fixed in latest 26.928.20755 release — cedric_chee · 2026-09-30
- Hands-On Comparison of Top Gen AI Frameworks for Go in 2026: Genkit, Eino, ADK Go and More — rseroter · 2026-09-30
- mitsuhiko: use any llama.cpp model with Pi as a discount classifier — mitsuhiko · 2026-09-30
- Obsidian Mind: 4.7k-star project gives Claude Code, Codex and Gemini agents persistent memory — tom_doerr · 2026-09-30
- Ambion 0.4.0 Released, Focusing on Simplified Core Abstractions — andreisavu · 2026-09-30
- Upcoming talk: 'Spring AI: There and Back Again' on building AI apps with Spring — therealdanvega · 2026-09-30