Qwen3.8-27B hits 268 tok/s on single Blackwell GPU with full context
EAccelerate_42 · x · 2026-08-24
A developer released an optimization recipe for Qwen3.8-27B on a single RTX PRO 5000 Blackwell (48GB). Using SGLang, DFlash2 block-16, and Triton attention, it achieves 267.8 tok/s (310.7 peak) with the full 262K context, doubling the EAGLE baseline. Reproducible scripts are provided.
More from coding & agent
- Agent Ran Game Server for 30 Days: Permission Split Made It Safe — RudimentarioUY · 2026-08-24
- Convert sessions between Claude and Codex without burning tokens on summaries — Effective_Farmer2542 · 2026-08-24
- Dad tests Claude Code Remote Control: Coding in small gaps — daniel_mac8 · 2026-08-24
- Practical Discussion: Integrating Local LLMs into Productive Workflows — ThomasAger · 2026-08-24
- Handling agent timeouts when actions might have succeeded — Real_KingZeotic · 2026-08-24
- Security Scan: A Third of Public MCP Servers Lack Safety Hints — Dear-Potential2625 · 2026-08-24