Qwen3.8-27B hits ~130 tok/s on Kaggle's free TPU with full 262k context

A-Rahim · reddit · 2026-09-04

The author spent a week getting Qwen3.8-27B (bf16, no quantization) running on Kaggle's free TPU v5e-8, exposed as an OpenAI-compatible endpoint: 130 tok/s single-stream with MTP (78 without), 540 tok/s across 8 streams, 10,300 tok/s prefill (105k-token prompt in 10s, 225k in 28s), full native 262,144 context, and 20 minutes from clicking run to a live endpoint via Cloudflare tunnel. Repo and Kaggle notebook are public.

Original post →

More from Infra

Infra channel →