Qwen3.8-27B hits ~130 tok/s on Kaggle's free TPU with full 262k context
A-Rahim · reddit · 2026-09-04
The author spent a week getting Qwen3.8-27B (bf16, no quantization) running on Kaggle's free TPU v5e-8, exposed as an OpenAI-compatible endpoint: 130 tok/s single-stream with MTP (78 without), 540 tok/s across 8 streams, 10,300 tok/s prefill (105k-token prompt in 10s, 225k in 28s), full native 262,144 context, and 20 minutes from clicking run to a live endpoint via Cloudflare tunnel. Repo and Kaggle notebook are public.
More from Infra
- Building apps with AI can cost 10,000x more energy than quick chatbot queries — shiringhaffary · 2026-09-04
- Taming a jet-engine local AI server with the motherboard's built-in BMC out-of-band fan control — HankYeomans · 2026-09-04
- Qwen 3.8 27B lands on Cerebras at 1,500 tokens per second — gibbonwalker · 2026-09-04
- DiffusionGemma-26B-A4B demo serves block-diffusion decoding at 800+ tok/s on one B200 — TheMoonMidas · 2026-09-04
- Databricks found $1.2M/year in wasted AI spend from 7 MCP-server bugs — matei_zaharia · 2026-09-04
- Neon-pioneered architecture underpins Lakebase, presented at VLDB — matei_zaharia · 2026-09-04