Tiny-Qwen update: PyTorch-native build for Qwen 3.8 27B model
No-Compote-6794 · reddit · 2026-08-15
The Tiny-Qwen repo was updated to support building Qwen 3.8 27B from scratch using PyTorch, with token-identical output to Hugging Face transformers. The codebase is cleaner, especially for linear attention. The update enables running the 27B model on 20GB+ memory with minimal overhead. It also includes a simple agentic harness providing CLI access.
More from Infra
- NVIDIA open-sources NeMo Switchyard for dynamic model routing in agent workflows — NVIDIAAI · 2026-08-15
- Vercel ranked as the world's fastest AI Gateway infrastructure — cramforce · 2026-08-15
- mcpp: Auto-generate MCP servers from C++ code via reflection — karurochari · 2026-08-15
- RTX 3090 gets 35 t/s on Qwen 3.8 27B — cviperr33 · 2026-08-15
- CME to launch futures contracts tracking Nvidia H100/B100 compute costs — AccBalanced · 2026-08-15
- Feedback: Serverless GPU capacity tight, placement times high — tobowers · 2026-08-15