Free Qwen3.8-27B endpoint launched: 262K context, vision, and tool calls
victormustar · x · 2026-08-15
Developer victormustar deployed a free public endpoint for Qwen3.8-27B that requires no token and is fully OpenAI-compatible.
Running on a single H200 GPU, the endpoint supports vision input, tool calling, and adjustable reasoning intensity. It features a 262K-token context window. The deployment costs roughly $5/hour, supports 50 concurrent requests, and achieves 60 tokens/s. The service will remain online for at least 72 hours.
More from Infra
- Intel explores HBM alternatives ZAM and XBM, production may take a decade — JOBhakdi · 2026-08-15
- RootCrak builds Web3 security infrastructure for autonomous agents — Thionne_WTZ · 2026-08-15
- Qwen3.8 Quantization: EXL3 Beats FP8 on Fidelity — malaiwah · 2026-08-15
- Why Do People Think GPUs Are Only Useful for 4-5 Years? — williamfalcon · 2026-08-15
- Idea: AI-Generated Mini Kernels for Bare Metal Server Deployment — _Stocko_ · 2026-08-15
- DwarfStar Aims to Run Frontier Models with Low Energy — antirez · 2026-08-15