Self-hosted Qwen 27B on RunPod hits only 17 tok/s generation, making OpenAI API hard to beat on cost per job
yeah280 · reddit · 2026-09-11
Testing Qwen3.8-27B GGUF Q80 on RunPod Serverless ($1.22/hr, 48GB VRAM) against OpenAI for 47k-token transcript analysis: prompt processing is fast (1073 tok/s) but generation crawls at 17.2 tok/s, meaning 8.7 minutes for 9k output tokens and $0.20-0.30 per video. OpenAI cost $5 per 1.6M tokens. llama.cpp ignores NextN/MTP speculative-decoding tensors. Poster weighs quantization, vLLM, better GPUs, and argues cost-per-completed-job is the right metric.
More from Infra
- Blackstone's Biggest AI Bet Is Compute, Backing Deals with Google, Nvidia, Anthropic — abhiadesai · 2026-09-11
- Frontier models now independently reach for speculative decoding and kernel optimization on InferenceBench — maksym_andr · 2026-09-11
- A Beginner-Friendly Guide to Budget Multi-GPU Local LLM Setups — lblblllb · 2026-09-11
- Chinese Nvidia challenger Enflame jumps 179% in Shanghai debut, raises $910M — pstAsiatech · 2026-09-11
- Qwen3.8 Flash Next hits 49 tok/s locally on 2x RTX 3090 with FlashNext llama.cpp fork — whiteh4cker · 2026-09-11
- Qdrant lines up three free community events with 4-hour vector tech stream — qdrant_engine · 2026-09-11