Self-hosted Qwen 27B on RunPod hits only 17 tok/s generation, making OpenAI API hard to beat on cost per job

yeah280 · reddit · 2026-09-11

Testing Qwen3.8-27B GGUF Q80 on RunPod Serverless ($1.22/hr, 48GB VRAM) against OpenAI for 47k-token transcript analysis: prompt processing is fast (1073 tok/s) but generation crawls at 17.2 tok/s, meaning 8.7 minutes for 9k output tokens and $0.20-0.30 per video. OpenAI cost $5 per 1.6M tokens. llama.cpp ignores NextN/MTP speculative-decoding tensors. Poster weighs quantization, vLLM, better GPUs, and argues cost-per-completed-job is the right metric.

Original post →

More from Infra

Infra channel →