Qwen 27B hits nearly 100 tokens/s on a local RTX 4090 via Ollama and Pinokio
cocktailpeanut · x · 2026-08-16
User zast57 ran a Qwen 3-series 27B model (written as "Qwen 3.28:27b" in the original post) locally on an RTX 4090 through Ollama + OpenWebUI, deployed via Pinokio. Measured generation speeds:
- Proofreading a text: 90.26 response tokens/s
- Generating Connect 4 game code: 96.77 response tokens/s
He was impressed by how fast it runs. Pinokio creator cocktailpeanut retweeted the benchmark.
More from Infra
- Test Shows Qwen 3.8 27B Q8 Beats Q6 Speed with MTP Enabled on Apple Silicon — jcmyang · 2026-08-16
- NInfer Engine Hits 880 tok/s on RTX 5090 with Qwen3.8-27B — Ond7 · 2026-08-16
- Nvidia reportedly investing $3B in SB Energy to back OpenAI data centers — rohanpaul_ai · 2026-08-16
- Seeed Unveils reComputer RK3576 Edge AI Module — ___Mufasaa · 2026-08-16
- Agent Capacity Planning Guide: Avoiding production surprises — blaizedsouza · 2026-08-16
- How GPU Architecture and Memory Bandwidth Dictate LLM Inference Speed — blaizedsouza · 2026-08-16