Qwen 27B hits nearly 100 tokens/s on a local RTX 4090 via Ollama and Pinokio

cocktailpeanut · x · 2026-08-16

User zast57 ran a Qwen 3-series 27B model (written as "Qwen 3.28:27b" in the original post) locally on an RTX 4090 through Ollama + OpenWebUI, deployed via Pinokio. Measured generation speeds:

He was impressed by how fast it runs. Pinokio creator cocktailpeanut retweeted the benchmark.

Original post →

More from Infra

Infra channel →