Rust Inference Engine Paddock Benchmarks: 3.5x Faster TTFT Than vLLM on Qwen

saltexx · reddit · 2026-08-22

A new inference engine written from scratch in Rust, Paddock, has been released with benchmarks showing significant performance gains over vLLM and llama.cpp. In a Qwen3.6-27B benchmark with 32 concurrent clients, Paddock achieved a TTFT of 697ms, compared to 2.5s for vLLM and 6.9s for llama.cpp. It won 11 of 13 scenarios against vLLM and all 13 against llama.cpp.

Key Features & Architecture:

Caveats:

Original post →

More from Infra

Infra channel →