RTX 3090 mini-bench: ninfer cuts TTFT from 3.4s to 29ms, prompt processing ~76x faster

milkipedia · reddit · 2026-09-11

A rigorous RTX 3090 mini-benchmark compares the author's ninfer-3090 stack against llama.cpp on two Qwen models, with careful methodology (3 repeats, token-weighted medians, nonce-based cache busting, untimed warmups, server-side timing).

Results:

The headline: ninfer collapses time-to-first-token from seconds to tens of milliseconds and boosts prompt throughput by 1–2 orders of magnitude, transforming long-prompt interactivity; generation speed is a wash. The author published the full 7-prompt test set (400–12,900 tokens), configs (Ubuntu 24.04, GGUF IQ4XS/Q4KM baselines), and invites reproduction and critique.

Original post →

More from Infra

Infra channel →