8 video VLMs benchmarked on a single RTX 3090: TTFT and throughput compared
SkyLordOmega · reddit · 2026-09-19
As part of a NIST/TREC visual QA submission, the author measured inference speed of 8 video VLMs of varying parameter sizes and architectures on a single RTX 3090, covering two tasks: returning ten ranked textual answers per video query, and ranking candidate answers by likelihood.
Results are given in two tables plus an environment photo. Caveats the author stresses:
- TTFT includes the vision encoder and server-side video decode, so it's not apples-to-apples with text LLM prefill;
- Concurrency = 1, no batched inference;
- Throughput says nothing about quality — two of the models produced output rejected in the final submission.
The author notes this is an anecdotal reference rather than a rigorous benchmark.
More from Infra
- Hyperscaler depreciation to hit $255B in 2026, $581B by 2029 — analysts — luisdans · 2026-09-19
- Running an M5 Max MacBook in low-power mode full time to trade tok/s for fan-less silence — joshwhiton · 2026-09-19
- Open-source inference rise pressures closed AI infra margins, analyst argues — AccBalanced · 2026-09-19
- RTX 4090 upgrade turns into a two-week ComfyUI freezing nightmare with no fix in sight — Jimbo_1995 · 2026-09-19
- Tesla AI5 chip enters trial production on Samsung's 2nm Texas fab, mass output by 2027 — XFreeze · 2026-09-19
- Data center water debate: producing one gallon of milk takes 628 gallons of water — dfinke · 2026-09-19