Is T4 Really 170x Slower Than A100?
Future-Structure-296 · reddit · 2026-07-15
While running a point-tracking model, the author found the NVIDIA T4 to be roughly 170x slower than the A100: processing half a video took 0.5 seconds on the A100 and 85 seconds on the T4.
The model settings included: 47 frames, 256×256, batch=1, pure FP32, with an architecture featuring local 4D correlation volume and a transformer for temporal context. The author ruled out common issues like disabled GPUs, models not loaded onto the card, ineffective cudnn.benchmark, and slow performance across two separate T4 machines. They suspect the bottleneck is related to the architecture or operator characteristics and plan to start with profiling to locate the issue.
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11