Local vision models manage ~3Hz tracking; real-time agent-level speed still a year out

yacineMTB · x · 2026-09-08

Citing a LearnOpenCV test, Sentdex notes GPT-6 Astra Ultra tracking a tennis ball took 11m49s and 7.87M tokens (7.5M cached) end-to-end. Local models (GLM 5.3 Flash, Qwen Next Flash, DSV4F-vision) currently hit only 3Hz for tracking+intelligence; he estimates Astra-level agents at 10-30Hz with <200ms latency are a year or less away. Key bottleneck: inference latency and context token cost, not per-frame capability.

Original post →

More from Models

Models channel →