Local vision models manage ~3Hz tracking; real-time agent-level speed still a year out
yacineMTB · x · 2026-09-08
Citing a LearnOpenCV test, Sentdex notes GPT-6 Astra Ultra tracking a tennis ball took 11m49s and 7.87M tokens (7.5M cached) end-to-end. Local models (GLM 5.3 Flash, Qwen Next Flash, DSV4F-vision) currently hit only 3Hz for tracking+intelligence; he estimates Astra-level agents at 10-30Hz with <200ms latency are a year or less away. Key bottleneck: inference latency and context token cost, not per-frame capability.
More from Models
- Apodex 1.1 Agent Team hits 63.3% pass rate on FrontierScience-Research, paper lands on Papers with Code — NielsRogge · 2026-09-08
- Gemini Pro runs research task for nearly 5 hours with barely any progress — teortaxesTex · 2026-09-08
- "Infinite Wikipedia That Talks Back": Using Smart LLMs to Get Deliberately Lost — signulll · 2026-09-08
- Astra Tries to Reproduce EUV Scattering Paper Computationally, Fails — teortaxesTex · 2026-09-08
- Command Code ships Muse Spark 1.3 Max with new reasoning effort parameter — alexandr_wang · 2026-09-08
- Jensen Huang confirms GPT-6 Astra trained on 100K+ Grace Blackwell NVL72 systems — rohanpaul_ai · 2026-09-08