Opinion: AI Video Processing Is Too Costly, Needs Native Vision Tools

JoelMahon · reddit · 2026-08-01

The author argues that current multimodal LLMs remain clumsy and expensive at video processing, typically extracting frames at fixed intervals (e.g., 1 fps) into image tokens. This approach completely ignores human vision's mechanism of "central focus + low-cost peripheral motion detection," leading to wasted compute and low efficiency.

The author believes that instead of just waiting for AGI, the industry should prioritize improving models' native tool-calling capabilities and harness frameworks. For instance, simple object tracking shouldn't require expensive full-video parsing or ad-hoc Python scripts but should be built-in as low-cost native tools. They even suggest models could automatically save user-verified ad-hoc tools for global reuse.

Original post →

More from AGI Musings

AGI Musings channel →