Real-time vision model runs expression, object detection and finger counting under 1 second

LinusEkenstam · x · 2026-09-24

Yohei Nakajima (creator of BabyAGI) demoed a real-time vision model running three tasks simultaneously — facial expression recognition, object detection, and finger counting — all in parallel with end-to-end latency under 1 second. Linus Ekenstam shared it calling it "bonkers". The demo highlights how far low-latency, multi-task visual understanding has come.

Original post →

More from Multimodal

Multimodal channel →