glance-vlm speedlab goes open source: MLX 8-bit cuts local camera VLM latency 27.6%

natesiggard · x · 2026-09-24

Yohei Nakajima open-sourced glance-vlm speedlab, a measurement-first latency study that turns any webcam into multiple locally running live AI detectors (emotion, count, object).

Key results (Apple M5, 32GB)

Methodological takeaway: optimize realized model work, not configuration labels or proxy token counts—several attractive changes produced no safe end-to-end win.

Original post →

More from Infra

Infra channel →