460M VLM Cuts First-Token Latency to 0.3s on iPhone Using Only 64 Visual Tokens

BTA_Labs · reddit · 2026-08-05

VisionPsy-Nano-460M-Flash is a 460M parameter Vision-Language Model (VLM) optimized for on-device deployment. It achieves massive inference speedups on smartphones through a radical mechanism: compressing visual tokens down to just 64.

The author notes important caveats: the 0.3s metric is only TTFT, it currently relies on a patched llama.cpp fork, context is limited to 8K, and it is designed primarily for single-image queries.

Original post →

More from Infra

Infra channel →