Mk1.5 adds video localization, expands into audio and audio-visual understanding

AkshatS07 · x · 2026-09-25

Version Mk1.5 expands perceptual understanding across modalities and fidelity: it adds video localization and extends into audio and audio-visual understanding. The team also shows improved egocentric understanding, reasoning, and cross-embodiment applications, with demos available on their blog.

Related event: Mk1.5 Real-Time Multimodal Reasoning Model Launches with 2-5x Lower Latency(4 posts)→

Original post →

More from Multimodal

Multimodal channel →