Mk1.5 Adds Video Localization, Audio and Audio-Visual Understanding

AkshatS07 · x · 2026-09-25

Mk1.5 expands perceptive understanding across modalities and fidelity: new video localization plus audio and audio-visual understanding. To address the pain of real-world agents struggling to fit world knowledge into model capacity, Mk1.5 is trained to reason over tools and external knowledge, cleanly factoring world knowledge from reasoning — leveraging search and reverse image search to answer complex knowledge queries efficiently.

Related event: Mk1.5 Real-Time Multimodal Reasoning Model Launches with 2-5x Lower Latency(4 posts)→

Original post →

More from Multimodal

Multimodal channel →