Mk1.5 Adds Video Localization, Audio and Audio-Visual Understanding
AkshatS07 · x · 2026-09-25
Mk1.5 expands perceptive understanding across modalities and fidelity: new video localization plus audio and audio-visual understanding. To address the pain of real-world agents struggling to fit world knowledge into model capacity, Mk1.5 is trained to reason over tools and external knowledge, cleanly factoring world knowledge from reasoning — leveraging search and reverse image search to answer complex knowledge queries efficiently.
Related event: Mk1.5 Real-Time Multimodal Reasoning Model Launches with 2-5x Lower Latency(4 posts)→
More from Multimodal
- Runway MCP lands in ChatGPT plugin directory, generating full ads from a single prompt — runwayml · 2026-09-26
- Runway MCP brings Gen-4.5 video and image generation into Claude chats — runwayml · 2026-09-26
- Higgsfield Genjutsu Goes Viral for Faking Private-Jet Lifestyles — CurieuxExplorer · 2026-09-26
- One Prompt, No Skills: Opus 5.5 Generates Motion Design Video With Fully Synthesized Music — prasenx · 2026-09-26
- Kelsey Allen's 3DSPA Brings Human-Like Physical Realism Scoring to Generative Video Models — VectorInst · 2026-09-26
- Suno v6 turns choreography into music: dancer Braylon Browner debuts 'Take Me Back' — suno · 2026-09-26