RadixArk's Miles uses a shared post-training design for VLMs and diffusion models

ying11231 · x · 2026-09-04

RadixArk published a blog post explaining how its Miles model supports cross-modal learning: since the real world is multimodal, models must learn across modalities to understand and recreate it. The team describes a shared post-training design that lets vision-language models (VLMs) and diffusion models learn under a unified framework. Full details are in the linked blog.

Original post →

More from Multimodal

Multimodal channel →