RadixArk's Miles uses a shared post-training design for VLMs and diffusion models
ying11231 · x · 2026-09-04
RadixArk published a blog post explaining how its Miles model supports cross-modal learning: since the real world is multimodal, models must learn across modalities to understand and recreate it. The team describes a shared post-training design that lets vision-language models (VLMs) and diffusion models learn under a unified framework. Full details are in the linked blog.
More from Multimodal
- HKUST-GZ and Tencent open-source VibeWorlding, an agent that builds interactive 3D worlds from natural language — jiqizhixin · 2026-09-04
- Single-pass 18-second AI video with Wan2.2 int8 workflow, full params shared — r0ni · 2026-09-04
- Midjourney comic trick: give every color one job — coral for pressure, turquoise for the eye — tisch_eins · 2026-09-04
- Seedance 2.5 prompt recreates authentic early-2000s DV vlog footage — 'captured, not generated' — eyishazyer · 2026-09-04
- AI video of a city moving to shade one orchid — full prompt shared — umesh_ai · 2026-09-04
- Fair-fight video test: HappyHorse 1.1 beats Kling 3.0 on character consistency under identical prompts — kimmonismus · 2026-09-04