IJCAI 2026 tutorial: building unified multimodal models from tokenizers to training

jindong_wang92 · x · 2026-08-15

ML researcher Jindong Wang announced a 3.5-hour tutorial on Unified Multimodal Models at IJCAI 2026, taking place August 16 in Bremen, Germany. His student will present it in person, with hybrid attendance via Zoom, an expected audience of 100–300, and room HS2010 at the conference site.

The tutorial is structured around three central questions:

The agenda traces the evolution of multimodal AI from isolated expertise to unified models, moves through a rigorous UMM definition, architectures and representation, and ends with practical training recipes.

Original post →

More from Research

Research channel →