IJCAI 2026 tutorial: building unified multimodal models from tokenizers to training
jindong_wang92 · x · 2026-08-15
ML researcher Jindong Wang announced a 3.5-hour tutorial on Unified Multimodal Models at IJCAI 2026, taking place August 16 in Bremen, Germany. His student will present it in person, with hybrid attendance via Zoom, an expected audience of 100–300, and room HS2010 at the conference site.
The tutorial is structured around three central questions:
- How to model? A systematic taxonomy of UMM architectures — External Expert Integration, Modular Joint Modeling, and End-to-End Unified Modeling — with trade-off analysis across autoregressive, diffusion, and hybrid approaches.
- How to represent? The "unified tokenizer" debate: continuous representations (e.g., CLIP) vs. discrete tokens (e.g., VQ-VAE), plus hybrid encodings balancing semantic understanding with generative fidelity.
- How to train? The full training lifecycle, from constructing interleaved image-text data and unified pre-training objectives to post-training alignment methods such as DPO and GRPO.
The agenda traces the evolution of multimodal AI from isolated expertise to unified models, moves through a rigorous UMM definition, architectures and representation, and ends with practical training recipes.
More from Research
- SPD Research: Robot Dexterity via Simulation Pre-training and Minimal Real Data — CSProfKGD · 2026-08-15
- 4 Key Reasoning Strategies Behind LLMs: From CoT to Tree of Thoughts — goyalshaliniuk · 2026-08-15
- MIT Professor: Algebraic Structure Predicts Transformer Length Generalization — ProfBuehlerMIT · 2026-08-15
- New Book 'Imbalanced Data' Debunks Common Myths in Classification Models — Al_Grigor · 2026-08-15
- Meta paper reveals Chinchilla scaling law blind spot, Skaling cuts error — rohanpaul_ai · 2026-08-15
- University Research Team Recruits ComfyUI Users and Creators for Interviews — Lopsided_State_8621 · 2026-08-15