Peking University Unveils UniMotion: A Unified Multimodal Framework for Human Motion
机器之心 · wechat · 2026-08-15
Researchers from Peking University and collaborators proposed UniMotion, integrating continuous human motion as an independent modality into a Unified Multimodal Model (UMM). The paper has been accepted by ECCV 2026.
- Architecture: Built on Show-o21.5B, using discrete tokens for text and continuous latent representations for Motion and RGB, coordinated via Hybrid Attention and modality-routed LoRA.
- Innovations: Introduces CMA-VAE for motion representation aligned with visual semantics (removable at inference) and LRA (Latent Reconstruction Alignment) for self-supervised pre-calibration.
- Capabilities: Unifies seven tasks including motion understanding, generation, prediction, editing, and visual human pose recovery.
- Performance: Shows competitive results on datasets like HumanML3D, demonstrating better semantic execution in complex motion generation compared to specialized baselines.
More from Multimodal
- Seedance2.5 Update: Now Supports 1080p Generation — CurieuxExplorer · 2026-08-15
- Cartesia releases Sonic 3.6 with enhanced Hindi support and new Urdu/Odia languages — nevrekaraishwa2 · 2026-08-15
- Biologist Sokrypton adds 3D and cartoon style rendering to protein viz tool — sokrypton · 2026-08-15
- AI Tool Used to Build 3D Walkthrough of Skara Brae Dwellings — NathanWilbanks_ · 2026-08-15
- AI Generates Ultra-Realistic Dogfight Video of J-35 vs Rafale — SimplyAnnisa · 2026-08-15
- ComfyUI 0.33 Breaks H3 Motion Context, Fix Released and Upstream Changes Explained — Sad_Berry_4621 · 2026-08-15