HuST Lab's Multimodal Flow: Fully Continuous Unified Language-Vision Generation
hustvl · hf · 2026-10-02
- Motivation: Existing unified multimodal models either discretize both text and images (quantization bottleneck) or mix discrete language with continuous images (modality-specific objectives/sampling). Fully continuous modeling avoids both trade-offs.
- Method: Text blocks and images are organized as ordered continuous hyperchunks; a shared chunk-causal flow backbone learns a single vector field via Flow Matching, with joint attention for cross-modal interaction and modality-specific FFNs.
- Results: MF-1 at 0.6B/1.2B/1.6B scales, pretrained on only 150B tokens, scores 82.8 avg on GenEval+DPG-Bench and 75.3 on VQAv2/MMBench/POPE, competitive with models trained on far more data and beating hybrid/discrete baselines under matched budgets. Code and models released on GitHub.
More from Multimodal
- Two white ducks having a 'very serious little conversation' in AI video — misovalko · 2026-10-02
- Tencent Hunyuan's Tex-Zero Trains 3D Texture Generation Without Any 3D Assets — Tencent-Hunyuan · 2026-10-02
- One prompt gets Claude Fable 5.5 to render a AAA-grade 3D Rube Goldberg machine — imjustnewatai · 2026-10-02
- Seedance 2.5 Video Goes Viral for Ultra-Realistic Early-2000s DV Camcorder Look — SimplyAnnisa · 2026-10-02
- ComfyUI Clips + sync.so Lip-Sync: Edit Before or After Processing? — Positive_Society1876 · 2026-10-02
- Dolly or Zoom? Prompting AI Video for Camera-Only Motion on a Static Car — Classic-Jellyfish-26 · 2026-10-02