Intern Lumina U2 unifies image, video, and 3D understanding in one diffusion LLM
bdsqlsz · x · 2026-09-05
The Intern team released Intern Lumina U2, a multi-codebook diffusion LLM unifying language, image, video, and 3D understanding with image generation and editing.
- Benchmarks (preliminary): ChartQA 86.52, CharXiv-DQ 83.65, HallusionBench 62.15, MMMU-Pro 36.13 — mostly beating InternVL-UL, LLaDA2.0-Uni, and LLaDA-o; Video-MME 51.26 on video; GenEval 0.81 and DPG 87.10 on generation, slightly behind some rivals.
- Method: a shared MoE Diffusion LLM backbone with discrete VQ levels for visual tokens; spatial positions processed in parallel while codebooks decode sequentially.
Page, code, and model are available; full technical report is upcoming.
More from Models
- OpenAI ships GPT-6 Astra with three new agent-building API updates — gabrielchua · 2026-09-05
- Simon Willison's pelican test shows GPT-6 Astra beats GPT-5.6 at every reasoning level — Simon Willison · 2026-09-05
- Meta publicly releases Muse Spark 1.3 max with stronger coding and agentic performance — EdwardSun0909 · 2026-09-05
- GPT-6 Astra hits 66% on ARC-AGI-3, near-100% with custom harness at ~$360 per game — AccBalanced · 2026-09-05
- 'First 17 seconds expose the huge deficiency in AI creative tools' — GPT-6 Astra demo critiqued — plopesresearch · 2026-09-05
- Astra's $20 plan fits only ~1.5 uses per 5 hours despite token-saving claims — oran_ge · 2026-09-05