NUS's MaLiang-Harness exposes the Program-to-Visual gap in code-driven image/video generation
NationalUniversityofSingapore · hf · 2026-09-30
NUS introduces MaLiang-Harness, a unified framework for MLLM-driven visual generation via executable programs. Key points:
- Defines the Program-to-Visual (P2V) gap: a program can execute correctly while violating requested composition, appearance, or motion.
- Framework organizes generation into a persistent process of construction, inspection, and revision, with three mechanisms: PEG state (persistent programs + task context), TGP (links edits to rendered evidence), REV (revision-aware editing and verification).
- Benchmarked 11 closed-source MLLMs on MaLiang-IBench and 4 on MaLiang-VBench. GPT-6-Astra hits 100% generation success, with 96.0% of image and 76.9% of video tasks meeting all quality thresholds.
- Notably finds that general capability scores poorly predict visual generation performance—similarly scored models differ substantially.
Code: github.com/gulucaptain/MaLiang-Harness
More from Multimodal
- AutoRef open-sourced: harness optimization for agentic multi-reference image generation — NunyaBuzor · 2026-09-30
- Claude Opus 5.5 video imagines the moment Girl with a Pearl Earring was painted — justin_hart · 2026-09-30
- Reddit user shares AI-generated sci-fi short film Erebus-9: The Hiveborn Archives — himeonisama · 2026-09-30
- Teaching Claude Opus to Annotate Speech Timing for ElevenLabs Voiceovers Proves Tricky — RileyRalmuto · 2026-09-30
- Alibaba DAMO unveils WorldAttention for efficient interactive video world models — Alibaba-DAMO-Academy · 2026-09-30
- Thinking Reward Model: rubric-first scoring sets open-source visual generation reward SOTA — Xuehai Bai · 2026-09-30