DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
Yupeng Xie, Zhenyang Wang, Liangwei Wang, Jiayi Zhu, Zhouan Shen, Yuyu Luo
cs.HC, cs.AI, cs.DB, cs.MA
2026-09-27
DataMagic turns raw tables into narrated data videos via a declarative spec plus generate-then-orchestrate agents: quality 2.13→3.89/5, exec rate >95%, user task time down 79.7%.
Data videos combine animated charts, voice narration, and synchronized motion into one narrative medium, common in business reporting and data journalism. Producing them takes three skill sets at once: data analysis, narrative design, and video editing. Existing tools each cover a slice. Automatic visualization systems (DeepEye, HAIChart) output static charts with no narrative or motion. Authoring tools (WonderFlow, Data Playwright) start from charts that are already made, so raw data never enters the pipeline. Pixel-level video models (Sora, Veo3) generate end to end but hallucinate numbers and cannot trace a visual element back to its source record.
What about asking an LLM to write the full rendering code in one shot? The paper tested this. On 109 real-world samples, GPT-5 direct generation scores 2.13/5 on quality, and execution success rates across four models range from 48.62% to 86.24%, with a typical failure being a crashed render leaving a blank screen. Two problems drive this: there is no unified representation for charts, narration, animation, and their timing; and the design space of insights, chart types, and narrative orderings is too large for greedy scene-by-scene generation.
DataMagic, an open-source system from HKUST (Guangzhou), stacks a declarative intermediate representation on a two-stage agent pipeline.
DVSpec encodes a video as a JSON scene sequence, each scene a self-contained five-tuple (id, type, content, narration, animation) carrying its own chart configuration and data slice. Two mechanisms fix long-standing pain points:
The Generate-then-Orchestrate strategy handles the search problem. Stage one: a Story Planner decomposes the query into orthogonal subtasks, a Data Manager generates Python code to extract a micro-dataset per subtask, and a Visual Designer designs charts in parallel, yielding a pool of candidate scenes (narration and animation are deliberately left empty here). Stage two: a Narration Director selects scenes by insight value and query coverage, orders them under a narrative pattern (Freytag's Pyramid by default, plus inverted-pyramid, comparison, and other structures), and writes narration through a sliding context window over adjacent scenes; an Animation Coordinator binds entities mentioned in narration to visual elements and sets trigger indices. Output is constrained to 60-120 seconds with at most 7 initial scenes.
DVSpec doubles as the shared state behind three interaction modes: canvas manipulation, script editing, and natural-language commands. Every edit updates fields of one target scene and re-renders only that scene.
109 samples from 60 real-world datasets (36 from T2R-bench, 24 from DAComp-DA, up to 150,000 rows), scored 1-5 on five dimensions by Gemini-2.5-Pro as judge, validated against three human experts on 60 videos (Pearson r = 0.91).
| Method | Exec rate | Avg. score |
| GPT-5 direct | 86.24% | 2.13 |
| Claude-Sonnet-4 direct | 84.40% | 2.18 |
| DataMagic (GPT-5) | 95.41% | 3.38 |
| DataMagic (Claude-Sonnet-4) | 98.17% | 3.89 |
The largest gains land exactly where direct generation is worst: Animation from 1.84 to 4.03 on average (+120%), Narrative from 2.00 to 3.43 (+72%). Direct-generation failures cluster in three classes: render crashes with blank screens, design defects such as truncation and occlusion, and audio-visual desynchronization ("blind narration"); the declarative binding eliminates the third class by design. Swapping the weakest model, DeepSeek-V3.2, into the framework lifts its execution rate from 48.62% to 96.33%, evidence the framework is model-agnostic. The cost is time: about 57 seconds for direct generation versus roughly 176 seconds end to end, a threefold slowdown traded for quality and stability.
Ablations: removing the Story Planner drops the average to 3.44, removing global orchestration to 3.54. Even without orchestration, Animation holds at 3.95 (down only 10%), confirming that audio-visual sync lives in the DVSpec declaration layer and does not depend on global planning.
User study (N=12, within-subjects, baseline a conversational workflow on the same Claude Sonnet 4.5): task time falls from 39.2 to 8.0 minutes (-79.7%), five of six NASA-TLX workload dimensions drop significantly, and narrative coherence rates 6.6 vs 2.8 on a 7-point scale. Seven of twelve participants overran the 40-minute target under the baseline; all twelve finished within it with DataMagic.
For BI and reporting teams this is an open-source path from raw tables to finished videos (React plus Remotion, with DVSpec reusable on its own). The broader lesson matches a pattern repeating across tasks that demand precision and traceability: a structured intermediate representation plus role-divided agents beats end-to-end generation. Because DVSpec serves as both the agents' interface and the human's editable state, automated output can be refined locally instead of regenerated wholesale, which black-box generation cannot offer. As a paper it is an HCI and visualization systems contribution; nothing here trains a model.
Author-stated: single-table inputs only, no joins; bound to one renderer (Remotion); animations follow structured templates, a conservative motion language that trades expressiveness for reliability; all user intervention happens after generation, with pre-generation plan confirmation left to future work. Appendix failure cases show the seams: extreme value distributions clip a pie chart out of view, and in one case the narration says 40.6% while the on-screen label reads 40.9%, so cross-modal consistency is not airtight.
Reasons for caution: scoring relies on an LLM judge, validated against experts on only 60 samples, so dimension-level robustness needs wider verification; the user study covers 12 graduate students; the baseline is a conversational LLM workflow rather than existing authoring tools, justified because those tools require pre-made charts; and the evaluation set was assembled by the authors from two benchmarks, in a field that still lacks a shared public benchmark for cross-system comparison.