One prompt: Claude Code assembles its own video pipeline with Z-Image, Minimax H3 and Music
dkackman11 · reddit · 2026-09-08
A Reddit user shares how a single prompt to Claude Code (running Sonnet) produced a 30-second scored video, with the model autonomously assembling a pipeline that existed nowhere as a template.
- Setup: an MCP server exposing self-describing generation tasks and workflows, plus custom skills explaining how to use it with LTX-2 and Minimax H3 (including H3's prompting skill)
- Sonnet decided on its own to: generate a reference image with Z-Image, block out six clips with Minimax H3 Ref2VA, compose a score with Minimax Music, then concatenate clips and overlay the audio
- The author's key point: Sonnet assembled the pipeline from parts based purely on what the tools said they could do
- Code for the MCP server, skills, and workflow engine is open-sourced on GitHub (diffusers-workflow)
More from coding & agent
- Xiaomi launches invite-only beta of MiMo Desktop, an agent that ships finished work — teortaxesTex · 2026-09-08
- Coval Raises $28M, Launches First Self-Improving Voice Agents With Phonely's Alma Model — massimosgrelli · 2026-09-08
- EU compliance checks as prepaid MCP tools for agents, from €0.10 per call — jithox_AI · 2026-09-08
- Same prompt, four models, one motorcycle: Blender MCP comparison crowns Astra — a10ondr · 2026-09-08
- Dev wires pre-push hook to run tests, SSH-sign a receipt as a git note — fforres · 2026-09-08
- 21 days to OpenAI DevDay: Codex team tracking multiple launches in 28-page deck — reach_vb · 2026-09-08