Band Builds Audio-Reactive WebGL + Local SD 1.5 Pipeline for Live Improv Music Video
XploitXploit · reddit · 2026-10-08
Frame, a duo from Buenos Aires, shared their fully local open-source workflow for an audio-reactive live music video: Demucs stems + librosa onsets drive a headless three.js scene, composited via green screen, then passed through SD 1.5 img2img with LCM-LoRA (10 steps) and a Depth Anything V2 ControlNet at 1024px, strength 0.59, guidance 2.1, fixed seed. Flicker is controlled by 25% frame blending, later refined with optical-flow blending.
More from Multimodal
- Musk Hypes Grok Bot After User Generates Full Explainer Video From a 15-Second Prompt — elonmusk · 2026-10-08
- AI pop singer Claudia open-sourced: downloadable soul via MCP server, skills and character sheets — Promptmethus · 2026-10-08
- UltraText Bench: bilingual dense text rendering eval, Qwen-Image drops from 86.50 to 42.86 at L3 — Westlake-University · 2026-10-08
- VIEScore2 unifies image eval with spatially grounded defect localization, beating Gemini-3-Flash — TIGER-Lab · 2026-10-08
- AI-generated feature 'A Woman Asleep' enters major film festival's main competition — lmoroney · 2026-10-08
- vLLM-Omni technical report: a unified serving runtime for omni-modal generation — vllm_project · 2026-10-08