LLM Steers Compounding Video Hallucinations to the Beat: A Controlled Drift Experiment
ART-ficial-Ignorance · reddit · 2026-09-04
The author chained Gemma 12b, LTX 2.3, an audio-reactive LoRA and Wan2GP into a controlled hallucination experiment: the LLM holds the full video structure and music energy map but writes each 6.9s clip prompt only after 'seeing' the previous clip's final frame, steering accumulated artifacts back toward the planned arc. Using a stationary camera on one morphing object made errors compound faster than the previous infinite-hallway run, yielding a black ceramic form that develops its own visual grammar before destabilizing into a network-like surface. Next step: let the LLM control LoRA strength and other settings.
More from Multimodal
- Saudi-backed HUMAIN launches Arabic-first voice suite with ASR, TTS and voice agents — HUMAIN · 2026-09-04
- CAT-Flow: Training-Free Adaptive Steps Cut Flow Matching Generation by Up to 40% — chaumian · 2026-09-04
- Text-to-4DGS demo shows text prompts driving dynamic 4D Gaussian scenes — solo_solipsist · 2026-09-04
- Rebuilding an entire campus from a single photo: blogger uses Atlas on UMD's CS building — gowthami_s · 2026-09-04
- Oscar-nominated writer pens one scene, eight artists turn it into AI films for HUMANCENTRIC — egeberkina · 2026-09-04
- Quiver AI teases Arrow 2.0 with bold retro monster illustration 'made entirely of code' — stuffyokodraws · 2026-09-04