UChicago Open-Sources Scalable Text-to-Video Pipeline: 1-Hour Video in 25 Minutes

Federal_Effect_3791 · reddit · 2026-08-05

A team of PhD students from the University of Chicago has developed and open-sourced a scalable text-to-video pipeline. The system features automatic stitching, audio, and music generation, allowing users to drag in an entire book and generate a full video with one click.

Currently, it takes about 25 minutes to generate a 1-hour video. The technology has been patented and implemented in UChicago's Humanities department. Originally aimed at making literature "watchable," the team is now seeking community feedback to explore more potential use cases.

Related event: UChicago Open-Sources Long-Form Text-to-Video Pipeline(2 posts)→

Original post →

More from Multimodal

Multimodal channel →