Testing Multimodal Models for Stand-up Comedy Gen: Natural but Flawed

Ror_Fly · x · 2026-08-12

A developer shared hands-on tests of generating 30-second stand-up comedy clips using a multimodal model. Pros: The character exhibits natural human mannerisms—blinking, pausing, and changing tones—making the dialogue feel significantly less robotic. Cons: The model struggled with prop consistency (like a hat), suffered from odd lip-syncing and chopped words, and remained too expensive to iterate even when generating in 480p.

Related event: Creators Test Fully AI-Generated Stand-Up Comedy Using Claude and Seedance(3 posts)→

Original post →

More from Multimodal

Multimodal channel →