ACTx486 turns real podcast video into interactive synthetic conversation
karinanguyen · x · 2026-09-24
Research demo ACTx486 blends real footage (Elon Musk on The Joe Rogan Experience) with generated synthetic faces, voices and speech: tap the video to ask questions, interrupt, switch languages, or dive into generated scenes like life on Mars.
Capabilities include real-time diagrams timed to the host's speech, personalized answers based on a viewer profile with cards sent to your phone, and overlaying new worlds while the original dialogue keeps playing in sync.
The team stresses statements in generated segments were never made by the people depicted, the project is unaffiliated with the show, and their write-up covers risks and implications for the future of media.
Related event: ACTx486 Turns Videos into Real-Time Conversational Agents(4 posts)→
More from Multimodal
- Claude Opus 5.5 generates a music video in ~1 shot: "not good, but not without interest" — NathanpmYoung · 2026-09-24
- Gemini 3.8 Flash TTS and Flash-Lite TTS land on Merge Gateway, top Hume voice quality index — shensi · 2026-09-24
- LemonSlice Launches Character World Model-1, a Real-Time Interactive Avatar Model — mhdfaran · 2026-09-24
- Reverse workflow: unpack a reference image with Extract Prompt, then remix it — JaynitMakwana · 2026-09-24
- Unverified claim: 'GPT-6 Astra' builds full video projects via Codex + Dreamina CLI — JaynitMakwana · 2026-09-24
- MiniMax H3 roundup: video VAE 2.2x faster encoding, music model under 12GB — optimisticalish · 2026-09-24