A Full Gemini Workflow: Audio Transcription, Kinetic Subtitles, and 60fps Rendering
AI_Andrew · x · 2026-09-24
A detailed walkthrough of using Gemini to turn a raw screen recording into a polished launch video: transcribe audio with Gemini, apply word-level kinetic subtitle styling, cut in announcement B-roll at the right timestamps, and hardware-accelerated render a 1920×1080/60fps master (18.6 Mbps MP4, faststart, AAC stereo) with a mirrored copy in the source folder.
More from Multimodal
- Artificial Analysis launches TTS leaderboard with new Pronunciation Robustness Benchmark across 95 models — ArtificialAnlys · 2026-09-24
- Google Gemini 3.8 Flash TTS heard in examples: shorthand expansion and contextual pronunciation — ArtificialAnlys · 2026-09-24
- Sample audio: Gemini 3.8 Flash TTS contextually appropriate pronunciation demos — ArtificialAnlys · 2026-09-24
- Gemini 3.8 Flash TTS tops Artificial Analysis pronunciation benchmark at 89.5% — ArtificialAnlys · 2026-09-24
- Runway CEO declares video the final interface alongside real-time video UI demo — _AustinCalvert_ · 2026-09-24
- ACTx486 turns real podcast video into interactive synthetic conversation — karinanguyen · 2026-09-24