FULL STORY

Seedance 2.5: From Leaks to Global Launch

ByteDance's Seedance 2.5 went through leaks, delays, and previews before its official global launch in late July, introducing native 30-second long video generation and advanced multimodal controls.

2026-07-07 ~ 2026-08-02 · 20 episodes · 302 posts

Episode 1 · CapCut Integrates Seedance 2.5 for Cinematic AI Video (2026-07-07, 3 posts)

ByteDance's Seedance 2.5 video model has been integrated into CapCut, delivering stunning cinematic quality. It can generate 30-second clips using 50 reference materials, drastically simplifying post-production workflows for creators.

Episode 2 · Seedance 2.5 Spotted in Jimeng with 180s Video Support Before Delay (2026-07-09, 9 posts)

The Seedance 2.5 video generation model was originally planned for an imminent release but faced a last-minute delay. Prior to the delay, the model had appeared in the Jimeng UI for testing, sparking widespread community interest due to its massively expanded video generation length.

Key Details and Capability Upgrades

Several users discovered traces of Seedance 2.5 within the Jimeng interface, introducing a new mode called "Ultra-Long Video (Beta)." This mode increases the maximum video generation length from the standard 30 seconds to up to 180 seconds (3 minutes). According to @aziz4ai, this leap in duration signifies a shift in the core bottleneck of video generation, suggesting that the bigger challenge now is having a compelling story to tell.

Reversal and Release Delay

Although users like @koltregaskes initially speculated that an official release was imminent based on the UI exposure—and some even claimed it had begun rolling out globally with early access—the situation quickly reversed. @bdsqlsz and @koltregaskes later noted that multiple sources confirmed Seedance 2.5 was delayed at the last minute, and notification emails had already been sent out. Currently, no specific reasons for the delay or a new launch timeline have been officially provided.

Episode 3 · ByteDance's Seedance 2.5 Leaked: Supports 30s Video Generation (2026-07-12, 5 posts)

ByteDance's upcoming video generation model, Seedance 2.5, experienced a series of API and demo leaks across multiple channels between July 12 and 13, signaling an imminent release. The most notable improvement is the extension of single-generation video duration to 30 seconds, alongside significant upgrades to multimodal inputs, positioning ByteDance at the forefront of the video generation duration race.

Key Details

According to API information shared by @matchaman11, Seedance 2.5 supports a combination of up to 30 images, 10 videos, and 10 audio clips as inputs, with output video duration increased to 30 seconds. Demos forwarded by @testingcatalog and @legitapi both featured a single-take generation of a "cyberpunk hacker robot working in front of multiple monitors," highlighting the model's ability to produce a 30-second video in one go. @WPHero noted that while this is based on early testing, the model's ability to stably generate 30-second clips demonstrates a solid push toward longer video formats.

Community Reactions

On July 13, @markk reported that Seedance 2.5 is "about to be released," framing these leaks as an official preview rather than isolated examples. Although most of the information comes from reposts, the concentrated timeline and mutually corroborating details strongly suggest that ByteDance is gearing up to launch a new version focused on extended video duration and enhanced multimodal capabilities.

Episode 4 · Seedance 2.5 Debuts With a Focus on Realism in Complex Video Scenes (2026-07-15, 10 posts)

Volcano Engine released Seedance 2.5 on July 15 and used the making-of of the short film Chasing the Wind to present its latest video-generation capabilities. Why it drew attention is that both the official post and subsequent commentary converged on one theme: realism, especially in scenes that usually expose AI video weaknesses, such as multi-person movement, crowd shots, sports actions, and object physics.

Officially disclosed updates

In its own post, Volcano Engine said Seedance 2.5 improves multi-person consistency, making simultaneous movement and ensemble shots less likely to break. It also said the model supports more flexible secondary editing, including changing only local parts of a video instead of regenerating everything.

Demos and outside reactions

Several reposts pointed to a football-themed demo built around classic Michael Owen moments. @nikolamr64990 argued the model sets a new bar for physical realism in AI video and said crowd characters now feel like real individuals rather than decorative filler. @ZabihullahAtal highlighted more natural motion, heavier-feeling objects such as the ball, and more consistent spectators. @sanchoyai said professional sports actions are usually where AI fails most visibly, but judged this demo to have no obvious weak spot. @heyabusiddik and @UncannyHarry also emphasized stable rendering in complex football scenes with many people moving at once.

Unconfirmed capability claims and limits

Some additional posts, especially @aziz4ai and a reposted preview, claimed Seedance 2.5 can generate 30-second clips, extend sequences to 3 minutes, and accept up to 50 references across images, video, and audio. However, those details were not fully laid out in the official post included in this cluster. Likewise, claims such as “strongest” or “close to live-action film” remain subjective judgments from posters rather than conclusions backed by systematic benchmarks or failure-case analysis.

Episode 5 · Seedance 2.5 Showcases 30-Second Single-Take AI Video (2026-07-20, 3 posts)

The Seedance 2.5 AI video model shifts the focus from pure image quality to single-take consistency by generating 30-second continuous shots. Tests highlight its strong motion fluidity and potential to revolutionize cinematic storytelling.

Episode 6 · ByteDance Previews Dreamina Seedance 2.5 with Enhanced Control (2026-07-26, 3 posts)

ByteDance's AI platform Dreamina is rolling out Seedance 2.5 globally, focusing on controllable consistency rather than just visual quality. The upgraded model supports up to 50 multimodal references and generates continuous single-take videos up to 30 seconds long.

Episode 7 · ByteDance's Dreamina Launches Aggressive Pricing and Multimodal Upgrades for Seedance 2.0 (2026-07-27, 26 posts)

ByteDance's Dreamina platform has introduced aggressive subsidized pricing and multimodal feature upgrades for its Seedance 2.0 video generation model, making it the most affordable option currently available. New users who purchase their first monthly membership can access a discounted rate of $0.083 per second, costing roughly $4.98 for a 60-second video. The model also features significant improvements in multimodal input support and generation controllability, greatly lowering the barrier to high-quality video creation.

Confirmed

  • 要点 Based on cross-comparison data provided by multiple authors, the generation cost for Dreamina Advanced's Seedance 2.0 is $0.083 per second, translating to about $4.98 for a 60-second 720p video. The pricing page also lists tiers including Free ($0/month), Basic ($9/month), and Standard ($21/month).
  • 要点 Functionally, the model supports simultaneous input of four modalities: image, video, audio, and text, with up to 12 files per generation. It shows enhancements in character and object consistency, text typography consistency, as well as detail and pacing control. It also supports an "intelligent continuation" feature to maintain a coherent storyline. Cross-modal reference capabilities can extract creative elements like camera movements and specific visual effects.
  • 要点 Dreamina integrates tools like Seedance 2.0, Seedream 5.0 Lite, and AI Agents into a single platform, allowing users to complete image generation, video generation, scene editing, and fine-tuning within one workflow.

Unconfirmed

  • 要点 Although the official pricing is highly attractive, standard single-generation costs remain significant. Without an effective free or low-cost trial mechanism, high exploration costs might still keep some creators outside the paywall.

Why it matters

  • 要点 This pricing strategy significantly lowers the financial barrier to high-quality video generation, prompting several creators to advise canceling subscriptions on other platforms and switching to Dreamina to save money.
  • 要点 To address trial-and-error costs, multiple authors suggest a "fast-then-fine" workflow: use Seedance 2.0 mini for quick, low-cost brainstorming, and switch to the full version for the final render once the concept is set, achieving the optimal balance between cost and quality.

6 more related posts →

Episode 8 · ByteDance Seedance 2.5 Launches Tomorrow, Supports 50 Assets & 30s Video (2026-07-28, 11 posts)

ByteDance's Dreamina Seedance 2.5 video generation model is set to officially launch on July 31. The new version supports mixing up to 50 multimodal reference assets (images, videos, etc.) in a single generation and produces videos over 30 seconds long. It will be powered by BytePlusGlobal and integrated into third-party platforms including Runway, Pika, Magnific, and ImagineArt. @ytjessie called the teaser video 'the best version yet,' while @AIwithGhotai and @eyishazyer noted that this update shifts AI video from prompt-based output to a more controllable, multimodal reference workflow.

Confirmed

  • Release date: According to a preview by Qiao Zhengqing, head of Dreamina industry operations, in a WeChat group, Seedance 2.5 is expected to launch on July 31 (Friday). @bdsqlsz and @xiaohu also confirmed the date.
  • Core capabilities: Single generation can mix up to 50 reference images, videos, or other multimodal assets, supporting up to 30 seconds of continuous video.
  • Platform integration: Besides ByteDance's own Dreamina, the model is powered by BytePlusGlobal and will soon be available on Runway, Pika, Magnific, and ImagineArt.

Why it matters

  • Workflow upgrade: @AIwithGhotai and @eyishazyer pointed out that this update is not just about improving output quality but moving AI video from prompt-based generation to a controllable, multimodal reference workflow.
  • Industry reception: @ytjessie described the teaser video as 'the best version yet,' indicating high expectations for quality.
  • Ecosystem expansion: Seedance 2.5, as a foundational video model, is quickly being adopted by multiple major third-party creative tools, showcasing ByteDance's ambition in multimodal and video generation ecosystems.

Episode 9 · ByteDance Seedance 2.0: Low Price, Real Physics, Short-Drama Costs Slashed (2026-07-29, 5 posts)

ByteDance's Seedance 2.0 video generation model on Dreamina has drawn industry attention with a low price of $0.083/sec ($4.98 for 60s) for new users. Tests show strong physical realism, and it's already used in AI short-drama production, significantly cutting compute costs. The model could reshape video creation cost structures, making it worth attention for creators and the industry.

Confirmed

  • Price advantage: Multiple creators confirmed the $0.083/sec new-user price, one of the lowest among 2.0 models.
  • Core capabilities: Reviews by @CodeByPoonam and @eyishazyer confirm realistic physics, multi-person lip-sync, dialogue, ambient sound, and cinematic multi-shot with coherent continuity.
  • Commercial adoption: SCMP reports Chinese studios widely use Seedance for AI comic short-dramas, cutting compute costs for an 80-min episode to 10k yuan, with the market projected at 40 billion yuan this year.

Why it matters

  • Cost restructuring: The low price lowers financial barriers for video creation, offering a cost-effective tool for independent creators.
  • Workflow innovation: @CodeByPoonam shared an efficient workflow: generate base video with Seedance then post-process, demonstrating practical value.

Episode 10 · MiniMax H3 Video Model Tests Impress with Native 2K and Commercial Quality (2026-07-30, 59 posts)

MiniMax's latest video generation model H3 has entered early testing. Hands-on tests by multiple creators and developers show it reaches top-tier performance on several metrics, with capabilities considered on par with or even surpassing Seedance 2.0, demonstrating high commercial value.

Confirmed

  • Core specs: H3 generates videos up to 15 seconds, with native resolution up to 2560×1440 (2K). As an all-in-one model, it supports up to 12 cross-modal reference inputs (including images, audio, video, each up to 15 seconds).
  • Multimodal generation: Introduces a new Omni Reference workflow, breaking the traditional text-only limitation. Users can combine product images, reference videos, and audio for generation; even 5 static images can produce coherent short films or title sequences.
  • Generation capabilities: Multiple creators confirm H3's excellent complex prompt understanding, executing intricate camera movements (e.g., one-shot high-speed flythrough). It performs well in physical motion naturalness, action coherence, text rendering clarity, and high-fidelity speech with lip sync. It also incorporates color, lighting, and composition intent into generation logic rather than mechanically reproducing elements.

Why it matters

  • Commercial-grade quality: In side-by-side tests with identical prompts, multiple reviewers note H3's output surpasses Seedance 2.0 in product interaction dynamics, transition smoothness, and narrative pacing, approaching high-end commercial ads.
  • Workflow innovation: H3 can separate background and characters for replacement, and supports direct output without post-processing. Combining multimodal references, high-quality output, and low barrier to entry makes it a highly practical AI video tool.

39 more related posts →

Episode 11 · ByteDance Releases Seedance 2.5: Native 30s Video and Multimodal Control (2026-07-30, 60 posts)

On July 31, ByteDance's Jimeng (Dreamina) officially released Seedance 2.5, a new video generation model focusing on long-form storytelling and precise multimodal control. It natively generates 30-second videos and offers a long-video mode extendable to 3 minutes. The product is now live on Jimeng and Higgsfield, significantly improving AI video creation workflows.

Confirmed

  • Long video generation: Native generation up to 30 seconds, with multi-round extension to create coherent stories up to 3 minutes. The model can organize multiple logically related shots in a single generation.
  • Multimodal and control: Supports up to 50 multimodal reference inputs (images, videos, audio). Users can precisely control motion trajectories, edit individual elements, and use interactive frame-level editing, including white-box control and green-screen workflows.
  • Platform and cost: Live on Jimeng domestically and internationally, API via BytePlus soon, integrated into Higgsfield. @venturetwins reports cost is about double that of Seedance 2.

Unconfirmed

  • Resolution discrepancy: Official claims native 4K, but beta testers like @op7418 report only 720P support.

Why it matters

  • The multimodal reference input and fine-grained control address core pain points in AI video generation. The strong long-form narrative capability could greatly optimize workflows. Testers like @EXM7777 praise quality as "$200 million movie" level; @mementomori2344323 and @Med1Ai commend physical weight, facial detail, complex interactions (explosions, fire), and seamless morphing. @FellMentKE highlights educational potential.

40 more related posts →

Episode 12 · Google Rolls Out Free Gemini Video Generation and Editing (2026-07-30, 3 posts)

Google announced that users can now create up to 10 videos for free using Gemini until August 2026. The company also introduced Gemini Omni, a new tool offering advanced video editing features like object replacement and spatial interaction.

Episode 13 · Seedance 2.0 Workflows Tested: Creating Commercial-Grade Videos at Low Cost (2026-07-30, 8 posts)

Recently, multiple creators tested and shared AI video generation workflows based on the Seedance 2.0 model, demonstrating its powerful potential and extremely low application costs in music MVs, dynamic blockbusters, and commercial ads. Tests show that, with appropriate prompts and auxiliary tools, this model can produce high-quality video content that meets commercial standards at a very low cost.

Confirmed

  • 要点 Creators @SimplyAnnisa and @eyishazyer respectively demonstrated the experience of using Seedance 2.0 on the Sjolt.ai platform to generate a 15-second fitness Vlog. By meticulously setting parameters such as DV 16mm handheld camera angle, natural camera shake, delayed focus, and dim lighting, they generated highly realistic video footage.
  • 要点 @techhalla shared an image-to-video workflow based on the Magnific platform. First, the Nano Banana Pro model is used to generate a reference image that locks in the overall style and vibe, and then switches to the Seedance 2.0 animation model, combining 3 prompts and rapid cuts to generate a visually striking dynamic blockbuster.
  • 要点 @techhalla also demonstrated a technique for making dynamic band videos: extracting single-member screenshots from the original video as a base, setting the tone with reference images, and using rapid-cut prompts to achieve a coherent dynamic effect.
  • 要点 Creator @FellMentKE pointed out after testing Dreamina Seedance 2.0 that the realism and cinematic quality of AI videos significantly improve when environmental elements like lighting, textures, and backgrounds exhibit natural motion.
  • 要点 A creator showcased a music MV titled "SmileyCard" produced entirely using the Seedance 2 model, verifying the model's ability to independently complete the full-process of long-form video creation.
  • 要点 A user shared a hardcore advertising generation solution, utilizing a combination of Seedance and GPT Image 2, spending only 15 minutes and $5 to generate a coherent product ad featuring complex movements, multiple actions, and comedic scenes.

Why it matters

  • 要点 These practical cases demonstrate that Seedance 2.0, combined with other models or tools, can produce high-quality video content at extremely low costs and high speeds, significantly lowering the barrier to creating complex video ads and music MVs.

Episode 14 · MiniMax Releases Multimodal Model H3 with 2K Native Stereo Video (2026-07-30, 18 posts)

MiniMax officially released its multimodal generation model H3, which unifies understanding and generation of text, images, video, and audio, supporting native stereo audio output and up to 15-second 2K resolution (24fps) video. H3 is now available on Hailuo web, MiniMax API, and platforms like fal and Topview, and will soon be open-sourced with open weights. Its competitive pricing (under one-third of mainstream models) has drawn community attention.

Confirmed

  • Multimodal and output specs: H3 supports multimodal context understanding and can directly output videos with native stereo sound, up to 15 seconds at 2K resolution (24fps), with commercial-grade visual quality. Author @量子位 noted that the model breaks the limitation of traditional video models that only generate raw footage, integrating editing logic, typography, transitions, background music, and visual effects end-to-end to directly output 2K finished videos.
  • 12-Asset Reference: Allows users to combine up to 9 images, 3 video clips, and 3 audio clips as references, precisely locking character appearance, motion trajectories, and audio features.
  • Commercial-grade capabilities: Excels in precise text and brand rendering, video-to-video motion transfer (V2V Motion Transfer), etc., targeting commercial scenarios like advertising and e-commerce.
  • Platforms and pricing: The model is available on fal, Topview, and Hailuo. Author @赛博禅心 revealed that its API price is less than one-third of mainstream models; author @angrypenguinPNG also noted its cost is a fraction while matching Seedance's capabilities.

Unconfirmed

  • Open-source details and timing: Official and multiple creators confirm the model will be open-sourced soon, but the exact timeline is not fully set. Author @cocktailpeanut mentioned that officials said they would release weights "in compliance with laws and regulations" in the coming days; some netizens joked about putting it on BitTorrent due to impatience.

Why it matters

The release of H3 marks further maturity in multimodal fusion for video generation models. The combination of native stereo audio and high-quality 2K visuals, along with fine-grained multi-reference control, significantly lowers the barrier for commercial video production. Its competitive pricing and upcoming open-source release are expected to quickly capture market share in the AI video generation space.

Episode 15 · MiniMax H3 Model Coming Soon with Open Weights, Video and Multimodal Capabilities Spark Buzz (2026-07-30, 7 posts)

MiniMax officially revealed on Hugging Face that its next-generation H3 model is coming soon, with weights to be open-sourced shortly. The official account also showcased generation results on Hailuo AI, impressing the community. Early tests indicate the H3 video model supports native 2560×1440 (2K) resolution and up to 15-second video generation, fully model-generated without post-editing. Additionally, H3 emphasizes native multimodal understanding, allowing users to freely combine up to 12 reference materials (mixing video, text, image, and audio). Community discussions highlight that if H3 can run locally in environments like ComfyUI on an RTX 3060-class GPU, it would greatly benefit developers with limited compute resources. Furthermore, MiniMax's Hugging Face page also updated information on the M3 model, showcasing its capabilities in sparse attention and mathematical proof generation.

Confirmed

  • MiniMax officially stated on Hugging Face that the H3 model is coming soon and weights will be open-sourced quickly.
  • The official account showcased actual generation results based on the H3 model on the Hailuo AI platform, drawing community amazement.
  • Information about the MiniMax M3 model was also revealed, demonstrating its capabilities in sparse attention and mathematical proof generation.

Unconfirmed

  • Video model specs: According to early hands-on feedback, the H3 video model supports native 2560×1440 (2K) resolution and up to 15-second video generation, fully model-generated without post-editing.
  • Native multimodal capability: Some developers report that H3 focuses on native multimodal understanding, allowing users to freely combine up to 12 reference materials (supporting mixed video, text, image, and audio inputs).
  • Generation quality: Testers generated highly cinematic visuals with minimal prompts, and anime-style animation generation was impressive.

Why it matters

  • If H3 can run locally in environments like ComfyUI on an RTX 3060-class GPU, it would greatly benefit developers with limited compute resources, lowering the technical barrier.
  • Its strong multimodal mixed-input capability is seen as an important step toward benchmarking against industry frontiers in the multimodal domain.

Episode 16 · Gemini Omni Transforms Video Generation and Editing (2026-07-30, 3 posts)

Google's Flow studio integrated with Gemini Omni is revolutionizing video workflows via natural language. Users can now easily generate videos, change backgrounds, and auto-dub silent clips, showcasing the model's powerful creative potential.

Episode 17 · MiniMax H3 Tops Video Editing Chart, Announces Open Weights (2026-07-31, 7 posts)

According to the latest evaluation data from Artificial Analysis, MiniMax's latest video model H3 ties with Google's Gemini Omni Flash for first place in the audio-inclusive video editing category (Elo 1130 vs 1121), ranks second globally in text-to-video (with audio), and top three in image-to-video. MiniMax has officially confirmed plans to release the model weights under a community license, making it one of the few leading video models to announce open-sourcing.

Confirmed

  • In Artificial Analysis's video editing leaderboard (with audio), MiniMax-H3 (open-source) and Google's Gemini Omni Flash tie for first place by a narrow margin (Elo 1130 vs 1121)
  • Tied for second globally in text-to-video (with audio), top three in image-to-video
  • Model supports multimodal input, generating 5-15 second 24fps clips with native audio
  • Model supports native 2K resolution and stereo sound generation, with rich video editing features
  • 2K video priced at $7.8
  • MiniMax officially confirms plans to release model weights under a community license

Why it matters

  • H3 is one of the few leading video models to announce open-sourcing weights, potentially lowering the barrier to entry in video generation
  • Native video editing capability breaks the limitation of generation from scratch, supporting direct editing with text, images, etc.
  • The $7.8 price for 2K video is lower than most competitors, potentially attractive for commercial applications

Episode 18 · Topview Launches MiniMax H3 at 30% of Seedance Price (2026-07-31, 9 posts)

Video creation platform Topview announced a partnership with MiniMax (Hailuo AI) to launch the MiniMax H3 video generation model, becoming one of the first platforms to integrate it. The new model emphasizes native multimodal input and high-resolution output, with disruptive pricing aimed at providing high-volume creators with a more cost-effective option.

Confirmed

  • MiniMax H3 supports native 2K resolution output and can generate clips up to 15 seconds long.
  • The model understands text, images, audio, and video simultaneously, enabling cross-modal creative control and native audio-video output.
  • Pricing: H3 costs 70% less than Seedance 2.0, i.e., only 30% of its price.
  • Topview offers Ultra annual users 60 days of unlimited generation.

Why it matters

  • By integrating MiniMax H3, Topview gives creators a new option with solid resolution and duration at a significantly lower cost than competitors, potentially lowering the barrier and expense of high-quality video creation.

Episode 19 · Seedance 2.5 Launches with 30s Native Generation and Local Editing (2026-07-31, 51 posts)

Seedance 2.5 video generation model has launched on Higgsfield and other platforms, with a 14-day unlimited generation event. The model focuses on long-video generation and detail realism, supporting native 30-second output and local editing. Multiple creators have tested it and believe it breaks through the cinematic quality bottleneck of AI video, excelling in physics simulation and multilingual processing, setting a new benchmark for AI filmmaking.

Confirmed

  • Core feature upgrades: According to @Flovaai and others, Seedance 2.5 supports native 30-second long video generation (no stitching), allows up to 50 multimodal reference materials for consistency, and supports precise local editing of specific segments without regenerating the whole video. @AlchainHust and @aziz4ai also verified its excellent performance in 30-second direct output with smooth camera movement and precise motion graphics execution.
  • Visual detail breakthroughs: Multiple creators reported outstanding performance in complex scenes. @Med1Ai noted excellent 3D depth preservation in foggy scenes; @heyabusiddik emphasized frame-level consistency of complex clothing and stitching textures during camera zoom, and successful physically accurate mirror reflections.
  • Lighting and motion optimization: @tischeins demonstrated deep shadows and detail retention in low-light environments. Lip sync, hand details, and natural body movements have been significantly improved, with almost no common AI artifacts. @techhalla verified high quality in complex camera movements (e.g., time freeze and reverse), and @johncoogan found seamless switching between four languages during generation.

Unconfirmed

  • Version differences: @TomLikesRobots mentioned in testing that the model currently supports up to 720p resolution, audio/video reference materials up to 15 seconds, and has stricter safety review than SD 2.0, with imperfections in instruction following. Whether these limitations exist in the upcoming full release remains to be verified.

Why it matters

The release of Seedance 2.5 further raises expectations for AI-generated UGC content. @socialwithaayan found through same-prompt comparison that version 2.5 significantly surpasses 2.0 in texture clarity, motion smoothness, and physical weight. The model not only enhances physical realism in dynamic scenes but also greatly optimizes creator workflows with local editing and multi-reference features.

31 more related posts →

Episode 20 · MiniMax H3 Video Model Nears Release with ComfyUI Support (2026-08-01, 2 posts)

The upcoming MiniMax H3 open-source video model has been deeply optimized by ComfyUI to run on consumer hardware. Its workflow is already accessible in specific running nodes for users to test.