Text
Describe the subject, action, environment, camera, lighting, style, and sound.
Seedance 2.0
Create multimodal AI video with Seedance 2.0. Combine text, images, video, and audio references to direct subjects, motion, camera language, visual style, and sound—then refine your result with native audio, video editing, and extension.
Released February 2026 • Multimodal Audio-Video Model

Seedance 2.0
Seedance 2.0
Start with text, an image, or supported reference media. Seedance 2.0 is designed for workflows where creative direction comes from more than a prompt alone.
Describe the subject, action, environment, camera, lighting, style, and sound.
Use images to establish characters, products, scenes, composition, or visual direction.
Reference motion, camera behavior, performance, effects, pacing, or an existing sequence.
Guide voice, rhythm, sound effects, ambience, music, or other audio direction.
Seedance 2.0 is the full general-purpose model of the 2.0 family, built for creators who need stronger quality and control without moving every project into a longer production-oriented workflow.
Combine written instructions with images, video, and audio instead of forcing every creative decision into one prompt.
Use source material to communicate subjects, composition, movement, camera behavior, visual effects, and sound more precisely.
Create visuals and synchronized audio within one generation workflow.
Use supported video editing when a strong result only needs a targeted change.
Extend an existing video while directing the next action, camera movement, or story beat.
Choose standard Seedance 2.0 when output quality and broad multimodal control matter more than the speed advantage of Fast or the efficiency advantage of Mini.
Seedance 2.0 can move from simple text-driven creation to complex reference-based audiovisual workflows.

Multimodal Video
Multimodal Video
Use different media to define the subject, environment, motion, camera language, visual direction, and sound.

Image to Video
Image to Video
Keep the visual foundation while directing action, expression, camera movement, atmosphere, and audio.

Audio + Video
Audio + Video
Generate dialogue, ambience, effects, music, and visual action as part of the same audiovisual workflow.
Seedance 2.0 is ByteDance Seed’s multimodal audio-video generation model released in February 2026. It introduced a unified creation architecture that can understand text, images, video, and audio together.
Use text, images, video, and audio as part of the same creative workflow.
Seedance 2.0 can work with up to 9 images, 3 video clips, and 3 audio clips together with natural-language instructions.
The model supports high-quality multi-shot audiovisual output up to 15 seconds.
Audio and video are generated together, including support for dialogue, effects, ambience, music, and dual-channel audio workflows.
Direct changes to an existing sequence instead of regenerating the full creative idea from the beginning.
Continue a video with additional action and connected shots based on new instructions.
Seedance 2.0 improves motion stability, multi-subject interaction, physical plausibility, instruction following, and controllability compared with the previous generation.
A key advantage of Seedance 2.0 is that different types of input can solve different creative problems.
Best for
Story, action, environment, camera, style, and creative intent.
Use text to explain what happens over time. A useful prompt defines the subject, main action, setting, framing, camera movement, visual treatment, and important audio direction.
Best for
Characters, products, locations, composition, costume, and visual style.
Use images when exact appearance matters more than a verbal description. Tell the model what each image represents and which details should remain recognizable.
Best for
Movement, performance, camera language, pacing, transitions, and effects.
A video reference can demonstrate motion or cinematography that would be difficult to describe precisely with words.
Best for
Voice, music, rhythm, ambience, dialogue, and sound direction.
Use audio when timing or performance depends on sound. Clearly identify what the audio should control in the generated result.
More references are not automatically better. Give each reference a clear role and avoid conflicting creative direction.
Reference media helps turn an abstract prompt into a more specific creative brief. Define what each source contributes, then use the prompt to explain how those elements should work together over time.
Seedance 2.0 treats sound as part of the scene rather than an afterthought.
Generate spoken performance together with the character and scene.
Match environmental or action sounds to what happens on screen.
Build environmental sound that supports the visual setting.
Use music or musical direction when it contributes to the scene.
Coordinate visual action with speech, rhythm, effects, or other sound cues.
Native audio can reduce separate production steps, but generated speech, music, effects, and synchronization may still need iteration.
Use an existing sequence as a reference and describe what should change. Editing is useful when the overall video works but a subject, action, scene detail, or story element needs revision.
Use editing when
Continue a sequence beyond its current ending while carrying forward the subject, scene, motion, and narrative direction.
Use extension when
The models share related workflows but optimize for different stages of video creation.
Best for
Full general-purpose multimodal creation
Choose it when
You want strong output quality, native audio, reference control, editing, and extension in the standard 2.0 workflow.
Main trade-off
More model than necessary for lightweight drafts or speed-first iteration.
Best for
Rapid iteration
Choose it when
Testing prompts, concepts, revisions, and multiple creative directions quickly.
Main trade-off
Optimizes for speed instead of maximizing final-generation quality.
Best for
Efficient higher-volume generation
Choose it when
Creating drafts, social variations, prototypes, and repeated content where efficiency matters.
Main trade-off
Less production headroom than the full 2.0 model.
Best for
Longer production-focused creation
Choose it when
You need up to 30-second storytelling, stronger references, more precise editing, extension, or advanced production control.
Main trade-off
Its additional capability is unnecessary for many standard multimodal tasks.
Choose Seedance 2.0 when you need the complete multimodal 2.0 workflow without specifically optimizing for speed, lower-cost volume, or the longer production workflow of 2.5.

Cinematic Scenes
Direct camera movement, atmosphere, action, visual style, and synchronized sound in a connected scene.

Character Performance
Combine character references, movement, dialogue, expression, and camera direction in one audiovisual workflow.

Product Advertising
Use product and environment references to explore commercial scenes, demonstrations, reveals, and branded concepts.

Reference-Driven Video
Combine subjects, scenes, camera references, motion examples, and audio to communicate a detailed creative brief.

Previsualization
Test storyboards, blocking, camera language, action, pacing, and sound before committing to a larger production workflow.

Social & Campaign Creative
Develop polished concepts for social content, campaign visuals, short narratives, and creative advertising.

01
Use the standard model when you need the full multimodal workflow and prioritize quality and control over Fast or Mini optimization.

02
Begin with text, an image, or the reference workflow currently available in the generator.

03
Upload supported images, videos, or audio when those sources communicate the subject, environment, movement, camera, or sound more clearly than text.

04
Explain which reference defines the character, which defines the scene, which controls movement, and which contributes audio or style.
05
Review subject consistency, motion, composition, camera behavior, audio, and instruction following. Refine the prompt or references, or use editing and extension when the base result already works.
Do not upload several files without explaining what each one should contribute.
Avoid references that show conflicting character, product, costume, or environment details unless the difference is intentional.
References establish context. The prompt should still explain the action, progression, camera behavior, and intended ending.
Use visual references for appearance and a motion/video reference for movement when those directions come from different sources.
A clear sequence with connected actions is easier to direct than many unrelated events competing inside the same clip.
If most of the result works, use editing or extension when available instead of repeatedly rebuilding the entire concept.
Explore different ways to combine visual direction, motion, sound, references, and storytelling with Seedance 2.0.

Multimodal

Cinematic

Character

Product

Animation

Advertising
Longer production
For 30-second storytelling, stronger reference control, advanced editing, extension, and production-focused workflows.
View modelFaster iteration
For rapid prompt testing, creative variations, and workflows where generation speed is the priority.
View modelEfficient generation
For drafts, prototypes, repeated creation, and higher-volume workflows.
View modelAudio + video
For an earlier Seedance workflow focused on native audiovisual generation, dialogue, and character performance.
View modelCombine text, images, video, and audio to direct complete audiovisual scenes with multimodal references, native sound, editing, and extension.