Create AI videos with the Wan model family on DoMax AI. Generate from text or images, add synchronized audio, work with character and video references, build multi-shot stories, continue existing clips, and move into longer multimodal production workflows with newer Wan models.
Text to Video•Image to Video•Native Audio•Reference to Video•Video Editing
Choose the Wan model that matches your workflow, then start from text, an image, or the reference options currently supported by that model. Use the live generator for the exact duration, resolution, audio, upload, and generation settings available on DoMax AI.
Longer multimodal production
Choose Wan 3.0
Advanced video control
Choose Wan 2.7
Reference-driven storytelling
Choose Wan 2.6
Straightforward audio-video generation
Choose Wan 2.5
Loading your workspace…
Which Wan AI Model Should You Use?
Start with the workflow you need. Each Wan generation expanded the amount of storytelling, references, audio, and video control available to creators.
Wan 3.0
Best for longer all-in-one creation
Up to 30-second model-level video generation
Text, image, video, and audio workflows
Up to 20 multimodal reference materials
Can use documents and web-page content as references at the model level
Native dialogue, background music, and sound effects
Video editing and extension
First-frame and first/last-frame control
Choose it when
You need a longer audiovisual sequence, many reference materials, or a workflow that combines generation, reference control, editing, and continuation inside one model.
You need strong 1080p-capable video creation with editing, continuation, frame control, or more advanced reference workflows without requiring the 30-second 3.0 workflow.
Improved motion and instruction following over earlier Wan generations
Choose it when
You need a simpler text-to-video or image-to-video workflow with synchronized audio and do not require newer multi-reference, editing, or long-form controls.
Wan AI is Alibaba’s family of generative visual models. The video line has evolved from straightforward text and image generation into a broader audiovisual system with synchronized sound, multi-shot storytelling, character references, video editing, continuation, and multimodal creative control.
01
Text to Video
Describe a scene, action, camera movement, atmosphere, and sound to generate a new video.
02
Image to Video
Use an image as the visual foundation and direct how the subject, camera, environment, and story develop over time.
03
Audio + Video
Wan 2.5 and newer video generations support synchronized audiovisual workflows.
04
Multi-Shot Storytelling
Newer Wan models can build connected shots instead of limiting every generation to one isolated camera view.
05
Reference-Driven Video
Use images or videos to guide characters, products, props, environments, movement, or creative direction in supported models.
06
Editing + Continuation
Later generations can modify an existing video or continue a sequence when the useful result should not be regenerated from the beginning.
Compare Wan AI Models
Use this table for a quick decision, then open the dedicated version page when you need detailed generation settings and workflows.
Model
Best for
Video length
Key difference
Main trade-off
Wan 3.0
Longer multimodal production
Up to 30 seconds at the model level
All-in-One generation, up to 20 multimodal references, native audio, editing, extension, and richer input types
More capability than necessary for simple short clips
Wan 2.7
Advanced short-form control
Up to 15 seconds
First/last-frame control, continuation, reference video, stronger performance, and video editing
Introduced synchronized audio into the modern Wan text-to-video and image-to-video workflow
Main trade-off
Fewer reference, editing, and long-form controls
How Wan AI Has Evolved
Illustrative footage · Not verified Wan output
September 2025
Wan 2.5
From
Silent or more limited earlier video workflows
To
Synchronized audiovisual generation
text-to-video with synchronized sound
image-to-video with synchronized sound
up to 10-second generation
improved instruction following
improved motion
improved visual quality
Illustrative footage · Not verified Wan output
December 2025
Wan 2.6
From
Single-clip audiovisual generation
To
Multi-shot and reference-driven storytelling
up to 15-second generation
intelligent multi-shot scheduling
better voice generation
more stable multi-speaker dialogue
reference-to-video
character and object references
multi-character performance
Illustrative footage · Not verified Wan output
April 2026
Wan 2.7
From
Reference-driven video
To
Broader generation, continuation, and editing control
stronger cinematic performance
more expressive dramatic and action scenes
first-frame generation
first-and-last-frame control
video continuation
stronger reference consistency
video editing
mixed image/video reference workflows
Illustrative footage · Not verified Wan output
August 2026
Wan 3.0
From
Multiple specialized workflows
To
All-in-One multimodal video creation
native generation up to 30 seconds
one model for multiple video-generation workflows
up to 20 multimodal reference materials
text, image, video, audio, document, and webpage-related inputs
native dialogue, BGM, and sound effects
editing and video extension
stronger long-form creative control
Choose Wan AI by Workflow
Up to 30-second video
Wan 3.0
All-in-One multimodal creation
Wan 3.0
Many reference materials
Wan 3.0
Text-to-video with native sound
Wan 2.5, 2.6, 2.7, or 3.0
Image-to-video with audio
Wan 2.5 or newer
Multi-shot storytelling
Wan 2.6, 2.7, or 3.0
Character or object reference-to-video
Wan 2.6, 2.7, or 3.0
Multi-character reference video
Wan 2.6, 2.7, or 3.0
First and last frame control
Wan 2.7 or 3.0
Video continuation
Wan 2.7 or 3.0
Instruction-based video editing
Wan 2.7 or 3.0
A simpler 5–10 second workflow
Wan 2.5
The newest model is not automatically necessary for every video. Match the model to your required duration, reference complexity, audio, editing, and level of production control.
See What Wan AI Can Create
Wan supports workflows from text prompts to reference-heavy audiovisual storytelling. These shared preview clips are illustrative, not verified Wan model outputs.
Illustrative footage · Not verified Wan output
Text to Video
Turn a Scene Description into Video
Define the subject, action, environment, camera movement, atmosphere, and audio direction.
Illustrative footage · Not verified Wan output
Image to Video
Bring a Still Image into Motion
Use an image to establish the visual starting point while the prompt directs movement, camera behavior, sound, and scene progression.
Illustrative footage · Not verified Wan output
Multi-Shot
Build Connected Shots
Develop a sequence with multiple camera views and connected narrative beats.
Reference
Result
Reference to Video
Direct with Characters and Objects
Use supported reference images or videos to guide subjects, products, props, scenes, appearance, or performance.
Illustrative footage · Not verified Wan output
Video Continuation
Continue an Existing Sequence
Extend a useful clip with another action or story beat instead of creating an unrelated video.
Illustrative footage · Not verified Wan output
Video Editing
Change an Existing Video
Use newer Wan workflows to modify visual elements, style, environment, movement, or other parts of an existing sequence.
Wan AI Video Creation Workflows
01
Text-to-Video
Create a new video directly from a written creative brief.
02
First-Frame Image-to-Video
Use a still image to establish the opening frame and visual identity of the generated clip.
03
First + Last Frame
Supported newer models can use both an opening and target ending frame to give the sequence clearer visual direction.
04
Reference-to-Video
Use characters, objects, scenes, video clips, or other supported media as creative references.
05
Video Continuation
Continue from an existing clip while maintaining the visual context and directing the next sequence.
06
Video Editing
Use instruction-based workflows to change an existing video rather than rebuilding the entire concept from zero.
Why Creators Use Wan AI
01
Audiovisual Creation
Later Wan models generate sound and video together, supporting dialogue, ambience, music, and sound effects.
02
Multi-Shot Storytelling
Build a sequence with connected camera changes instead of limiting every result to a single shot.
03
Reference Control
Use visual references when exact characters, products, objects, scenes, or motion are difficult to communicate with text alone.
04
1080p-Class Video
The current Wan video family includes model-level 1080p generation options across the versions shown on this page.
05
Workflow Choice
Choose between simpler audiovisual generation, reference-heavy production, advanced editing, or longer all-in-one multimodal creation.
What Should You Know Before Choosing a Wan Model?
Before you choose
Complex multi-character interaction, hands, exact anatomy, detailed object geometry, and highly constrained choreography can still require multiple generations.
Longer videos introduce more opportunities for subject, motion, audio, and scene continuity to drift.
Reference material can conflict when characters, products, environments, camera language, or audio cues point in different directions.
Native audio reduces separate production steps but does not guarantee perfect dialogue, music, effects, voice consistency, or synchronization.
First/last-frame control guides the sequence but does not guarantee that every intermediate frame will follow a predetermined path.
Video editing is generative rather than a frame-by-frame replacement for traditional compositing or editing software.
More advanced models may consume more resources than necessary for simple tests or short clips.
Duration, resolution, reference limits, audio settings, file formats, credits, and pricing depend on the selected model and current DoMax AI integration. Treat the live generator as the source of truth.
How to Use Wan AI on DoMax AI
01
Choose a Wan model
Match the model to your required duration, reference workflow, audio, editing, and level of creative control.
02
Choose how to start
Begin from text, an image, or another reference workflow currently supported by the selected model.
03
Describe the video
Define the subject, action, scene, camera movement, shot progression, visual style, dialogue, music, or sound when relevant.
04
Add references when useful
Use supported images, videos, audio, or other source material when the creative direction is easier to show than describe.
05
Generate and refine
Review motion, subject consistency, shot transitions, reference fidelity, audio, pacing, and overall story progression. Refine a focused part of the brief or use editing and continuation when supported.
Wan AI FAQ
What is Wan AI?
Wan AI is Alibaba’s family of generative visual models. Its video models support workflows including text-to-video, image-to-video, synchronized audio, multi-shot storytelling, reference-to-video, video continuation, editing, and multimodal creative control depending on the selected version.
Which Wan AI versions are available on DoMax AI?
This Wan AI page currently includes Wan 3.0, Wan 2.7, Wan 2.6, and Wan 2.5. Use the live model selector as the source of truth for current availability.
What is the latest Wan video model?
As of September 2026, Wan 3.0 is the newest major Wan video generation included on this page. Alibaba Cloud released Wan 3.0 in August 2026 with an All-in-One workflow, model-level generation up to 30 seconds, multimodal references, native audio, editing, and extension.
Which Wan model should I use?
Use Wan 3.0 for longer and more complex multimodal production, Wan 2.7 for advanced 15-second generation with frame control, continuation, and editing, Wan 2.6 for reference-driven multi-character storytelling, and Wan 2.5 for a simpler audiovisual text-to-video or image-to-video workflow.
Can Wan AI generate video from text?
Yes. Text-to-video is supported across Wan 2.5, 2.6, 2.7, and 3.0. Available duration, resolution, audio, and other controls vary by model.
Can Wan AI turn an image into a video?
Yes. Wan supports image-to-video workflows. Newer generations expand image-driven creation with first-and-last-frame control, video continuation, audio, references, and other multimodal capabilities.
Does Wan AI generate audio?
Yes. Wan 2.5 introduced synchronized audio-video generation into the versions shown here, and later models expand native dialogue, music, effects, and reference-based audiovisual workflows.
Can Wan AI use character or product references?
Yes. Reference-to-video became a major workflow in Wan 2.6 and was expanded in Wan 2.7 and Wan 3.0. Depending on the model, references can help preserve characters, products, props, scenes, visual identity, motion, or voice.
Can Wan AI edit or extend existing video?
Newer Wan models support these workflows. Wan 2.7 includes instruction-based video editing and video continuation, while Wan 3.0 expands editing and extension inside its All-in-One workflow.
How long can Wan AI videos be?
Model-level duration depends on the version. Wan 2.5 supports 5- or 10-second video, Wan 2.6 and 2.7 support up to 15 seconds, and Wan 3.0 supports generation up to 30 seconds. The live DoMax AI generator should be used to confirm the exact duration options available here.
Explore the Wan AI Model Family
Longer All-in-One Creation
Wan 3.0
For up to 30-second generation, rich multimodal references, native audio, editing, continuation, and more complete audiovisual storytelling.