Menu
Home
Explore
AI Generators
Video
Image
Models
Library
My Creations
Get MoreCredits
HomeAI ModelsWan AI
Products
  • AI Video Generator
  • AI Image Generator
  • AI Audio Generator
AI Models
  • Seedance AI
  • Kling AI
  • Wan AI
  • Nano Banana
  • GPT Image
Resources
  • Explore
  • Help Center
Company & Legal
  • About Us
  • Contact Us
  • Terms of Service
  • Privacy Policy
  • Refund Policy

© 2026 DoMax AI. All rights reserved.

  1. Home
  2. AI Models
  3. Wan AI

Wan AI

Wan AI Video Generator

Create AI videos with the Wan model family on DoMax AI. Generate from text or images, add synchronized audio, work with character and video references, build multi-shot stories, continue existing clips, and move into longer multimodal production workflows with newer Wan models.

Text to Video•Image to Video•Native Audio•Reference to Video•Video Editing

Updated September 2026

Start CreatingCompare Wan Models

Wan AI

Illustrative footage · Not verified Wan output

Create with Wan AI

Choose the Wan model that matches your workflow, then start from text, an image, or the reference options currently supported by that model. Use the live generator for the exact duration, resolution, audio, upload, and generation settings available on DoMax AI.

Longer multimodal production

Choose Wan 3.0

Advanced video control

Choose Wan 2.7

Reference-driven storytelling

Choose Wan 2.6

Straightforward audio-video generation

Choose Wan 2.5

0 characters
Resolution
Ratio
Duration3s2s30s
Loading your workspace…

Which Wan AI Model Should You Use?

Start with the workflow you need. Each Wan generation expanded the amount of storytelling, references, audio, and video control available to creators.

Wan 3.0

Best for longer all-in-one creation

  • Up to 30-second model-level video generation
  • Text, image, video, and audio workflows
  • Up to 20 multimodal reference materials
  • Can use documents and web-page content as references at the model level
  • Native dialogue, background music, and sound effects
  • Video editing and extension
  • First-frame and first/last-frame control

Choose it when

You need a longer audiovisual sequence, many reference materials, or a workflow that combines generation, reference control, editing, and continuation inside one model.

View Wan 3.0

Wan 2.7

Best for advanced 15-second video control

  • Up to 15-second text and image-driven video
  • Native audio-video synchronization
  • First-frame and first/last-frame generation
  • Video continuation
  • Reference-to-video workflows
  • Instruction-based video editing
  • Stronger performance and cinematic shot changes

Choose it when

You need strong 1080p-capable video creation with editing, continuation, frame control, or more advanced reference workflows without requiring the 30-second 3.0 workflow.

View Wan 2.7

Wan 2.6

Best for reference-driven multi-character video

  • Up to 15-second generation
  • Multi-shot storytelling
  • Synchronized audio and video
  • Improved dialogue and voice generation
  • Reference-to-video
  • Multi-character reference workflows

Choose it when

You want reference-driven character or object videos, multi-person performance, dialogue, and multi-shot narrative generation.

View Wan 2.6

Wan 2.5

Best for straightforward audio-video generation

  • Text-to-video
  • Image-to-video
  • Synchronized audio generation
  • 5- or 10-second model-level video
  • 480p, 720p, and 1080p model-level output
  • Improved motion and instruction following over earlier Wan generations

Choose it when

You need a simpler text-to-video or image-to-video workflow with synchronized audio and do not require newer multi-reference, editing, or long-form controls.

View Wan 2.5

What Is Wan AI?

Wan AI is Alibaba’s family of generative visual models. The video line has evolved from straightforward text and image generation into a broader audiovisual system with synchronized sound, multi-shot storytelling, character references, video editing, continuation, and multimodal creative control.

01

Text to Video

Describe a scene, action, camera movement, atmosphere, and sound to generate a new video.

02

Image to Video

Use an image as the visual foundation and direct how the subject, camera, environment, and story develop over time.

03

Audio + Video

Wan 2.5 and newer video generations support synchronized audiovisual workflows.

04

Multi-Shot Storytelling

Newer Wan models can build connected shots instead of limiting every generation to one isolated camera view.

05

Reference-Driven Video

Use images or videos to guide characters, products, props, environments, movement, or creative direction in supported models.

06

Editing + Continuation

Later generations can modify an existing video or continue a sequence when the useful result should not be regenerated from the beginning.

Compare Wan AI Models

Use this table for a quick decision, then open the dedicated version page when you need detailed generation settings and workflows.

ModelBest forVideo lengthKey differenceMain trade-off
Wan 3.0Longer multimodal productionUp to 30 seconds at the model levelAll-in-One generation, up to 20 multimodal references, native audio, editing, extension, and richer input typesMore capability than necessary for simple short clips
Wan 2.7Advanced short-form controlUp to 15 secondsFirst/last-frame control, continuation, reference video, stronger performance, and video editingShorter and less unified than the 3.0 workflow
Wan 2.6Reference-driven storytellingUp to 15 secondsMulti-shot narrative, stronger voice generation, multi-character reference-to-videoFewer editing and continuation options than 2.7
Wan 2.5Straightforward audiovisual generation5 or 10 secondsIntroduced synchronized audio into the modern Wan text-to-video and image-to-video workflowFewer reference, editing, and long-form controls

Wan 3.0

Best for

Longer multimodal production

Video length

Up to 30 seconds at the model level

Key difference

All-in-One generation, up to 20 multimodal references, native audio, editing, extension, and richer input types

Main trade-off

More capability than necessary for simple short clips

Wan 2.7

Best for

Advanced short-form control

Video length

Up to 15 seconds

Key difference

First/last-frame control, continuation, reference video, stronger performance, and video editing

Main trade-off

Shorter and less unified than the 3.0 workflow

Wan 2.6

Best for

Reference-driven storytelling

Video length

Up to 15 seconds

Key difference

Multi-shot narrative, stronger voice generation, multi-character reference-to-video

Main trade-off

Fewer editing and continuation options than 2.7

Wan 2.5

Best for

Straightforward audiovisual generation

Video length

5 or 10 seconds

Key difference

Introduced synchronized audio into the modern Wan text-to-video and image-to-video workflow

Main trade-off

Fewer reference, editing, and long-form controls

How Wan AI Has Evolved

Illustrative footage · Not verified Wan output

September 2025

Wan 2.5

From

Silent or more limited earlier video workflows

To

Synchronized audiovisual generation

  • text-to-video with synchronized sound
  • image-to-video with synchronized sound
  • up to 10-second generation
  • improved instruction following
  • improved motion
  • improved visual quality
Illustrative footage · Not verified Wan output

December 2025

Wan 2.6

From

Single-clip audiovisual generation

To

Multi-shot and reference-driven storytelling

  • up to 15-second generation
  • intelligent multi-shot scheduling
  • better voice generation
  • more stable multi-speaker dialogue
  • reference-to-video
  • character and object references
  • multi-character performance
Illustrative footage · Not verified Wan output

April 2026

Wan 2.7

From

Reference-driven video

To

Broader generation, continuation, and editing control

  • stronger cinematic performance
  • more expressive dramatic and action scenes
  • first-frame generation
  • first-and-last-frame control
  • video continuation
  • stronger reference consistency
  • video editing
  • mixed image/video reference workflows
Illustrative footage · Not verified Wan output

August 2026

Wan 3.0

From

Multiple specialized workflows

To

All-in-One multimodal video creation

  • native generation up to 30 seconds
  • one model for multiple video-generation workflows
  • up to 20 multimodal reference materials
  • text, image, video, audio, document, and webpage-related inputs
  • native dialogue, BGM, and sound effects
  • editing and video extension
  • stronger long-form creative control

Choose Wan AI by Workflow

Up to 30-second video

Wan 3.0

All-in-One multimodal creation

Wan 3.0

Many reference materials

Wan 3.0

Text-to-video with native sound

Wan 2.5, 2.6, 2.7, or 3.0

Image-to-video with audio

Wan 2.5 or newer

Multi-shot storytelling

Wan 2.6, 2.7, or 3.0

Character or object reference-to-video

Wan 2.6, 2.7, or 3.0

Multi-character reference video

Wan 2.6, 2.7, or 3.0

First and last frame control

Wan 2.7 or 3.0

Video continuation

Wan 2.7 or 3.0

Instruction-based video editing

Wan 2.7 or 3.0

A simpler 5–10 second workflow

Wan 2.5

The newest model is not automatically necessary for every video. Match the model to your required duration, reference complexity, audio, editing, and level of production control.

See What Wan AI Can Create

Wan supports workflows from text prompts to reference-heavy audiovisual storytelling. These shared preview clips are illustrative, not verified Wan model outputs.

Illustrative footage · Not verified Wan output

Text to Video

Turn a Scene Description into Video

Define the subject, action, environment, camera movement, atmosphere, and audio direction.

Illustrative footage · Not verified Wan output

Image to Video

Bring a Still Image into Motion

Use an image to establish the visual starting point while the prompt directs movement, camera behavior, sound, and scene progression.

Illustrative footage · Not verified Wan output

Multi-Shot

Build Connected Shots

Develop a sequence with multiple camera views and connected narrative beats.

Reference

Result

Reference to Video

Direct with Characters and Objects

Use supported reference images or videos to guide subjects, products, props, scenes, appearance, or performance.

Illustrative footage · Not verified Wan output

Video Continuation

Continue an Existing Sequence

Extend a useful clip with another action or story beat instead of creating an unrelated video.

Illustrative footage · Not verified Wan output

Video Editing

Change an Existing Video

Use newer Wan workflows to modify visual elements, style, environment, movement, or other parts of an existing sequence.

Wan AI Video Creation Workflows

01

Text-to-Video

Create a new video directly from a written creative brief.

02

First-Frame Image-to-Video

Use a still image to establish the opening frame and visual identity of the generated clip.

03

First + Last Frame

Supported newer models can use both an opening and target ending frame to give the sequence clearer visual direction.

04

Reference-to-Video

Use characters, objects, scenes, video clips, or other supported media as creative references.

05

Video Continuation

Continue from an existing clip while maintaining the visual context and directing the next sequence.

06

Video Editing

Use instruction-based workflows to change an existing video rather than rebuilding the entire concept from zero.

Why Creators Use Wan AI

01

Audiovisual Creation

Later Wan models generate sound and video together, supporting dialogue, ambience, music, and sound effects.

02

Multi-Shot Storytelling

Build a sequence with connected camera changes instead of limiting every result to a single shot.

03

Reference Control

Use visual references when exact characters, products, objects, scenes, or motion are difficult to communicate with text alone.

04

1080p-Class Video

The current Wan video family includes model-level 1080p generation options across the versions shown on this page.

05

Workflow Choice

Choose between simpler audiovisual generation, reference-heavy production, advanced editing, or longer all-in-one multimodal creation.

What Should You Know Before Choosing a Wan Model?

Before you choose
  • Complex multi-character interaction, hands, exact anatomy, detailed object geometry, and highly constrained choreography can still require multiple generations.
  • Longer videos introduce more opportunities for subject, motion, audio, and scene continuity to drift.
  • Reference material can conflict when characters, products, environments, camera language, or audio cues point in different directions.
  • Native audio reduces separate production steps but does not guarantee perfect dialogue, music, effects, voice consistency, or synchronization.
  • First/last-frame control guides the sequence but does not guarantee that every intermediate frame will follow a predetermined path.
  • Video editing is generative rather than a frame-by-frame replacement for traditional compositing or editing software.
  • More advanced models may consume more resources than necessary for simple tests or short clips.
  • Duration, resolution, reference limits, audio settings, file formats, credits, and pricing depend on the selected model and current DoMax AI integration. Treat the live generator as the source of truth.

How to Use Wan AI on DoMax AI

DoMax AI product interface: Choose a Wan model

01

Choose a Wan model

Match the model to your required duration, reference workflow, audio, editing, and level of creative control.

DoMax AI product interface: Choose how to start

02

Choose how to start

Begin from text, an image, or another reference workflow currently supported by the selected model.

DoMax AI product interface: Describe the video

03

Describe the video

Define the subject, action, scene, camera movement, shot progression, visual style, dialogue, music, or sound when relevant.

DoMax AI product interface: Add references when useful

04

Add references when useful

Use supported images, videos, audio, or other source material when the creative direction is easier to show than describe.

05

Generate and refine

Review motion, subject consistency, shot transitions, reference fidelity, audio, pacing, and overall story progression. Refine a focused part of the brief or use editing and continuation when supported.

Wan AI FAQ

What is Wan AI?
Wan AI is Alibaba’s family of generative visual models. Its video models support workflows including text-to-video, image-to-video, synchronized audio, multi-shot storytelling, reference-to-video, video continuation, editing, and multimodal creative control depending on the selected version.
Which Wan AI versions are available on DoMax AI?
This Wan AI page currently includes Wan 3.0, Wan 2.7, Wan 2.6, and Wan 2.5. Use the live model selector as the source of truth for current availability.
What is the latest Wan video model?
As of September 2026, Wan 3.0 is the newest major Wan video generation included on this page. Alibaba Cloud released Wan 3.0 in August 2026 with an All-in-One workflow, model-level generation up to 30 seconds, multimodal references, native audio, editing, and extension.
Which Wan model should I use?
Use Wan 3.0 for longer and more complex multimodal production, Wan 2.7 for advanced 15-second generation with frame control, continuation, and editing, Wan 2.6 for reference-driven multi-character storytelling, and Wan 2.5 for a simpler audiovisual text-to-video or image-to-video workflow.
Can Wan AI generate video from text?
Yes. Text-to-video is supported across Wan 2.5, 2.6, 2.7, and 3.0. Available duration, resolution, audio, and other controls vary by model.
Can Wan AI turn an image into a video?
Yes. Wan supports image-to-video workflows. Newer generations expand image-driven creation with first-and-last-frame control, video continuation, audio, references, and other multimodal capabilities.
Does Wan AI generate audio?
Yes. Wan 2.5 introduced synchronized audio-video generation into the versions shown here, and later models expand native dialogue, music, effects, and reference-based audiovisual workflows.
Can Wan AI use character or product references?
Yes. Reference-to-video became a major workflow in Wan 2.6 and was expanded in Wan 2.7 and Wan 3.0. Depending on the model, references can help preserve characters, products, props, scenes, visual identity, motion, or voice.
Can Wan AI edit or extend existing video?
Newer Wan models support these workflows. Wan 2.7 includes instruction-based video editing and video continuation, while Wan 3.0 expands editing and extension inside its All-in-One workflow.
How long can Wan AI videos be?
Model-level duration depends on the version. Wan 2.5 supports 5- or 10-second video, Wan 2.6 and 2.7 support up to 15 seconds, and Wan 3.0 supports generation up to 30 seconds. The live DoMax AI generator should be used to confirm the exact duration options available here.

Explore the Wan AI Model Family

Longer All-in-One Creation

Wan 3.0

For up to 30-second generation, rich multimodal references, native audio, editing, continuation, and more complete audiovisual storytelling.

Explore Wan 3.0

Advanced Video Control

Wan 2.7

For 15-second generation, first/last-frame control, continuation, reference workflows, cinematic performance, and video editing.

Explore Wan 2.7

Reference-Driven Storytelling

Wan 2.6

For multi-shot video, synchronized audio, multi-character reference workflows, dialogue, and character-focused creation.

Explore Wan 2.6

Straightforward Audio + Video

Wan 2.5

For text-to-video or image-to-video with synchronized audio in a simpler short-form workflow.

Explore Wan 2.5

Find the Right Wan AI Model

Choose the balance of duration, audio, references, editing, and storytelling control that fits your next AI video.

Start CreatingCompare Wan Models