Menu
Home
Explore
AI Generators
Video
Image
Models
Library
My Creations
Get MoreCredits
HomeAI ModelsWan AIWan 2.7
Products
  • AI Video Generator
  • AI Image Generator
  • AI Audio Generator
AI Models
  • Seedance AI
  • Kling AI
  • Wan AI
  • Nano Banana
  • GPT Image
Resources
  • Explore
  • Help Center
Company & Legal
  • About Us
  • Contact Us
  • Terms of Service
  • Privacy Policy
  • Refund Policy

© 2026 DoMax AI. All rights reserved.

Plans before generating for higher coherence
Updated June 2026

Wan 2.7 AI Video Generator

Alibaba Tongyi Lab's most controllable AI video model. Its planning-first mode maps the composition before generation, which usually means higher coherence and fewer artifacts. It also brings four video workflows for editing, extension, and reference-driven generation.

By Alibaba Tongyi LabOpen-source familyFour video workflowsUp to 9 reference imagesNo subscription
Planning Mode4 WorkflowsThousand-FacePrecise ColorLong Text
Try Wan 2.7Buy Credits
0 characters
Resolution
Ratio
Duration3s2s15s
Loading your workspace…

Model angle

The Thinking AI video model for higher-control professional workflows.

Why it exists

Planning-first generation, workflow depth, and open-source Wan heritage set it apart.

Planning-mode claims on this page are described as of June 2026. The feature improves coherence but does not eliminate every artifact, and faster standard generation is still the better choice for quick iteration.

Wan 2.7 Workflow Previews

These shared demo clips illustrate possible workflows. They are not verified Wan 2.7 outputs or actual planning-mode comparisons; prompts are conceptual ideas.

AllPlanning A/BFramesVideo ExtendVideo ReferenceVideo EditingColor-ControlledText-to-Video
text-to-videoIllustrative footage

Luxury lobby with layered foreground, midground, and background motion planned through Thinking Mode.

Related prompt idea: wan27-sample-2

text-to-videoIllustrative footage

Same luxury lobby prompt without Thinking Mode for side-by-side coherence comparison.

Related prompt idea: wan27-sample-1

text-to-videoIllustrative footage

Crowded train platform with multiple unique faces demonstrating Thousand-Face Realism.

text-to-videoIllustrative footage

Coffee campaign scene controlled by exact palette values to demonstrate brand-color precision.

framesIllustrative footage

Sunrise frame to sunset frame transition over a mountain ridge using the Frames workflow.

framesIllustrative footage

Closed flagship store to launch-night crowd transition anchored by first and last frames.

video-extendIllustrative footage

Extend a character exit shot with consistent camera follow and preserved wardrobe detail.

video-extendIllustrative footage

Continue a product reveal clip while holding material reflections and studio light continuity.

video-referenceIllustrative footage

Reference a handheld fashion film, then restage the same motion in a monochrome gallery editorial.

video-referenceIllustrative footage

Borrow the pacing of a POV city walk and apply it to a premium automotive reveal scene.

video-editingIllustrative footage

Convert a summer beach scene into winter with snow, colder grading, and heavier wind.

video-editingIllustrative footage

Change daytime retail footage into a premium night launch with neon reflections and darker atmosphere.

text-to-videoIllustrative footage

Long in-video signage and product typography rendered inside a modern retail environment.

text-to-videoIllustrative footage

Another Thinking Mode A/B scene: crowded kitchen choreography with complex reflections and motion beats.

Related prompt idea: wan27-sample-15

text-to-videoIllustrative footage

Same kitchen choreography prompt without Thinking Mode for side-by-side comparison.

Related prompt idea: wan27-sample-14

What is Wan 2.7?

Wan 2.7 is Alibaba Tongyi Lab's most advanced AI video and image generation model, released in April 2026. It builds on the open-source Wan family, whose earlier versions earned unusual global momentum through public-weight distribution, Hugging Face adoption, and benchmark visibility across Chinese AI video discussions.
The defining feature is Thinking Mode. Before generating, the model first interprets the prompt and plans the composition: where the subject should sit, how the camera should behave, what the light should emphasize, and how motion should unfold. That planning step is why it can often produce more coherent output with fewer common artifact patterns than direct generation workflows.
This release also expands the family into a broader professional workflow system. Rather than behaving like a single text-to-video endpoint, it is framed around four production modes: Frames, Video Extend, Video Reference, and Video Editing. That turns the model into more than a generator; it becomes a practical post-generation and reference-driven creative tool.
Other headline additions include Thousand-Face Realism, precise color control for brand work, long-text rendering, and broader reference support. On DoMax, the model runs via hosted cloud inference, so users get the benefit of the family's open-source provenance without downloading weights or managing local GPUs. For the underlying family context, see the Wan AI hub.

Why Creators Choose This Release

Planning-first generation

This release is framed around a planning-first workflow. Instead of generating immediately, it interprets prompt intent and decides composition logic before rendering the final clip.

Four video workflows

Frames, Video Extend, Video Reference, and Video Editing give this release unusually broad workflow coverage for teams that do more than raw text-to-video generation.

Thousand-Face Realism

This feature targets one of AI video's recurring weaknesses: cloned-looking faces. It aims for stronger character differentiation across multi-person scenes.

Precise color control

Use HEX-style palette instructions for brand work that requires exact, repeatable color direction rather than approximate prompt-only interpretation.

Long-text rendering

It is positioned for stronger in-video text generation, which matters for signage, interface shots, retail scenes, and ad creative with visible typography.

Open-source family heritage

It inherits the family’s open-source credibility. On DoMax you use hosted inference, but the architectural lineage remains a real differentiator versus closed competitors.

Thinking Mode: Plan Before You Generate

Thinking Mode is the most important reason this page exists separately from every other AI video model page on DoMax. Instead of treating the prompt as a direct instruction to render immediately, this release inserts a planning step. In practical terms, the model acts less like a generator reacting instantly and more like a senior creative system deciding how the scene should be structured before it starts drawing frames.
That difference matters because many AI video artifacts come from under-planned scenes: characters appear in awkward places, limbs drift, shadows contradict the light source, or the camera moves in ways that don't feel logically motivated. Thinking Mode does not eliminate every mistake, and this page should not overclaim that it does. But for complex scenes, layered blocking, and client-grade compositions, it changes the success rate enough to matter.
The practical rule is simple. Use this mode for final deliverables, complex multi-element scenes, branded client work, multi-character setups, and any scene where coherence matters more than speed. Skip it when you are still exploring prompts, running budget-sensitive batches, or testing rough ideas where a faster standard pass is enough. Honest tradeoff: planning-first generation is slower and should be treated as a premium option, not the default for every job.
The A/B comparisons in the sample gallery are the proof section for this claim. They show the same scene idea with planning on and off. That is the right way to explain the feature: not as vague marketing language, but as a visible difference in scene stability, composition logic, and artifact reduction as of June 2026.

Without Planning vs With Planning

Generation approachDirect generationPlanning step first, then generation
LatencyFaster and better for prompt testingSlower and better for final-quality output
Typical use caseIteration and rough conceptingProfessional scenes and deliverables
Artifact resistanceStandard artifact riskUsually fewer coherence and anatomy issues
Budget fitLower total cost for batchesHigher cost justified by better structure
This planning feature should be understood as a professional control, not a universal default. It is best when the scene is complicated enough that planning materially improves the result. For quick tests, simple single-subject shots, or budget-limited iteration, the standard mode is still the more sensible workflow.

The Four Video Workflows

Frames

Anchor the first and last visual state, then let the model generate the transition between them. This is useful for scene bridges, day-to-night changes, or controlled visual continuity across clips.

Video Extend

Upload an existing clip and continue it. The value here is continuity: the model keeps the established visual language, pacing, and motion direction instead of forcing a full regeneration from zero.

Video Reference

Transfer motion rhythm, style cues, or compositional structure from a reference clip into a new scene. This is a strong workflow when a creative team already has a target feel they want to echo.

Video Editing

Modify an existing video with instructions instead of rebuilding the scene from scratch. For example, change season, weather, time of day, or palette while preserving the base shot structure.

Professional Realism and Color Precision

Thousand-Face Realism is this release's answer to one of the most recognizable AI-video weaknesses: multiple different characters who somehow still look like cousins of the same synthetic face. For scenes with casts, crowd energy, or different customer personas, this matters because repetition breaks immersion immediately.
Precise color control matters for a completely different reason: brand safety. Many image and video models can get close to the requested palette. This one is positioned for teams that need something stricter. If your brand palette is built around exact greens, warm neutrals, or a signature launch color, prompt-level palette references become part of the creative workflow rather than an afterthought.
Together, these two features push the model away from generic prompt novelty and toward professional use: one is about believable human differentiation, the other about repeatable visual governance. That combination is why it makes sense for campaigns, pre-vis, fashion, retail, and color-critical content rather than just experimental clips.

How to Use Wan 2.7

1

Choose a workflow and write a prompt

Pick from standard generation, Frames, Video Extend, Video Reference, or Video Editing. Use reference images or palette language if the shot depends on continuity or strict art direction.

2

Enable Thinking Mode if quality matters most

Turn it on when the scene is complicated, client-facing, or visually dense. Keep it off when you are still exploring, testing, or optimizing for speed and budget.

3

Generate and download

Run the job, review the result, and download the MP4. If the first pass is close but not final, refine the prompt, palette instructions, or workflow selection before running the next version.

Try it now

Wan 2.7 Prompt Examples

These examples are organized around where the model actually differentiates: planning- optimized prompts, palette-controlled brand work, multi-character realism, and workflow-specific prompting patterns.

Thinking Mode optimized

text-to-videoPlanning on

Luxury lobby with layered depth

A luxury hotel lobby at sunrise with three layers of motion: guests crossing the foreground, staff resetting flowers near the desk, and light shifting across a polished marble floor. Cinematic camera push-in, stable composition, high coherence.

text-to-videoPlanning on

Complex kitchen choreography

A busy open kitchen where two chefs plate dishes while a third passes behind them carrying copper pans. Warm tungsten light, stainless reflections, and controlled camera drift from left to right.

text-to-videoPlanning on

Multi-character commuter scene

Five commuters waiting on a rainy train platform, each with distinct face shapes, coats, and posture. Neon station reflections in puddles, a train entering in the far background, restrained handheld camera motion.

text-to-videoPlanning on

Museum fashion film

A fashion model walks slowly through a modern art museum while sunlight cuts across white walls and a second figure appears briefly in the background. Elegant pacing, clean geometry, and stable shadow direction.

Color-controlled brand work

text-to-videoPlanning on

Coffee campaign palette

A modern coffee shop interior. Brand palette: #5C3A1E, #F5E6D3, #2D8C4E. Soft morning light, handcrafted ceramic cups, shallow depth of field, warm cinematic commercial style.

text-to-videoPlanning on

Sportswear launch film

Athletic apparel launch scene. Strict palette: #111827, #22C55E, #F8FAFC. Urban rooftop at blue hour, reflective glass, premium movement, and precise product color retention.

text-to-videoStandard mode

Cosmetics hero close-up

Luxury skincare bottle on rippling black water. Brand colors: #EADBC8, #C79B6C, #161616. Macro beauty lighting, slow camera push, premium reflections, exact palette control.

text-to-videoPlanning on

Tech keynote stage

Conference keynote stage using only palette #0F172A, #38BDF8, #FFFFFF. Minimalist product screen graphics, crisp spot lighting, premium reveal pacing, and strict brand color discipline.

Multi-character realism

text-to-videoPlanning on

Six faces, one boardroom

Six executives around a glass conference table, each with clearly distinct facial structure, eye shape, age, and styling. Bright city skyline through windows, subtle gestures, natural business energy.

text-to-videoPlanning on

Street market ensemble

A crowded evening street market with multiple vendors and shoppers. Each face should feel unique, not cloned. Steam rising from food stalls, lantern light, layered depth, documentary realism.

Video workflows

video-referencePlanning on

Reference-driven car reveal

Use the reference video motion, but restage the scene as a premium electric car reveal inside a concrete gallery with dramatic white key light and deep graphite flooring.

video-editingStandard mode

Winterize the beach scene

Edit the uploaded beach footage into winter: remove bright summer light, add grey sky, cold wind, and falling snow while keeping camera movement and composition intact.

video-extendStandard mode

Extend the character exit

Continue the uploaded clip with the main character walking past the frame edge, camera following for two more beats, while preserving lighting, wardrobe, and movement rhythm.

video-referencePlanning on

Editorial style transfer

Transfer the pacing and compositional rhythm of the reference footage into a monochrome fashion editorial scene with silver reflections and clean gallery architecture.

Frame anchored

framesPlanning on

Sunrise to sunset bridge

Generate the transition from the provided sunrise frame to the provided sunset frame while preserving the same mountain ridge and moving the light naturally through the day.

framesPlanning on

Store opens for launch

Use the first frame of a closed flagship store and the last frame of the same space full of launch visitors. Generate a polished opening transition with elegant customer flow.

Wan 2.7 vs Wan 2.6 vs Wan 2.5: Choose Your Wan

This release is not meant to make Wan 2.6 or Wan 2.5 irrelevant. The more honest framing is that the family now covers three different jobs: the premium planning-first workflow, the role-playing and multi-shot middle tier, and the affordable audio-visual baseline.
Release positionWan 2.6 and 2.5 are earlier workflow milestonesLatest professional flagship in the family
Thinking ModeNot the core story for older versionsYes — the defining feature
Role-playing and storyboardingWan 2.6 is still the clearest storyboarding releaseStill relevant, but no longer the headline
Audio-visual heritageWan 2.5 was the first audio-visual WanInherited, but not the main reason to choose it
Cost profileLower-cost routes for iteration and volumePremium workflow with higher compute expectations
Use Wan 2.7 for professional work that needs planning-led coherence, precise palette control, workflow breadth, or stronger realism. Use Wan 2.6 when role-playing and multi-shot storyboarding are the main job. Use Wan 2.5 when you want the most affordable audio-visual Wan for higher-volume production.

Wan 2.7 vs Sora 2, Veo 3.1, Kling 3.0, and Seedance 2.0

This section stays intentionally conservative. Competitive AI video specs change quickly, and this page should not fabricate numbers it cannot verify. What can be said safely is that this model's differentiator is architectural and workflow-based, not just a generic claim of being "better."
ArchitectureMost major competitors on DoMax are closed-source familiesPart of an open-source family with hosted inference on DoMax
Planning-first generationPlanning-style generation is not the main public positioning of other production modelsThinking Mode is the core story here as of June 2026
Workflow breadthSome competitors emphasize generation quality, iteration, or lip-syncThis release explicitly centers Frames, Extend, Reference, and Editing workflows
Best fitKling 3.0 is strong for 4K and multilingual lip-sync; Seedance 2.0 is strong for motion and prompt adherenceStrongest when controllability, palette precision, and planning-first coherence matter most
The honest recommendation is to compare the same prompt across families. DoMax supports Wan AI, Kling AI, and Seedance AI, which makes side-by-side workflow testing much more useful than relying on stale listicle rankings alone.

What Creators Make with Wan 2.7

Brand-color-accurate ad creatives

Use palette-controlled prompting when the campaign must stay close to approved color systems.

Multi-character films

Thousand-Face Realism is especially relevant when a scene needs distinct-looking people rather than one repeated facial template.

Professional pre-vis

Planning-first generation makes this release a better candidate for complex pre-visualization and concept development where coherence matters.

Content restructuring

Use the Video Editing workflow when you need to revise footage rather than generate an entirely new composition every time.

Longer sequences through extension

Video Extend and Frames are both relevant when a creator wants to turn short clips into a longer continuous sequence.

Reference-led style work

Video Reference is helpful when a team already has a motion or look target and wants the new scene to follow that visual logic.

Frame-anchored transitions

Frames workflow is suited to transitions where the starting image and ending image both matter and the path between them must feel intentional.

Readable in-video text

Long-text rendering matters for retail scenes, interface-style shots, signage, packaging, or ads that put readable text on-screen.

Wan 2.7 Pricing

Wan 2.7 is the premium tier of the Wan family on DoMax at about 8 credits per second before any workflow-specific or planning-mode overhead. The model earns that premium by offering planning-first generation, broader workflow depth, and more professional creative controls. A common cost-saving pattern is to explore prompts in Wan 2.5 or Wan 2.6 first, then move the winning concept into Wan 2.7 for the final pass.
Buy Credits
PackPriceCredits~Seconds
Starter$15300~38 seconds
Plus$491,500~188 seconds
Ultra$993,500~438 seconds
Business$29912,000~1,500 seconds

Frequently Asked Questions about Wan 2.7

Find quick answers about availability, capabilities, supported inputs, generation settings, credits, downloads, and commercial use.

What is Wan 2.7?
Wan 2.7 is Alibaba Tongyi Lab's most advanced Wan release, launched in April 2026. It introduced Thinking Mode prompt planning, Thousand-Face Realism, precise color control, long-text rendering, and four formal video workflows inside one model family.
What is Thinking Mode in Wan 2.7?
Thinking Mode is Wan 2.7's signature architectural feature. Instead of generating immediately, the model first interprets the prompt and plans composition, camera logic, lighting, and motion structure. That extra planning step usually improves coherence and reduces common artifact issues.
When should I use Thinking Mode?
Use Thinking Mode for final deliverables, complex multi-element scenes, client work, brand-critical content, and multi-character setups. Skip it for quick prompt testing, cheaper batch iteration, or simple scenes where speed matters more than polished coherence.
What are the four Wan 2.7 video workflows?
Wan 2.7 is positioned around four workflow types: Frames, Video Extend, Video Reference, and Video Editing. Together they make the model useful not only for new generation, but also for continuation, reference-driven creation, and instruction-based modification.
What is Thousand-Face Realism?
Thousand-Face Realism is Wan 2.7's answer to the usual 'AI same-face' problem. It aims to produce more distinct facial structures, expressions, and identity separation across different characters so scenes do not feel like clones of one template face.
How does color control work in Wan 2.7?
Wan 2.7 supports prompt-level color direction with palette language and HEX-style brand references. That matters for ad creative and brand campaigns where 'approximately the right color' is not enough and a campaign palette must stay controlled.
Is Wan 2.7 open source?
Wan belongs to an open-source family whose weights are publicly available through the broader Wan ecosystem. On DoMax, however, Wan 2.7 runs through hosted cloud inference. The open-source value on this page is architectural provenance and community validation, not a downloadable local package from DoMax.
How is Wan 2.7 different from Wan 2.6?
Wan 2.6 emphasized role-playing, multi-shot storyboarding, and style transfer. Wan 2.7 keeps the broader professional positioning but adds Thinking Mode, stronger realism control, precise color handling, and more explicit video editing workflows.
How is Wan 2.7 different from Wan 2.5?
Wan 2.5 was the first audio-visual Wan and remains the more affordable family entry point. Wan 2.7 is the premium professional option built for coherence, workflow flexibility, and higher-control output rather than lowest-cost generation.
Is Wan 2.7 free to try?
DoMax can offer free credits to test Wan 2.7. After that, usage runs on paid credits. Because Wan 2.7 is positioned as the most advanced Wan workflow, it should be treated as a premium model rather than a free-unlimited tool.

Try Other AI Video Models on DoMax

Wan 2.6

Use the workflow-rich sibling when role-playing and multi-shot storyboarding matter more than planning-first generation.

Wan 2.5

Use the most affordable Wan when you want high-volume audio-visual generation with lower operating cost.

Kling 3.0

Compare Wan 2.7 with a premium closed-source alternative focused on 4K, AI Director, and multilingual lip-sync.

Wan AI

Return to the Wan family hub to compare all supported Wan versions in one place.

Start Creating with Wan 2.7 Today

Use planning mode when coherence matters, pick the workflow that matches the job, and move from rough concept to professional output inside one hosted Wan 2.7 workflow.

Planning modeFour workflowsOpen-source heritage
Try Wan 2.7 FreeBuy Credits
Hugging Face Wan page·Alibaba Tongyi Lab