Luxury lobby with layered foreground, midground, and background motion planned through Thinking Mode.
Related prompt idea: wan27-sample-2
Alibaba Tongyi Lab's most controllable AI video model. Its planning-first mode maps the composition before generation, which usually means higher coherence and fewer artifacts. It also brings four video workflows for editing, extension, and reference-driven generation.
Model angle
The Thinking AI video model for higher-control professional workflows.
Why it exists
Planning-first generation, workflow depth, and open-source Wan heritage set it apart.
These shared demo clips illustrate possible workflows. They are not verified Wan 2.7 outputs or actual planning-mode comparisons; prompts are conceptual ideas.
Luxury lobby with layered foreground, midground, and background motion planned through Thinking Mode.
Related prompt idea: wan27-sample-2
Same luxury lobby prompt without Thinking Mode for side-by-side coherence comparison.
Related prompt idea: wan27-sample-1
Crowded train platform with multiple unique faces demonstrating Thousand-Face Realism.
Coffee campaign scene controlled by exact palette values to demonstrate brand-color precision.
Sunrise frame to sunset frame transition over a mountain ridge using the Frames workflow.
Closed flagship store to launch-night crowd transition anchored by first and last frames.
Extend a character exit shot with consistent camera follow and preserved wardrobe detail.
Continue a product reveal clip while holding material reflections and studio light continuity.
Reference a handheld fashion film, then restage the same motion in a monochrome gallery editorial.
Borrow the pacing of a POV city walk and apply it to a premium automotive reveal scene.
Convert a summer beach scene into winter with snow, colder grading, and heavier wind.
Change daytime retail footage into a premium night launch with neon reflections and darker atmosphere.
Long in-video signage and product typography rendered inside a modern retail environment.
Another Thinking Mode A/B scene: crowded kitchen choreography with complex reflections and motion beats.
Related prompt idea: wan27-sample-15
Same kitchen choreography prompt without Thinking Mode for side-by-side comparison.
Related prompt idea: wan27-sample-14
This release is framed around a planning-first workflow. Instead of generating immediately, it interprets prompt intent and decides composition logic before rendering the final clip.
Frames, Video Extend, Video Reference, and Video Editing give this release unusually broad workflow coverage for teams that do more than raw text-to-video generation.
This feature targets one of AI video's recurring weaknesses: cloned-looking faces. It aims for stronger character differentiation across multi-person scenes.
Use HEX-style palette instructions for brand work that requires exact, repeatable color direction rather than approximate prompt-only interpretation.
It is positioned for stronger in-video text generation, which matters for signage, interface shots, retail scenes, and ad creative with visible typography.
It inherits the family’s open-source credibility. On DoMax you use hosted inference, but the architectural lineage remains a real differentiator versus closed competitors.
| Generation approach | Direct generation | Planning step first, then generation |
|---|---|---|
| Latency | Faster and better for prompt testing | Slower and better for final-quality output |
| Typical use case | Iteration and rough concepting | Professional scenes and deliverables |
| Artifact resistance | Standard artifact risk | Usually fewer coherence and anatomy issues |
| Budget fit | Lower total cost for batches | Higher cost justified by better structure |
Anchor the first and last visual state, then let the model generate the transition between them. This is useful for scene bridges, day-to-night changes, or controlled visual continuity across clips.
Upload an existing clip and continue it. The value here is continuity: the model keeps the established visual language, pacing, and motion direction instead of forcing a full regeneration from zero.
Transfer motion rhythm, style cues, or compositional structure from a reference clip into a new scene. This is a strong workflow when a creative team already has a target feel they want to echo.
Modify an existing video with instructions instead of rebuilding the scene from scratch. For example, change season, weather, time of day, or palette while preserving the base shot structure.
Pick from standard generation, Frames, Video Extend, Video Reference, or Video Editing. Use reference images or palette language if the shot depends on continuity or strict art direction.
Turn it on when the scene is complicated, client-facing, or visually dense. Keep it off when you are still exploring, testing, or optimizing for speed and budget.
Run the job, review the result, and download the MP4. If the first pass is close but not final, refine the prompt, palette instructions, or workflow selection before running the next version.
These examples are organized around where the model actually differentiates: planning- optimized prompts, palette-controlled brand work, multi-character realism, and workflow-specific prompting patterns.
A luxury hotel lobby at sunrise with three layers of motion: guests crossing the foreground, staff resetting flowers near the desk, and light shifting across a polished marble floor. Cinematic camera push-in, stable composition, high coherence.
A busy open kitchen where two chefs plate dishes while a third passes behind them carrying copper pans. Warm tungsten light, stainless reflections, and controlled camera drift from left to right.
Five commuters waiting on a rainy train platform, each with distinct face shapes, coats, and posture. Neon station reflections in puddles, a train entering in the far background, restrained handheld camera motion.
A fashion model walks slowly through a modern art museum while sunlight cuts across white walls and a second figure appears briefly in the background. Elegant pacing, clean geometry, and stable shadow direction.
A modern coffee shop interior. Brand palette: #5C3A1E, #F5E6D3, #2D8C4E. Soft morning light, handcrafted ceramic cups, shallow depth of field, warm cinematic commercial style.
Athletic apparel launch scene. Strict palette: #111827, #22C55E, #F8FAFC. Urban rooftop at blue hour, reflective glass, premium movement, and precise product color retention.
Luxury skincare bottle on rippling black water. Brand colors: #EADBC8, #C79B6C, #161616. Macro beauty lighting, slow camera push, premium reflections, exact palette control.
Conference keynote stage using only palette #0F172A, #38BDF8, #FFFFFF. Minimalist product screen graphics, crisp spot lighting, premium reveal pacing, and strict brand color discipline.
Six executives around a glass conference table, each with clearly distinct facial structure, eye shape, age, and styling. Bright city skyline through windows, subtle gestures, natural business energy.
A crowded evening street market with multiple vendors and shoppers. Each face should feel unique, not cloned. Steam rising from food stalls, lantern light, layered depth, documentary realism.
Use the reference video motion, but restage the scene as a premium electric car reveal inside a concrete gallery with dramatic white key light and deep graphite flooring.
Edit the uploaded beach footage into winter: remove bright summer light, add grey sky, cold wind, and falling snow while keeping camera movement and composition intact.
Continue the uploaded clip with the main character walking past the frame edge, camera following for two more beats, while preserving lighting, wardrobe, and movement rhythm.
Transfer the pacing and compositional rhythm of the reference footage into a monochrome fashion editorial scene with silver reflections and clean gallery architecture.
Generate the transition from the provided sunrise frame to the provided sunset frame while preserving the same mountain ridge and moving the light naturally through the day.
Use the first frame of a closed flagship store and the last frame of the same space full of launch visitors. Generate a polished opening transition with elegant customer flow.
| Release position | Wan 2.6 and 2.5 are earlier workflow milestones | Latest professional flagship in the family |
|---|---|---|
| Thinking Mode | Not the core story for older versions | Yes — the defining feature |
| Role-playing and storyboarding | Wan 2.6 is still the clearest storyboarding release | Still relevant, but no longer the headline |
| Audio-visual heritage | Wan 2.5 was the first audio-visual Wan | Inherited, but not the main reason to choose it |
| Cost profile | Lower-cost routes for iteration and volume | Premium workflow with higher compute expectations |
| Architecture | Most major competitors on DoMax are closed-source families | Part of an open-source family with hosted inference on DoMax |
|---|---|---|
| Planning-first generation | Planning-style generation is not the main public positioning of other production models | Thinking Mode is the core story here as of June 2026 |
| Workflow breadth | Some competitors emphasize generation quality, iteration, or lip-sync | This release explicitly centers Frames, Extend, Reference, and Editing workflows |
| Best fit | Kling 3.0 is strong for 4K and multilingual lip-sync; Seedance 2.0 is strong for motion and prompt adherence | Strongest when controllability, palette precision, and planning-first coherence matter most |
Use palette-controlled prompting when the campaign must stay close to approved color systems.
Thousand-Face Realism is especially relevant when a scene needs distinct-looking people rather than one repeated facial template.
Planning-first generation makes this release a better candidate for complex pre-visualization and concept development where coherence matters.
Use the Video Editing workflow when you need to revise footage rather than generate an entirely new composition every time.
Video Extend and Frames are both relevant when a creator wants to turn short clips into a longer continuous sequence.
Video Reference is helpful when a team already has a motion or look target and wants the new scene to follow that visual logic.
Frames workflow is suited to transitions where the starting image and ending image both matter and the path between them must feel intentional.
Long-text rendering matters for retail scenes, interface-style shots, signage, packaging, or ads that put readable text on-screen.
| Pack | Price | Credits | ~Seconds |
|---|---|---|---|
| Starter | $15 | 300 | ~38 seconds |
| Plus | $49 | 1,500 | ~188 seconds |
| Ultra | $99 | 3,500 | ~438 seconds |
| Business | $299 | 12,000 | ~1,500 seconds |
Find quick answers about availability, capabilities, supported inputs, generation settings, credits, downloads, and commercial use.
Use the workflow-rich sibling when role-playing and multi-shot storyboarding matter more than planning-first generation.
Use the most affordable Wan when you want high-volume audio-visual generation with lower operating cost.
Compare Wan 2.7 with a premium closed-source alternative focused on 4K, AI Director, and multilingual lip-sync.
Return to the Wan family hub to compare all supported Wan versions in one place.
Use planning mode when coherence matters, pick the workflow that matches the job, and move from rough concept to professional output inside one hosted Wan 2.7 workflow.