Reference performer turned into a keynote speaker on a dark stage with clear continuity of face, voice, and delivery style.
Reference: Keynote talent reference
Alibaba Tongyi Lab's narrative AI video model. Wan 2.6 was the first domestic model with role-playing: upload a reference video, extract a character's appearance and voice cues, then generate consistent multi-shot narrative scenes. Built for short films, drama, and ad production.
Model retired
This model has been retired. Please use our latest supported model from the same brand for new generations.
Use Wan 3.0Model angle
The role-playing and storyboard Wan for narrative production.
Honest fit
Best when character continuity, multi-shot scenes, and story structure matter more than the latest flagship controls.
These shared demo clips illustrate workflow ideas, not verified Wan 2.6 outputs. The prompts and labels describe concepts rather than the properties of the clips.
Reference performer turned into a keynote speaker on a dark stage with clear continuity of face, voice, and delivery style.
Reference: Keynote talent reference
Reference character re-staged into a neon city walk at night while preserving identity and natural gait.
Reference: Street character reference
Talent reference used for a premium skincare spokesperson scene with direct-to-camera delivery.
Reference: Spokesperson reference
Sunny kitchen narrative auto-expanded into a breakfast mini film with a stable lead character.
Fishing village opening sequence generated from one narrative prompt.
Coffee brand story rendered as a four-shot commercial sequence from one prompt.
Anime reference transferred onto a downtown crossing sequence while preserving the new scene content.
Includes style reference workflow
Vintage monochrome style reference applied to a new fashion editorial scene.
Includes style reference workflow
Rain-soaked commuter platform with restrained handheld realism and synchronized ambient audio.
Luxury watch macro reveal with premium reflections and ad-style pacing.
Street food stall still image animated into a warm night-market reel.
Portrait still image expanded into a fashion corridor walk with subtle motion.
Reference character used in a candlelit palace monologue scene for a dramatic period-style performance.
Reference: Drama reference
Detective office sequence rendered from one prompt with stable story logic across shots.
Wan 2.6 was the first domestic Chinese-developed AI video model to support role-playing, which is the key heritage claim that still distinguishes this release inside the family.
This release can expand one narrative prompt into a sequence of connected shots, making it much more useful for short films, ad storytelling, and mini-scene construction.
Give the model a style reference video and it can generate new content that follows that visual language, pacing, or treatment rather than starting from a blank aesthetic.
Wan 2.6 extended the family beyond Wan 2.5’s 10-second limit, which gave it enough room for a more meaningful narrative beat rather than just a quick visual loop.
It keeps synchronized audio in the workflow, so voice, ambient sound, and scene timing remain part of one generation process rather than a separate dubbing step.
DoMax hosts the inference, but Wan 2.6 still inherits the broader Wan family’s open-source credibility and research visibility, which closed model families cannot claim in the same way.
Pick text-to-video, image-to-video, role-playing, multi-shot storyboard, or style transfer depending on whether your priority is identity consistency, narrative structure, or aesthetic reference.
Set aspect ratio, duration, and audio language. For role-playing, use a clear 3–30 second reference video that shows the person’s face, voice, and natural motion as clearly as possible.
Run the generation, review the clip, and download the MP4. Multi-shot, role-playing, and style-transfer jobs may need one or two prompt refinements before they lock into the exact narrative beat you want.
These prompts are grouped around the three reasons to choose Wan 2.6: role-playing, multi-shot narrative structure, and style transfer, with a few supporting cinematic and social examples for broader usage.
[character] gives a high-energy keynote on creativity in a dark theater, spotlight on stage, audience visible in soft blur, confident hand gestures and clear delivery.
[character] walks through a rainy Tokyo street at night, neon signs reflecting on wet pavement, cinematic over-the-shoulder coverage and natural stride.
[character] introduces a luxury skincare product in a clean studio, direct-to-camera delivery, elegant gestures, premium soft-box lighting.
[character] stands in a candlelit palace hallway delivering a short dramatic monologue, rich costume texture, slow push-in camera move, emotional intensity.
[character] sits for a creator interview in a warm studio, speaking naturally toward an off-camera host, subtle hand motion and believable posture.
Wide shot: sunny kitchen at morning. Medium shot: young chef whisking batter. Close-up: batter pouring into a pan. Medium shot: chef plates pancakes and smiles at camera.
Wide shot: sunrise over a small fishing harbor. Medium shot: fisherman opens a wooden door holding gear. Close-up: hands tying line to a reel. Tracking shot: walking toward the boat through morning mist.
Wide shot: singer alone in a long red corridor. Medium shot: walking toward camera while lip-syncing. Close-up: hand trailing along the wall. Wide shot: turning into a bright performance room.
Wide shot: small cafe opens at dawn. Medium shot: barista grinds beans. Close-up: espresso pouring into a cup. Medium shot: customer takes first sip by the window.
Wide shot: detective office at night. Medium shot: investigator reading a file under desk lamp. Close-up: fingers tracing a suspect photo. Medium shot: phone rings and the detective turns.
Use the uploaded style reference to render a crowded downtown crosswalk in a polished anime style, bold lighting, expressive motion, and clean line work.
Apply the reference video style to a monochrome fashion editorial inside a stone gallery, preserving vintage film grain, flicker, and soft contrast.
A commuter stands alone on a rain-soaked train platform at blue hour, soft reflections in puddles, restrained handheld motion, cinematic realism.
Macro product film of a luxury watch on black stone, controlled highlights, shallow depth of field, elegant camera orbit, premium ad finish.
Animate a still image of a street food stall into a vertical social reel with steam, customer movement, and warm tungsten night lighting.
Vertical fitness reel, early morning rooftop workout, high-energy pacing, confident movement, clean branded athleisure look.
| Release position | Wan 2.5 is the affordable audio-visual baseline | Wan 2.6 is the narrative middle tier; Wan 2.7 is the premium flagship |
|---|---|---|
| Role-playing | Not part of Wan 2.5's story | Wan 2.6 introduced it and Wan 2.7 inherits it |
| Multi-shot storyboarding | Single-shot oriented baseline | Introduced in 2.6 and carried forward by 2.7 |
| Premium controls | Not the focus | Wan 2.7 adds Thinking Mode, stronger realism, and color precision beyond 2.6 |
| Cost fit | Lowest family cost | Mid-tier cost that matches narrative production workflows |
| Role-playing context | Sora 2 arrived first globally | Wan 2.6 was first domestically and made the feature central to its release story |
|---|---|---|
| Multi-shot storytelling | Kling 3.0 has AI Director and its own multi-shot angle | Wan 2.6 centers storyboard-like narrative prompting rather than flagship 4K positioning |
| Style transfer | Not every competitor markets this workflow equally | Wan 2.6 makes style transfer part of its feature identity |
| Best fit | Some alternatives lean into 4K, lip-sync, or prompt adherence | Wan 2.6 is strongest when character continuity and short-form narrative structure matter most |
Use role-playing to keep the same person credible across multiple scenes and prompts.
Multi-shot storyboard prompting makes it easier to generate a beat with progression instead of one isolated frame idea.
Reference yourself once, then produce different scenarios without filming every location or setup.
Capture a consenting spokesperson once and explore multiple campaign scenarios around the same identity.
A storyboard-oriented workflow is more suitable for multi-beat editorial sequences than a pure single-shot model.
Style transfer can keep multiple episodes or campaign cuts inside the same aesthetic language.
Use role-playing plus multi-shot coverage to rough out scenes before production locks casting or locations.
Keep the same character identity while changing setting, framing, or narrative context for different markets.
| Pack | Price | Credits | ~Seconds |
|---|---|---|---|
| Starter | $15 | 300 | ~50 seconds |
| Plus | $49 | 1,500 | ~250 seconds |
| Ultra | $99 | 3,500 | ~583 seconds |
| Business | $299 | 12,000 | ~2,000 seconds |
Find quick answers about availability, capabilities, supported inputs, generation settings, credits, downloads, and commercial use.
Move up to the flagship when you want Thinking Mode, stronger realism controls, and broader premium workflows.
Drop to the budget Wan when you need affordable audio-visual volume more than narrative feature depth.
Compare Wan 2.6 against a premium closed-source alternative centered on 4K, AI Director, and multilingual lip-sync.
Compare Wan 2.6 against a motion-focused alternative when prompt adherence and cinematic polish matter more than role-playing.
Use role-playing for character continuity, multi-shot prompting for story beats, and style transfer for aesthetic consistency inside one hosted Wan 2.6 workflow.