A founder presents a premium coffee machine with natural English narration and realistic steam ambience.
The first Kling model with native audio-visual generation. Generate cinematic 1080p video with synchronized audio, Motion Control reference workflows, and Elements multi-image consistency on affordable DoMax credits.
A founder presents a premium coffee machine with natural English narration and realistic steam ambience.
These shared demo clips illustrate workflow ideas. They are not verified Kling 2.6 outputs, and their prompt labels do not describe the actual clip properties.
A founder presents a premium coffee machine with natural English narration and realistic steam ambience.
A chef introduces a noodle recipe in Mandarin while kitchen sounds and warm ambient music play naturally underneath.
Animate a fantasy forest still with drifting fog, subtle birdsong, and a gentle cinematic push-in.
Reference dancer motion transferred into a samurai training sequence in a misty bamboo courtyard.
Reference camera drift restaged as an electric coupe sliding through a rain-lit tunnel.
Hand motion from a plating reference transferred into a Parisian pastry glazing scene.




Four reference images preserve the same red-haired explorer, coat texture, brass compass, and dusk palette across the generated observatory scene.




Character, outfit, environment, and style references are combined into one coherent luxury fashion corridor scene.


A lighthouse keeper lights a lantern while the start frame anchors the empty stairwell and the last frame anchors the glowing lamp against the storm outside.


A runner exits a subway tunnel into sunrise, using explicit start and end frame anchors to bridge into the next narrative clip.
Vertical product reel with English dialogue, close-up unboxing, and clean rhythmic edit energy.
Vertical food reel with Mandarin narration, sizzling wok audio, and quick punch-in edits for social distribution.
The prompt library here is organized around the workflows that make Kling 2.6 worth choosing: native audio dialogue, Motion Control, Elements reference packs, cinematic singles, and polished vertical social output.
A barista in a sunlit cafe says in natural English, 'Your cappuccino is ready,' while milk foam swirls into a heart pattern and soft room ambience fills the background.
In a warm kitchen, a chef looks at camera and says in Mandarin, '今天我们做一道很简单的早餐,' while chopping herbs and hearing light utensil sounds in the background.
A founder presents a matte-black smart bottle on a clean studio table and says in English, 'Hydration reminders that actually feel premium,' with subtle synth branding audio underneath.
A fashion creator on a neon-lit rooftop speaks in Mandarin about the look of the day while wind ambience, city noise, and soft electronic music sit under the dialogue.
Kling 2.6 was the first Kling model to generate synchronized voice, sound effects, ambience, and video together in one pass, which remains its clearest heritage advantage.
Upload a 3–30 second reference video and transfer its motion structure into a new scene, making 2.6 unusually useful for motion-matched recreations and camera transfer workflows.
Use up to 4 reference images to preserve characters, outfits, environments, or style language across multiple generations without re-explaining every detail in the prompt.
Anchor the start and end frames of a generation so multiple 10-second clips can connect into a longer sequence with better continuity than independent generations.
Kling 2.6 runs at a lower credit-per-second rate than Kling 3.0, which makes it a better fit for 1080p workflows and reference-heavy experimentation.
Months of community usage have made Kling 2.6 one of the better-documented audio-visual video models, especially for reference-driven workflows that creators want to repeat reliably.
Reference dancer motion transferred into a samurai training sequence in a misty bamboo courtyard.
Reference camera drift restaged as an electric coupe sliding through a rain-lit tunnel.
Hand motion from a plating reference transferred into a Parisian pastry glazing scene.
Four reference images preserve the same red-haired explorer, coat texture, brass compass, and dusk palette across the generated observatory scene.
Character, outfit, environment, and style references are combined into one coherent luxury fashion corridor scene.
A lighthouse keeper lights a lantern while the start frame anchors the empty stairwell and the last frame anchors the glowing lamp against the storm outside.
A runner exits a subway tunnel into sunrise, using explicit start and end frame anchors to bridge into the next narrative clip.
Use text-to-video for direct prompts, image-to-video for still-image animation, Motion Control for reference-driven motion transfer, or Elements when you need up to four image references for consistency.
Pick aspect ratio, duration from 3 to 10 seconds, and audio language if you want spoken output. Kling 2.6 runs at 1080p and up to 48 FPS.
Click Generate. Kling 2.6 usually renders a synchronized audio-visual clip in roughly 60–120 seconds. Download the MP4 or keep iterating with references.
| Released | Dec 3, 2025 | Feb 5, 2026 |
|---|---|---|
| Max resolution | 1080p | 1080p |
| Max duration | 10s | 15s |
| Native audio | Yes | Yes |
| Audio languages | English, Chinese | English, Chinese, Japanese, Korean, Spanish |
| Multi-shot AI Director | No | Yes |
| Motion Control reference video | Yes | Not a primary workflow |
| Elements references | Up to 4 images | Different memory-style workflow |
| First/Last Frame | Yes | Not the page's lead workflow |
| Credit cost | Lower | Higher |
| Max resolution | 1080p | Check latest source |
|---|---|---|
| Native audio | Yes | Varies by model |
| Audio languages | English, Chinese | Varies by model |
| Motion Control reference video | Yes | Not typically a lead feature |
| Elements multi-image | Yes, up to 4 | Varies by model |
| First/Last Frame anchors | Yes | Varies by model |
| Max duration | 10s | Varies by model |
Produce affordable 1080p launch clips, explainers, and short ads with spoken audio baked into the generation instead of patched in afterward.
Create Chinese and English social clips with native voice generation for regional campaigns, creator updates, and product explainers.
Elements references make it practical to keep the same lead character, outfit, and visual palette across multiple clips in a campaign or narrative sequence.
Motion Control is ideal when the motion itself matters most and you want to restage it with a different character, brand, environment, or look.
First/Last Frame control helps turn multiple 10-second generations into one longer scene with stronger continuity from clip to clip.
Use native dialogue and ambient sound for vertical short-form content that would otherwise require a second audio-editing pass after video generation.
Build music-led clips with synchronized sound design, stylized character action, and motion references drawn from existing choreography or performance video.
Directors and agencies can block motion, mood, and continuity quickly before a production shoot locks creative decisions and budget.
| Pack | Price | Credits | ~Seconds |
|---|---|---|---|
| Starter | $15 | 300 | ~60 seconds |
| Plus | $49 | 1,500 | ~300 seconds |
| Ultra | $99 | 3,500 | ~700 seconds |
| Business | $299 | 12,000 | ~2,400 seconds |
Find quick answers about availability, capabilities, supported inputs, generation settings, credits, downloads, and commercial use.
Move up to longer clips, multi-shot storyboarding, and broader language support when your project needs the flagship workflow.
Compare a top ByteDance cinematic model when prompt adherence and a different motion aesthetic matter for your test set.
Use a proven production-stable Seedance option when you want a different cost-to-quality balance for audio-visual output.
See the full Kling model family on DoMax and compare 2.6 with 3.0 in one place.
The first audio-visual Kling is still one of the most workflow-friendly options on DoMax for reference-driven 1080p video.