Kuaishou's Kling AI team designed its fourth-generation model, Kling 4.0, for people who want to direct a shot rather than gamble on one prompt. You pin keyframes on a timeline of up to 30 seconds, give every reference file a job, and the model fills in the motion, camera work and sound between those anchors.
Coming soon
This workspace will run Kling 4.0 soon
Generation switches on here as soon as Kling 4.0 is connected. You can already shape the same idea with the models below and bring the prompt over later.
Start a project with
Kling 4.0
Keyframe-driven video model by Kling AI
Model profile
The numbers that decide how much a single Kling 4.0 generation can hold, from clip length to the number of files it can learn from.
| Item | Kling 4.0 |
|---|---|
| Made by | Kling AI, the video platform of Kuaishou |
| Inputs | Text, images, video clips, subjects and voice references |
| Clip length | 3 to 30 seconds per generation |
| Timeline control | Up to 10 keyframes |
| References | Up to 15 in total: 10 images, 5 videos, 7 subjects |
| Image quality | Up to 4K; 10-bit HDR at 1080p and 4K |
| Aspect ratios | 16:9, 9:16, 1:1 and 21:9 ultrawide |
| Sound | Two-channel stereo with tighter lip sync |
| Text and speech | On-screen text and dialogue in nine languages |
| Prompt length | Up to 8,000 tokens |
| Editing | Up to 5 existing clips, 30 seconds in total |
| Extension | A sequence can be extended to about two minutes |
Kling 3.0 already produced cinematic shots with native sound. Kling 4.0 keeps that base and mainly adds room on the timeline and finer control over it.
Length
Kling 3.0:3–15 secondsKling 4.0:3–30 seconds
A full scene with a setup, an action and a payoff fits into one take instead of two clips stitched together.
Frame control
Kling 3.0:First and last frameKling 4.0:Up to 10 keyframes
The middle of the shot can be pinned as well, such as a turn of the head or the moment a product is revealed.
Image quality
Kling 3.0:Up to 4KKling 4.0:Up to 4K, 10-bit HDR at 1080p and 4K
Skies and skin tones grade without banding, and highlights keep their detail.
Aspect ratios
Kling 3.0:16:9, 9:16, 1:1Kling 4.0:Adds 21:9 ultrawide
Cinema-style widescreen compositions come straight out of the model, with no cropping afterwards.
Sound
Kling 3.0:Native audioKling 4.0:Stereo audio, more precise lip sync
Lines land on the mouth movement, and ambience has left and right space.
A keyframe is a still image pinned to a moment of the clip. Kling 4.0 animates the path from one keyframe to the next, so the frames you choose become the skeleton of the shot. A typical four-point plan looks like this:
Opening frame
Set the place, the light and the main subject. This is the first image a viewer sees.
First change
Pin the moment the action starts: a character stands up, a lid lifts, a car pulls away.
Turning point
Fix the frame where the shot changes direction, such as a reveal or a new camera angle.
Final frame
Lock the ending, so the clip lands on a composition you can cut from or extend.
Two keyframes are enough for a simple move. Add more only where a moment must look exactly as planned: every extra keyframe narrows what the model can invent between them.
Omni Reference lets one task carry several kinds of source material at once. Each kind has its own ceiling inside the shared total of 15.
| Input | Up to | Best used for |
|---|---|---|
| Images | 10 | Character turnarounds, product angles, locations, storyboards and style frames |
| Videos | 5, 30 s combined | Motion, choreography, camera paths and pacing to follow |
| Subjects | 7 | Recurring people, mascots or products that must look the same throughout |
| Voice | Supported | The tone and timbre of a speaking character |
State the role of each file in the prompt, for example “image 2 is the jacket, video 1 is the camera move”. References without a stated role compete with each other.
Flash is the lighter edition of the same model generation. It gives up length, resolution and keyframes in exchange for quicker results.
For finished pieces you plan to publish
For drafts, tests and high volume
A practical routine: try several ideas quickly in Flash, then rebuild the one you keep in the full model and add keyframes where it matters.
Each can return a full 30-second take in one go. Where they part ways is the way you tell them what to do.
| Comparison | Kling 4.0 | Seedance 3.0 |
|---|---|---|
| Longest single generation | 30 seconds | 30 seconds |
| Main way to steer | Keyframes pinned on a timeline | A reference stack plus a written prompt |
| Reference capacity | 15 inputs: 10 images, 5 videos, 7 subjects | 50 inputs: 30 images, 10 videos, 10 audio files |
| Audio as input | Voice tone references | Up to 10 audio files |
| Strongest at | Exact timing of key moments, HDR delivery | Large multimodal briefs, scene planning with camera blocking |
Pick Kling 4.0 when specific moments have to appear at specific times. Pick Seedance 3.0 when the brief is a large collection of images, clips and sound that the model should learn from.
Long prompts work when every part has a fixed place. Write them in this order and keep each part to a phrase or two.
1. Subject
A cyclist in a yellow rain jacket
2. Action
brakes, steps off the bike and looks up
3. Setting
on a wet cobblestone street at dusk
4. Camera
a low tracking shot that rises into a crane move
5. Light
neon signs reflected in the puddles
6. Timing
0–10 s riding, 10–20 s stopping, 20–30 s looking up
7. Sound
rain, tyre hiss, a distant tram bell
8. References
subject 1 is the cyclist, image 3 is the street
0–8 s: steam rises from a ceramic cup on a sunlit wooden counter, slow push-in. 8–16 s: a barista's hand slides the cup forward and the focus pulls to the latte art. 16–24 s: a window seat, a woman takes the first sip and smiles. 24–30 s: wide shot of the café at golden hour. Sound: grinder hum, quiet acoustic guitar, a cup set down on its saucer.
A drone glides over terraced rice fields at sunrise, mist lying in the valleys, one farmer walking along a ridge. The camera descends slowly and finishes at eye level beside the farmer. Birdsong and wind only, no music.
A kitchen at night under warm practical light. Subject 1, an older man, asks: "Did you finish it?" Subject 2, a teenage girl, holds up a sketchbook and answers: "Almost." Medium two-shot, slight handheld movement, quiet stereo room tone.
Image 1 is the sneaker, image 2 is the studio backdrop. The sneaker turns once on a glossy turntable while a soft light sweeps across the mesh; the camera ends on a three-quarter hero angle. A soft electronic pulse underneath.
The longer timeline and keyframe control pay off most in work where timing and consistency are part of the brief.
Short answers on length, picture quality, control, sound and how Kling 4.0 sits next to other video models.
It is the fourth major video model released by Kling AI, the generative video team at Kuaishou. Give it text, still images, short clips and subject references, and it returns a 3 to 30 second clip with sound, following up to 10 keyframes you place on the timeline.
Thirty seconds per generation, with three seconds as the minimum. For anything longer, extend the finished clip; extensions can stretch one scene to roughly two minutes.
Up to 4K. Both the 1080p and 4K outputs carry 10-bit HDR, so sunsets and skin tones survive grading without visible steps. The lighter Flash edition stops at 720p in 8-bit SDR.
You place up to 10 still images at chosen moments of the clip. The model treats them as fixed points and generates the motion between them.
Use two keyframes for a simple start-to-end move, and add more only where a moment has to look exactly as planned.
Up to 15 in total:
Voice references can also set the tone of a speaking character.
Flash is tuned for speed rather than finish. It generates up to 20 seconds at up to 720p in 8-bit SDR and does not take first/last frame or multi-keyframe input. Switch to the full model when you need the 30-second length, keyframes, the full reference budget, 4K or HDR.
Yes. It produces two-channel stereo audio together with the picture, including dialogue, ambience and effects, with more accurate lip sync than earlier Kling models.
Characters can speak in nine languages, with room for different accents and dialects. The model can also render on-screen text in nine languages, as well as emoji.
Yes. Its editing mode takes up to five existing clips, 30 seconds in total, and can change a subject, background, expression, camera angle or line of dialogue from a text, image or subject instruction.
Yes. Besides 16:9, 9:16 and 1:1, Kling 4.0 adds a 21:9 ultrawide frame for cinema-style compositions.
Prompts can run to 8,000 tokens, which is enough for a timed script with camera notes, dialogue and sound cues. Order matters more than length: list subject, action, setting, camera, light, timing and sound in the same sequence every time.
Yes, if you need clips longer than 15 seconds, control over moments in the middle of a shot, HDR output or a 21:9 frame. For short, single-action clips Kling 3.0 remains a solid choice.
Pick Kling 4.0 when exact timing matters, because its keyframes fix what happens when. Pick Seedance 3.0 when you want to steer the result with a large set of references; it accepts up to 50 images, videos and audio files in one task.
Register the person as a subject reference, show them in the opening keyframe, and refer to them by the same label throughout the prompt. One person, one reference: two different photos of the same character pull the result in two directions.