Kling Elements: Consistent Characters Across Clips
Kling Elements lets you upload reference images of a subject and reuse that subject across generations so it looks the same each time. Give it a few clean, well-lit images of the same character in the same outfit from different angles, then write prompts that describe the action and scene rather than the character's appearance, and Kling keeps the identity while changing everything else.
Kling Elements lets you upload reference images of a subject (a person, a product, a pet, a vehicle) and reuse that subject across generations so it looks the same each time. You give Kling a few clear pictures, then refer to that subject in your prompt. It is the most reliable way in Kling to build a series of clips around one character without the face, outfit or object drifting between shots, and it is the feature to reach for before you plan a multi-shot story, an ad with a recurring presenter or a product video with several angles.
What Elements does and does not do
In Kling's current interface, Elements is a reference feature. You upload one or more images of a subject, and Kling uses them to preserve that subject's identity when it generates video. You can then write a prompt that places the subject in a new setting, doing a new action, with a different camera. The scene is free; the subject is anchored.
It helps to be clear about the limits:
- It is not a trained model. Elements steers the generation towards your references; it does not memorise every freckle or stitch. Expect strong resemblance rather than a pixel match.
- It preserves the subject, not the scene. Background, lighting and colour grade still come from your prompt, which is what you want for a series.
- Extreme angles and occlusion weaken it. A face seen from behind, or half hidden by a hand, gives the model less to hold on to.
- Clothing is part of the identity. If your references show a red jacket, the model treats the jacket as part of the subject. Changing outfit between clips means new references or very explicit prompting.
Treat it as a strong steer that you back up with good prompting, not a lock you can rely on blindly.
Preparing reference images
The quality of your references decides the quality of the consistency. Spend time here and you save credits later.
- Use the same subject in the same outfit. Every image should show the identical character, hairstyle, clothing and accessories. Mixed references produce a blend.
- Vary the angle, not the look. A front view, a three-quarter view and a profile give the model a sense of the shape. Keep expression neutral or mild.
- Keep backgrounds plain. A simple, uncluttered background stops the model confusing scenery with subject.
- Match the lighting. Soft, even light in all references. Dramatic shadows in one image and flat light in another weaken the identity.
- Fill the frame. The subject should be large and sharp. A tiny figure in a wide shot gives the model almost nothing.
- One subject per image. No other people, pets or hero objects in the references.
Where do the references come from? Most people generate them in an image model with a locked style and a consistent description, then pick the three or four that match best. If you work in Midjourney, the Midjourney consistent characters guide shows how to produce a matching set. The same principles apply in any image generator: identical description, identical style settings, change only the angle.
Writing prompts with an Element
Once the subject is set up, your prompt changes shape. Describe the action, the scene and the camera. Do not re-describe the character in detail; the references already do that, and a long appearance description can pull the result away from them.
- Refer to the subject the way the interface expects. The exact method of naming or inserting an element in the prompt may differ between versions, so follow the on-screen hint in the current interface.
- Keep outfit words consistent. If the references show a yellow raincoat, say "in her yellow raincoat" rather than "in a coat". Consistency in words supports consistency in pixels.
- One subject per clip to begin with. Two elements in one shot is possible but harder; master single-subject clips first.
- Favour moderate motion. Slow turns, walking, gestures and expressions preserve identity better than running, spinning or fast head movements.
- Name the camera. "Medium shot, static camera" keeps the face at a readable size. Very wide shots shrink the subject and reduce resemblance.
For the general prompt structure, see the Kling prompting guide.
Settings at a glance
| Setting | What it does | Start with |
|---|---|---|
| Element references | Images that define the subject's identity | Three or four clean images, same outfit, different angles |
| Model version | Which Kling model generates the clip; Elements support depends on version | The latest version that offers Elements in the current interface |
| Mode | Standard is faster; Professional holds faces and detail more steadily | Standard to test the subject, Professional for keeper clips |
| Duration | 5 or 10 seconds | 5 s; identity drifts more over longer clips |
| Negative prompt | Faults to avoid | blur, distortion, deformed face, extra limbs, different person, text |
| Camera movement | Preset moves on supported models | Static or a slow push-in |
Elements may not be available on every model version or plan. Check the current interface and pricing page.
Fixing drift
- The face changes mid-clip. Shorten to 5 seconds, slow the motion, keep the face larger in frame and switch to Professional mode.
- The outfit colour shifts. Name the colour in the prompt and add the wrong colour to the negative prompt.
- Proportions look off. Your references probably mix lens types or angles too aggressively. Replace the outlier image.
- Two subjects swap features. Generate each subject in its own clip, or give each a distinct silhouette and colour so the model can tell them apart.
- Elements is not enough for a hero shot. Combine approaches: generate a still of the character in the exact pose you need, then use it as the start frame in image-to-video. The image-to-video guide covers that route.
Try it
Element: [your character, three or four reference images, same outfit]
Prompt: She walks slowly through a sunlit market, glancing at a stall of oranges and smiling. Medium shot, slow tracking shot from the side, warm morning light, shallow depth of field, natural colours.
Negative prompt: blur, distortion, deformed face, extra limbs, different person, text.
Mode: Standard. Duration: 5 s.
Generate the same prompt three times and compare the faces. If they match, change the scene to an evening street and run again; if the identity holds across both, your references are good enough to build a series on.