KlingConsistent characters

Kling Elements: Consistent Characters Across Clips

By Marco Cavazzana · Updated · 5 min read
ShareEmail
Short answer

Kling Elements lets you upload reference images of a subject and reuse that subject across generations so it looks the same each time. Give it a few clean, well-lit images of the same character in the same outfit from different angles, then write prompts that describe the action and scene rather than the character's appearance, and Kling keeps the identity while changing everything else.

Kling Elements lets you upload reference images of a subject (a person, a product, a pet, a vehicle) and reuse that subject across generations so it looks the same each time. You give Kling a few clear pictures, then refer to that subject in your prompt. It is the most reliable way in Kling to build a series of clips around one character without the face, outfit or object drifting between shots, and it is the feature to reach for before you plan a multi-shot story, an ad with a recurring presenter or a product video with several angles.

What Elements does and does not do

In Kling's current interface, Elements is a reference feature. You upload one or more images of a subject, and Kling uses them to preserve that subject's identity when it generates video. You can then write a prompt that places the subject in a new setting, doing a new action, with a different camera. The scene is free; the subject is anchored.

It helps to be clear about the limits:

  • It is not a trained model. Elements steers the generation towards your references; it does not memorise every freckle or stitch. Expect strong resemblance rather than a pixel match.
  • It preserves the subject, not the scene. Background, lighting and colour grade still come from your prompt, which is what you want for a series.
  • Extreme angles and occlusion weaken it. A face seen from behind, or half hidden by a hand, gives the model less to hold on to.
  • Clothing is part of the identity. If your references show a red jacket, the model treats the jacket as part of the subject. Changing outfit between clips means new references or very explicit prompting.

Treat it as a strong steer that you back up with good prompting, not a lock you can rely on blindly.

Preparing reference images

The quality of your references decides the quality of the consistency. Spend time here and you save credits later.

  1. Use the same subject in the same outfit. Every image should show the identical character, hairstyle, clothing and accessories. Mixed references produce a blend.
  2. Vary the angle, not the look. A front view, a three-quarter view and a profile give the model a sense of the shape. Keep expression neutral or mild.
  3. Keep backgrounds plain. A simple, uncluttered background stops the model confusing scenery with subject.
  4. Match the lighting. Soft, even light in all references. Dramatic shadows in one image and flat light in another weaken the identity.
  5. Fill the frame. The subject should be large and sharp. A tiny figure in a wide shot gives the model almost nothing.
  6. One subject per image. No other people, pets or hero objects in the references.

Where do the references come from? Most people generate them in an image model with a locked style and a consistent description, then pick the three or four that match best. If you work in Midjourney, the Midjourney consistent characters guide shows how to produce a matching set. The same principles apply in any image generator: identical description, identical style settings, change only the angle.

Writing prompts with an Element

Once the subject is set up, your prompt changes shape. Describe the action, the scene and the camera. Do not re-describe the character in detail; the references already do that, and a long appearance description can pull the result away from them.

  • Refer to the subject the way the interface expects. The exact method of naming or inserting an element in the prompt may differ between versions, so follow the on-screen hint in the current interface.
  • Keep outfit words consistent. If the references show a yellow raincoat, say "in her yellow raincoat" rather than "in a coat". Consistency in words supports consistency in pixels.
  • One subject per clip to begin with. Two elements in one shot is possible but harder; master single-subject clips first.
  • Favour moderate motion. Slow turns, walking, gestures and expressions preserve identity better than running, spinning or fast head movements.
  • Name the camera. "Medium shot, static camera" keeps the face at a readable size. Very wide shots shrink the subject and reduce resemblance.

For the general prompt structure, see the Kling prompting guide.

Settings at a glance

SettingWhat it doesStart with
Element referencesImages that define the subject's identityThree or four clean images, same outfit, different angles
Model versionWhich Kling model generates the clip; Elements support depends on versionThe latest version that offers Elements in the current interface
ModeStandard is faster; Professional holds faces and detail more steadilyStandard to test the subject, Professional for keeper clips
Duration5 or 10 seconds5 s; identity drifts more over longer clips
Negative promptFaults to avoidblur, distortion, deformed face, extra limbs, different person, text
Camera movementPreset moves on supported modelsStatic or a slow push-in

Elements may not be available on every model version or plan. Check the current interface and pricing page.

Fixing drift

  • The face changes mid-clip. Shorten to 5 seconds, slow the motion, keep the face larger in frame and switch to Professional mode.
  • The outfit colour shifts. Name the colour in the prompt and add the wrong colour to the negative prompt.
  • Proportions look off. Your references probably mix lens types or angles too aggressively. Replace the outlier image.
  • Two subjects swap features. Generate each subject in its own clip, or give each a distinct silhouette and colour so the model can tell them apart.
  • Elements is not enough for a hero shot. Combine approaches: generate a still of the character in the exact pose you need, then use it as the start frame in image-to-video. The image-to-video guide covers that route.

Try it

Element: [your character, three or four reference images, same outfit]
Prompt: She walks slowly through a sunlit market, glancing at a stall of oranges and smiling. Medium shot, slow tracking shot from the side, warm morning light, shallow depth of field, natural colours.
Negative prompt: blur, distortion, deformed face, extra limbs, different person, text.
Mode: Standard. Duration: 5 s.

Generate the same prompt three times and compare the faces. If they match, change the scene to an evening street and run again; if the identity holds across both, your references are good enough to build a series on.

Frequently asked questions

How many reference images does Kling Elements need?
A small set of clean images works best: three or four showing the same subject in the same outfit from different angles. More images do not help if they are inconsistent, and a single image gives the model too little information about the subject's shape.
Can Kling Elements keep a product or pet consistent, not just a person?
Yes. Elements works for any subject you can photograph or generate clearly, including products, pets and vehicles. Use the same preparation rules: plain backgrounds, even lighting, the subject large in frame and no other objects in the references.
Why does my character still look different between Kling clips?
Usually the references are inconsistent, the motion is too fast, or the face is too small in frame. Tighten the references, slow the action, use a medium or close shot and switch to Professional mode. For a critical shot, use a matching still as the start frame in image-to-video instead.
DreamdriveComing to Chrome
Join the waiting list