KlingImage-to-videoParameters & settings
Kling Image to Video Guide: Start Frames, Settings, Tips
To make a video from an image in Kling, switch to image-to-video, upload your still as the start frame, write a short prompt that describes motion rather than appearance, choose Standard or Professional mode and a 5 or 10 second duration, then generate. The image fixes the composition and look; the prompt only has to say what moves.
Image-to-video in Kling starts from a picture you already have and animates it for 5 or 10 seconds. You upload the still as the start frame, describe the motion you want in the prompt, pick a mode and duration, and the model works out how the scene should move. Because the first frame is fixed, you keep control of composition, character design and colour, which is why most working creatives prefer this route over text-to-video for anything that has to match a brand or a storyboard.
How Kling image-to-video works
Kling (made by Kuaishou, at klingai.com) offers two ways to generate video: text-to-video and image-to-video. In image-to-video the uploaded image becomes the first frame of the clip, so the output inherits its framing, palette and subject. The prompt no longer needs to describe what the scene looks like. Its job is to describe what happens: who moves, how the camera moves and what changes over the clip.
You can also add an end frame. With both a start and an end image, Kling generates the transition between them, which is useful for reveals, product turns and simple morphs. That works well when the two frames share a camera position and subject, for example a closed box and the same box open. It works poorly when the frames differ in composition, lighting or lens, because the model has to invent a jump that ends up looking like a cut or a smear. Generate the end frame in the same image model and session as the start frame and change only the thing that should move.
Several model versions are available in the current interface (1.0, 1.5, 1.6, 2.0, 2.1 and 2.5 Turbo at the time of writing), and newer versions generally follow motion prompts more faithfully. The model list changes often, so check which version is selected before you generate.
Step by step: your first clip
- Prepare the still. Use a clean, well-lit image at the aspect ratio you want in the final video. Crop before uploading rather than hoping the model will reframe. Leave a little room around the subject so motion has somewhere to go.
- Upload it as the start frame. Switch to image-to-video and drop the image into the start frame slot.
- Write a motion prompt. One or two sentences about the subject's movement and the camera. For example: "The woman turns her head slowly towards the window and smiles. Slow push-in, shallow depth of field."
- Add a negative prompt. List what you do not want, such as "blur, distortion, extra limbs, text, flicker".
- Choose mode and duration. Standard mode for drafts, Professional for the final take. Start with 5 seconds; move to 10 only when the 5-second version proves the motion works.
- Generate, review, iterate. Watch the clip at full size, then step through the first and last second frame by frame, where artefacts most often appear. Change one thing per retry so you learn what each change does.
Writing the motion prompt
The most common mistake is pasting the original image prompt back in. Kling already sees the picture. Instead:
- Name the subject and one action. "The cyclist pushes off and rides out of frame to the left." One action per clip is far more reliable than three.
- Describe the camera separately. "Static camera", "slow dolly in", "gentle handheld drift". Kling also has camera movement controls in the interface for some models; if you use them, keep the text description consistent with what you selected.
- Set the pace. Words like "slowly", "gently" and "subtle" calm the model down. Without them, motion tends to be too fast for a 5-second clip.
- Say what should stay still. "Background remains static" or "face stays in frame" reduces unwanted drift.
For a deeper look at prompt order and vocabulary, see the Kling prompting guide.
Settings at a glance
| Setting | What it does | Start with |
|---|---|---|
| Model version | Selects which Kling model generates the clip; newer versions follow prompts more closely | The latest standard version for quality, a Turbo version for quick tests |
| Mode | Standard is faster and cheaper; Professional is higher quality, slower and uses more credits | Standard for drafts, Professional for the final take |
| Duration | Length of the clip | 5 s, then 10 s once the motion is right |
| Start frame | The image the clip begins on | A clean still at the final aspect ratio |
| End frame | Optional image the clip ends on | Leave empty unless you need a specific destination |
| Negative prompt | Things to avoid | blur, distortion, extra limbs, text, flicker |
| Creativity / Relevance | On earlier models, a slider between following the prompt literally and giving the model freedom (similar to CFG) | Middle; nudge towards Relevance if the model ignores you |
| Camera movement | Preset camera moves on supported models | None, then one simple move |
Credit costs depend on mode, duration and version. Check the platform's current pricing page rather than relying on figures from older guides.
Common problems and fixes
- The subject warps or melts. Too much motion was requested. Simplify to one action, add "subtle" and try Professional mode.
- Nothing moves. The prompt was descriptive but not active. Use a verb and name the camera move.
- The face changes. Keep the face large in frame, ask for "face stays consistent", and avoid fast head turns. For a recurring person, see Kling Elements for consistent characters.
- The end frame arrives early and the clip freezes. Use a shorter duration or make the two frames more different so there is enough change to fill the time.
- Text or logos in the still smear. Video models struggle with lettering. Remove it from the still and add it back in your editor.
Try it
Start frame: a ceramic coffee cup on a wooden table by a window, morning light.
Prompt: Steam rises gently from the cup. Soft sunlight shifts across the table as a cloud passes. Slow dolly in towards the cup, shallow depth of field, static background.
Negative prompt: blur, distortion, flicker, text, extra objects.
Mode: Standard. Duration: 5 s.
Change one variable at a time: swap "slow dolly in" for "static camera" to see how much of the motion is camera versus scene, then move to Professional mode for the version you keep.