Runway Act-Two: Performance Capture for AI Characters
Runway Act-Two takes a driving video of a real performance and transfers the facial expression, speech movement and, unlike its predecessor Act-One, body and hand motion onto a character you supply as an image or video. Film yourself delivering the performance, upload it alongside the character, and the character acts it out.
Act-Two is the Runway tool that lets a character you have designed act out a performance you have filmed. You record yourself, or an actor, delivering a line or a gesture; you upload that clip as the driving video; you supply a character as an image or a video; and Runway maps the performance onto the character. It is the quickest route from a storyboard still to a talking, emoting shot, and this guide covers how to shoot for it, how to prepare the character and how to avoid the mistakes that break the illusion.
Act-One versus Act-Two
Act-One was Runway’s first performance transfer tool. It took the facial performance from a driving video and applied it to a character image: expression, eye direction and mouth movement. Act-Two is its successor and widens the capture to include body and hand motion, so a shrug, a pointed finger or a lean into the camera now carries across as well as the face. Act-Two also accepts a character video as the target, not only a still, which means you can give an already-moving shot a new performance.
In practice, Act-Two is the one to use for anything with gesture or posture, and the one most people will find in the current interface. If you are working on a tight head-and-shoulders talking shot, the two behave similarly; the body capture simply matters less.
Shooting the driving video
The driving video is the performance. Everything Runway can transfer has to be visible in it, so treat the shoot seriously even if it is just you and a phone.
- Frame to match the output. If the character image is a medium shot from the waist up, film yourself from the waist up. Mismatched framing forces the tool to guess what the unseen parts are doing.
- Light the face evenly. A window or a soft lamp in front of you, no hard shadows across the mouth or eyes. The tool reads expression from the face; shadows hide it.
- Keep the camera still. Use a tripod or prop the phone. Camera shake in the driving video becomes confusing motion in the output.
- Perform slightly larger than life. Subtle micro-expressions tend to flatten in transfer. Give the eyebrows and the mouth a little more than feels natural.
- Keep hands in frame and uncrossed. If you want hand motion to transfer, the hands need to be clearly visible against a contrasting background, not overlapping the face or each other.
- Record clean audio. The audio from the driving video is what you will usually keep, so avoid echoey rooms and background noise.
Keep takes short. A performance of a few seconds per shot is easier to review, cheaper to generate and easier to cut. Runway charges credits per second of generated video, so long rambling takes cost more and give you less usable material. Check the current pricing page for Act-Two costs.
Preparing the character
The character can be a still image or a video. For a still, the same rules that apply to any consistent-character work apply here: a clear face, a neutral or mild expression, soft light, and framing that matches the driving video. A character looking straight at the camera is the easiest starting point, because the performance is usually delivered to camera too. Profiles and extreme angles are harder.
If your character comes from a wider project, generate the still with Gen-4 References so that it matches the other shots. For a video target, pick a clip with gentle, slow motion; a character already spinning around leaves little room for the new performance.
Avoid characters with heavy occlusion around the mouth such as thick beards, masks or hands on the chin, and avoid very stylised faces with no visible mouth. The transfer needs features it can move.
Settings at a glance
| Setting | What it does | Start with |
|---|---|---|
| Driving video | Your filmed performance: face, body and hands in Act-Two | A short, well-lit, tripod take framed like the character |
| Character input | The image or video that receives the performance | A front-facing still with a mild expression |
| Aspect ratio | Shape of the output clip | The ratio of your character image |
| Expression or motion controls | Where offered, adjusts how strongly the performance is applied; labels vary | Defaults, then adjust after one test |
| Audio | Whether the driving clip’s sound is kept | Keep it, you will sync to it |
The exact set of sliders and toggles in the Act-Two panel has changed across updates, so treat the table as a map of what to look for rather than the precise labels you will see.
Try it
Driving video: 6 second tripod take, waist-up, soft window light, delivered to camera. Line: "You said the ferry left at nine. It is ten past." Slight head tilt on the second sentence, one open-palm gesture on "ten past".
Character: front-facing medium shot, neutral expression, plain background, same waist-up framing.
Act-Two is driven by the footage rather than a text prompt, so the thing to copy here is the shot plan. Change the line and the gesture, keep the framing match, and if the mouth sync looks soft, re-shoot with your face a little closer to the camera and the light a little flatter.
Common problems and fixes
- Mouth movement does not match the words. The driving video’s face is probably too small in frame or in shadow. Re-shoot closer and brighter.
- Hands do odd things. They were overlapping or leaving frame in the driving clip. Keep them separated, visible and slower.
- The character’s identity slips during big head turns. Reduce the range of head movement in your performance, or use a character video rather than a still.
- The result looks stiff. Your performance was too subtle. Act a little bigger and let your shoulders move.
- Background warps. Prefer character images with plain or softly blurred backgrounds.
Keep the driving clip, the character file and the output together for every shot. When a client asks for the same performance on a different character, you only need to swap one input.