Field Notes · July 9, 2026

Motion-Matching: Shoot It Cheap, Generate It Right

The single biggest difference between AI footage that looks like a slot-machine pull and AI footage that looks like your movie is control. Prompt-only generation asks the model to imagine everything — blocking, camera, timing. Sometimes it nails it. Often it doesn’t, and you burn takes finding out.

Here’s the professional move: shoot the shot first. On your phone, in a backyard, with whoever’s around. You’re not shooting footage — you’re shooting motion. Then extract a pose skeleton and depth map from that clip and hand it to the generator as a video reference. The output follows your blocking, your camera, your timing — with the character, wardrobe, and world swapped to whatever the scene needs.

A free tool we like for this is TheoreticallyPose by Theoretically Media — it runs entirely in your browser (Chrome or Edge), nothing uploads anywhere, and it exports exactly the control video you need. Drag your clip in, hit Track, then Bake, then Export. Feed the result to your generation as a reference alongside a prompt like: “use video 1 as motion and depth” — then describe the world.

Three field lessons: One — fight scenes, falls, and anything with body contact benefit most; that’s exactly the stuff prompt-only generation fails. Two — keep takes short; a 6–10 second motion reference holds better than 15. Three — the model occasionally ignores the reference entirely. Don’t fight it — pull another take. (On Greenlit Dark, failed generations never charge, and multiple takes are one dropdown.)

Your actors’ real faces on those motion-matched bodies? That’s the part that needs signed releases and a whitelist — which is the part we do. Start with your script.

← All Field Notes  ·  Submit a script