Subject + action
Name the actor or object and use a physical verb: the woman turns left, the cup slides right, or the fabric lifts in the wind.
How to write video to prompt · practical tutorial
Break the reference into shots, write only visible and audible facts, give each motion a clear owner and direction, then adapt the frames and text to the AI video model you will use.
Direct answer
Who or what is in the shot? What physical action happens? Where does it happen? How is the shot framed? What does the camera do? How fast and in which direction does each movement happen? What dialogue, music, ambience, or effect must be heard?
Write one continuous shot at a time. When a reference video cuts to another angle or scene, start another prompt. If you use image-to-video, attach the opening frame and remove visual details the image already communicates. If you use a model with first-and-last-frame control, attach both boundary frames and describe the transition between them.
Guide reviewed September 15, 2026. Veo and Runway model notes were checked against the official guides linked below.
Prompt anatomy
A prompt should remove uncertainty that matters to the shot. Do not list film terms because they sound professional. Name the physical result you need and leave out details already fixed by an input image.
Name the actor or object and use a physical verb: the woman turns left, the cup slides right, or the fabric lifts in the wind.
For text-to-video, state the location and framing: a close product shot on a red background, or a wide street view at night.
Keep them separate. The object can move while the camera stays locked, or the camera can track a subject that moves through the scene.
Use concrete relationships: left to right, slowly toward the lens, after the hand reaches the cup, or during the final two seconds.
Include them when text must create the look. If a first-frame image already shows the palette, lighting, and material, let the image carry those details.
Quote exact speech and name the audible event. Keep sound in edit notes when the selected generation workflow does not use it.
Reference video workflow
Watch first, write second. The prompt should come from a checked shot list, not from a single impression of the whole video.
Write the start and end time for each continuous shot. A title card, close-up, wide view, or new location begins a separate generation task.
Record the subject, action, scene, object direction, camera behavior, visible text, and sound. Mark uncertain causes as guesses instead of placing them in the prompt.
Decide whether the model will receive text only, one opening image, or opening and ending images. This determines which visual facts still belong in the text.
Put the main physical change first. Add camera or environmental movement next. Use order words only when two actions truly happen in sequence.
If direction is wrong, fix direction. If appearance is wrong, replace the frame. If timing is wrong, simplify the sequence or split the shot.
Trim each result to the source timecode, rebuild exact titles as overlays, restore sound where necessary, and place the generated shots in order.
From vague to testable
“Masterpiece,” “4K,” and “cinematic” do not explain the movement in a reference clip. A testable prompt tells you what to compare after generation.
The concrete version creates three checks: the cup stays centered, its size changes, and the camera does not move. If one check fails, you know which instruction to revise.
Model-specific editing
Reuse the same observed facts, then change how frames, motion, audio, and timing are delivered. The model interface is part of the prompt.
| Workflow | Keep in the text | Put outside the text |
|---|---|---|
| Text to video | Subject, action, scene, composition, camera, style, and supported sound cues | Shot order and trimming notes for the editor |
| Runway image to video | Subject motion, environmental motion, camera motion, direction, speed, and timing | Appearance already visible in the START image; exact text and sound for the edit |
| Veo first + last frames | The continuous transition, visual direction, dialogue, ambience, and effects | Opening and ending composition carried by the two images; shot assembly notes |
| Unknown model | A complete reviewed brief with clear facts and uncertainty | Unsupported model controls until you verify the target interface |
Model sources: Google's Veo 3.1 guide and Runway's image-to-video prompting guide. Confirm the controls in the model interface you use because model behavior and supported inputs can change.
Before generating
Read the prompt while replaying its shot. Delete any claim you cannot point to on the timeline. Then check that every noun has one job and every movement has an owner.
The prompt does not cross a cut or combine unrelated locations, angles, or title cards.
The first sentence states the physical change that must happen during the shot.
Words such as left, forward, closer, clockwise, slowly, or suddenly remove ambiguity when the source proves them.
The prompt does not call a moving object a pan or mistake digital scaling for a physical camera move.
The opening, reference, and ending images are not treated as interchangeable screenshots.
Unverified lens, hardware, seed, original prompt, lighting equipment, and render settings remain outside the final prompt.
Specific questions
For text-to-video, start with subject, physical action, scene, composition, camera movement, visual style, and sound when the model supports it. For image-to-video, let the image carry appearance and use the text mainly for subject, scene, and camera motion.
Usually no. Treat each continuous shot as one generation. State the important action order and timing, but split a clip at cuts or major scene changes instead of forcing several unrelated events into one prompt.
Use the shortest prompt that removes the important ambiguity. A simple image-to-video motion may need one or two sentences. Text-to-video or a shot with dialogue may need a longer structured brief. Length alone does not improve control.
Follow the model you selected. Runway recommends positive phrasing, such as “the locked camera remains still,” rather than “no camera movement.” Veo exposes a separate negative-prompt control in some API workflows, so do not mix model rules blindly.
You can reuse the same reviewed facts, but the final text should change. Veo can use first and last frames plus audio direction. Runway image-to-video uses the starting image for appearance and benefits from a shorter motion-focused prompt.
Play the source at the matching timecode and check subject, action, environment, object direction, camera behavior, visible text, dialogue, music, and effects. Keep guesses outside the prompt until you verify them.