How to write video to prompt · practical tutorial

How to Write Video to Prompt Instructions

Break the reference into shots, write only visible and audible facts, give each motion a clear owner and direction, then adapt the frames and text to the AI video model you will use.

Direct answer

A useful video prompt answers seven concrete questions

Who or what is in the shot? What physical action happens? Where does it happen? How is the shot framed? What does the camera do? How fast and in which direction does each movement happen? What dialogue, music, ambience, or effect must be heard?

Write one continuous shot at a time. When a reference video cuts to another angle or scene, start another prompt. If you use image-to-video, attach the opening frame and remove visual details the image already communicates. If you use a model with first-and-last-frame control, attach both boundary frames and describe the transition between them.

Guide reviewed September 15, 2026. Veo and Runway model notes were checked against the official guides linked below.

Prompt anatomy

Write What the Model Must Control

A prompt should remove uncertainty that matters to the shot. Do not list film terms because they sound professional. Name the physical result you need and leave out details already fixed by an input image.

Subject + action

Name the actor or object and use a physical verb: the woman turns left, the cup slides right, or the fabric lifts in the wind.

Scene + composition

For text-to-video, state the location and framing: a close product shot on a red background, or a wide street view at night.

Camera + object motion

Keep them separate. The object can move while the camera stays locked, or the camera can track a subject that moves through the scene.

Direction + speed + timing

Use concrete relationships: left to right, slowly toward the lens, after the hand reaches the cup, or during the final two seconds.

Style and light

Include them when text must create the look. If a first-frame image already shows the palette, lighting, and material, let the image carry those details.

Dialogue and sound

Quote exact speech and name the audible event. Keep sound in edit notes when the selected generation workflow does not use it.

Reference video workflow

How to Write a Prompt From an Existing Video

Watch first, write second. The prompt should come from a checked shot list, not from a single impression of the whole video.

  1. 01

    Mark every cut

    Write the start and end time for each continuous shot. A title card, close-up, wide view, or new location begins a separate generation task.

  2. 02

    Describe only evidence

    Record the subject, action, scene, object direction, camera behavior, visible text, and sound. Mark uncertain causes as guesses instead of placing them in the prompt.

  3. 03

    Choose the generation workflow

    Decide whether the model will receive text only, one opening image, or opening and ending images. This determines which visual facts still belong in the text.

  4. 04

    Write one action sequence

    Put the main physical change first. Add camera or environmental movement next. Use order words only when two actions truly happen in sequence.

  5. 05

    Generate, compare, and change one field

    If direction is wrong, fix direction. If appearance is wrong, replace the frame. If timing is wrong, simplify the sequence or split the shot.

  6. 06

    Restore the edit

    Trim each result to the source timecode, rebuild exact titles as overlays, restore sound where necessary, and place the generated shots in order.

From vague to testable

Replace Praise Words With Visible Instructions

“Masterpiece,” “4K,” and “cinematic” do not explain the movement in a reference clip. A testable prompt tells you what to compare after generation.

The concrete version creates three checks: the cup stays centered, its size changes, and the camera does not move. If one check fails, you know which instruction to revise.

Model-specific editing

Do Not Send the Same Final Prompt to Every Model

Reuse the same observed facts, then change how frames, motion, audio, and timing are delivered. The model interface is part of the prompt.

WorkflowKeep in the textPut outside the text
Text to videoSubject, action, scene, composition, camera, style, and supported sound cuesShot order and trimming notes for the editor
Runway image to videoSubject motion, environmental motion, camera motion, direction, speed, and timingAppearance already visible in the START image; exact text and sound for the edit
Veo first + last framesThe continuous transition, visual direction, dialogue, ambience, and effectsOpening and ending composition carried by the two images; shot assembly notes
Unknown modelA complete reviewed brief with clear facts and uncertaintyUnsupported model controls until you verify the target interface

Model sources: Google's Veo 3.1 guide and Runway's image-to-video prompting guide. Confirm the controls in the model interface you use because model behavior and supported inputs can change.

Before generating

Video Prompt Review Checklist

Read the prompt while replaying its shot. Delete any claim you cannot point to on the timeline. Then check that every noun has one job and every movement has an owner.

One shot

The prompt does not cross a cut or combine unrelated locations, angles, or title cards.

One clear main action

The first sentence states the physical change that must happen during the shot.

Motion has direction

Words such as left, forward, closer, clockwise, slowly, or suddenly remove ambiguity when the source proves them.

Camera and object are separate

The prompt does not call a moving object a pan or mistake digital scaling for a physical camera move.

Frame role is explicit

The opening, reference, and ending images are not treated as interchangeable screenshots.

Unknown details stay unknown

Unverified lens, hardware, seed, original prompt, lighting equipment, and render settings remain outside the final prompt.

Specific questions

How to Write Video to Prompt FAQ

What is the basic structure of an AI video prompt?

For text-to-video, start with subject, physical action, scene, composition, camera movement, visual style, and sound when the model supports it. For image-to-video, let the image carry appearance and use the text mainly for subject, scene, and camera motion.

Should I describe every second of the video?

Usually no. Treat each continuous shot as one generation. State the important action order and timing, but split a clip at cuts or major scene changes instead of forcing several unrelated events into one prompt.

How long should a video prompt be?

Use the shortest prompt that removes the important ambiguity. A simple image-to-video motion may need one or two sentences. Text-to-video or a shot with dialogue may need a longer structured brief. Length alone does not improve control.

Should I use negative prompts?

Follow the model you selected. Runway recommends positive phrasing, such as “the locked camera remains still,” rather than “no camera movement.” Veo exposes a separate negative-prompt control in some API workflows, so do not mix model rules blindly.

Can I copy one prompt into Veo and Runway?

You can reuse the same reviewed facts, but the final text should change. Veo can use first and last frames plus audio direction. Runway image-to-video uses the starting image for appearance and benefits from a shorter motion-focused prompt.

How do I check whether a generated prompt is accurate?

Play the source at the matching timecode and check subject, action, environment, object direction, camera behavior, visible text, dialogue, music, and effects. Keep guesses outside the prompt until you verify them.