Opening frame
A real image captured at the start of the shot. Use it as the primary image input so the generated shot begins with the source composition.
Veo 3.1 workflow selected
Upload a reference clip and turn every shot into a Veo-ready prompt with real opening and ending frames, timing, visual direction, and reviewed sound notes.
Upload and analyze
Get prompts by shot, real start and end frames, and model-ready instructions in one analysis.
2. Upload your source video
Choose a short video to begin.
Report a problemYour source stays visible
Compare every generated shot description with the original video before you copy the prompt.
Direct answer
Split the reference at every cut. For each shot, use its actual opening image as Veo's first frame, its actual ending image as the last-frame constraint, and a prompt that describes what changes between them. Put reviewed dialogue, ambience, and sound effects in a separate audio section.
Generate each shot separately, trim it to the source timecode, and assemble the clips in order. This gives Veo one continuous action to solve at a time instead of asking it to reconstruct an edited multi-shot video in one generation.
Prepared for the next click
The page has already selected Veo. After analysis, every shot contains the exact pieces you need to move from the reference clip to Veo without guessing which frame or paragraph to use.
A real image captured at the start of the shot. Use it as the primary image input so the generated shot begins with the source composition.
A real image captured just before the next cut. Use it as the last-frame constraint when your Veo 3.1 workflow exposes that control.
Observed subject, physical action, setting, object movement, camera behavior, and visible text are written as one continuous shot.
Dialogue, music, ambience, and sound effects are kept apart from the visual paragraph so you can check or remove them before generation.
One shot at a time
Do not paste the complete export into one Veo request. The export includes production notes for you; the model only needs the frame pair and the prompt inside the marked copy section.
Play the source from the shot start to the shot end. Correct any mistaken action, camera movement, visible wording, dialogue, or sound before generating.
The start image establishes the opening composition. The end image constrains the final state before the edit cuts to the next shot.
The pack compares the source shot length with Veo’s available duration choices and tells you what to generate and how much to keep in the final edit.
Copy only the text between the prompt markers. Generate, compare the action and sound with the source, then revise the incorrect field instead of adding random style words.
Keep the source-length portion of each generation. Rebuild exact titles as editor overlays when generated text changes, then place the clips in the listed order.
Verified fixture pattern
The project's test fixture contains a yellow cup on a dark blue background, visible “MUSE” text, a scale change, and a short tone. The useful output is specific about those observed changes and stays silent about an unverified lens, render engine, or original model.
Official controls checked September 15, 2026
Google's current Veo 3.1 documentation says the model can generate native audio and can use a first image plus a last frame. Its prompt guide recommends concrete subject, action, scene, camera, composition, ambience, and audio cues. That is why this page preserves appearance and motion details instead of reducing the video to a one-line summary.
| Decision | Use this input | Reason |
|---|---|---|
| Opening composition matters | START image | The first image gives Veo the subject, layout, color, and exact state at the beginning of the shot. |
| Ending state matters | END image | Veo 3.1 can interpolate toward a last-frame constraint, which is useful for a controlled transition. |
| Speech or sound matters | Reviewed audio paragraph | Veo can create native audio, but generated sound is a new result and is not extracted audio from the reference file. |
| The source contains several cuts | Separate shot requests | A frame pair should describe one continuous shot. The edit, rather than a single generation, restores the cut sequence. |
Capability source: Google AI for Developers — Generate videos with Veo 3.1. Model controls can change, so check the Veo interface you use before spending credits.
Fit and limits
Use it when you have a short reference, want the same sequence of shots, and need a starting prompt plus real frame inputs. It is especially useful when a product, character, camera move, or transition must begin and end in recognizable states.
It cannot recover a proprietary prompt or guarantee an exact copy. A first-and-last-frame pair may constrain composition but still leave the path between those frames open to the model. Fast hand movement, readable typography, precise lip sync, and identity can require several attempts or manual finishing.
If only motion matters and one starting image already contains the complete look, compare the Runway video-to-prompt workflow, which produces shorter motion-focused text.
Specific questions
Choose Veo, upload a short MP4, MOV, or WebM, and review the detected shots. For each shot, download the opening and ending frames, copy only that shot prompt, and generate the shots separately before joining them in source order.
Veo 3.1 supports a first image plus a last-frame constraint. The pair tells the model how a shot should begin and end, while the text explains the motion, scene changes, and sound between those two states.
Veo 3.1 can generate native audio. Keep reviewed dialogue in quotation marks when exact wording matters, then describe sound effects and ambience plainly. Frame to Prompt separates audio direction so you can verify it against the source before copying.
A single prompt is a poor fit for unrelated cuts. Generate one continuous shot at a time, using that shot’s frames and text. Trim each result to the source timecode and assemble the generated clips in the listed order.
No. It analyzes your reference and prepares the frames, prompt, timing, and edit instructions. You still paste the prompt and upload the listed images inside the Veo interface you use.
No model can guarantee an exact reconstruction from a finished clip. Faces, readable text, fast actions, precise timing, and audio may drift. Use the extracted frames as constraints, check one shot at a time, and add exact titles in your editor when necessary.