Scene text with timecodes
Each shot receives a start time, end time, subject, setting, composition, action, and transition boundary.
Free · editable text output
Convert a short video into structured text by shot, including the scene, visible action, camera movement, timing, readable words, and reviewable sound notes.
Upload and analyze
Get prompts by shot, real start and end frames, and model-ready instructions in one analysis.
2. Upload your source video
Choose a short video to begin.
Report a problemYour source stays visible
Compare every generated shot description with the original video before you copy the prompt.
Direct answer
Upload a short video and the tool writes a separate text prompt for each detected shot. The result names the subject, scene, physical action, object movement, camera behavior, visible words, and sound cues, with timecodes that let you check every description against the source.
This page describes how the video looks and moves so the text can guide an AI video model. It is not a transcript generator, caption downloader, or promise to recover the original hidden prompt.
Page reviewed September 15, 2026. Upload limits and output fields match the live tool on this page.
Readable, editable output
The output is split into fields before it is assembled into a prompt. That gives you a place to correct one mistaken detail instead of editing a dense paragraph.
Each shot receives a start time, end time, subject, setting, composition, action, and transition boundary.
Subject movement, object movement, and camera movement are described separately, including direction and timing when the evidence supports them.
Readable on-screen text, dialogue, music, ambience, and effects appear as reviewable fields rather than being silently mixed into visual prose.
Edit the fields, choose a universal or model-specific format, copy the complete pack, or download the result as a text file.
Video in, checked text out
A useful text prompt describes change over time. Review direction, speed, and camera behavior carefully because one wrong verb can produce a different shot.
Choose an MP4, MOV, or WebM file no longer than 60 seconds and strictly below 20 MiB (20,971,520 bytes). A short clip is faster to check shot by shot.
Start with the subject, action, scene, camera, objects, visible text, and sound notes. The timecode tells you exactly where to compare each field with the source.
Remove style labels you cannot see, correct misread words, and separate camera motion from an object moving across a static frame.
Use Universal for a complete description or select an AI model format. Copy the full sequence, or work with one shot when you plan to generate it separately.
Choose the right kind of text
The phrase “video to text” can describe three jobs. This page is for a generation prompt, so it prioritizes visible action and camera behavior while keeping sound available for review.
| Output | What it records | Best next use |
|---|---|---|
| Video to text prompt | Shots, scene, action, camera, timing, visible text, and supported sound notes | Recreate or adapt the reference with an AI video generator. |
| Transcript | Spoken words in chronological order | Captions, meeting notes, or dialogue editing. |
| Short video description | A compact summary of what happens overall | Cataloging, search, or a quick human overview. |
Know what the text can prove
The output can describe frames and audible events in the uploaded file. It cannot reveal a private prompt, random seed, hidden reference image, model setting, or editor timeline that is no longer present in the render.
Speech and visible words can be missed or misread, especially when they are fast, quiet, stylized, or small. Review them at the listed timecode before using the prompt.
If your goal is an exact transcript or accessibility captions, use a dedicated transcription tool. If your goal is to reproduce the visual sequence with AI, use the structured prompt and frames returned here.
Specific questions
It is a written set of instructions describing a video’s shots, scene, subject, action, camera behavior, timing, visible words, and supported sound cues for use in an AI video workflow.
It can return reviewable dialogue and sound notes when supported, but it is not an exact transcription or caption service. Use a transcription tool when word-for-word speech is the main task.
Yes. Edit each shot field after analysis. The complete prompt and every model-specific format update from your corrected text.
Yes. After analysis, copy the complete prompt pack or download it as a plain text file.
Yes. The current public tool is free, needs no login or payment details, and has no ad gate. It permits up to three analyses per visitor each day, subject to a small site-wide capacity.
You can upload MP4, MOV, or WebM files no longer than 60 seconds and strictly below 20 MiB (20,971,520 bytes). Try an H.264 MP4 if another file will not preview in your browser.