By Genflow Editorial
Evidence checked: August 27, 2026
A text-to-video prompt should define one observable shot: what the viewer sees at the start, what changes, how the camera behaves, what must remain stable, and how the shot should end. It is not a screenplay, a mood board, and a list of every filmmaking word you know.
This guide turns that principle into a production record you can review and reuse. It draws on current first-party prompting guidance and Genflow's public Text to Video AI workflow. It does not present a new model benchmark or promise that the same wording will behave identically across models.
The short answer: write a shot contract
For a short generated clip, start with six fields:
- Frame: shot size, angle, and the main subject.
- Action: one primary change that can be seen over time.
- Camera: one movement, its direction, and its pace—or a locked camera.
- World: location, lighting, atmosphere, and visual treatment.
- Continuity: the identity, object, layout, or spatial relationship that must remain stable.
- End state: what should be true in the final moment.
Write the result as clear sentences:
[Frame and subject]. [Primary action]. [Camera behavior and pace]. [World]. [Continuity rule]. [End state].
Runway's current Text to Video Prompting Guide recommends clarity rather than a target word count and warns that overly long prompts can introduce conflicting requests. Google Cloud's Veo prompting guide uses a related structure: cinematography, subject, action, context, and style or ambience.
The six-field contract adds continuity and an end state because those fields make the output easier to approve. They state what cannot drift and where the shot is supposed to arrive.
Separate the creative brief from model settings
Do not force every production decision into prose. Keep two layers.
| Creative shot contract | Generation settings |
|---|---|
| Subject, action, camera, world, continuity, end state | Model and version |
| What the viewer should observe | Duration options supported by that model |
| Story and visual priorities | Aspect ratio and resolution |
| Acceptance requirements | Seed, reference inputs, audio controls, or safety settings when available |
This separation keeps the brief portable without pretending that controls are universal. Current Google AI video documentation, for example, lists capabilities and request parameters that depend on the selected model. Another provider may expose a different set.
If you hide a platform setting inside a long paragraph, a teammate may not know whether the model ignored it or the request never configured it. Save the prose and the settings together, but label them separately.
Turn an idea into one observable shot
Suppose the idea is:
A child finds a tiny glowing whale in a library and follows it into a storybook world.
That is a story premise, not one short shot. It contains a discovery, a reaction, a pursuit, a location transition, and a new world. Asking one generation to solve all of it makes review ambiguous.
Break the premise into beats:
- Establish the child alone in the library.
- Reveal the glowing whale between the books.
- Show the child's reaction.
- Follow the whale toward an opening book.
- Transition into the illustrated world.
Now choose one beat for the first test. For example, the reveal:
Medium-wide eye-level shot of a child standing between tall library shelves. A palm-sized blue whale made of soft light rises slowly from behind an open book. The camera makes a gentle four-second push-in. Warm reading lamps contrast with the whale's cool glow; dust moves lightly in the air. Preserve the child's yellow raincoat, the whale's scale, and the shelf geometry. End with the whale hovering at the child's eye level while the child remains still.
This is an untested starting brief. Its value is not that it guarantees a good clip. Its value is that the team can identify whether the subject, action, camera, continuity, and ending were followed.
Create the acceptance matrix before generation
Approval should not begin with “Do I like it?” Translate the shot contract into observable checks first.
| Field | Expected behavior | Blocking failure | Revision target |
|---|---|---|---|
| Frame | Medium-wide, child and nearby shelves readable | Crop hides the subject or changes the intended distance | Frame and subject placement |
| Action | One small whale rises from the book | Several whales appear or the whale performs unrelated actions | Primary action and quantity |
| Camera | Slow push-in for the shot | Orbit, whip pan, cut, or rapid zoom | Camera direction and pace |
| World | Warm lamps, cool whale glow, light dust | Lighting switches style or obscures the subject | Lighting and atmosphere |
| Continuity | Raincoat, whale scale, and shelves stay recognizable | Identity, scale, clothing, or geometry visibly drifts | Continuity rule or shot ambition |
| End state | Whale hovers at the child's eye level | The action has no readable final pose | Ending and timing |
The matrix distinguishes a requirement failure from a subjective preference. A stakeholder may prefer a more saturated blue, but an extra whale is a direct mismatch with the brief.
Give the camera one job
Camera words are useful when they describe a spatial path. “Cinematic” does not tell the viewpoint where to move. “Slow dolly push-in over four seconds” does.
Runway's camera prompting reference separates framing, angle, movement, and speed. For a first test, use one movement source of truth:
- Locked camera: use when subject action is the main event or exact layout matters.
- Push-in or dolly in: use when attention should narrow toward a subject or reveal.
- Pull-back: use when the environment or scale should become the new information.
- Pan or tilt: use when the camera needs to reveal something beside or above the starting frame.
- Tracking move: use when the camera should travel with a moving subject.
Avoid stacking an orbit, zoom, crane, and fast subject movement in the same diagnostic pass. If the output fails, you will not know which spatial instruction to simplify.
Define continuity as an acceptance requirement
Text-to-video starts without a fixed source frame, so the model has more freedom to invent the scene. That freedom is useful for concepts and background footage, but it can make exact identity or layout harder to preserve. Runway explicitly positions text-to-video as useful when exact character or scene consistency is not the priority.
A continuity rule does not guarantee preservation. It names the requirement the reviewer should check:
- preserve the subject's clothing and defining features;
- keep the product label plane facing the camera;
- retain the number and relative position of key objects;
- keep the room layout and screen direction coherent;
- maintain one visual treatment across the shot.
If exact identity, typography, packaging, or composition is mandatory, the production method may need a reference image, image-to-video, compositing, or a separate finishing pass. Do not respond to a method mismatch by endlessly adding adjectives.
Use one-variable revisions
When the first output fails, record expected versus observed behavior and change one failed field.
| Version | Field changed | Expected | Observed | Decision |
|---|---|---|---|---|
| v1 | Initial contract | Slow push-in; one whale; stable child | Example: camera orbits; two whales appear | Reject |
| v2 | Camera only | Locked camera; one whale; stable child | Record the actual result | Accept or revise |
| v3 | Action only, if needed | One whale rises and stops at eye level | Record the actual result | Accept or revise |
The observations in this table are placeholders for a real run, not claims that this article generated those results. Replace them with actual output notes.
Changing one field preserves information. If you rewrite the full prompt after every failure, you cannot tell whether the improvement came from the camera, action, timing, or a different random generation.
Three prompt starters with different risk profiles
These examples are newly written, untested briefs. Adapt them to the controls and limits of the model you select.
Low-complexity atmosphere shot
Wide locked shot of a quiet greenhouse just before sunrise. Condensation slides slowly down the nearest glass panes while leaves move gently in a light breeze. Pale blue morning light gradually warms. Preserve the greenhouse structure and the position of the central worktable. End as the first direct sunbeam reaches the table.
Why it is easier to diagnose: the camera is fixed, the environment supplies the motion, and the end state is visible.
Product reveal
Medium close-up of an unbranded ceramic bottle centered on a dark stone pedestal. A narrow band of light travels slowly from left to right across the bottle while the camera makes a restrained push-in. Fine mist stays behind the product. Preserve the bottle silhouette, cap, pedestal edge, and centered composition. End on a clean front-facing hero frame.
Production warning: if a real package needs exact readable text, use an approved product reference and plan to verify or composite the label. A text instruction alone is not proof of brand accuracy.
Character beat
Medium eye-level shot of a courier in a red jacket waiting under a bus shelter at night. The courier looks up as one paper ticket slips from their hand and moves across the wet pavement. The camera pans slowly with the ticket for two seconds, then stops. Preserve the jacket, bag, shelter, and direction of travel. End with the ticket resting beside a puddle in the foreground.
Why it is riskier: attention transfers from a person to a small moving object. Review screen direction, hand anatomy, object count, and the final foreground position separately.
Convert multiple shots into a workflow, not one giant prompt
For a sequence, give each beat its own contract and connect them with shared continuity fields. A compact shot record can include:
| Record | Purpose |
|---|---|
| Scene and shot ID | Keeps prompts aligned with the script or storyboard |
| Shot contract | Defines frame, action, camera, world, continuity, and ending |
| Model settings | Records the actual generation configuration |
| Shared continuity block | Identifies recurring character, object, palette, and location rules |
| Output and version | Preserves the result that was reviewed |
| Acceptance matrix | Shows which requirements passed or failed |
| Decision and owner | Records approval, rejection, and the next action |
An AI-generated shot list can provide a starting structure, but it does not make directorial decisions or maintain consistency automatically. Runway's AI shot list guide recommends human review and maps framing, angle, movement, subject action, and technical notes into prompt-ready language.
Save the reason a version won
Genflow's public Text to Video AI page describes a workflow that keeps scene intent, prompt structure, model choices, constraints, and review criteria together, then reuses that production path for new variants. The broader AI Video Generator page similarly treats prompts, references, models, and review context as a reusable system.
The reusable asset is therefore not the prompt sentence alone. Save:
- the original idea and selected shot beat;
- the six-field contract;
- provider, model version, duration, aspect ratio, and available controls;
- reference inputs and their rights or provenance;
- every meaningful prompt version;
- expected-versus-observed notes;
- accepted output and rejection reasons;
- the person or rule that approved the shot.
When you adapt the workflow to another product or character, preserve the structure but re-evaluate the continuity rules and risk. Reuse is a controlled starting point, not evidence that one prompt transfers unchanged.
Text-to-video prompt FAQ
How long should a text-to-video prompt be?
There is no universal ideal length. Include enough detail to specify one coherent shot and its acceptance requirements. Remove decorative wording that adds no observable instruction, and split conflicting actions into separate shots.
Should I include camera movement in every prompt?
No. A locked camera is a valid direction. Use camera movement when it reveals information or supports the intended beat, not because every clip needs motion in both the subject and viewpoint.
Can a prompt preserve the same character across many shots?
A prompt can name continuity requirements, but text alone does not guarantee identity. Use the reference and consistency controls supported by the selected model, keep shared attributes documented, and review every shot. Exact continuity may require a different generation method or post-production.
Should a full story go into one prompt?
Usually not for a diagnostic short clip. Break the story into observable beats, assign each beat a shot contract, and assemble the approved shots. This makes failure and revision easier to locate.
What should I change first when the output is wrong?
Compare the output with the acceptance matrix. Change the smallest failed field: framing, action, camera, world, continuity, or end state. If the problem depends on exact identity, text, or missing visual reference information, reconsider the production method.
Start a reusable text-to-video workflow
Open Genflow's Text to Video AI workflow. Select one story beat, write the six-field shot contract, and add the acceptance matrix before generating. Keep the prompt, settings, output, and review decision together so the next variant starts with evidence rather than memory.
Sources and creation method
Responsible publisher: Genflow Editorial, the team responsible for Genflow's public product documentation and blog. Corrections can be sent through Genflow Support.
First-party Genflow evidence: Text to Video AI and AI Video Generator, checked August 27, 2026.
Primary external guidance: Runway Text to Video Prompting Guide, Runway AI camera prompts, Runway AI shot list, Google Cloud Veo 3.1 prompting guide, and Google AI video documentation, checked August 27, 2026.
How this article was produced: AI assisted with search-result research, drafting, and organization. The retained evidence pack lists ten reviewed pages and the selected content gap. This run did not perform a controlled video-generation benchmark, so the prompt examples are labeled as untested starting briefs and no success-rate claim is made.
