How to Write AI Video Prompts: A Filmmaker's Guide to Better Generations

By Richard Garland · 2026-08-16

AI video prompting is often treated like a contest to see who can write the longest description. It isn't. A great video prompt doesn't describe everything imaginable — it directs a shot.

That distinction matters. If you're generating an image, describing what something looks like may be enough. Video introduces another dimension: time. Something has to happen. A person moves, a camera follows, wind pushes through a room, light changes, someone hesitates before answering a question.

Good AI video prompting is less about piling adjectives into a paragraph and more about communicating the same things a director communicates to a cinematographer, a performer and a crew. This guide shows you how. If you want the wider context first, it sits alongside our complete AI short film workflow.

The AI Video Prompting Framework

A useful framework for building a prompt is:

Subject → Action → Environment → Shot → Camera → Lighting → Visual Style → Motion → Audio

You won't need every element in every prompt. But understanding them gives you control.

1. Subject

Who or what are we looking at? Instead of "a woman," try "a woman in her early thirties wearing a weathered orange utility jacket." Instead of "a robot," try "a battered silver delivery robot with one flickering blue eye."

Give the model enough to establish the subject without burying it under irrelevant detail. Ask what visually defines them, and what actually matters to this shot. If a detail needs to stay consistent across multiple shots — clothing, hairstyle, age, a particular prop — it's worth establishing clearly.

2. Action

Now answer the most important video question: what happens?

Compare "a detective in a diner" with "a detective sits alone in a nearly empty diner, slowly stirring untouched coffee while watching someone outside through the rain-covered window." Now there's a shot.

Actions can be extremely subtle. Someone might slowly raise their eyes, hesitate before opening a door, tighten their grip on an object, turn toward a sound, stumble backward, or remain perfectly still while the environment moves around them. Don't mistake more action for better action — one meaningful action is usually more useful than six competing ones.

3. Environment

Environment gives the model spatial and atmospheric context. "A man walks" becomes "a man walks through an abandoned subway station partially flooded with ankle-deep water." Now the environment participates: water ripples around his shoes, lights reflect from the floor, the architecture establishes scale.

Think about location, time of day, weather, background activity, foreground objects and atmosphere — but don't turn the environment into a furniture inventory. Describe the details that matter visually.

4. Choose the Shot

Here's where prompting starts becoming filmmaking. Don't just tell the model what to see. Tell it how we're seeing it: extreme wide, wide, medium, medium close-up, close-up, extreme close-up, over-the-shoulder, two-shot, POV, low angle, high angle, top-down, macro.

The shot size changes the meaning. Consider "wide shot of a lone astronaut standing inside an enormous abandoned hangar" versus "extreme close-up of the astronaut's eyes reflecting the abandoned hangar." Same character, completely different storytelling — the first communicates scale and isolation, the second communicates reaction.

Ask what information the audience needs from this shot, then frame accordingly.

5. Direct the Camera

Camera movement is one of the most powerful and most frequently abused parts of a prompt. The useful vocabulary is small and precise:

  • Static / locked — the camera doesn't move
  • Pan — rotates horizontally; tilt — rotates vertically
  • Dolly in / out — physically moves toward or away from the subject
  • Tracking shot — travels with the subject; truck — moves sideways
  • Crane — moves vertically or through a sweeping elevated path
  • Arc / orbit — circles the subject
  • Handheld — introduces instability and immediacy
  • Aerial / drone — moves through the scene from above

These aren't interchangeable vocabulary words. They change how a shot feels, and movement should have a reason. If a character realizes she's being followed, "slow dolly inward as her expression changes from confusion to fear" gradually closes the space between us and her. But "static wide shot as she realizes someone is standing motionless behind her" creates tension precisely through stillness. Cinematic camera movement isn't automatically better than a locked camera — sometimes not moving is the directing choice.

6. Lighting Is Storytelling

Instead of adding "cinematic lighting" to every prompt, describe where the light comes from and what it's doing. Compare that phrase with "cold moonlight enters through the blinds while a warm desk lamp illuminates one side of his face." The second gives the model something concrete.

Think about source, direction, intensity, color temperature, contrast, shadows and practical lights:

Harsh fluorescent ceiling lights create pale green highlights and deep shadows beneath the eyes.
Soft morning sunlight diffuses through sheer curtains, creating low-contrast natural light.
Flashing red emergency lights intermittently illuminate the dark corridor.

Lighting shouldn't simply make the shot prettier. It should help establish the world.

7. Define Style Without Drowning in Adjectives

This is where prompts go off the rails: "cinematic, masterpiece, ultra cinematic, incredible, award-winning, stunning, breathtaking, 8K, hyperrealistic, professional cinematography." That isn't direction. It's enthusiasm.

Instead, define the visual language:

Restrained 1970s science-fiction aesthetic, practical production design, muted earth tones, subtle film grain.
Clean contemporary commercial photography, high-key lighting, crisp surfaces, controlled studio reflections.
Naturalistic documentary aesthetic, available light, handheld observational camera.

Specific aesthetic decisions communicate more than stacks of generic quality words.

8. Describe Motion, Not Just Objects

Video models have to understand how the world changes over time, so don't forget secondary motion. "A woman stands on a train platform" becomes far more alive as: "a woman stands motionless on an outdoor train platform while wind pushes her coat and loose hair sideways. Commuters pass behind her in soft motion blur. A train approaches in the distance."

The subject barely moves, but the shot is alive. Look for motion in hair, clothing, smoke, steam, rain, dust, foliage, crowds, reflections, shadows, vehicles, water and background characters. Environmental movement can make an otherwise simple generation feel dramatically more convincing.

9. Add Audio Intentionally

Modern AI video models increasingly support native or prompt-directed audio. Think in three layers.

Dialogue — what is said, and how? Consider emotion, volume, pace, vocal quality and pauses, not just the words. Sound effects — what would actually make sound here? "Soft electrical buzzing, rain hitting the metal roof, distant thunder." Music — if the shot needs it, describe its function: "sparse ambient synth score slowly increasing in tension."

But don't automatically add music. Sometimes "no music, only room tone and breathing" is the stronger choice.

Putting It Together

Let's build a prompt progressively.

Weak:

A woman in a space station.

We know the subject and approximate location. Not much else.

Better:

A woman sits alone inside an abandoned space station at night. She looks through a window at Earth.

We have an action now, but we're still leaving most of the filmmaking to the model.

Directed:

Medium close-up of a woman in her early thirties wearing a faded orange utility jacket, sitting alone inside the dark observation deck of an abandoned orbital station. She slowly raises her eyes toward a large window as Earth comes into view beyond the glass. The camera performs a subtle dolly inward. Cold blue light from Earth illuminates her face while dim amber instrument lights glow behind her. Restrained cinematic science-fiction realism, shallow depth of field, subtle natural movement. Quiet electrical hum, distant structural creaks, no music.

We've now directed the subject, the action, the environment, the shot size, the camera, the lighting, the style, the motion and the audio. The prompt isn't better because it's longer. It's better because the additional words have jobs.

One Shot, One Purpose

One of the easiest ways to break a generation is asking too much of it:

A man enters a bar, sits down, orders a drink, notices his ex-wife across the room, walks over to her, they argue, she throws her drink at him and he leaves.

That's not a shot. That's a scene. Break it up:

  • Shot 1 — Wide tracking shot following a tired man entering a dim hotel bar and walking toward an empty stool.
  • Shot 2 — Medium shot as he sits at the bar and quietly signals the bartender.
  • Shot 3 — Close-up. He suddenly stops moving as something across the room catches his attention.
  • Shot 4 — Over-the-shoulder shot revealing a woman seated at a corner table.

Now the model has manageable jobs. And more importantly, you have an edit.

Prompt for What the Audience Sees

Avoid relying too heavily on abstract backstory. "Marcus is devastated because his brother died ten years ago and this is the first time he's returned home since the funeral" may help establish context, but almost none of it is visible.

Translate emotion into performance instead: "Marcus stands silently in the doorway of his childhood bedroom. His shoulders remain tense. He reaches toward an old photograph on the desk, hesitates before touching it, then slowly lowers his hand."

AI video models generate pictures and sound. Give internal ideas external evidence.

Image-to-Video Prompting Is Different

When you start from an image, that frame has already established much of the shot — what the character looks like, what they're wearing, where they are, the composition, the color, the lighting and the visual style. You don't need to describe all of that again.

Concentrate instead on what changes after frame one. If your starting image shows a woman beneath a neon sign in the rain, don't re-describe her hair and jacket. Write:

She slowly looks over her shoulder as wind moves her hair and jacket. Rain continues falling around her. The camera gently pushes toward her face. Her expression changes from calm to concerned as she notices something behind the camera.

The image handles appearance. The prompt handles time. That's the division of labor.

Put the technique into practice.

Generate from text or a starting image with leading AI video models directly on RevaultAI.

Generate Video

Prompting Dialogue

Dialogue deserves restraint. If a shot contains dialogue, give the character enough time to actually perform it. Don't ask someone to deliver five sentences while running down a hallway, firing a weapon and changing expression six times in an eight-second clip.

Break the scene apart:

Medium close-up. The woman stares at the radio, barely breathing. After a short pause she quietly says: "I know that voice."

Then cut, and let another shot carry the response. Treat dialogue as a performance, not a caption.

Camera Movement Without Chaos

A common mistake is stacking every move at once: "dolly zoom tracking orbit crane shot cinematic camera movement." Pick one movement and know why you're using it.

  • For intimacy — slow dolly inward
  • For isolation — slow dolly backward, gradually revealing the enormous empty room around him
  • For energy — fast handheld tracking shot following alongside the runner
  • For revelation — camera slowly cranes upward, revealing thousands of people beyond the wall
  • For importance — slow controlled arc around the subject
  • For tension — locked camera, no movement at all

How Much Detail Is Too Much?

There's no magic prompt length. Ask instead whether every instruction is helping the shot. A fifty-word prompt can be excellent. A 250-word prompt can be excellent. A 250-word prompt can also be a confused pile of contradictions.

Watch for instructions fighting each other: static handheld camera, fast slow movement, bright low-key lighting, extreme close-up showing the entire city. More detail doesn't fix contradictory direction. Clarity beats volume.

Change One Thing at a Time

Suppose the generation is almost right, but the camera moves too aggressively. Don't rewrite the character, location, lighting, lens, action and style all at once. Change "fast dolly inward" to "extremely slow, subtle dolly inward" and generate again.

Treat prompting like an experiment and change variables deliberately. Otherwise, when the next generation improves, you won't know why.

Use References and Frames When Available

Text isn't your only directing tool. Depending on the model and workflow, you may be able to use starting images, ending images, character references, scene references, previous video, reference audio or seeds.

Use them when available. Trying to force exact visual continuity entirely through prose is inefficient — a picture communicates character appearance, costume, composition and production design instantly. When a model supports first-and-last-frame generation, the two images can establish where a shot begins and ends, leaving the model to generate the transition between them.

Model behavior varies here, so it's worth knowing which tool you're working in. Our comparison of Seedance 2.0, Veo 3.1, Kling 3.0 and Wan 2.6 covers where each one is strongest.

Advanced: Prompt With Time

For more complicated generations, think temporally. Instead of describing a collection of actions, specify their order:

0–3 seconds: Wide shot. A man stands alone beneath a streetlight while heavy snow falls around him. Camera remains locked.

3–6 seconds: He notices something offscreen and slowly turns his head.

6–10 seconds: The camera begins a subtle dolly inward as his expression changes from confusion to recognition.

This can be useful when the model responds well to structured prompting. But don't use timestamps merely because they look sophisticated — use them when timing actually matters.

Negative Prompts: Use Them Surgically

If your model supports negative prompting, use it to address recurring unwanted results rather than turning it into a giant superstition list copied from somewhere online. Start with the positive direction, then exclude specific problems: "no camera shake, no additional people entering frame."

Behavior and availability vary by model, so use the controls provided by the system you're generating with rather than assuming every model interprets negative prompts identically.

The Prompt Is Not Sacred

This might be the most useful advice in the guide. If a generation contains four incredible seconds and six broken ones, use the four seconds. Don't spend another twenty generations trying to make the original prompt produce a flawless ten-second shot just because that's what you first imagined.

AI filmmaking isn't a prompt-writing competition. The prompt is a production tool. The footage is what matters.

A Reusable Prompt Template

[Shot / composition] of [subject + important visual details] [performing a clear action] in [environment + relevant atmospheric details]. [Camera movement]. [Lighting direction and quality]. [Visual style]. [Secondary motion]. [Dialogue / sound / ambience / music if needed].

In practice:

Low-angle medium shot of a battered service robot standing in the doorway of an abandoned roadside diner at dawn. The robot slowly steps inside and looks around the empty room. Camera gently tracks backward as it approaches. Pale morning sunlight enters through dusty windows, creating long shadows across the floor. Restrained retro-futurist realism, weathered practical design, muted colors. Dust drifts through the light and a broken ceiling fan turns slowly overhead. Quiet wind outside, soft mechanical footsteps, distant electrical buzz, no music.

That's a prompt with a job.

A Quick Checklist

Before generating, ask:

  • Subject — is it obvious what we're looking at?
  • Action — does something clearly happen?
  • Environment — do we know where the action occurs?
  • Composition — have I chosen the right shot size?
  • Camera — should the camera move? If so, how and why?
  • Lighting — where is the light coming from?
  • Style — have I defined a visual language instead of stacking buzzwords?
  • Motion — what else moves in the scene?
  • Audio — what should we hear?
  • Timing — am I asking for too much within the clip?
  • Purpose — what does this shot contribute to the film?

You don't need to specify all eleven every time. But you should know the answers.

Frequently Asked Questions

How long should an AI video prompt be?

There's no ideal word count. A prompt should be long enough to communicate the shot clearly without introducing unnecessary or contradictory instructions. Simple shots may require only a sentence or two, while highly directed shots may benefit from more detail.

Should I say "cinematic" in AI video prompts?

You can, but the word alone provides relatively little direction. Shot size, camera movement, lighting, composition and visual style communicate what you mean by "cinematic" much more precisely.

Why doesn't my AI video follow my entire prompt?

You may be asking too much of one generation. Simplify the action, eliminate conflicting directions and consider breaking a complex sequence into several shots.

What's the best structure for an AI video prompt?

A useful starting framework is subject, action, environment, shot, camera, lighting, style, motion, audio. Adapt it rather than treating it as a rigid formula.

Is image-to-video better than text-to-video?

Neither is universally better. Text-to-video offers more freedom and exploration. Image-to-video gives you a strong visual starting point and greater control over composition, character appearance and style.

How do I make AI video look more cinematic?

Think like a filmmaker rather than adding the word "cinematic." Make deliberate decisions about framing, camera movement, lighting, blocking, depth, environmental motion, sound and — most importantly — what the shot is supposed to communicate.

Should every AI video prompt include camera movement?

No. A locked camera can be just as intentional as an elaborate tracking shot. Movement should serve the shot.

Better Prompting Is Better Directing

The most important shift is simple. Stop asking "how do I describe this image?" and start asking "how would I direct this shot?"

Where is the camera? What does the subject do? What moves in the background? Where does the light come from? What changes between the first frame and the last? What should we hear? And why does this shot exist?

As AI video models improve, they'll handle more of the technical work. That doesn't make direction less important — it makes taste, intention and decision-making more valuable. The goal isn't to write the world's most impressive prompt. It's to make a shot worth putting in your film.

Turn your prompt into a film.

Generate your next shot on RevaultAI, then submit the finished work to the Gallery for consideration.

Generate Video Submit Your Film

AI filmmaking is evolving quickly. Model capabilities and available tools can change over time, so always check the current capabilities of the tools in your workflow.