Writing shots

Writing a shot: what a video model actually acts on

A video prompt is a shot list, not a wish. Here is which parts of a sentence change the render, which parts are quietly ignored, and how to tell the difference.

The short answer

Write a shot the way a cinematographer would describe one: name the subject, say what the light is doing, say what the camera is doing, and stop. Abstract mood words ("epic", "beautiful", "cinematic") change almost nothing, while a camera move, a light source and a lens length change almost everything. If a render came out flat, the missing ingredient is nearly always the camera — a static description produces a static shot.

A prompt is a shot list

The most useful shift is to stop writing what you want to see and start writing what the camera is doing while it sees it. A video model is not illustrating a noun; it is producing a moving image, and the parts of your sentence that describe motion, light and lens are the parts it has the most to act on.

Four ingredients carry nearly all the weight:

  • Subject — who or what, and what they are doing. A verb beats an adjective.
  • Light — where it comes from and what it is like. "Late afternoon through a dusty window" is a direction; "beautiful lighting" is not.
  • Camera — the move. Push in, pull back, track alongside, hold still. Naming "hold still" is itself a choice and produces a better locked-off shot than saying nothing.
  • Lens and distance — close on a face, wide on a room. This is what decides whether the render feels intimate or observational, and it is the ingredient people leave out most often.

Everything else — mood words, genre names, quality adjectives — is decoration. It is not harmful, but a shot that is failing is almost never failing for lack of the word "cinematic".

What gets ignored

Some instructions read like directions but have nothing to attach to.

Negations. "No people in the shot" frequently produces people. The model is matching a description to what it knows, and the strongest signal in that sentence is the word people. Describe the empty street instead: "an empty street at dawn, no movement but the traffic lights changing."

Precise counts. "Seven birds" gets you some birds. If a number matters to the story, keep it small enough to be a composition rather than a count — one, two, a pair.

Text. Signs, labels and titles come out as convincing-looking nonsense. Write the shot so nothing has to be legible, or accept it as texture.

Timing within the shot. "After three seconds, the door opens" is not a control the model has. A five-second clip is one continuous idea. If your shot has a then in it, that is not one shot — it is two, and the second one is an extension.

Length, resolution and what they cost you

A five-second shot is enough for one gesture: a look, a turn, a door opening. A ten-second shot has room for a gesture and its consequence, but it also has twice as long to drift, and drift is what makes generated footage feel generated. Most sequences are better as three five-second shots than one long one, and they cost the same.

Resolution is a straight trade: at 768p a credit buys a second, at 480p it buys two. The practical rule is to work at 480p while you are finding the film — while you are still discovering what the scene is — and switch up once you know what you are shooting. A wrong shot at high resolution is the most expensive thing you can make.

Reading a bad render

When a shot comes back wrong, the fault is usually diagnosable from the render itself.

What you gotWhat was missing
A postcard: pretty, static, emptyA camera move, and a subject doing something
A face that morphs mid-clipToo long a duration for the amount of motion described
Something generic and glossySpecificity: a place, a time of day, a real light source
Not the thing you asked for at allCompeting instructions — cut the sentence in half

The last row is worth dwelling on. When two parts of a prompt disagree, the model does not split the difference; it picks one. A shot described as both "wide, aerial" and "close on her hands" will be one or the other, and which one is not something you get to choose after the fact.

A worked example

Start with what most people write:

A robot in a city at night.

Add a subject doing something:

A robot crouches to fix a broken streetlight in a city at night.

Add the light:

A robot crouches to fix a broken streetlight, the only warm light on a wet street, everything else in blue neon.

Add the camera and the lens:

Close on a robot's hands as it fixes a broken streetlight — the only warm light on a wet street, everything else blue neon. Slow push in, shallow focus, handheld.

The last version is not longer because more words are better. It is longer because each clause added a decision the model would otherwise have made for you, and the decisions it makes for you are, by definition, the average ones.

Common questions

Does prompt length help?

Up to a point. A sentence with a subject, a light and a camera move in it beats a paragraph of adjectives every time. Past roughly forty words the additions start competing with each other and the model resolves the conflict by ignoring some of them — usually the ones you cared about.

Why do my shots look like stock footage?

Because stock footage is what a description with no camera in it looks like. "A woman in a cafe" is a stock frame. "A woman in a cafe, handheld, drifting past her shoulder toward the window" is a shot.

Every clip made with The Director is generated. None of it was filmed, and none of it is a record of anything that happened.