Prompt craft · Published 2026-09-24

Most AI video prompts fail because they are written like image prompts.

An image prompt describes a frame. A video prompt has to describe a shot: what is in it, what moves, where the camera is, and how long any of that has to hold together. Here is a six-slot structure that covers all of it, the rewrites that show what changes, and the prompt patterns that reliably produce a clip you cannot use.

AIVideoAuditor desk · Craft guide, not a benchmark · Verify behaviour against your model's current docs
The six slots

Write every prompt in the same six slots, in the same order. The order matters more than the wording: put the element that must not drift at the front, because the opening of a prompt is the part a video model tends to hold most consistently across frames. Everything after it is progressively more negotiable.

Slot 1 · Subject

Who or what the shot is about, described concretely enough to be re-identifiable. Vague subjects drift the fastest.

"A grey-haired barista in a green apron"

Slot 2 · Action

Exactly one continuous action. Not a sequence. If you catch yourself writing then, you have written two generations.

"pouring steamed milk into a cup"

Slot 3 · Camera

Framing plus movement. Saying nothing here does not give you a neutral camera, it gives you whatever the model defaults to, which is usually a slow drift.

"medium shot, locked-off camera"

Slot 4 · Lighting

The single biggest lever on whether output reads as real or as rendered, and it costs you nothing to specify.

"warm window light from the left, soft shadows"

Slot 5 · Style

Medium and treatment. Pick one register and stop. Stacking cinematic, hyperreal, 8k and anime in one prompt makes them compete.

"shot on 35mm film, shallow depth of field"

Slot 6 · Duration

How long the shot needs to hold. Ask for the shortest clip that serves the edit. Coherence on current models degrades with length, so a 5-second ask is a materially easier ask than a 10-second one.

"5 seconds"

A grey-haired barista in a green apron, pouring steamed milk into a cup, medium shot, locked-off camera, warm window light from the left with soft shadows, shot on 35mm film with shallow depth of field, 5 seconds.

Three rewrites
Product shot

BEFORE "beautiful cinematic product video of a candle, 8k, hyperrealistic, amazing quality, trending"

AFTER "A cream pillar candle on a walnut table, flame flickering steadily, slow push-in on a 50mm lens, low warm key light from the right against a dark background, shot on 35mm film, 5 seconds."

The before is six quality adjectives and no shot. It names no camera, no light and no motion, so the model invents all three. The after gives the model one subject, one motion, one camera move and one lighting setup, which is the whole job. Quality words like 8k and amazing quality are doing nothing here that the film stock and lens references are not doing better.

Character shot

BEFORE "a woman walks into the office, sits down at her desk, opens her laptop and starts typing while talking to her colleague"

AFTER "A woman in a navy blazer typing at a laptop, seated at a desk, medium close-up, locked-off camera, cool overhead office light, natural documentary style, 5 seconds."

The before chains four actions and adds a second character into a few seconds of output. That is the single most reliable way to get a clip where the prompt is partly ignored, because the model has more instructions than it has frames to satisfy them. The after keeps one person doing one thing. Shoot the walk-in as its own generation and cut the two together.

Text on screen

BEFORE "a storefront with a sign that reads FRESH BREAD DAILY, morning light"

AFTER "A bakery storefront with an unlettered wooden sign above the door, morning light, static wide shot, 5 seconds."

Readable text in frame remains one of the weakest areas of current video models, and longer strings degrade faster than short ones. Asking for a blank sign and adding the lettering in your editor gets you a usable clip on the first generation instead of a reroll loop.

What to keep out of the prompt

Four patterns cost more credits than anything else, because they tend to fail in ways a reroll does not fix. They are worth designing around rather than prompting harder at.

  • Chained actions. Anything with then in it. One action per generation. See the per-model references for prompt adherence failure.
  • Readable text in frame. Signs, labels, screens, subtitles. Add them in post. Reference: text rendering failure.
  • Hand close-ups. Fingers at the centre of frame are still where anatomy breaks most visibly. Frame them further from camera. Reference: hand artifact failure.
  • Negative phrasing in the main prompt. No watermark and not blurry name the thing you do not want. Use the model's dedicated negative prompt field if it has one, and otherwise leave it out.
Adapting the structure per model

The six slots are model-agnostic, but which slots carry the most weight is not, and the models change often enough that anything specific here should be checked against current documentation before you rely on it. Two practical adjustments are stable enough to be worth stating:

  • Models that generate audio need a seventh slot for it, and it should be as literal as the rest. Describe the sound you want, not the mood.
  • Image-to-video changes what slot 1 is for. The reference image already fixes the subject, so spend slot 1 on what should stay locked and give the remaining budget to action and camera.

For prompts that real creators actually shipped, with the model each one ran on, see the prompt library. For the failure modes named above, with what each one looks like and how to describe it when you open a support ticket, see the failure reference.

Don't want to fight the tools? We'll produce your product video for you — scroll-stopping, platform-native, done in 2-3 days.

We'll produce it for you →