A cinematographer's eye: why your generated shots look dead

The tools produce shots that are technically stunning and cinematically lifeless. The difference is not the prompt. It is the eye behind it, and that eye can be learned.

AA Abdelilah Arahal
6 min read Updated 21 September 2026

You write: "a man walking down a city street at night, cinematic."

You get a clean, well lit, completely flat shot. Nothing wrong with it, and nothing alive in it.

The problem is not the tool. The problem is that you described what is in the shot and said nothing about how it is photographed. The difference between those two is the difference between a shopping list and a scene.

The vocabulary you are missing

Every shot in every film you have seen is a set of decisions. A director does not say "film the man". They specify:

Decision What it means Effect on the viewer
Shot size How much of the frame the subject fills Emotional proximity
Lens 24mm wide, 85mm long Distortion and background separation
Height Eye level, low, high Who holds power in the scene
Movement Locked, tracking, crane, handheld Stability or unease
Lighting Where the light comes from, how hard The entire mood
Depth of field What is sharp and what dissolves Where you look

That last row is the most powerful tool on the list and the least used in prompts. When the background dissolves, you force the viewer onto one face. When everything is sharp, you let them wander.

Before and after

Before: "a man walking down a city street at night, cinematic"

After: "medium close shot of a man in his mid forties walking toward camera, 85mm lens, shallow depth of field dissolving the street lights behind him into circles, camera retreating ahead of him steadily on a dolly, hard side light from a neon sign to his left leaving half his face in shadow, street wet after rain and reflecting the light, night"

The second is not merely longer. The second contains decisions, and every decision eliminates thousands of possibilities the model was otherwise free to choose from.

The template I use

Prompt
Write me a video shot description in the language of cinematography, in this order, as one continuous paragraph: 1. Shot size (wide, medium, close, detail) 2. Subject: who or what, and exactly what they are doing 3. Lens in millimetres, and depth of field 4. Camera movement, or explicitly locked off 5. Light source, direction and hardness 6. Time of day and weather 7. One concrete tactile detail that grounds the place The idea I want to shoot: [your idea in one sentence] Do not use words like "cinematic", "stunning" or "high quality". Replace every vague adjective with a specific technical decision.

That final constraint matters most. Forbid yourself vague adjectives and you are forced into decisions.

What the tools still cannot do

Honestly, and I say this from a background behind the camera:

Continuity across shots remains weak. The same character across five shots drifts in the face. This is improving fast, but it is not solved.

Performance is absent. An expression moving from doubt to understanding over two seconds is what makes a scene; models generate one expression rather than the transition between two.

Spatial logic stumbles. A reverse shot that should show the same room from the opposite side rarely matches.

Which is why the most successful use today is not "making a film" but isolated shots: an establishing shot, a transition, a background plate, a visual element inside an edit shot by humans.

This is exactly where I see the gap. Plenty of people watch a generated shot online and try to reproduce it, and end up with something striking for two seconds and usable nowhere.

What I work on with the people I train is the opposite: not a shot that impresses in a feed, but one that gets used, in an advert, in a film, in a pitch to a client. The difference is not the tool. It is that the second one was written in the vocabulary of the trade: a specific camera move, an aspect ratio chosen for a reason, a deliberate colour treatment, and a point of view you can justify.

Someone who combines prompting skill with that vocabulary does not produce clips for show. They produce material that sells.

What this means for you

If you make content: learn the six camera terms in the table above. One hour of study improves every prompt you write afterwards, permanently.

If you are a working cinematographer: your advantage is not operating the tool. It is knowing what to ask for. Years spent looking through a lens gave you the vocabulary everyone else lacks.

If you train creative teams: start with the language of cinema, not the tool's interface. Tools change every six months. The language of the lens is a century old and is not going to change.

In closing

Generation made executing a shot nearly free, and when execution becomes free, all the value moves to the decision. Knowing what to shoot is now worth more than knowing how to shoot it.

Try this today: take a prompt you wrote before and rewrite it with the template above. The difference in the result will tell you everything.

And if you want to master this on your own projects, that is what we do in the AI filmmaking course.

Common questions

Why do my generated shots look flat?
Because you described what is in the shot without specifying how it is photographed. Size, lens, depth of field, lighting and movement create the feeling, and without them the model produces the average of everything it has seen.
Which words should I avoid in video prompts?
Vague adjectives like "cinematic", "stunning" and "high quality". Replace them with technical decisions: 85mm, shallow depth of field, hard side light, locked off camera.
What is the single most useful thing to learn first?
Depth of field. It is the strongest tool for directing the viewer's eye and the least used in prompts: a dissolved background forces attention onto one face, while everything sharp leaves the viewer to wander.
Where do video generation tools still fail?
Continuity across multiple shots, performance that changes over time, and spatial logic in reverse angles. That is why isolated shots rather than complete scenes are the most successful use today.
Is camera skill still worth having?
More than before. When execution becomes free the value moves to the decision, and someone who spent years behind a lens owns the decision vocabulary that people who only know the tool lack.
Share

No comments yet

Leave a comment

Never published. Used only if I reply to you directly.

Comments are read before they appear.

Related reading

4 min read

Will AI replace designers?

Execution got cheap. Taste did not. The problem is that many designers were selling execution while believing they were selling taste.

Start with a short call

Fifteen minutes to understand what your team does and what you want to change. If training is not the right answer, I will say so.