A cinematographer's eye: why your generated shots look dead
The tools produce shots that are technically stunning and cinematically lifeless. The difference is not the prompt. It is the eye behind it, and that eye can be learned.
You write: "a man walking down a city street at night, cinematic."
You get a clean, well lit, completely flat shot. Nothing wrong with it, and nothing alive in it.
The problem is not the tool. The problem is that you described what is in the shot and said nothing about how it is photographed. The difference between those two is the difference between a shopping list and a scene.
The vocabulary you are missing
Every shot in every film you have seen is a set of decisions. A director does not say "film the man". They specify:
| Decision | What it means | Effect on the viewer |
|---|---|---|
| Shot size | How much of the frame the subject fills | Emotional proximity |
| Lens | 24mm wide, 85mm long | Distortion and background separation |
| Height | Eye level, low, high | Who holds power in the scene |
| Movement | Locked, tracking, crane, handheld | Stability or unease |
| Lighting | Where the light comes from, how hard | The entire mood |
| Depth of field | What is sharp and what dissolves | Where you look |
That last row is the most powerful tool on the list and the least used in prompts. When the background dissolves, you force the viewer onto one face. When everything is sharp, you let them wander.
Before and after
Before: "a man walking down a city street at night, cinematic"
After: "medium close shot of a man in his mid forties walking toward camera, 85mm lens, shallow depth of field dissolving the street lights behind him into circles, camera retreating ahead of him steadily on a dolly, hard side light from a neon sign to his left leaving half his face in shadow, street wet after rain and reflecting the light, night"
The second is not merely longer. The second contains decisions, and every decision eliminates thousands of possibilities the model was otherwise free to choose from.
The template I use
That final constraint matters most. Forbid yourself vague adjectives and you are forced into decisions.
What the tools still cannot do
Honestly, and I say this from a background behind the camera:
Continuity across shots remains weak. The same character across five shots drifts in the face. This is improving fast, but it is not solved.
Performance is absent. An expression moving from doubt to understanding over two seconds is what makes a scene; models generate one expression rather than the transition between two.
Spatial logic stumbles. A reverse shot that should show the same room from the opposite side rarely matches.
Which is why the most successful use today is not "making a film" but isolated shots: an establishing shot, a transition, a background plate, a visual element inside an edit shot by humans.
This is exactly where I see the gap. Plenty of people watch a generated shot online and try to reproduce it, and end up with something striking for two seconds and usable nowhere.
What I work on with the people I train is the opposite: not a shot that impresses in a feed, but one that gets used, in an advert, in a film, in a pitch to a client. The difference is not the tool. It is that the second one was written in the vocabulary of the trade: a specific camera move, an aspect ratio chosen for a reason, a deliberate colour treatment, and a point of view you can justify.
Someone who combines prompting skill with that vocabulary does not produce clips for show. They produce material that sells.
What this means for you
If you make content: learn the six camera terms in the table above. One hour of study improves every prompt you write afterwards, permanently.
If you are a working cinematographer: your advantage is not operating the tool. It is knowing what to ask for. Years spent looking through a lens gave you the vocabulary everyone else lacks.
If you train creative teams: start with the language of cinema, not the tool's interface. Tools change every six months. The language of the lens is a century old and is not going to change.
In closing
Generation made executing a shot nearly free, and when execution becomes free, all the value moves to the decision. Knowing what to shoot is now worth more than knowing how to shoot it.
Try this today: take a prompt you wrote before and rewrite it with the template above. The difference in the result will tell you everything.
And if you want to master this on your own projects, that is what we do in the AI filmmaking course.
Common questions
- Why do my generated shots look flat?
- Because you described what is in the shot without specifying how it is photographed. Size, lens, depth of field, lighting and movement create the feeling, and without them the model produces the average of everything it has seen.
- Which words should I avoid in video prompts?
- Vague adjectives like "cinematic", "stunning" and "high quality". Replace them with technical decisions: 85mm, shallow depth of field, hard side light, locked off camera.
- What is the single most useful thing to learn first?
- Depth of field. It is the strongest tool for directing the viewer's eye and the least used in prompts: a dissolved background forces attention onto one face, while everything sharp leaves the viewer to wander.
- Where do video generation tools still fail?
- Continuity across multiple shots, performance that changes over time, and spatial logic in reverse angles. That is why isolated shots rather than complete scenes are the most successful use today.
- Is camera skill still worth having?
- More than before. When execution becomes free the value moves to the decision, and someone who spent years behind a lens owns the decision vocabulary that people who only know the tool lack.
No comments yet
Leave a comment