Lesson 6 of 22
Why pulling the camera back put four faces on his shirt
Every pixel you add to the frame is a pixel the model gets to invent.
The first thing everyone gets wrong is framing, and they get it wrong in a predictable direction: too tight. Image models default to a face-filling crop. If your prompt doesn't state a distance, you get a headshot — cropped forehead, no headroom, no room behind the person. It reads as a video call, not a TikTok.
So you pull back. And that's where it gets expensive, because widening the frame doesn't just show more room — it hands the model more canvas to invent on.
What actually happened
We loosened a talking-head prompt from a tight crop to waist-up with the room visible. Two things appeared that nobody asked for. The actor's plain band tee came back with four printed faces across the chest — at a glance, in a three-second scroll, that reads as four extra people. And a single held hairbrush became two mirrored brushes, one in each hand.
Neither is a prompt-adherence failure. The model did what we asked — it filled the frame. We'd given it more frame and no rules about what was allowed in it.
The law
Wider framing = more hallucination canvas. Every framing loosen needs matching content bans in the same edit. Widen without banning and you're buying lottery tickets.
The bans that fixed it
- Only one person in frame — and say that this includes posters, prints, photographs and reflections. The model counts a printed face as a face.
- Plain clothing: no graphics, no text, no faces on any garment. This one clause kills the most common version.
- At most one held object. Hands are where duplication happens; naming the count stops the mirroring.
How to check it in ten seconds
Pull the first and last frame of every render and put them side by side. Then read the frame for content you didn't ask for: garment prints, reflections, duplicated hands, anything with a face on it. Both checks take less time than re-rolling a bad clip.
Takeaway
Every framing loosen needs matching content bans, written in the same edit. Otherwise you're gambling.
Questions this raises
Why does AI video default to such a tight crop?
Training data skews heavily toward portrait and headshot compositions, so with no distance specified the model regresses to the mean. The fix is an explicit distance in every prompt path, not a one-off on the shots where you notice it.
Should I just generate wide and crop in post?
It helps framing but not hallucination — the invented content is still in the frame you keep, and cropping costs resolution. Ban the content at generation time and you get both.
Course outline
Progress
0 of 22 lessons completed
Chapter 1
Start here
Chapter 2
The script
Chapter 3
The frame
Chapter 4
The clip
Chapter 5
The edit
Chapter 6
Judging it honestly