The single hardest unsolved problem in AI filmmaking isn't image quality — today's models render skin, light, and motion convincingly. It's continuity. Ask a generative model for the same character in eight different shots and, left unmanaged, you'll get eight subtly different people: a nose that shifts, a jacket that changes color, eyes that drift half a shade. An audience won't consciously catalog every discrepancy, but they will feel that something is off, and that feeling breaks trust in the story faster than almost any other flaw.
The Consistency Problem In AI Video
Generative video models don't have a persistent memory of a character the way a costume department or a casting choice does. Every generation is, in a real sense, starting over — reconstructing a face and a body from a text description and whatever reference material you give it. That means consistency isn't something the model does for you by default. It's something a director has to engineer, shot by shot, the same way a traditional production engineers continuity through wardrobe notes and script supervisors.
Reference Frames As The Anchor
The most reliable tool in the kit is the reference frame: one carefully generated and approved image that becomes the character's visual ground truth. Once that frame exists — face, proportions, wardrobe, palette all locked — it gets fed back into the model alongside every new prompt, so each subsequent shot is generated in reference to that identity instead of guessing at one from scratch. Treat the reference frame the way a traditional production treats a lookbook: nothing ships until it matches.
This is also where a lot of productions go wrong by rushing. Spending extra time perfecting the reference frame before generating a single sequence shot always pays for itself later — it's far cheaper to fix a face once than to fix it across twenty shots.
Locking Seeds and Style Tokens
Beyond reference frames, locking the model's seed value — the random starting point for generation — reduces unwanted variation between takes of the same shot. It's not a silver bullet on its own; a locked seed with a sloppy prompt still drifts. But paired with a consistent set of style tokens — the same lighting descriptors, the same lens language, the same color-grade vocabulary — in every prompt for a given character or scene, it meaningfully tightens the visual thread running through a sequence.
What We Learned Producing Pick Up Gerald
Pick Up Gerald, our fully AI-directed short film built on Google Veo, was as much an exercise in continuity discipline as it was in storytelling. With a single recurring lead character carrying the film across multiple locations and emotional beats, any drift in his face or wardrobe would have been immediately visible. The production leaned hard on locked reference frames for every costume and lighting state the character appears in, then cross-checked each new generated shot against that library before it was allowed into the edit.
The lesson that traveled forward into every project since: consistency isn't a generation-time problem you solve once. It's a pipeline discipline you maintain on every single shot, with a reference library that grows more authoritative as the production goes on. Treat your character's first approved frame as canon, and everything after it as a shot that has to earn its place next to that canon — not the other way around.


