← Blog

How to Get Consistent Characters in AI Video

September 17, 2026 · 6 min read · by the ClipCraft team

Consistent characters are the hardest ordinary thing in AI video. Any generator can give you a great five-second shot; ask for the same person in scene two and you get their taller cousin in different clothes. The cause is structural: every scene is an independent shoot, and the model has no memory of what it made for you a minute ago. The fix is equally structural, and it's the whole design of ClipCraft's Story Mode: define each character once, restate who's in every scene automatically, and optionally chain finished scenes so the next one can literally see the last.

Why characters drift between scenes

A text prompt underdetermines a face. "A cheerful cartoon banana" matches a thousand bananas, and the model picks a different one each run, so cross-scene identity has to come from somewhere stronger than adjectives: reference images, repeated character sheets, and video the model can extend. We learned the failure mode the embarrassing way. The first version of Story Mode matched characters to scenes by name, silently, and a scene that didn't literally say "BANANA" got no reference at all, which produced the proud moment of a banana protagonist turning into a strawberry mid-story. Version two replaced that guesswork with visible, tappable controls.

Build a cast once

Story Mode's cast section takes up to six characters, each with a named headshot, a "looks like" line, and a "sounds like" line, plus an optional voice sample (experimental, and real: it rides to the model as reference audio, with the three biggest talkers in a scene taking the slots). Headshots travel with every scene the character appears in as reference images, which is why Story Mode runs on the Seedance 2.0 family, currently the line that accepts multiple character references per shot. The character sheet is prepended to each scene's prompt before anything else, so identity survives even when a long scene description crowds the prompt limit.

Story Mode with the cast section, a three-scene bracket script, storyboard cards with per-scene costs and the continuity toggle
The whole flow on one screen: cast slots, the bracket script, a storyboard with per-scene costs, and the continuity toggle with its honest price note.

Tag who's in each scene

The storyboard preview builds one card per scene with tap-to-toggle character chips: auto-detected from the script when a name appears in the heading, a bracket note or a spoken line, and overridable when the detection guesses wrong. A scene with a cast defined but nobody selected gets an amber warning chip instead of a silent nothing. The UI's own instruction is the best one-line summary of the whole system: every scene a character appears in restates who they are, so scene four's banana is scene one's banana.

Scene continuity, and what it honestly costs

The strongest consistency tool is the 🔗 continuity toggle: each finished scene rides into the next as a reference video, so style, characters and voices carry over the way a TV episode carries its own look. It's off by default for a reason the UI states plainly: chained scenes bill the video-input rate, about 20 percent more per second (85 tokens a second instead of 70 at 480p on the default model), because feeding a video into a generation is priced higher by the provider. Scene one always bills the normal rate, and continuity can't be combined with a fixed end frame. Turn it on for the final render, not the drafts.

The script format, and the ChatGPT shortcut

Scripts use a bracket format that reads like a shooting script: a scene heading with direction and duration, NAME: "dialogue" lines that lip-sync, and bracketed camera notes. A Load-example button drops in a three-scene treehouse story, and on the default draft settings (Seedance 2.0 Mini at 480p) that example prices at 350 tokens a scene, 1,050 for the 15-second story, shown on the generate button before you commit. If writing scripts isn't your thing, the copy-prompt-for-ChatGPT button exports an instruction block, with your character names baked in, that tells the AI to name characters in every scene and re-describe props and locations each time, because scenes are independent shoots.

Stills drift too: the image-side version

The same principle covers images. In the AI Image Studio, attach a reference image of your character with the 📎 button and prompt the new scene; the models that accept references keep the face recognizable across generations. A neat loop for video people: pull a clean frame of your character with the frame extractor workflow and use it as the reference headshot for the next batch of scenes.

The AI Image Studio with a prompt box, an add-image reference button, size presets and the FLUX Schnell model selected
The Image Studio: the 📎 Add image button is where a character reference rides along with your prompt.

What to know before you spend tokens

Draft at the cheap default (Mini, 480p) until the story works, then re-run the keeper scenes at higher settings. Consistency comes from the references, not the resolution, so the cheap drafts tell you everything except how sharp it ends up.

The bigger picture: model shootouts like our AI video generator roundup tell you who renders best, but a character that survives eight scenes comes from workflow, and that's a thing you can set up in an afternoon. Start with a free account to explore, and the editing guide covers what happens after the scenes exist.

Give your story a cast that stays cast

Headshots, character sheets, per-scene tagging and continuity chaining, with per-scene prices shown before you spend. Scripted in brackets, stitched in your browser.

Open Story Mode