How to Get Consistent Characters in AI Video
September 17, 2026 · 6 min read · by the ClipCraft team
Consistent characters are the hardest ordinary thing in AI video. Any generator can give you a great five-second shot; ask for the same person in scene two and you get their taller cousin in different clothes. The cause is structural: every scene is an independent shoot, and the model has no memory of what it made for you a minute ago. The fix is equally structural, and it's the whole design of ClipCraft's Story Mode: define each character once, restate who's in every scene automatically, and optionally chain finished scenes so the next one can literally see the last.
Why characters drift between scenes
A text prompt underdetermines a face. "A cheerful cartoon banana" matches a thousand bananas, and the model picks a different one each run, so cross-scene identity has to come from somewhere stronger than adjectives: reference images, repeated character sheets, and video the model can extend. We learned the failure mode the embarrassing way. The first version of Story Mode matched characters to scenes by name, silently, and a scene that didn't literally say "BANANA" got no reference at all, which produced the proud moment of a banana protagonist turning into a strawberry mid-story. Version two replaced that guesswork with visible, tappable controls.
Build a cast once
Story Mode's cast section takes up to six characters, each with a named headshot, a "looks like" line, and a "sounds like" line, plus an optional voice sample (experimental, and real: it rides to the model as reference audio, with the three biggest talkers in a scene taking the slots). Headshots travel with every scene the character appears in as reference images, which is why Story Mode runs on the Seedance 2.0 family, currently the line that accepts multiple character references per shot. The character sheet is prepended to each scene's prompt before anything else, so identity survives even when a long scene description crowds the prompt limit.

Tag who's in each scene
The storyboard preview builds one card per scene with tap-to-toggle character chips: auto-detected from the script when a name appears in the heading, a bracket note or a spoken line, and overridable when the detection guesses wrong. A scene with a cast defined but nobody selected gets an amber warning chip instead of a silent nothing. The UI's own instruction is the best one-line summary of the whole system: every scene a character appears in restates who they are, so scene four's banana is scene one's banana.
Scene continuity, and what it honestly costs
The strongest consistency tool is the 🔗 continuity toggle: each finished scene rides into the next as a reference video, so style, characters and voices carry over the way a TV episode carries its own look. It's off by default for a reason the UI states plainly: chained scenes bill the video-input rate, about 20 percent more per second (85 tokens a second instead of 70 at 480p on the default model), because feeding a video into a generation is priced higher by the provider. Scene one always bills the normal rate, and continuity can't be combined with a fixed end frame. Turn it on for the final render, not the drafts.
The script format, and the ChatGPT shortcut
Scripts use a bracket format that reads like a shooting script: a scene heading with direction and duration, NAME: "dialogue" lines that lip-sync, and bracketed camera notes. A Load-example button drops in a three-scene treehouse story, and on the default draft settings (Seedance 2.0 Mini at 480p) that example prices at 350 tokens a scene, 1,050 for the 15-second story, shown on the generate button before you commit. If writing scripts isn't your thing, the copy-prompt-for-ChatGPT button exports an instruction block, with your character names baked in, that tells the AI to name characters in every scene and re-describe props and locations each time, because scenes are independent shoots.
Stills drift too: the image-side version
The same principle covers images. In the AI Image Studio, attach a reference image of your character with the 📎 button and prompt the new scene; the models that accept references keep the face recognizable across generations. A neat loop for video people: pull a clean frame of your character with the frame extractor workflow and use it as the reference headshot for the next batch of scenes.

What to know before you spend tokens
- The image and video generators need the Creator or Studio plan; galleries stay viewable on any plan.
- Voice samples are labeled experimental and mean it: voices get close, less reliably than faces do.
- Scenes generate sequentially and stitch in your browser; if a stitch ever fails, a send-to-timeline button lays every scene end to end in VideoCraft instead, in script order.
- There's no continuous music bed across scenes yet; scenes carry their own dialogue and effects, and a soundtrack is an edit-time add.
The bigger picture: model shootouts like our AI video generator roundup tell you who renders best, but a character that survives eight scenes comes from workflow, and that's a thing you can set up in an afternoon. Start with a free account to explore, and the editing guide covers what happens after the scenes exist.
Give your story a cast that stays cast
Headshots, character sheets, per-scene tagging and continuity chaining, with per-scene prices shown before you spend. Scripted in brackets, stitched in your browser.
Open Story Mode