Why the face drifts
Every generation is a fresh roll of the dice. Ask for “a 24-year-old woman with dark hair” twice and you get two different people, because nothing ties the second image to the first. Text alone cannot carry identity — a description narrow enough to fix a face would have to specify measurements no model reads reliably. The words that feel most descriptive to you (“striking”, “girl-next-door”, “high cheekbones”) are exactly the words that carry no geometry, so the model re-invents the geometry each time.
“Close enough” is the actual failure
The drift that kills accounts is not the obvious kind. A face that comes back visibly wrong gets deleted immediately. The dangerous one is the photo that is 90% right — same hair, same vibe, a jawline a few percent narrower — because you keep it. Do that fifty times and the fiftieth post shares no features with the first. Nobody can point at the photo where it changed, but the feed reads as several women, and a viewer scrolling a profile registers that instantly even when they cannot articulate it. Judge every keeper against the anchor file itself, never against the photo you generated an hour ago, or you are measuring drift against drift.
Rule 1 — build the anchor first
Generate 4–6 variants with no reference at all, then pick one or two where the face reads the same from different angles. Those become the anchor: the permanent identity of the character. Everything you make from then on references them. Spend real time here, because every later photo inherits whatever you settle for. At 10 tokens a photo, six anchor candidates cost 60 tokens — a tenth of the smallest monthly plan — which makes this the cheapest step in the entire pipeline and the one worth redoing three times if the face is not right. Regenerating a month of content because the anchor was mediocre is what costs money.
Rule 2 — never send more than two references
Three or more reference photos make the model average them, and the average of three faces is a fourth face. Two is the practical limit: one front, one at an angle. Past that the drift starts again, only slower and harder to notice. This is counter-intuitive — more references feel like more information, and more evidence of who this person is — but the model is not building a composite identity from them. It is blending pixels, and a blend of three faces belongs to none of the three.
Rule 3 — a fixed angle set, generated once
Front and over-the-shoulder, generated once from the anchor, cover most scenes. Reuse those files rather than regenerating the angle each time — a regenerated angle is a new roll of the dice, and it will not match the previous one exactly. Treat the angle set the way you would treat a photograph of a real person: a file you open, not a thing you re-derive. The moment you regenerate “the same three-quarter shot” instead of reusing the file, you have introduced a second anchor, and from then on half your output descends from each.
Rule 4 — keep the identity block identical
Write the identity part of the prompt once, word for word, and change only the scene around it: “Same woman as in the reference image. Keep her exact face: identical facial proportions, same eyes, nose, lips, jawline and eyebrows. Do not idealize, beautify or alter any of her features.” Rewriting that block between generations is the most common cause of a face that slowly becomes someone else. The instruction not to beautify matters more than it looks: left alone, image models drift toward a generic attractiveness average, so a character with a distinctive nose or an asymmetric smile loses exactly the features that made her recognisable.
Give every reference an explicit job
A reference cited without a stated purpose leaks. Attach a photo for the outfit and it will also push its lighting, its framing and its facial expression into the result. Write the role instead: “use @image1 for the face and hair only”, “@image2 is the outfit”. Our working rule is that the face comes from one image and one image only; later references may supply clothing or setting, but never features. And if a reference supplies the outfit, describe that outfit in the text as well — otherwise the model quietly pulls the clothing from the face photo, and you get the anchor's shirt in every scene.
Say what you want, not what you fear
Negations are for control instructions, not objects. “No subtitles”, “no watermark”, “no music” work as intended, because they switch off a whole production behaviour. But naming an object in order to exclude it — “no sunglasses”, “no jewellery” — tends to bring it into frame, because the word is in the prompt and the prompt is what the model is drawing from. Describe the positive state instead: “bare ears and neck”, “clear unobstructed eyes”. The same applies to faces: “not too made up” often returns heavy makeup, while “bare skin, no foundation, visible freckles” returns what you meant.
Consistency in video is a different problem
A clip is not one image, so the drift has somewhere new to hide: motion. Two camera moves in a single clip produce wobble and identity slippage, which is why a single deliberate move — or, for advertising, a locked-off static camera — holds a face far better than anything dynamic. When a character passes through several looks or scenes within one clip, the body position has to be pinned in the text as explicitly as the face: state that the stance and framing stay identical across every section, or the model reinterprets the pose at each transition and the person changes with it. Chaining clips works the same way as chaining photos — the first clip becomes the reference for the next, cited as ground truth for identity and look.
Change one variable at a time
When a generation misses, the reflex is to rewrite the prompt and try again. That teaches you nothing, because if the next one works you do not know which edit did it. Change one thing — the camera, the action, or the strength of the reference — and regenerate. Diagnosis costs one photo at a time, and a fault you actually understand stops recurring. This is the habit that separates people whose accounts stay consistent for months from people who rebuild their character every few weeks.
Do you need to train a model?
Training a LoRA on a set of images is the traditional answer to this problem, and it does work. It also costs hours per character, has to be redone whenever you want a meaningfully different look, and locks you into whichever base model you trained against. Reference-based generation from a disciplined anchor set reaches the consistency an audience can actually perceive, at a cost measured in single photos and a turnaround measured in seconds. Start here. If you eventually run a character at a volume where training pays for itself, the anchor set you built is exactly the training data you will need.