Generating content with two consistent characters at once is harder than a single-subject workflow — both identities need to hold steady, and the interaction between them needs to look physically coherent, especially in video.
Generate freeKeep a separate reference photo for each character and be explicit in the prompt about which description applies to which — this reduces the chance the model blends or confuses the two identities.
Two bodies in close proximity is where video models break down most often — limbs merging, hands passing through skin, contact points drifting as the clip advances. Templates specifically tuned for two-person scenes hold contact points and proportions coherent through motion, which a generic single-subject prompt usually doesn't.
Two-character scenes have more that can go wrong than single-subject ones, so starting from an existing template with a proven prompt structure for that specific interaction is more reliable than writing the whole scene from scratch.
This usually happens when the prompt doesn't clearly separate which description applies to which character, or when there's no reference photo for each. Being explicit about both, separately, fixes most of this.
Contact points between two bodies (limbs, hands) are where generic video prompts most often glitch or merge as the clip progresses — templates tuned specifically for two-person scenes handle this more reliably.
Yes — every character in the generation must be a fictional, AI-generated adult, always shown as an adult. Real, identifiable people or anyone under 18 are banned by the content rules.











