OCDevel AI Video Generation Podcast

Character Consistency: Sheets, References, and When Multi-Reference Beats a LoRA

This tutorial climbs the next rung: keeping one character looking like the same person across multiple shots. We start with why drift happens (a video generator is stateless, so it re-derives a plausible face on every call) and the "second-clip identity drift" wall, [documented as almost never random](https://www.runcomfy.com/trainer/ai-toolkit/wan-2-2-i2v-character-consistency-lora). Then the four anchors, weakest to strongest: - **Character sheets** built in an image model: turnaround, expression set, neutral lighting, plain background, one full-height shot ([Higgsfield Soul ID guide](https://scribehow.com/page/Higgsfield_Soul_ID_The_Best_Tool_for_AI_Character_Consistency_in_2026__i1nfbuF-TcalH-r-LeNQgg)). Tools include [Nano Banana Pro](https://wavespeed.ai/blog/posts/google-nano-banana-pro-complete-guide-2026/), [FLUX.2 Pro](https://selfielab.me/blog/flux-2-pro-multi-reference-character-sheets-guide-20260307), Seedream 4.5, and [Ideogram Character](https://blog.fal.ai/introducing-ideogram-character/). - **Single reference / start frame** (image-to-video), plus no-training adapters PuLID and IPAdapter ([LoRA vs references](https://zsky.ai/blog/lora-training-guide)). [Runway Gen-4 reportedly hits 95%+ from one reference](https://selfielab.me/blog/runway-gen-4-character-consistency-guide-20260215). - **Native multi-reference**, the episode's thesis: [Runway Gen-4 References](https://replicate.com/runwayml/gen4-image), [Veo 3.1 Ingredients to Video](https://blog.google/innovation-and-ai/products/veo-updates-flow/), [Kling Elements](https://kling.ai/blog/kling-ai-3-0-multi-reference-inpainting-guide), [Seedance 2.0 Omni Reference](https://vicsee.com/blog/seedance-2-0-omni-reference), and [Midjourney Omni-Reference](https://selfielab.me/blog/midjourney-v7-omni-reference-character-mastery-20260215). - **Trained character LoRA** on [fal.ai](https://fal.ai/models/fal-ai/flux-lora-fast-training) or [Replicate](https://replicate.com/replicate/fast-flux-trainer/train): roughly fifteen to thirty varied images, two to five dollars a run, base-model lock-in. Decision rule: default to multi-reference; train a LoRA only for high-volume, exact-lock, stable-base work. Plus pitfalls (outfit drift, lighting, identity bleed, reference quality), provenance ([SynthID and C2PA](https://blog.google/innovation-and-ai/products/identifying-ai-generated-media-online/), the [EU AI Act and SB 942](https://www.eyesift.com/ai-image-detection-2026-c2pa-content-credentials-synthid-watermarks-diffusion-fingerprints-deepfake/)), real-person rights ([NO FAKES Act](https://en.wikipedia.org/wiki/No_Fakes_Act)), and benching it yourself on the [Video Arena leaderboard](https://huggingface.co/spaces/ArtificialAnalysis/Video-Generation-Arena-Leaderboard).