CIVITAI / Workflows

LTX 2.3 Three-Image Reference Video Workflow

This workflow is designed for LTX 2.3 three-image reference video generation, giving creators a controlled way to turn multiple visual references into a coherent cinematic video. Instead of relying on only one source image, this workflow uses three separate reference images to guide the final result, making it more practical for character consistency, product presentation, scene control, and short-form AI video production. The core idea is multi-reference visual anchoring. A single image often cannot provide enough information for a stable video. It may show the character clearly, but not the product. It may show the lighting, but not the intended camera angle. It may show the scene, but not the subject identity. By using three reference images, this workflow gives the model more visual context. One image can define the main character or subject, the second image can define the product, object, clothing, or key design element, and the third image can provide the background, mood, color palette, or scene atmosphere. This makes the workflow especially useful for commercial-style video generation. For example, creators can use it to build AI influencer clips, beauty product showcases, fashion previews, character-driven advertisements, cinematic product reveals, short social media videos, and Civitai / RunningHub demonstration assets. The prompt can then act as the director, telling the model how the three references should be combined and how the action should develop over time. The workflow is based on an LTX 2.3 video generation route, using image reference guidance, prompt conditioning, video latent creation, sampling, decoding, and final video export. In a typical use case, the reference images are resized and prepared before being passed into the video generation stage. The model then uses these images as guide signals while following the written prompt. This helps the final output stay closer to the intended visual design instead of drifting into a random text-only result. The strength of this workflow is not just reference fusion, but reference fusion for video. In image generation, a reference mismatch may only affect one frame. In video generation, that mismatch can become flicker, identity drift, unstable clothing, object deformation, or inconsistent backgrounds. By giving the workflow three visual anchors, creators can improve the chance that the

LTXV 2.3 #character
在 Civitai 查看原始条目
LTX 2.3 Three-Image Reference Video Workflow

公开版本

v1.0

LTXV 2.3