CIVITAI / Workflows
LTX 2.3 Image and Text Video 10S Similarity Preservation Workflow
Watch the full video first if you want to understand how this LTX 2.3 image-and-text video workflow works in practice. The video shows how one reference image can be combined with text control, how the 10-second similarity system keeps the subject stable, and how to run the full workflow online without rebuilding a complex local ComfyUI environment. This ComfyUI workflow is designed for LTX 2.3 image-reference video generation with text-controlled motion and 10-second likeness preservation. Its main purpose is to let creators start from one image, describe the desired action or camera movement with text, and generate a controlled video while keeping the original subject, composition, and visual identity more stable across the clip. The workflow is built around the LTX 2.3 distilled 1.1 generation route. It uses the LTX 2.3 video checkpoint, Gemma3 fp8 text encoder, LTX Audio VAE, LTXVConditioning, LTXVImgToVideoConditionOnly, LTXVPreprocess, Image_Resize_longsize, LTX2_NAG, ManualSigmas, CFGGuider, SamplerCustomAdvanced, LTXVLatentUpsampler, LTXVConcatAVLatent, LTXVSeparateAVLatent, tiled decoding, and final video output. This makes the workflow more structured than a basic one-pass image-to-video graph. The image side provides the visual anchor. The reference image is resized, prepared through LTXVPreprocess, and injected into the generation process through LTXVImgToVideoConditionOnly. This helps the model preserve the character, object, scene, lighting, clothing, and composition from the original image. The text prompt then controls the motion direction, expression, camera movement, atmosphere, and cinematic behavior. The key update is the 10-second similarity preservation system. The workflow uses similarity and anchor-style guidance during the later stages, especially around the latent upscaling and HD refinement process. This helps reduce common image-to-video issues such as face drift, hairstyle changes, clothing inconsistency, subject deformation, background collapse, and unwanted identity changes. For creators making character videos, this is one of the most important improvements. The generation process is divided into three stages. The first stage builds the initial composition and motion base. The second stage performs latent-space upscaling while keeping stronger similarity control and weak anchor stability. The third stage applies final hig
公开版本
LTXV 2.3