CIVITAI / Workflows
LTX 2.3 Multi-Image Reference OmniNFT + Relay Video Fusion Workflow
Watch the full video first if you want to understand how this LTX 2.3 multi-image reference workflow works in practice. The video shows how multiple images are fused into one video generation pipeline, why OmniNFT + Relay control matters, and how to launch the workflow online without rebuilding a complex local ComfyUI environment. This ComfyUI workflow is designed for LTX 2.3 multi-image reference video generation, using OmniNFT, Relay-style prompt control, and the distilled 1.1 model route to create more controllable image-to-video results. The main purpose of this workflow is to let creators use several reference images at the same time instead of relying on only one starting frame. This makes it much more suitable for character consistency, multi-angle visual guidance, object continuity, environment reference, and cinematic video generation. The workflow is built around the LTX 2.3 distilled 1.1 video route. It uses a Gemma3-based LTX text encoder, LTX video VAE, LTX Audio VAE, LTXVConditioning, LTX2_NAG for stronger negative guidance, and LTXVAddGuideMulti for multi-image reference control. The workflow also uses ManualSigmas, CFGGuider, RandomNoise, SamplerCustomAdvanced, LTXVConcatAVLatent, LTXVSeparateAVLatent, LTXVLatentUpsampler, tiled VAE decoding, and CreateVideo output. The key node is LTXVAddGuideMulti. This node allows multiple images to act as visual guides across the video timeline. Each guide can be assigned a frame index and strength value, so the workflow can control when a reference image becomes important and how strongly it affects the output. One image can define the main character, another can define the scene, another can provide clothing or object details, and another can guide a later-frame direction. The workflow uses a three-stage rendering structure. The first stage focuses on initial composition and motion foundation. The second stage handles latent continuation and latent-space expansion. The third stage performs high-resolution refinement after the latent upscaler. This staged approach is more stable than forcing the whole video into one single pass. Compared with ordinary image-to-video workflows, this graph gives creators stronger control over visual continuity. A normal single-image workflow may struggle with stable identity, clothing consistency, background logic, or multi-reference storytelling. This multi-image wor

Public versions
LTXV 2.3