CIVITAI / Workflows

LTX 2.3 Four-Image Reference Audio-Driven Video Workflow

This workflow is designed for LTX 2.3 four-image reference audio-driven video generation. It combines multiple visual references with audio-aware video latent routing, making it suitable for creators who want a more controlled cinematic video instead of a random text-only result. The main purpose is to use four reference images as visual anchors, then guide the video generation with prompt structure, temporal motion planning, and audio-related conditioning so the final output feels more coherent, more rhythmic, and more production-ready. The workflow is built around the LTX 2.3 video generation system, using LTX video and audio latent components, LTX VAE decoding, Gemma-style text conditioning, custom sampler routes, manual sigma control, and final video export. Compared with a simple image-to-video workflow, this setup is more advanced because it does not depend on only one image. It allows the user to provide multiple reference images that can define different parts of the final video: character identity, product appearance, clothing or pose, background atmosphere, color tone, camera style, and visual direction. The four-image reference structure is the most important visual control layer. In practical use, Image 1 can define the main subject, Image 2 can provide the product or object, Image 3 can guide the scene or environment, and Image 4 can provide the final mood, style, or lighting reference. This gives LTX 2.3 more visual information to work with, reducing the chance of identity drift, unstable product appearance, or inconsistent scene design. For product videos, AI influencer clips, fashion showcases, beauty ads, music-video style shots, and short-form commercial content, this kind of multi-reference structure is much more useful than single-image generation. The audio-driven part makes this version different from the normal four-image reference workflow. The graph includes audio VAE routing, audio latent connection, LTXVConcatAVLatent, LTXVSeparateAVLatent, and LTXVAudioVAEDecode-style processing, allowing the video pipeline to carry audio information through the generation and export process. This makes the workflow suitable for videos where rhythm, performance, presentation timing, music atmosphere, or spoken content matters. It is not just a silent image animation pipeline; it is structured for video output with audio-aware handling. The wor

LTXV 2.3 #character
在 Civitai 查看原始条目
LTX 2.3 Four-Image Reference Audio-Driven Video Workflow

公开版本

v1.0

LTXV 2.3