CIVITAI / Workflows
LTX 2.3 Audio-Reactive Animation Workflow
Watch the full video first if you want to understand how this LTX 2.3 audio-reactive animation workflow works in practice. The video shows how one image and one audio track can be connected into a staged animation pipeline, how the video length follows the audio duration, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 audio-reactive animation generation. Its main purpose is to take a source image and an audio file, then generate a video clip whose duration and visual rhythm are organized around the audio input. Instead of creating a silent image-to-video clip and manually matching it to music later, this workflow brings audio into the generation structure from the beginning, making it more suitable for music animation, MV fragments, sound-driven visual clips, and social media video production. The workflow is built around the LTX 2.3 video generation route. It uses image reference preparation, audio duration detection, automatic frame calculation, LTXVImgToVideoConditionOnly, LTXVConditioning, CFGGuider, ManualSigmas, SamplerCustomAdvanced, LTXVLatentUpsampler, AV latent combination, tiled VAE decoding, and CreateVideo output. The graph also includes VRAM purge tools, fps control, latent size checking, image resizing, mask handling, and multi-stage refinement. The audio side is one of the most important parts of this workflow. The input audio is measured through an Audio Duration node, then converted into a frame count through a math expression. This keeps the generated video length aligned with the audio length and reduces manual calculation errors. The workflow uses 24 fps logic and LTX-friendly temporal length rules, so the video can follow a cleaner generation structure instead of using arbitrary frame counts. The image side provides the visual identity. The source image is resized and prepared, then injected into the video process through LTXVImgToVideoConditionOnly. This allows the generated animation to preserve the original character, object, scene, or visual style while still producing motion. The same image reference can be reused across later refinement stages, helping the workflow maintain continuity after latent upscaling. The generation pipeline uses a three-stage structure. The first stage builds the initial animation and base composition. The se

公开版本
LTXV 2.3