CIVITAI / Workflows

For-Loop LTX 2.3 Long MV Auto Generation Workflow

This ComfyUI workflow is designed for LTX 2.3 long MV generation, audio-driven video creation, and for-loop style automatic music video production. Unlike a simple image-to-video workflow that only generates one short clip, this workflow focuses on longer video output by combining audio duration detection, automatic frame calculation, image-to-video conditioning, audio-video latent processing, latent upscaling, multi-stage sampling, and final video assembly. The workflow is built around LTX 2.3, using ltx-2.3-22b-dev as the main video model, Gemma 3 12B as the text encoder, LTX 2.3 spatial upscaler for latent enhancement, and motion/control LoRA support for stronger video consistency. It can take image input, audio input, prompts, and automatically calculate the number of frames needed for the video. The frame logic follows the LTX-compatible 8n+1 rule, helping users avoid frame-count errors when matching video duration to music or narration. A key part of this workflow is the automatic duration system. The audio duration is read, converted into frame length, and aligned with the required LTX frame structure. This makes the workflow more practical for MV production because users do not need to manually calculate every segment. The workflow also uses LTXVConditioning, LTXVImgToVideoConditionOnly, LTXVConcatAVLatent, and LTXVSeparateAVLatent to connect image guidance, audio-video latent logic, and video generation. The workflow is structured for long-form generation. Instead of forcing the whole MV into one single heavy render, it uses a staged process. The first stage creates the base motion and visual direction. Later stages can continue, refine, upscale, and improve the latent video result. This makes it easier to build longer music videos, character MVs, digital idol clips, cinematic visual loops, and stylized AI video segments. The workflow also includes LTXVLatentUpsampler for higher-quality output. This allows the video to be generated more efficiently at a manageable stage first, then enhanced later through latent upscaling and additional refinement. This is useful for balancing speed, quality, and VRAM usage. Final output is handled through VHS_VideoCombine, which combines the generated frames with the audio track into a finished MP4 video. This makes the workflow suitable for actual publishing, not just frame preview. It can be used for YouTube,

LTXV 2.3 #character
View original on Civitai
For-Loop LTX 2.3 Long MV Auto Generation Workflow

Public versions

v1.0

LTXV 2.3