CIVITAI / Workflows

LTX 2.3 Video Re-Speaking OmniNFT + Relay Lip-Sync Replacement Workflow

Watch the full video first if you want to understand how this LTX 2.3 video re-speaking workflow works in practice. The video shows how an existing talking video can be guided by a new audio track, how the lip-sync control pipeline is organized, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 video re-speaking, word replacement, and audio-driven lip-sync generation. The main purpose of this workflow is to take an existing talking video or character reference and regenerate the mouth movement so it can match a new audio track. Instead of creating a completely new character from scratch, the workflow focuses on preserving the original person, framing, motion style, and video identity while changing the spoken content. The workflow is built around the LTX 2.3 distilled 1.1 route. It uses ltx-2.3-22b-dev-dare-ties-distilled-1.1 as the main checkpoint, Gemma3 fp8 text encoding, LTX Audio VAE, LipDub IC LoRA, VBVR / I2V stabilization LoRA, LTXVAudioVAEEncode, LTXVSetAudioRefTokens, LTXAddVideoICLoRAGuide, LTXVCropGuides, ManualSigmas, CFGGuider, SamplerCustomAdvanced, LTXVLatentUpsampler, LTXVAudioVAEDecode, and final video output. The key control module is LipDub IC LoRA. This module helps the model focus on mouth movement, speaking rhythm, and audio-related facial changes. The workflow encodes the new audio into audio latent space, then uses LTXVSetAudioRefTokens to inject audio reference tokens into the conditioning. This allows the rendering stages to follow the replacement speech instead of only producing generic motion. The video reference is handled through LTXAddVideoICLoRAGuide. This node injects the original video or visual reference into the generation process, helping the output preserve the face, camera angle, clothing, background, and overall identity. The workflow uses this guide across multiple rendering stages, so the character does not drift too far while the mouth movement is updated. The generation process is divided into three stages. The first stage creates the base lip-sync motion and audio-aligned structure. The second stage inherits audio tokens from the first stage and performs latent-space refinement. The third stage uses the previous output as a stronger reference and applies final high-resolution refinement. This staged structure is import

LTXV 2.3 #character
View original on Civitai
LTX 2.3 Video Re-Speaking OmniNFT + Relay Lip-Sync Replacement Workflow

Public versions

v1.0

LTXV 2.3