CIVITAI / Workflows

Title: LTX-2 Lip Sync Workflow

Description: LTX-2 Lip Sync Workflow is a ComfyUI workflow designed for audio-driven lip sync video generation, talking character animation, and image-to-video portrait performance using LTX-2. Instead of only creating a silent motion clip from an image, this workflow brings audio into the generation process and lets the video latent and audio latent work together, making it suitable for creating short speaking videos, AI presenters, dialogue clips, digital human previews, character voice performances, and social media talking-head content. The workflow is built around the LTX-2 19B Dev FP8 checkpoint, using both the main video model and the dedicated LTX audio VAE pipeline. The audio file is encoded into an audio latent, then combined with the video latent through an audio-video latent workflow. This design allows the model to use the input audio as part of the generation condition, instead of treating the audio as something added after the video is finished. The result is a more direct audio-to-mouth-motion relationship, which is important for lip sync, speech rhythm, facial timing, and natural talking performance. The core logic of the workflow is image + audio to lip-synced video. You provide a source image as the visual identity reference and an audio file as the speech or singing reference. The image is used to initialize the character appearance and video layout, while the audio latent guides the speaking rhythm. The workflow then generates a video where the character can appear to talk along with the provided audio. A key part of this workflow is the LTXVAudioVAEEncode stage. The input audio is processed by the LTX audio VAE and converted into an audio latent. This audio latent is then passed into the later video generation stage through LTXVConcatAVLatent, where it is combined with the video latent. After sampling, LTXVSeparateAVLatent is used to separate the final video latent from the audio-video latent structure. This gives the workflow a clear audio-video pipeline: load audio, encode audio, combine audio with video latent, sample, separate video latent, then decode or upscale the final result. The workflow also uses image-to-video logic through LTXVImgToVideoInplace. This helps preserve the source image identity and composition while allowing the generated frames to move. For portrait images, this is especially useful because the face, clot

ZImageTurbo #character
在 Civitai 查看原始条目
Title:  LTX-2 Lip Sync Workflow

公开版本

v1.0

ZImageTurbo