CIVITAI / Workflows
LTX 2.3 Distill 1.1 + VBVR 240K | High-Probability Digital Human Workflow
V This workflow is designed for high-probability LTX 2.3 digital human generation, built around LTX 2.3 Distill 1.1 and VBVR 240K enhancement. Its main purpose is to create a more reliable image-to-video talking-person or digital-avatar result, where the character keeps a stable identity, controlled facial motion, cleaner body movement, and stronger final visual quality. The workflow uses an LTX 2.3 video generation structure with ltx-2.3-22b-distilled-1.1, distilled LoRA support, VBVR-style image-to-video enhancement, Gemma-based LTX text encoding, LTX video VAE, LTX audio VAE routing, NAG enhancement, IC LoRA motion-track control, spatial latent upscaling, custom sampling, tiled decoding, and final video export. This makes it more production-oriented than a simple first-frame animation workflow. The core advantage of this setup is probability and stability. Digital human generation is not only about making a still image move. The model must preserve the face, avoid identity drift, maintain the original clothing and composition, keep the speaking performance believable, and avoid random body motion. This workflow is designed to improve those weak points by using the input image as a strong visual anchor, then reinforcing the generation through LTXVImgToVideoConditionOnly, LTXVPreprocess, audio/video latent routing, and multiple refinement stages. The workflow includes a dedicated audio path. Audio can be encoded through LTXVAudioVAEEncode and connected into the video latent process, allowing the output to behave more like a digital human video rather than a silent image animation. This is useful for AI presenters, talking avatars, product explanation videos, character narration, short drama dialogue, virtual influencer clips, and commercial-style social media content. NAG enhancement is another important part of the workflow. It helps strengthen generation control and reduce unwanted drift during sampling. For digital human videos, this is especially useful because even small errors in the face, mouth, hands, or camera motion can make the result feel unstable. The workflow also uses a motion-track control LoRA to guide the movement more deliberately, helping the character perform with more controlled motion instead of random animation. The pipeline is also staged for better final quality. It first builds the base video from the input image and conditi

公开版本
LTXV 2.3