CIVITAI / Workflows
VBVR Digital Human | Stable Talking Avatar Workflow
This workflow is designed for VBVR digital human video generation, focusing on stable talking-avatar animation from a single reference image and an audio input. Its main purpose is to help creators generate a more controlled digital human result where the character keeps the same face, framing, camera angle, clothing, and background while performing natural speaking motion. The workflow is built around an LTX 2.3 video generation pipeline with VBVR I2V LoRA enhancement, LTX audio / video latent routing, Gemma-style text encoding, LTX video VAE, LTX audio VAE, NAG enhancement, IC LoRA motion-track control, spatial latent upscaling, multi-stage sampling, tiled decoding, and final video export. Compared with a basic image-to-video workflow, this setup is more suitable for talking-avatar production because it combines visual anchoring, audio routing, and controlled motion guidance in one graph. The core idea is simple but important: keep the shot stable and make the woman speak. In digital human generation, the biggest problem is often not whether the image can move, but whether it moves too much. A weak workflow may change the face, zoom the camera out, alter the clothing, deform the mouth, create extra hands, or shift the background. This workflow is designed to reduce those problems by keeping the camera and scene steady while concentrating motion on the face, mouth, head, and subtle body performance. VBVR is used here as an image-to-video consistency and motion-control booster. It helps the model follow the source image more closely and reduces random drift during generation. This is especially important for digital human videos because the first frame usually defines the person’s identity. If the generated video loses that identity after a few seconds, the result becomes unusable for avatar content, product narration, AI presenters, or character dialogue. The workflow also includes an audio latent route. The audio is encoded through LTXVAudioVAEEncode, connected into the audio/video latent structure, and later separated and decoded for final output. This makes the workflow more than a silent animation setup. It is designed for speaking-person videos where the final result needs both visual motion and usable audio-video export. Another important part is the use of NAG and IC LoRA motion-track control. NAG helps stabilize generation guidance, while the m

Public versions
LTXV 2.3