CIVITAI / Workflows
LTX 2.3 Single-Person Digital Human OmniNFT + Relay Audio-Driven Workflow
Watch the full video first if you want to understand how this LTX 2.3 single-person digital human workflow works in practice. The video shows how one character image and one audio file can be turned into an audio-driven talking video, how the staged rendering structure improves stability, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 single-person digital human video generation, using OmniNFT, Relay-style prompt control, and the distilled 1.1 model route to create an audio-driven talking character video from a still image. The main purpose of this workflow is to make single-character digital human production easier, more repeatable, and more suitable for real creator use. The workflow starts with one character image. The image is resized and prepared before entering the LTX video pipeline. This image becomes the main identity reference for the digital human, controlling the face, clothing, framing, visual style, and overall composition. The workflow then uses LTXVImgToVideoConditionOnly to inject the image into the video latent process, allowing the model to preserve the original subject while generating motion. The audio side is also important. The workflow includes Audio Duration detection and a SimpleMath frame calculation system. The audio length is read automatically, then converted into an LTX-compatible frame count. This helps reduce manual frame-count mistakes and keeps the generated video length closer to the input audio. The workflow also uses LTXVEmptyLatentAudio, LTXVConcatAVLatent, and LTXVSeparateAVLatent to connect audio latent logic with video latent generation. The core generation process is divided into three stages. The first stage establishes the character motion, composition, and basic video structure. The second stage performs latent-space refinement and continuation. The third stage applies high-resolution refinement after latent upscaling. This three-stage structure is more stable than a simple one-pass render because each stage has a clearer purpose: build motion, improve structure, then polish quality. The workflow also includes LTX2_NAG and a universal negative prompt structure. These are used to reduce common digital human problems such as identity drift, face distortion, broken mouth shapes, unstable lip movement, flicker, frame ji

Public versions
LTXV 2.3