CIVITAI / Workflows
LongCat Avatar Single-Image Looping Digital Human Workflow
Watch the full video first if you want to understand how this LongCat Avatar workflow works in practice. The video shows how one reference image and one audio track can be turned into a longer talking-avatar video, how loop continuation is handled, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LongCat Avatar single-image looping digital human generation. Its main purpose is to take one character image and one driving audio file, then generate a talking-avatar video that can extend beyond a short first segment through loop-based continuation. Instead of repeatedly creating disconnected talking-head clips, this workflow builds a more continuous production route for long-form avatar narration, virtual hosting, and audio-driven digital human content. The workflow is built around LongCat-Avatar-15_bf16.safetensors as the main avatar model. It also uses LongCat-Avatar-15_dmd_distill_lora_rank128_bf16.safetensors as the DMD acceleration LoRA, WanVideoWrapper generation nodes, WanVideo VAE, WanVideo scheduler, Whisper large v3 encoder, and LongCat Avatar embed extension. The audio is processed through Whisper to extract speech features, allowing the generated avatar to follow the timing, rhythm, and mouth movement of the input voice. The visual side is driven by a single reference image. The image is resized into the target video format, encoded into a WanVideo latent, and used as the identity and appearance anchor for the avatar. This reference image controls the face, clothing, framing, lighting, background, and general visual style. Because the workflow focuses on one image, it is easier to keep the character stable than a multi-shot switching setup. The key node is WanVideoLongCatAvatarExtendEmbeds. It combines the previous latent, audio embeds, reference image latent, frame count, overlap setting, and continuation logic into a LongCat-compatible conditioning structure. The workflow uses a segmented generation design, with a 93-frame segment and 13-frame overlap. The first stage generates the initial speaking segment, then the loop continuation stage uses previous frames and audio progress to extend the video while preserving visual continuity. This is important for digital human production. A normal single-image talking-avatar workflow may work for a short clip, but it o

公开版本
Wan Video 2.2 T2V-A14B