CIVITAI / Workflows

LongCat Avatar Multi-Image Shot-Switching Digital Human Workflow

Watch the full video first if you want to understand how this LongCat Avatar workflow works in practice. The video shows how multiple reference images can be organized into a talking-avatar pipeline, how shot switching and loop extension are handled, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LongCat Avatar multi-image shot-switching digital human generation. Its main purpose is to turn several reference images and one driving audio track into a longer talking-avatar video with controllable visual changes. Instead of using only one image for a single static talking head, this workflow builds a reference image pool and prompt pool so the avatar can switch between different prepared visual states while keeping the talking rhythm and audio-driven mouth movement. The workflow is built around LongCat-Avatar-15_bf16.safetensors as the main avatar model, LongCat-Avatar DMD LoRA as the distilled acceleration layer, WanVideoWrapper generation nodes, WanVideo VAE, Whisper large v3 encoder, LongCat Avatar embed extension, and a segmented sampling / looping structure. The audio is loaded first and passed through Whisper, which extracts speech features for mouth movement and speaking behavior. This makes the workflow suitable for audio-driven digital human videos rather than ordinary silent image-to-video animation. The image side is organized as a multi-reference system. The workflow includes up to ten image input groups. Each image is resized to a unified 1280×720 canvas, then encoded into a LongCat-compatible latent. These images can represent different characters, outfits, backgrounds, camera angles, or visual states. The workflow also includes an image index switch, allowing the user to select which reference image enters the generation path. The prompt side is also modular. The graph contains GPT-5-based reverse prompt generation nodes and a prompt pool. Each image can have a corresponding LongCat talking-avatar prompt describing identity, appearance, scene relationship, camera framing, speaking behavior, lip-sync, subtle head movement, natural facial expression, hand gestures, and continuity locks. This makes the workflow more practical than manually writing every avatar prompt from scratch. The generation structure uses a first-stage render plus loop extension. The firs

Wan Video 2.2 T2V-A14B #character
在 Civitai 查看原始条目
LongCat Avatar Multi-Image Shot-Switching Digital Human Workflow

公开版本

v1.0

Wan Video 2.2 T2V-A14B