CIVITAI / Workflows

Two-Person InfiniteTalk Native Loop Long-Duration Workflow

This ComfyUI workflow is designed for two-person InfiniteTalk native looping, dual-speaker talking video generation, and long-duration audio-driven character interaction. The main goal of this workflow is to generate a two-character dialogue video from a start frame and audio input, then continue the output through a native loop structure so creators can extend the video duration more naturally across multiple segments. Unlike a simple single-person talking-head workflow, this graph is built for two speakers. It uses two speaker regions, two audio encoder outputs, character masks, InfiniteTalk multi-speaker model patching, previous-frame continuation, and repeated video generation stages. This makes it suitable for AI dialogue scenes, two-person digital human videos, interview-style content, virtual host conversations, short drama dialogue, character interaction videos, product explanation conversations, and long-form AI video narration. The workflow is built around the Wan 2.1 InfiniteTalk multi-speaker pipeline. It uses a Wan video model route, UMT5 text encoder, Wan VAE, wav2vec2 audio encoder, InfiniteTalk multi-speaker model patch, start image input, two character masks, two audio encoder outputs, sampler control, continuation frames, and CreateVideo / SaveVideo output nodes. The central generation module is WanInfiniteTalkToVideo, which receives the model, InfiniteTalk model patch, positive and negative conditioning, VAE, audio features, start image, previous frames, speaker masks, width, height, video length, motion frame count, and audio scale. The key feature of this workflow is dual-speaker native continuation. In many AI video workflows, a two-person scene is difficult to maintain because the model may not know which person should speak, which mouth should move, or how to keep both characters stable. This workflow solves that problem by using speaker-specific masks and audio encoder outputs. Character 1 and Character 2 can each have their own mask region, allowing the model to understand where each speaking area is located. The workflow includes instructions for drawing masks in the ComfyUI MaskEditor. Users upload the start frame, open the image in MaskEditor, then draw the mask for Character 1. The same process is repeated for Character 2. These masks are important because they define the active speaker regions. Without clear masks, the mode

LTXV2 #character
在 Civitai 查看原始条目
Two-Person InfiniteTalk Native Loop Long-Duration Workflow

公开版本

v1.0

LTXV2