CIVITAI / Workflows
SCAIL-2 Two-Person Reference Editing Long-Video Workflow
Watch the full video first if you want to understand how this SCAIL-2 two-person reference editing workflow works in practice. The video shows how two reference characters can replace two selected people in a driving video, while keeping left-right identity alignment, synchronized motion, original scene structure, lighting coherence, and long-video continuity more stable. This ComfyUI workflow is designed for SCAIL-2 two-person biological reference editing. Its main purpose is to take a two-person reference image and use it to replace both selected people in a two-person driving video. Unlike a simple two-person pose-driving workflow, this version is explicitly configured for character replacement. The reference image provides the two replacement identities, while the driving video provides the motion, timing, body interaction, camera rhythm, and original scene context. The workflow is built around wan2.1_14B_SCAIL_2_fp8_scaled.safetensors as the main SCAIL-2 model. It also uses WAN VAE, UMT5 WAN text encoding, CLIP Vision, SAM3 tracking, SCAIL2ColoredMask, WanSCAILToVideo, SamplerCustom, VAEDecode, ForLoop continuation, overlap-frame trimming, ColorTransfer, final video combining, and original audio restoration. A multi-LoRA chain is preserved to improve motion quality, character stability, and final visual consistency. The most important switch in this workflow is replacement_mode=true. This tells the SCAIL route to perform two-person skeleton guidance with reference character replacement. The positive prompt focuses on replacing both selected people with two reference characters, following two-person pose guidance, keeping left and right identity alignment, preserving the original scene structure, maintaining natural synchronized motion, coherent lighting, and smooth temporal consistency. The negative prompt is also built for two-person failure cases. It suppresses bad video quality, flicker, only one person being replaced, missing second person, wrong identity order, identity swap, identity drift, deformed bodies, distorted faces, extra limbs, missing hands, blur, and low-quality output. These problems are especially common in two-person editing because the model has to preserve both bodies, both identities, and their interaction at the same time. The workflow uses strict 512×896 alignment. Both the reference image and the driving video are resized

公开版本
Wan Video 2.2 T2V-A14B