CIVITAI / PUBLIC LIBRARY

AI-KSK Civitai 模型与工作流

集中检索 AI-KSK 在 Civitai 公开发布的模型、LoRA 与工作流,查看版本、基础模型和官方来源。

385

公开条目

23

基础模型方向

60732

累计下载

FLUX.2 Timestep Distillation Eight-Step Reference Image Workflow
Workflows 2026-06-06

FLUX.2 Timestep Distillation Eight-Step Reference Image Workflow

Watch the full video first if you want to understand how this FLUX.2 timestep distillation workflow works in practice. The video shows how FLUX.2 can run through a compact eight-step generation route, how optional reference images can be connected through latent conditioning, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for FLUX.2 timestep distillation image generation and reference-guided image editing. Its main purpose is to reduce the sampling cost of FLUX.2 while keeping the workflow clean, practical, and easy to test. Instead of using a heavy full-step FLUX.2 generation route, this workflow uses a distilled eight-step model path, making it more suitable for fast prompt testing, image editing experiments, reference-based generation, and online workflow demonstrations. The workflow is built around flux2_distilled_8step_fp8mixed.safetensors as the main diffusion model. It uses mistral_3_small_flux2_fp8.safetensors as the Flux2 text encoder, flux2-vae.safetensors as the VAE, Flux2Scheduler for step scheduling, BasicGuider for the core guidance route, KSamplerSelect with Euler sampling, RandomNoise, EmptyFlux2LatentImage, SamplerCustomAdvanced, VAEDecode, and SaveImage. The whole graph is compact, which makes the logic easier to understand than a large multi-stage image workflow. The most important idea is timestep distillation. The workflow uses an eight-step distilled FLUX.2 model, meaning the sampling process is compressed into a much smaller number of steps while still trying to preserve usable image quality, prompt following, composition, and detail. This is useful when creators need speed, repeated testing, or fast comparison between prompts and reference images. The prompt side uses a structured JSON-style editing prompt. In the uploaded workflow, the example task focuses on single-image food editing: transforming a pasta dish into a Michelin-style lobster truffle tagliatelle while preserving the bowl angle, tabletop composition, shallow depth of field, and realistic food photography base. This shows that the workflow is not only for pure text-to-image generation, but also for controlled image editing and commercial-style visual transformation. The reference image system is optional. The graph includes ReferenceLatent nodes and VAEEncode nodes. If the ReferenceLat

Wan Video 2.2 T2V-A14B 57 downloads
查看公开资料
Workflows 2026-06-03

LTX 2.3 Image and Text Video 10S Similarity Preservation Workflow

Watch the full video first if you want to understand how this LTX 2.3 image-and-text video workflow works in practice. The video shows how one reference image can be combined with text control, how the 10-second similarity system keeps the subject stable, and how to run the full workflow online without rebuilding a complex local ComfyUI environment. This ComfyUI workflow is designed for LTX 2.3 image-reference video generation with text-controlled motion and 10-second likeness preservation. Its main purpose is to let creators start from one image, describe the desired action or camera movement with text, and generate a controlled video while keeping the original subject, composition, and visual identity more stable across the clip. The workflow is built around the LTX 2.3 distilled 1.1 generation route. It uses the LTX 2.3 video checkpoint, Gemma3 fp8 text encoder, LTX Audio VAE, LTXVConditioning, LTXVImgToVideoConditionOnly, LTXVPreprocess, Image_Resize_longsize, LTX2_NAG, ManualSigmas, CFGGuider, SamplerCustomAdvanced, LTXVLatentUpsampler, LTXVConcatAVLatent, LTXVSeparateAVLatent, tiled decoding, and final video output. This makes the workflow more structured than a basic one-pass image-to-video graph. The image side provides the visual anchor. The reference image is resized, prepared through LTXVPreprocess, and injected into the generation process through LTXVImgToVideoConditionOnly. This helps the model preserve the character, object, scene, lighting, clothing, and composition from the original image. The text prompt then controls the motion direction, expression, camera movement, atmosphere, and cinematic behavior. The key update is the 10-second similarity preservation system. The workflow uses similarity and anchor-style guidance during the later stages, especially around the latent upscaling and HD refinement process. This helps reduce common image-to-video issues such as face drift, hairstyle changes, clothing inconsistency, subject deformation, background collapse, and unwanted identity changes. For creators making character videos, this is one of the most important improvements. The generation process is divided into three stages. The first stage builds the initial composition and motion base. The second stage performs latent-space upscaling while keeping stronger similarity control and weak anchor stability. The third stage applies final hig

LTXV 2.3 198 downloads
查看公开资料
Anima Edit Image Editing and Local Inpainting Workflow
Workflows 2026-05-31

Anima Edit Image Editing and Local Inpainting Workflow

Watch the full video first if you want to understand how this Anima Edit workflow works in practice. The video shows how Anima Base can be used for image editing, local inpainting, outfit change, and reference-guided transformation, while keeping the workflow online and easy to test without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for Anima image editing and local transformation. Its main purpose is to take an existing image, define the edit direction through prompt and mask logic, then regenerate the target area while preserving the original image structure as much as possible. Instead of using Anima only for pure text-to-image generation, this workflow turns Anima into a practical editing pipeline for changing clothes, adjusting appearance, modifying a scene, testing character edits, and comparing different edit LoRA routes. The workflow is built around anima_baseV10.safetensors as the main model. It uses qwen_3_06b_base.safetensors as the text encoder and qwen_image_vae.safetensors as the VAE. The input image is loaded through LoadImage, then resized to a controlled 1024×1024 working canvas. This gives the editing process a stable base resolution and avoids random size mismatch problems. A key part of the workflow is the IC-style side-by-side editing structure. AILab_ICLoRAConcat combines the original image and edit target into a left-right layout, producing both an image and a base mask. This structure helps the model understand that the right side should become the edited result while the left side provides the original reference. The workflow also uses InpaintModelConditioning to prepare positive conditioning, negative conditioning, latent image, image pixels, and mask data for local inpainting. The workflow includes Anima-LLLite inpainting v2 through AnimaLLLiteApply. This module injects lightweight inpainting guidance into the Anima model, helping the model follow the image and mask more accurately during local edits. It is especially useful when the target is not to redraw the whole image, but to change a selected area while keeping identity, pose, framing, and background logic more stable. The graph also includes two edit LoRA routes: AnimaEditV1 and lora_edit_ZeroTwo. These are loaded through LoraLoaderModelOnly and connected into separate generation branches. This makes the workflow useful for testin

Anima 576 downloads
查看公开资料
Anima HiresFix 1.5x High-Resolution Refinement Workflow
Workflows 2026-05-31

Anima HiresFix 1.5x High-Resolution Refinement Workflow

Watch the full video first if you want to understand how this Anima HiresFix workflow works in practice. The video shows how Anima Base can generate a 1024 image first, then use a 1.5x latent refinement stage to improve detail, structure, and final image quality without making the workflow overly complex. This ComfyUI workflow is designed for Anima Base 1.0 high-resolution refinement using a classic HiresFix-style two-stage generation route. Its main purpose is to create a stable 1024×1024 base image first, then upscale the latent to 1536×1536 and run a second controlled refinement pass. Compared with a simple one-pass Anima generation workflow, this structure gives the final image more detail, cleaner linework, better texture density, and stronger overall polish. The workflow is built around anima_baseV10.safetensors as the main model. After loading the model, ModelSamplingAuraFlow applies a shift value of 3, preparing the Anima model for the intended sampling behavior. The text side uses qwen_3_06b_base.safetensors as the text encoder with the stable_diffusion type setting. The output is decoded with qwen_image_vae.safetensors, which keeps the workflow aligned with the Anima / Qwen Image VAE ecosystem. The first generation stage uses an EmptyLatentImage at 1024×1024 with batch size 1. The positive prompt defines the image direction, while the negative prompt suppresses low-quality output, weak score tags, blurry artifacts, JPEG artifacts, and unwanted artist-name artifacts. The first KSampler is configured as the base sample stage, using 34 steps, CFG 4.5, er_sde sampler, simple scheduler, and denoise 1. This stage is responsible for building the main composition, character structure, pose, lighting, and global visual identity. After the base latent is generated, the workflow sends it into LatentUpscale. The upscale method is bicubic, and the target size is 1536×1536, which is exactly a 1.5x increase from the original 1024 base. This is important because it improves the working resolution without jumping too aggressively into an unstable size. A moderate 1.5x latent upscale is often easier to control than a larger upscale when the goal is refinement rather than complete regeneration. The second KSampler performs the HiresFix refinement stage. It uses 18 steps, CFG 4.5, er_sde sampler, simple scheduler, and denoise 0.35. The lower denoise value means t

Anima 412 downloads
查看公开资料
FLUX.2 Dev PiD Direct 4K Image Generation Workflow
Workflows 2026-05-29

FLUX.2 Dev PiD Direct 4K Image Generation Workflow

Watch the full video first if you want to understand how this FLUX.2 Dev + PiD workflow works in practice. The video shows how a 1024 base image can be generated with FLUX.2 Dev, how the PiD 2K-to-4K enhancement lane improves the final image, and how to launch the workflow online without building a local ComfyUI environment. This ComfyUI workflow is designed for FLUX.2 Dev direct 4K image generation using PiD as the high-resolution enhancement stage. Its main purpose is to create a strong FLUX.2 Dev base image first, capture the correct latent state during sampling, and then send that latent into a PiD refinement pipeline for a cleaner and more detailed final output. The workflow is built around flux2_dev_fp8mixed.safetensors as the main UNet model. It uses mistral_3_small_flux2_fp8.safetensors as the Flux2 text encoder, flux2-vae.safetensors as the VAE, and EmptyFlux2LatentImage for the base canvas. The base resolution is controlled through separate width and height primitive nodes, both set to 1024, giving the workflow a clean 1024×1024 starting point before the 2K-to-4K PiD pass. The prompt is handled through PiDTextPrompt. This node sends the same text into the positive prompt encoder and also provides the caption for PiD preparation. In this workflow, the prompt is written for a high-end commercial product photograph, with a luxury skincare bottle, wet black stone, cyan and gold rim lighting, crisp reflections, tiny water droplets, clean product label text, and premium advertising composition. This makes the workflow especially suitable for testing detail, typography, packaging clarity, lighting control, and product-style realism. The main FLUX.2 generation stage uses CLIPTextEncode, FluxGuidance, and PiDKSamplerCapture. FluxGuidance is set to 4, while the sampler route uses 50 steps, CFG 4, Euler sampler, simple scheduler, denoise 1.0, and capture_step 45. PiDKSamplerCapture outputs both the final native latent and a PiD latent. The native latent is decoded through the Flux2 VAE and saved as the baseline result, while the PiD latent is sent into the enhancement lane. The PiD lane is configured with the Flux2 backbone, 2kto4k checkpoint type, scale 4, auto download enabled, and cleanup after prepare enabled. PiDSample then performs the high-resolution enhancement pass with 4 PiD steps, CFG scale 1.0, fixed seed, aggressive cleanup, and sequential_b

Flux.2 D 244 downloads
查看公开资料
FLUX.1 Dev PiD Direct 4K Image Generation Workflow
Workflows 2026-05-29

FLUX.1 Dev PiD Direct 4K Image Generation Workflow

Watch the full video first if you want to understand how this FLUX.1 Dev + PiD workflow works in practice. The video shows how a 1024 base image can be generated and then pushed into a PiD-based 4K enhancement lane, how the native baseline compares with the enhanced result, and how to launch the workflow online without building a local ComfyUI setup. This ComfyUI workflow is designed for FLUX.1 Dev 4K image generation using PiD as the high-resolution enhancement stage. Its main purpose is to create a strong FLUX.1 Dev base image first, capture the correct intermediate latent state, and then send it into a PiD 2K-to-4K refinement pipeline for sharper, richer, and more detailed final output. The workflow is built around flux1-dev-fp8.safetensors as the main checkpoint. It uses a 1024×1024 base latent, CLIP text encoding, FluxGuidance, PiDTextPrompt, PiDKSamplerCapture, PiDPrepare, PiDSample, PiDFinalize, VAEDecode, and SaveImage. The graph also includes a native FLUX output lane and a PiD enhanced output lane, making it useful for direct comparison between the standard generation result and the 4K-enhanced version. The first part of the workflow is the FLUX.1 Dev base generation. The prompt is written into PiDTextPrompt and then passed into the positive prompt encoder. FluxGuidance is set to 3.5, while CFG is set to 1.0, which is important because FLUX Dev does not use the negative prompt in the same way as older CFG-heavy models. The workflow note also makes this clear: for FLUX Dev and Schnell, CFG should stay at 1.0 because negative prompting is ignored when CFG is 1.0. The sampling stage uses PiDKSamplerCapture instead of a normal KSampler. This node generates the native FLUX latent while also capturing a PiD-ready latent at the selected capture step. In this workflow, the setup uses 28 steps, Euler sampler, simple scheduler, denoise 1.0, and capture_step 26. This is a practical setting because it allows the base image to develop enough structure before the PiD stage receives the captured latent. After the base latent is captured, PiDPrepare converts it into a PiD-compatible preparation object. The workflow is configured for the FLUX backbone, 2kto4k PiD checkpoint type, scale 4, auto download enabled, and cleanup after prepare enabled. Then PiDSample runs the high-resolution PiD pass with 4 PiD steps, CFG scale 1.0, fixed seed, aggressive cleanup, an

Flux.1 D 126 downloads
查看公开资料
SD3.5 Large PiD Direct 4K Image Generation Workflow
Workflows 2026-05-29

SD3.5 Large PiD Direct 4K Image Generation Workflow

Watch the full video first if you want to understand how this SD3.5 Large + PiD workflow works in practice. The video shows how a 1024 base image can be generated with SD3.5 Large, how the PiD enhancement stage pushes it toward 4K output, and how to launch the workflow online without building a local ComfyUI environment. This ComfyUI workflow is designed for SD3.5 Large direct 4K image generation using PiD as the high-resolution enhancement stage. Its main purpose is to create a clean SD3.5 Large base image first, capture the correct latent state during sampling, and then send that latent into a PiD 2K-to-4K refinement pipeline. The result is a more detailed and more production-oriented output than a simple native image generation pass. The workflow is built around sd3.5_large_fp8_scaled.safetensors as the main checkpoint. It uses a 1024×1024 EmptySD3LatentImage as the base canvas, CLIPTextEncode for positive and negative conditioning, PiDTextPrompt for unified prompt and caption handling, PiDKSamplerCapture for SD3 latent generation and PiD latent capture, PiDPrepare for preparing the captured latent, PiDSample for the PiD refinement pass, PiDFinalize for final image conversion, VAEDecode for native baseline output, and SaveImage for both comparison outputs. The first lane is the native SD3.5 Large generation route. The prompt is written into PiDTextPrompt, then sent into the positive prompt encoder and also used as the caption for PiD preparation. The negative prompt suppresses common image problems such as deformation, blur, low resolution, bad anatomy, extra fingers, JPEG artifacts, and watermarks. This makes the workflow suitable for both artistic generation and cleaner comparison testing. The key generation node is PiDKSamplerCapture. Unlike a normal KSampler, this node produces the final native latent while also capturing a PiD-ready latent at a selected capture step. In this workflow, the sampling setup uses 25 steps, CFG 4.4, Euler sampler, sgm_uniform scheduler, denoise 1.0, and capture_step 23. The captured PiD sigma is automatically sent into PiDPrepare, which helps the PiD stage receive the correct sampling-state information instead of relying on manual sigma guessing. After capture, PiDPrepare is configured with the SD3 backbone, 2kto4k PiD checkpoint type, scale 4, auto download enabled, and cleanup after prepare enabled. PiDSample then p

Other 102 downloads
查看公开资料
LongCat Avatar Single-Image Looping Digital Human Workflow
Workflows 2026-05-28

LongCat Avatar Single-Image Looping Digital Human Workflow

Watch the full video first if you want to understand how this LongCat Avatar workflow works in practice. The video shows how one reference image and one audio track can be turned into a longer talking-avatar video, how loop continuation is handled, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LongCat Avatar single-image looping digital human generation. Its main purpose is to take one character image and one driving audio file, then generate a talking-avatar video that can extend beyond a short first segment through loop-based continuation. Instead of repeatedly creating disconnected talking-head clips, this workflow builds a more continuous production route for long-form avatar narration, virtual hosting, and audio-driven digital human content. The workflow is built around LongCat-Avatar-15_bf16.safetensors as the main avatar model. It also uses LongCat-Avatar-15_dmd_distill_lora_rank128_bf16.safetensors as the DMD acceleration LoRA, WanVideoWrapper generation nodes, WanVideo VAE, WanVideo scheduler, Whisper large v3 encoder, and LongCat Avatar embed extension. The audio is processed through Whisper to extract speech features, allowing the generated avatar to follow the timing, rhythm, and mouth movement of the input voice. The visual side is driven by a single reference image. The image is resized into the target video format, encoded into a WanVideo latent, and used as the identity and appearance anchor for the avatar. This reference image controls the face, clothing, framing, lighting, background, and general visual style. Because the workflow focuses on one image, it is easier to keep the character stable than a multi-shot switching setup. The key node is WanVideoLongCatAvatarExtendEmbeds. It combines the previous latent, audio embeds, reference image latent, frame count, overlap setting, and continuation logic into a LongCat-compatible conditioning structure. The workflow uses a segmented generation design, with a 93-frame segment and 13-frame overlap. The first stage generates the initial speaking segment, then the loop continuation stage uses previous frames and audio progress to extend the video while preserving visual continuity. This is important for digital human production. A normal single-image talking-avatar workflow may work for a short clip, but it o

Wan Video 2.2 T2V-A14B 82 downloads
查看公开资料
LongCat Avatar Multi-Image Shot-Switching Digital Human Workflow
Workflows 2026-05-28

LongCat Avatar Multi-Image Shot-Switching Digital Human Workflow

Watch the full video first if you want to understand how this LongCat Avatar workflow works in practice. The video shows how multiple reference images can be organized into a talking-avatar pipeline, how shot switching and loop extension are handled, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LongCat Avatar multi-image shot-switching digital human generation. Its main purpose is to turn several reference images and one driving audio track into a longer talking-avatar video with controllable visual changes. Instead of using only one image for a single static talking head, this workflow builds a reference image pool and prompt pool so the avatar can switch between different prepared visual states while keeping the talking rhythm and audio-driven mouth movement. The workflow is built around LongCat-Avatar-15_bf16.safetensors as the main avatar model, LongCat-Avatar DMD LoRA as the distilled acceleration layer, WanVideoWrapper generation nodes, WanVideo VAE, Whisper large v3 encoder, LongCat Avatar embed extension, and a segmented sampling / looping structure. The audio is loaded first and passed through Whisper, which extracts speech features for mouth movement and speaking behavior. This makes the workflow suitable for audio-driven digital human videos rather than ordinary silent image-to-video animation. The image side is organized as a multi-reference system. The workflow includes up to ten image input groups. Each image is resized to a unified 1280×720 canvas, then encoded into a LongCat-compatible latent. These images can represent different characters, outfits, backgrounds, camera angles, or visual states. The workflow also includes an image index switch, allowing the user to select which reference image enters the generation path. The prompt side is also modular. The graph contains GPT-5-based reverse prompt generation nodes and a prompt pool. Each image can have a corresponding LongCat talking-avatar prompt describing identity, appearance, scene relationship, camera framing, speaking behavior, lip-sync, subtle head movement, natural facial expression, hand gestures, and continuity locks. This makes the workflow more practical than manually writing every avatar prompt from scratch. The generation structure uses a first-stage render plus loop extension. The firs

Wan Video 2.2 T2V-A14B 78 downloads
查看公开资料
Stable Audio 3 Sound Asset Generation Workflow
Workflows 2026-05-26

Stable Audio 3 Sound Asset Generation Workflow

Watch the full video first if you want to understand how this Stable Audio 3 workflow works in practice. The video shows how a simple text idea can be expanded into a structured audio prompt, how different sound categories affect the result, and how to launch the workflow online without building the full ComfyUI audio environment locally. This ComfyUI workflow is designed for Stable Audio 3 sound asset generation. Its main purpose is to turn text descriptions into usable audio assets, including music tracks, instrument loops, sound effects, one-shot samples, ambience, cinematic hits, UI sounds, game audio elements, and production-ready creative sound material. Instead of only generating random audio from a short prompt, this workflow adds a category-aware prompt expansion layer so the final Stable Audio prompt becomes more precise and more suitable for the target audio type. The workflow is built around the Stable Audio 3 Medium Base route. It uses stable_audio_3_medium_base.safetensors as the main checkpoint, t5gemma_b_b_ul2.safetensors as the Stable Audio text encoder, and a separate text-generation route for intelligent prompt rewriting. The generation pipeline uses CLIPTextEncode for positive and negative conditioning, EmptyLatentAudio for defining the target audio duration, KSampler for latent audio sampling, and VAEDecodeAudio to decode the generated latent into an actual audio waveform. The most important design in this workflow is the optional reprompt system. Users can input a short idea, then decide whether to enable prompt expansion. When reprompt is enabled, the workflow uses a category-aware prompt template. The available categories include Music, Instrument, SFX, and One-shot. Each category has different prompt rules. Music prompts focus on genre, instruments, layers, rhythm, mood, BPM, and track length. Instrument prompts focus on playing technique, timbre, production texture, BPM, and loop or stem length. SFX prompts focus on sound source, material texture, spatial environment, movement, attack, decay, and duration. One-shot prompts focus on short isolated audio samples such as hits, stabs, plucks, drum sounds, impacts, or short sound design elements. This structure makes the workflow more practical than a simple text-to-audio setup. In ordinary audio generation, a vague prompt like “dark cinematic sound” may produce inconsistent results.

LTXV 2.3 198 downloads
查看公开资料
LTX 2.3 Audio-Reactive Animation Workflow
Workflows 2026-05-26

LTX 2.3 Audio-Reactive Animation Workflow

Watch the full video first if you want to understand how this LTX 2.3 audio-reactive animation workflow works in practice. The video shows how one image and one audio track can be connected into a staged animation pipeline, how the video length follows the audio duration, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 audio-reactive animation generation. Its main purpose is to take a source image and an audio file, then generate a video clip whose duration and visual rhythm are organized around the audio input. Instead of creating a silent image-to-video clip and manually matching it to music later, this workflow brings audio into the generation structure from the beginning, making it more suitable for music animation, MV fragments, sound-driven visual clips, and social media video production. The workflow is built around the LTX 2.3 video generation route. It uses image reference preparation, audio duration detection, automatic frame calculation, LTXVImgToVideoConditionOnly, LTXVConditioning, CFGGuider, ManualSigmas, SamplerCustomAdvanced, LTXVLatentUpsampler, AV latent combination, tiled VAE decoding, and CreateVideo output. The graph also includes VRAM purge tools, fps control, latent size checking, image resizing, mask handling, and multi-stage refinement. The audio side is one of the most important parts of this workflow. The input audio is measured through an Audio Duration node, then converted into a frame count through a math expression. This keeps the generated video length aligned with the audio length and reduces manual calculation errors. The workflow uses 24 fps logic and LTX-friendly temporal length rules, so the video can follow a cleaner generation structure instead of using arbitrary frame counts. The image side provides the visual identity. The source image is resized and prepared, then injected into the video process through LTXVImgToVideoConditionOnly. This allows the generated animation to preserve the original character, object, scene, or visual style while still producing motion. The same image reference can be reused across later refinement stages, helping the workflow maintain continuity after latent upscaling. The generation pipeline uses a three-stage structure. The first stage builds the initial animation and base composition. The se

LTXV 2.3 139 downloads
查看公开资料
LTX 2.3 Multi-Image Reference OmniNFT + Relay Video Fusion Workflow
Workflows 2026-05-25

LTX 2.3 Multi-Image Reference OmniNFT + Relay Video Fusion Workflow

Watch the full video first if you want to understand how this LTX 2.3 multi-image reference workflow works in practice. The video shows how multiple images are fused into one video generation pipeline, why OmniNFT + Relay control matters, and how to launch the workflow online without rebuilding a complex local ComfyUI environment. This ComfyUI workflow is designed for LTX 2.3 multi-image reference video generation, using OmniNFT, Relay-style prompt control, and the distilled 1.1 model route to create more controllable image-to-video results. The main purpose of this workflow is to let creators use several reference images at the same time instead of relying on only one starting frame. This makes it much more suitable for character consistency, multi-angle visual guidance, object continuity, environment reference, and cinematic video generation. The workflow is built around the LTX 2.3 distilled 1.1 video route. It uses a Gemma3-based LTX text encoder, LTX video VAE, LTX Audio VAE, LTXVConditioning, LTX2_NAG for stronger negative guidance, and LTXVAddGuideMulti for multi-image reference control. The workflow also uses ManualSigmas, CFGGuider, RandomNoise, SamplerCustomAdvanced, LTXVConcatAVLatent, LTXVSeparateAVLatent, LTXVLatentUpsampler, tiled VAE decoding, and CreateVideo output. The key node is LTXVAddGuideMulti. This node allows multiple images to act as visual guides across the video timeline. Each guide can be assigned a frame index and strength value, so the workflow can control when a reference image becomes important and how strongly it affects the output. One image can define the main character, another can define the scene, another can provide clothing or object details, and another can guide a later-frame direction. The workflow uses a three-stage rendering structure. The first stage focuses on initial composition and motion foundation. The second stage handles latent continuation and latent-space expansion. The third stage performs high-resolution refinement after the latent upscaler. This staged approach is more stable than forcing the whole video into one single pass. Compared with ordinary image-to-video workflows, this graph gives creators stronger control over visual continuity. A normal single-image workflow may struggle with stable identity, clothing consistency, background logic, or multi-reference storytelling. This multi-image wor

LTXV 2.3 263 downloads
查看公开资料
LTX 2.3 Video Extension OmniNFT + Relay Vertical Widening Workflow
Workflows 2026-05-25

LTX 2.3 Video Extension OmniNFT + Relay Vertical Widening Workflow

Watch the full video first if you want to understand how this LTX 2.3 video extension workflow works in practice. The video shows how a vertical video can be expanded into a wider frame, why OmniNFT + Relay guidance matters, and how to launch the workflow online without rebuilding the full ComfyUI setup locally. This ComfyUI workflow is designed for LTX 2.3 video extension, vertical-to-wide frame expansion, and guided video outpainting. The main purpose of this workflow is to take a narrow or vertical video-style input and expand it into a wider composition while keeping the original subject, motion, and visual identity as stable as possible. Instead of simply stretching the image or cropping the video, this workflow uses LTX 2.3 generation to synthesize new side areas and create a more natural wide-frame result. The workflow is built around the LTX 2.3 distilled 1.1 route. It uses ltx-2.3-22b-dev-dare-merged-distilled-1.1.safetensors as the main checkpoint, Gemma-based text encoding, LTX Audio VAE, LTXVConditioning, LTXAddVideoICLoRAGuide, LTXVCropGuides, LTXVConcatAVLatent, LTXVSeparateAVLatent, ManualSigmas, CFGGuider, SamplerCustomAdvanced, VAEDecodeTiled, and VHS_VideoCombine. The graph also includes image resizing, color correction, image switching, and image concatenation nodes, which are important for comparing and assembling the expanded output. The key idea is guided extension. The source frame or reference image is first prepared through resizing and LTXVPreprocess. Then LTXAddVideoICLoRAGuide injects the visual guide into the LTX generation process, helping the model preserve the original content while expanding beyond the initial frame boundary. LTXVCropGuides helps manage the guided area so the model can focus on the extension region instead of freely changing the whole image. The workflow also uses audio-video latent logic. Empty video latent and empty audio latent are created, then combined through LTXVConcatAVLatent before sampling. After generation, LTXVSeparateAVLatent separates the video and audio latent streams again. This makes the workflow compatible with LTX 2.3 audio-video generation structure and final video output. Compared with ordinary video resizing, this workflow does not only change the canvas size. It generates new visual content for the expanded region. Compared with basic image outpainting, it works in a video pipeline

LTXV 2.3 139 downloads
查看公开资料
LTX 2.3 Video Re-Speaking OmniNFT + Relay Lip-Sync Replacement Workflow
Workflows 2026-05-25

LTX 2.3 Video Re-Speaking OmniNFT + Relay Lip-Sync Replacement Workflow

Watch the full video first if you want to understand how this LTX 2.3 video re-speaking workflow works in practice. The video shows how an existing talking video can be guided by a new audio track, how the lip-sync control pipeline is organized, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 video re-speaking, word replacement, and audio-driven lip-sync generation. The main purpose of this workflow is to take an existing talking video or character reference and regenerate the mouth movement so it can match a new audio track. Instead of creating a completely new character from scratch, the workflow focuses on preserving the original person, framing, motion style, and video identity while changing the spoken content. The workflow is built around the LTX 2.3 distilled 1.1 route. It uses ltx-2.3-22b-dev-dare-ties-distilled-1.1 as the main checkpoint, Gemma3 fp8 text encoding, LTX Audio VAE, LipDub IC LoRA, VBVR / I2V stabilization LoRA, LTXVAudioVAEEncode, LTXVSetAudioRefTokens, LTXAddVideoICLoRAGuide, LTXVCropGuides, ManualSigmas, CFGGuider, SamplerCustomAdvanced, LTXVLatentUpsampler, LTXVAudioVAEDecode, and final video output. The key control module is LipDub IC LoRA. This module helps the model focus on mouth movement, speaking rhythm, and audio-related facial changes. The workflow encodes the new audio into audio latent space, then uses LTXVSetAudioRefTokens to inject audio reference tokens into the conditioning. This allows the rendering stages to follow the replacement speech instead of only producing generic motion. The video reference is handled through LTXAddVideoICLoRAGuide. This node injects the original video or visual reference into the generation process, helping the output preserve the face, camera angle, clothing, background, and overall identity. The workflow uses this guide across multiple rendering stages, so the character does not drift too far while the mouth movement is updated. The generation process is divided into three stages. The first stage creates the base lip-sync motion and audio-aligned structure. The second stage inherits audio tokens from the first stage and performs latent-space refinement. The third stage uses the previous output as a stronger reference and applies final high-resolution refinement. This staged structure is import

LTXV 2.3 140 downloads
查看公开资料
LTX 2.3 Text to Video OmniNFT + Relay Three-Stage No-Subtitle Workflow
Workflows 2026-05-25

LTX 2.3 Text to Video OmniNFT + Relay Three-Stage No-Subtitle Workflow

Watch the full video first if you want to understand how this LTX 2.3 text-to-video workflow works in practice. The video shows how a clean prompt can be turned into a complete video clip, how the three-stage rendering structure improves stability, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 text-to-video generation, using OmniNFT, Relay-style prompt control, and the distilled 1.1 model route to create clean video outputs from text prompts. The main purpose of this workflow is to make text-to-video generation more controllable, more stable, and more suitable for publishing, especially when users want a no-subtitle, no-watermark, no-extra-text output. The workflow is built around the LTX 2.3 distilled 1.1 generation route. It uses an LTX 2.3 checkpoint, Gemma3-based text encoding, LTXVConditioning, EmptyLTXVLatentVideo, LTXVEmptyLatentAudio, LTX2_NAG negative guidance, ManualSigmas, CFGGuider, SamplerCustomAdvanced, LTXVLatentUpsampler, VAEDecodeTiled, LTXVAudioVAEDecode, and CreateVideo. The graph also includes seed control, fps control, universal negative prompting, VRAM management, and audio-video latent handling. The core idea is to generate video from text while maintaining stronger control over structure, motion, and final image quality. The positive prompt defines the subject, action, camera movement, lighting, environment, atmosphere, and cinematic direction. The negative prompt is designed to suppress common LTX video problems, including low quality, flicker, unstable perspective, identity drift, broken anatomy, subtitles, captions, UI overlays, logos, watermarks, unreadable text, and unwanted audio artifacts. The workflow uses a three-stage rendering structure. The first stage focuses on initial composition and motion foundation. It creates the base video latent and establishes the main visual direction. The second stage performs latent-space upscaling and refinement, allowing the workflow to improve structure and detail without rebuilding the whole video from scratch. The third stage applies final high-resolution polish, using another controlled sampling pass before tiled VAE decoding and video assembly. Compared with ordinary text-to-video workflows, this graph is more production-oriented. A simple one-pass T2V workflow may be fast, but it often s

LTXV 2.3 121 downloads
查看公开资料
LTX 2.3 Image to Video OmniNFT + Relay One-Image Film Workflow
Workflows 2026-05-25

LTX 2.3 Image to Video OmniNFT + Relay One-Image Film Workflow

Watch the full video first if you want to understand how this LTX 2.3 image-to-video workflow works in practice. The video shows how one image can be turned into a complete video clip, how the staged rendering pipeline improves stability, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 image-to-video generation, using OmniNFT, Relay-style prompt control, and the distilled 1.1 model route to turn a single image into a finished video clip. The main purpose of this workflow is to make one-image video generation more stable, more controllable, and more production-ready than a basic image-to-video graph. The workflow starts from a single input image. The image is resized and prepared through Image_Resize_longsize and LTXVPreprocess, then passed into LTXVImgToVideoConditionOnly as the main visual condition. This allows the source image to guide the video identity, composition, subject placement, and overall visual style while still giving the LTX model enough freedom to generate motion. The workflow is built around the LTX 2.3 distilled 1.1 route. It uses LTX video conditioning, LTX audio-video latent logic, an LTX Audio VAE, LTX2_NAG negative guidance, a universal negative prompt, and a three-stage rendering pipeline. The graph also includes Seed Everywhere, fps control, EmptyLTXVLatentVideo, LTXVEmptyLatentAudio, LTXVConcatAVLatent, LTXVSeparateAVLatent, ManualSigmas, CFGGuider, SamplerCustomAdvanced, LTXVLatentUpsampler, VAEDecodeTiled, and final video output. The key generation structure is divided into three stages. The first stage focuses on initial composition and base motion. It uses the image condition to establish the main character or scene and generate the first stable video latent. The second stage performs latent-space expansion, reconditioning, and continuation, helping the result gain more structure and detail. The third stage performs high-resolution refinement after latent upscaling, making the final output cleaner and more suitable for publishing. Compared with ordinary image-to-video workflows, this graph is more structured. A simple one-pass workflow may create motion but often struggles with identity drift, weak detail, inconsistent lighting, or unstable composition. This version uses staged sampling, manual sigma control, image conditioning, neg

LTXV 2.3 187 downloads
查看公开资料
LTX 2.3 Single-Person Digital Human OmniNFT + Relay Audio-Driven Workflow
Workflows 2026-05-25

LTX 2.3 Single-Person Digital Human OmniNFT + Relay Audio-Driven Workflow

Watch the full video first if you want to understand how this LTX 2.3 single-person digital human workflow works in practice. The video shows how one character image and one audio file can be turned into an audio-driven talking video, how the staged rendering structure improves stability, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 single-person digital human video generation, using OmniNFT, Relay-style prompt control, and the distilled 1.1 model route to create an audio-driven talking character video from a still image. The main purpose of this workflow is to make single-character digital human production easier, more repeatable, and more suitable for real creator use. The workflow starts with one character image. The image is resized and prepared before entering the LTX video pipeline. This image becomes the main identity reference for the digital human, controlling the face, clothing, framing, visual style, and overall composition. The workflow then uses LTXVImgToVideoConditionOnly to inject the image into the video latent process, allowing the model to preserve the original subject while generating motion. The audio side is also important. The workflow includes Audio Duration detection and a SimpleMath frame calculation system. The audio length is read automatically, then converted into an LTX-compatible frame count. This helps reduce manual frame-count mistakes and keeps the generated video length closer to the input audio. The workflow also uses LTXVEmptyLatentAudio, LTXVConcatAVLatent, and LTXVSeparateAVLatent to connect audio latent logic with video latent generation. The core generation process is divided into three stages. The first stage establishes the character motion, composition, and basic video structure. The second stage performs latent-space refinement and continuation. The third stage applies high-resolution refinement after latent upscaling. This three-stage structure is more stable than a simple one-pass render because each stage has a clearer purpose: build motion, improve structure, then polish quality. The workflow also includes LTX2_NAG and a universal negative prompt structure. These are used to reduce common digital human problems such as identity drift, face distortion, broken mouth shapes, unstable lip movement, flicker, frame ji

LTXV 2.3 100 downloads
查看公开资料
Anima Base 1.0 Enhancement Node A/B Comparison Workflow
Workflows 2026-05-22

Anima Base 1.0 Enhancement Node A/B Comparison Workflow

Watch the full video first if you want to understand the workflow logic quickly.This ComfyUI workflow is designed for Anima Base 1.0 enhancement node comparison, A/B testing, and practical model behavior analysis. Instead of only showing a finished generation result, this workflow is built like a controlled experiment. Each group contains an A baseline and a B experimental branch, making it easier to judge whether a specific node actually improves quality, stability, speed, or low-step performance. The workflow uses anima_baseV10.safetensors as the base model, qwen_3_06b_base.safetensors as the Qwen Image CLIP text encoder, and qwen_image_vae.safetensors as the VAE. The testing rule is strict: within each A/B group, the model file, CLIP, VAE, positive prompt, negative prompt, seed, resolution, sampler, scheduler, steps, and CFG are kept the same. The A branch is the clean baseline. The B branch only adds the target test node. This makes the comparison much more meaningful than casual prompt testing. The workflow is divided into two major categories. The D series is designed for distilled / Turbo low-step testing, using steps=8 and CFG=1. This section is useful for testing Turbo LoRA, NAGuidance, and CFGNorm under fast generation conditions. Turbo LoRA is tested to see whether it can compensate for low-step quality loss. NAGuidance is tested to see whether negative concepts can still affect generation when CFG is very low. CFGNorm is tested to observe whether it improves color, edge stability, and composition under Turbo-style sampling. The N series is designed for non-distilled standard CFG workflows, using steps=40 and CFG=4. This section tests CFGZeroStar, RescaleCFG, and SageAttention. CFGZeroStar is useful for observing changes in classifier-free guidance behavior. RescaleCFG is tested to see whether it can reduce overexposure, over-saturation, color damage, or CFG-related image collapse. SageAttention is treated mainly as a performance and attention acceleration test rather than a style enhancement node. The correct result should stay close to the baseline while improving speed or memory behavior. This workflow is especially useful for creators who want to understand which nodes are actually worth using in Anima Base production. It helps separate real improvement from psychological bias. By comparing A and B under controlled conditions, users can de

Anima 122 downloads
查看公开资料
For-Loop LTX 2.3 Long MV Auto Generation Workflow
Workflows 2026-05-18

For-Loop LTX 2.3 Long MV Auto Generation Workflow

This ComfyUI workflow is designed for LTX 2.3 long MV generation, audio-driven video creation, and for-loop style automatic music video production. Unlike a simple image-to-video workflow that only generates one short clip, this workflow focuses on longer video output by combining audio duration detection, automatic frame calculation, image-to-video conditioning, audio-video latent processing, latent upscaling, multi-stage sampling, and final video assembly. The workflow is built around LTX 2.3, using ltx-2.3-22b-dev as the main video model, Gemma 3 12B as the text encoder, LTX 2.3 spatial upscaler for latent enhancement, and motion/control LoRA support for stronger video consistency. It can take image input, audio input, prompts, and automatically calculate the number of frames needed for the video. The frame logic follows the LTX-compatible 8n+1 rule, helping users avoid frame-count errors when matching video duration to music or narration. A key part of this workflow is the automatic duration system. The audio duration is read, converted into frame length, and aligned with the required LTX frame structure. This makes the workflow more practical for MV production because users do not need to manually calculate every segment. The workflow also uses LTXVConditioning, LTXVImgToVideoConditionOnly, LTXVConcatAVLatent, and LTXVSeparateAVLatent to connect image guidance, audio-video latent logic, and video generation. The workflow is structured for long-form generation. Instead of forcing the whole MV into one single heavy render, it uses a staged process. The first stage creates the base motion and visual direction. Later stages can continue, refine, upscale, and improve the latent video result. This makes it easier to build longer music videos, character MVs, digital idol clips, cinematic visual loops, and stylized AI video segments. The workflow also includes LTXVLatentUpsampler for higher-quality output. This allows the video to be generated more efficiently at a manageable stage first, then enhanced later through latent upscaling and additional refinement. This is useful for balancing speed, quality, and VRAM usage. Final output is handled through VHS_VideoCombine, which combines the generated frames with the audio track into a finished MP4 video. This makes the workflow suitable for actual publishing, not just frame preview. It can be used for YouTube,

LTXV 2.3 166 downloads
查看公开资料
Anima Base Image-to-Image + ControlNet Face & Hand Repair Workflow
Workflows 2026-05-17

Anima Base Image-to-Image + ControlNet Face & Hand Repair Workflow

This workflow is designed for Anima Base image-to-image generation with ControlNet-style structure guidance, face refinement, and hand repair. Its main purpose is to take an existing image as the visual reference, regenerate it through Anima Base, preserve the original composition and style direction, and then automatically improve the most error-prone areas: the face, eyes, hands, and fingers. Unlike a basic image-to-image workflow, this setup is not only a simple restyle or redraw process. It combines Anima Base generation, image scaling, latent encoding, prompt-guided reconstruction, optional ControlNet / LLLite-style guidance, face detection, SAM-assisted refinement, hand detection, hand segmentation, FaceDetailer repair, and final preview / export logic. This makes it more suitable for creators who want a complete anime image polishing pipeline instead of a one-pass redraw. The workflow uses anima_baseV10.safetensors as the main Anima Base model route, qwen_3_06b_base.safetensors as the text encoder, and qwen_image_vae.safetensors as the VAE. It also includes image_scale_pixel_v2, VAEEncode, VAEDecode, CLIPTextEncode, NAGuidance, AnimaLLLiteApply, AIO_Preprocessor, FaceDetailer, SAMLoader, UltralyticsDetectorProvider, and multiple preview nodes. The structure shows that this workflow is built for controlled regeneration plus automatic detail correction. The image-to-image section first receives the input image, scales it to a suitable working resolution, encodes it into latent space, and uses the prompt to guide the Anima Base redraw. This helps preserve the main layout while giving the model enough freedom to improve detail, color, lighting, anime rendering quality, and character styling. The positive prompt route defines the target anime key visual style, while the negative prompt suppresses common failures such as low quality, blurry faces, bad anatomy, extra fingers, malformed hands, duplicated characters, cropped bodies, text, and watermark artifacts. A key part of this workflow is the ControlNet-style guidance section. The workflow includes a depth preprocessor route and Anima LLLite application logic. This can help preserve the structural relationship of the input image, such as pose, depth, silhouette, body placement, and large composition. In image-to-image generation, this kind of control is important because a pure prompt-based redraw can

Anima 513 downloads
查看公开资料
Anima Base Text-to-Image + ControlNet Face & Hand Repair Workflow
Workflows 2026-05-17

Anima Base Text-to-Image + ControlNet Face & Hand Repair Workflow

This workflow is designed for Anima Base text-to-image generation with ControlNet-style structure guidance, face refinement, and hand repair. Its main purpose is to let creators start from a pure text prompt, generate an anime-style image through Anima Base, use structure control to improve composition stability, and then automatically refine the most failure-prone areas: the face, eyes, hands, and fingers. Unlike a basic text-to-image workflow, this setup is not only a single-pass image generator. It combines Anima Base generation, Qwen image text encoding, empty latent creation, prompt-guided sampling, NAG guidance, ControlNet / Anima LLLite-style structure control, latent upscaling, second-pass refinement, face detection, SAM-assisted facial repair, hand detection, hand segmentation, FaceDetailer repair, preview nodes, and final image output. This makes it more suitable for creators who want a complete anime image production pipeline rather than a simple prompt-to-picture graph. The workflow uses anima_baseV10.safetensors as the main Anima Base model route, qwen_3_06b_base.safetensors as the text encoder, and qwen_image_vae.safetensors as the VAE. It also includes CLIPSetLastLayer, EmptyLatentImage, CLIPTextEncode, NAGuidance, AnimaLLLiteApply, AIO_Preprocessor, ClownsharKSampler_Beta, LatentUpscaleBy, VAEDecode, FaceDetailer, SAMLoader, UltralyticsDetectorProvider, and multiple preview nodes. This structure shows that the workflow is built for text-to-image generation plus controlled refinement. The first generation stage starts from an empty latent image, meaning the image is created from text rather than from an existing input image. The positive prompt describes the target anime key visual, such as an adult anime beauty, celestial fantasy scene, long platinum hair, luminous skin, floating palace balcony, colossal sky dragon, clouds, wind, sunlight, and cinematic scale contrast. The negative prompt suppresses low quality, blurry output, JPEG artifacts, low resolution, and other weak-image problems. A key part of the workflow is the ControlNet-style guidance route. The workflow includes a reference image path, DepthAnything preprocessing, and Anima LLLite application. This allows the creator to use structure guidance while still generating from text. In practical terms, this helps the output keep stronger pose, depth, silhouette, layout, and spatial

Anima 375 downloads
查看公开资料
IC Edit HD Video Restoration & Enhancement Workflow
Workflows 2026-05-14

IC Edit HD Video Restoration & Enhancement Workflow

This workflow is designed for IC Edit-style high-definition video restoration and enhancement, built on an LTX 2.3 video upscale / repair pipeline. Its main purpose is to take an existing low-quality or compressed video, preserve the original composition and motion, and rebuild the final output with cleaner details, reduced artifacts, better texture, and a more stable high-definition look. Unlike a simple video upscaler or sharpening filter, this workflow is closer to a generative restoration pipeline. It does not only enlarge the frame or add artificial sharpness. Instead, it uses LTX 2.3, IC LoRA video upscale models, video frame extraction, prompt-guided enhancement, latent reconstruction, tiled VAE decoding, and final video recombination to improve the clip while keeping the original structure intact. This makes it useful when the goal is not to create a new video from scratch, but to repair and polish an existing result. The workflow uses ltx-2.3-22b-dev-dare-ties-distilled-1.1 as the core model route, with LTXVAudioVAELoader, CheckpointLoaderSimple, LTXAVTextEncoderLoader, LTXVConditioning, LTXAddVideoICLoRAGuide, LTXVImgToVideoConditionOnly, VAEEncodeForInpaint, SamplerCustomAdvanced, VAEDecodeTiled, VHS_LoadVideo, and VHS_VideoCombine. It also loads IC LoRA upscale models such as ltx2.3-ic-video-upscale-general and ltx2.3-video-upscale-v2, which shows that the workflow is specifically tuned for video quality restoration rather than normal image-to-video generation. A key part of this workflow is the positive enhancement prompt. The workflow asks the model to enhance the input video to clean high-definition quality, remove compression artifacts, noise, blur, jagged edges, color blocks, motion smearing, local dirt, and frame flickering, while reconstructing clearer facial details, skin texture, hair strands, clothing fabric, and background structure. At the same time, it explicitly preserves the original composition, character identity, motion rhythm, camera movement, lighting atmosphere, and color style. This is the correct direction for repair work: improve quality, but do not rewrite the video. The negative prompt is also practical. It suppresses blur, oversaturation, pixelation, low resolution, grain, distortion, noise, compression artifacts, JPEG artifacts, glitches, watermark, text, logo, signature, copyright marks, subtitles, distorted sound

LTXV 2.3 116 downloads
查看公开资料
LTX 2.3 Dual Digital Human | IC Edit No-Subtitle Dialogue Workflow
Workflows 2026-05-14

LTX 2.3 Dual Digital Human | IC Edit No-Subtitle Dialogue Workflow

This workflow is designed for LTX 2.3 dual-person digital human dialogue generation, with IC Edit-style control and a strong focus on clean subtitle-free output. Its main purpose is to take a two-person reference image or character scene, generate a controlled dialogue-style video, and keep the final result clean without unwanted subtitles, fake captions, random text, watermark-like marks, overlays, or UI-style artifacts appearing on the screen. Compared with a single-person digital human workflow, this setup is more demanding because it needs to maintain two character identities at the same time. A good dual-person dialogue video must preserve left-right placement, facial consistency, clothing, body proportion, camera framing, background stability, and interaction logic. If the workflow is not controlled well, the two characters may swap positions, merge faces, duplicate body parts, drift away from the original image, or create random mouth movement that does not match the intended dialogue structure. The workflow uses LTX 2.3 as the main video generation backbone, with LTX video VAE, LTX audio VAE, image resizing, image-to-video conditioning, LTXVConditioning, LTXVImgToVideoConditionOnly, LTXVConcatAVLatent, LTXVSeparateAVLatent, SamplerCustomAdvanced, ManualSigmas, latent upscaling, tiled VAE decoding, audio decoding, CreateVideo, SaveVideo, and VRAM cleanup logic. This makes it a more complete production workflow rather than a simple one-pass image animation graph. The core generation design follows a staged rendering structure. The first stage builds the base motion, character presence, camera structure, and dialogue performance from the reference image and prompt conditioning. Later stages continue from the generated latent result with lower sigma values, refining motion stability, facial detail, clothing texture, background consistency, and final visual quality. This staged approach is especially useful for two-person digital human scenes because both subjects need to remain coherent across the full video. A key feature of this workflow is its dual-character control direction. The workflow is built for restrained dialogue performance rather than chaotic motion. The ideal output should show two people facing the camera or interacting naturally, with subtle head movement, mouth movement, facial expression changes, small hand gestures, and stable bod

LTXV 2.3 103 downloads
查看公开资料
IC Edit Subtitle & Watermark Removal Video Cleanup Workflow
Workflows 2026-05-14

IC Edit Subtitle & Watermark Removal Video Cleanup Workflow

This workflow is designed for IC Edit-style subtitle and watermark removal, built on an LTX 2.3 video restoration pipeline. Its main purpose is to take an existing video with hardcoded subtitles, captions, random AI text, logo overlays, signatures, platform watermarks, or semi-transparent marks, and reconstruct a cleaner video result while preserving the original motion, subject identity, camera movement, lighting, and scene composition. Unlike a simple blur, crop, mosaic, or overlay method, this workflow uses a generative restoration approach. It does not just cover the unwanted text area. Instead, it uses LTX 2.3, subtitle-removal IC LoRA, watermark-removal IC LoRA, video frame extraction, prompt-guided reconstruction, latent inpainting-style processing, tiled VAE decoding, and final video recombination to rebuild the hidden background more naturally. The workflow uses ltx-2.3-22b-dev-dare-ties-distilled-1.1 as the main model route, with LTXVAudioVAELoader, CheckpointLoaderSimple, LTXAVTextEncoderLoader, LTXVConditioning, LTXAddVideoICLoRAGuide, LTXVImgToVideoConditionOnly, VAEEncodeForInpaint, SamplerCustomAdvanced, LTXVCropGuides, VAEDecodeTiled, VHS_LoadVideo, and VHS_VideoCombine. It also loads two dedicated IC LoRA models: ltx2.3-ic-subtitles-remove-general-lora and ltx2.3-ic-watermark-remove-general-lora. This makes the workflow focused on cleanup and reconstruction rather than ordinary image-to-video generation. The positive prompt is specifically written for removal tasks. It asks the model to remove subtitles, captions, hardcoded text, AI-generated garbled text, platform watermarks, logo overlays, signatures, and semi-transparent marks. At the same time, it asks the model to restore the underlying image using surrounding visual context while preserving facial features, body shape, object boundaries, lighting, texture continuity, camera motion, and scene composition. This is the correct logic for video cleanup: erase the unwanted layer, but do not destroy the original scene. The negative prompt suppresses common restoration failures such as blur, oversaturation, pixelation, low resolution, grain, distortion, noise, compression artifacts, JPEG artifacts, glitches, watermark, text, logo, signature, copyright marks, subtitles, distorted sound, saturated sound, and loud audio artifacts. This helps reduce the chance that the workflow removes one tex

LTXV 2.3 82 downloads
查看公开资料