CIVITAI / Workflows
LTX 2.3 Text to Video OmniNFT + Relay Three-Stage No-Subtitle Workflow
Watch the full video first if you want to understand how this LTX 2.3 text-to-video workflow works in practice. The video shows how a clean prompt can be turned into a complete video clip, how the three-stage rendering structure improves stability, and how to launch the workflow online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for LTX 2.3 text-to-video generation, using OmniNFT, Relay-style prompt control, and the distilled 1.1 model route to create clean video outputs from text prompts. The main purpose of this workflow is to make text-to-video generation more controllable, more stable, and more suitable for publishing, especially when users want a no-subtitle, no-watermark, no-extra-text output. The workflow is built around the LTX 2.3 distilled 1.1 generation route. It uses an LTX 2.3 checkpoint, Gemma3-based text encoding, LTXVConditioning, EmptyLTXVLatentVideo, LTXVEmptyLatentAudio, LTX2_NAG negative guidance, ManualSigmas, CFGGuider, SamplerCustomAdvanced, LTXVLatentUpsampler, VAEDecodeTiled, LTXVAudioVAEDecode, and CreateVideo. The graph also includes seed control, fps control, universal negative prompting, VRAM management, and audio-video latent handling. The core idea is to generate video from text while maintaining stronger control over structure, motion, and final image quality. The positive prompt defines the subject, action, camera movement, lighting, environment, atmosphere, and cinematic direction. The negative prompt is designed to suppress common LTX video problems, including low quality, flicker, unstable perspective, identity drift, broken anatomy, subtitles, captions, UI overlays, logos, watermarks, unreadable text, and unwanted audio artifacts. The workflow uses a three-stage rendering structure. The first stage focuses on initial composition and motion foundation. It creates the base video latent and establishes the main visual direction. The second stage performs latent-space upscaling and refinement, allowing the workflow to improve structure and detail without rebuilding the whole video from scratch. The third stage applies final high-resolution polish, using another controlled sampling pass before tiled VAE decoding and video assembly. Compared with ordinary text-to-video workflows, this graph is more production-oriented. A simple one-pass T2V workflow may be fast, but it often s

公开版本
LTXV 2.3