CIVITAI / Workflows
Bernini-R Three-Image Reference Cinematic Video Workflow
Watch the full video first if you want to understand how this Bernini-R three-image reference video workflow works in practice. The video shows how multiple reference images can be combined into one cinematic video generation pipeline, how the prompt enhancement system rewrites the visual concept, and how the final video can be generated online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for Bernini-R three-image reference video generation. Its main purpose is to take three visual references and use them together as the foundation for a finished short video. Compared with a single-image image-to-video workflow, this graph gives the model more visual material to understand subject identity, supporting elements, atmosphere, scene structure, and cinematic direction. The workflow is built around the Bernini-R high-noise and low-noise model route. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, VAEDecode, CreateVideo, SaveVideo, and PathchSageAttentionKJ. The model chain includes both LightX2V acceleration LoRA and UnifiedReward-Flex LoRA for the high-noise and low-noise branches, helping the workflow stay more efficient while improving visual quality and coherence. The reference section is the core of this workflow. Three LoadImage nodes provide three separate image references. Each image is processed through image_scale_pixel_v2, then combined through BatchImagesNode. These batched images enter BerniniPromptEnhancer and BerniniConditioning as the multi-reference visual condition. This allows the workflow to treat the first image as the main subject, the second image as a secondary presence or object, and the third image as another visual element, environment cue, or story component. The prompt system is also important. BerniniPromptEnhancer is used to build a Bernini-specific instruction with r2v reference-to-video logic. Then RHLLMChatNode rewrites the instruction into a more complete video prompt. The output is cleaned through StringReplace nodes, removing the JSON wrapper before sending the rewritten prompt into CLIPTextEncode. This makes the workflow more practical because the user can start from a rough idea and let the system expand it in

公开版本
Wan Video 2.2 T2V-A14B