CIVITAI / Workflows

Bernini-R Single-Image Reference Cinematic Video Workflow

Watch the full video first if you want to understand how this Bernini-R single-image reference video workflow works in practice. The video shows how one reference image can be expanded into a finished cinematic video, how the prompt enhancement chain converts a rough idea into a stronger Bernini instruction, and how to run the full workflow online without rebuilding a local ComfyUI environment. This ComfyUI workflow is designed for Bernini-R single-image reference video generation. Its main purpose is to take one reference image, or an expandable batch of reference images, and generate a complete video clip from it. Unlike a pure text-to-video workflow, this graph uses the reference image as the visual anchor, so the final video can preserve subject identity, style direction, visual mood, and composition logic more effectively. The workflow is built around the Bernini-R high-noise and low-noise model structure. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, PathchSageAttentionKJ, VAEDecode, CreateVideo, and SaveVideo. The model chain also includes LightX2V-style LoRA support and UnifiedReward-Flex LoRA support for both high-noise and low-noise routes, helping the final result stay more efficient, coherent, and visually polished. The reference image side is flexible. The workflow includes multiple LoadImage nodes, image scaling nodes, and BatchImagesNode. This means the graph can be used as a single-image reference workflow, but it can also expand into multi-reference input when needed. The image is scaled and prepared before entering BerniniConditioning, where it becomes the visual condition for the generated video. The prompt side is one of the strongest parts of this workflow. BerniniPromptEnhancer is used to build a Bernini-specific prompt structure. In the uploaded graph, the task type is set around r2v / reference-to-video logic, and the example prompt describes an epic cinematic fantasy scene in a collapsing floating holy city above the clouds. The prompt is then passed into RHLLMChatNode, which rewrites the instruction into a more complete video-generation prompt. After that, StringReplace nodes clean the JSON wrapper, and the final rewritten text is automatically connected into

Wan Video 2.2 T2V-A14B #character
在 Civitai 查看原始条目
Bernini-R Single-Image Reference Cinematic Video Workflow

公开版本

v1.0

Wan Video 2.2 T2V-A14B