CIVITAI / Workflows

Bernini-R Image-to-Video Source Image Animation Workflow

Watch the full video first if you want to understand how this Bernini-R image-to-video workflow works in practice. The video shows how one source image can be turned into a dynamic video, how the prompt enhancement chain expands the motion instruction, and how the final result can be generated online without rebuilding a local ComfyUI setup. This ComfyUI workflow is designed for Bernini-R image-to-video generation. Its main purpose is to take a still image as the starting visual condition, then generate a short video based on a text instruction. Compared with pure text-to-video generation, this workflow gives the model a concrete visual anchor. The source image provides the subject, composition, framing, environment, and initial visual identity, while the prompt controls the action, reaction, camera behavior, atmosphere, and scene progression. The workflow is built around the Bernini-R high-noise and low-noise model route. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, VAEDecode, CreateVideo, SaveVideo, and PathchSageAttentionKJ. The model chain also includes LightX2V LoRA and UnifiedReward-Flex LoRA for both high-noise and low-noise stages, helping improve generation efficiency, motion coherence, and final visual quality. The source image section is the foundation of the workflow. LoadImage imports the starting image, then image_scale_pixel_v2 prepares the image size and alignment before sending it into the Bernini-R conditioning structure. This makes the workflow suitable for animating portraits, character images, trackside scenes, product photos, concept images, and cinematic still frames. The prompt creation section is also important. BerniniPromptEnhancer is set to the i2v task type. The user can write a simple instruction, and the workflow converts it into a Bernini-specific image-to-video prompt. RHLLMChatNode then rewrites the task into a more detailed cinematic instruction. The output is cleaned through StringReplace nodes, removing the JSON wrapper before sending the final prompt into CLIPTextEncode. In the uploaded example, the source image is animated into a trackside scene where a woman reacts with extreme surprise as an F1 car and a black truck race past her, creating smok

Wan Video 2.2 T2V-A14B #character
View original on Civitai
Bernini-R Image-to-Video Source Image Animation Workflow

Public versions

v1.0

Wan Video 2.2 T2V-A14B