CIVITAI / Workflows

Bernini-R Text-to-Image Visual Generation Workflow

Watch the full video first if you want to understand how this Bernini-R text-to-image workflow works in practice. The video shows how a short text idea can be expanded into a detailed visual prompt, then processed through the Bernini-R dual-model route to generate a polished image-style result inside a clean ComfyUI pipeline. This ComfyUI workflow is designed for Bernini-R text-to-image generation. Its main purpose is to turn a simple creative idea into a finished visual result without requiring a source image, source video, reference image, or reference video. Unlike reference-based Bernini workflows, this version is focused on pure text input. The user writes the concept, and the workflow handles prompt enhancement, LLM rewriting, Bernini conditioning, sampling, decoding, and final output. The workflow is built around the Bernini-R high-noise and low-noise model structure. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, VAEDecode, CreateVideo, SaveVideo, and PathchSageAttentionKJ. Even though this workflow is set up as a text-to-image task, it still uses the Bernini generation pipeline and video-compatible output structure, which makes it flexible for single-frame visual generation or short output testing. The prompt section is one of the strongest parts of the workflow. BerniniPromptEnhancer is set to the t2i task type. The user can enter a rough visual idea, and the enhancer builds a Bernini-specific prompt structure. In the uploaded example, the concept is a Van Gogh Sunflowers-inspired realistic fractured cubism image. RHLLMChatNode then rewrites the short idea into a much more detailed visual instruction. The output is cleaned through StringReplace nodes, removing the JSON wrapper before sending the final prompt into CLIPTextEncode. The generation section uses BerniniConditioning with a 1280×720 setup and very short length logic, making it suitable for image-style generation and concept testing. The first KSamplerAdvanced stage handles the high-noise construction phase, where the main composition, subject, geometry, color, and structure are created. The second KSamplerAdvanced stage handles low-noise refinement, improving visual polish, texture, detail consistency, and final image

Wan Video 2.2 T2V-A14B #character
View original on Civitai
Bernini-R Text-to-Image Visual Generation Workflow

Public versions

v1.0

Wan Video 2.2 T2V-A14B