CIVITAI / PUBLIC LIBRARY

AI-KSK Civitai Models and Workflows

Search AI-KSK public Civitai models, LoRAs, and workflows with versions, base models, and source links.

385

public entries

23

base model families

60732

total downloads

Krea2 Dual-Phase Rendering RAW Draft Turbo Finishing Workflow
Workflows 2026-07-04

Krea2 Dual-Phase Rendering RAW Draft Turbo Finishing Workflow

Watch the full video first if you want to understand how this Krea2 dual-phase rendering workflow works in practice. The video demonstrates a two-stage Krea2 production chain that uses a RAW model to build the initial image and a Turbo model to refine and complete the final result. This ComfyUI workflow is designed for Krea2 RAW drafting and Turbo finishing. Unlike a simple one-pass Krea2 workflow, this version separates the render into two phases. The first phase focuses on structure, composition, and early visual formation. The second phase uses a Turbo model to tighten the image, improve the final look, and create a more polished output. The first stage loads krea2_raw_fp8_scaled.safetensors as the RAW base model. This RAW route is passed through Power Lora Loader from rgthree, with multiple LoRA files enabled at strength 0.7. The active LoRA stack includes QJ_Krea2_Lora_E15_mk2, Detailer-KREA2, Neo_Rococo_Krea2_v1, dora, krea2_neondrip, krea2-mikkoph, and krea2_Enhancer. This gives the RAW stage a strong stylized foundation before the image is handed to the second phase. The second stage loads krea2_turbo_bf16.safetensors as the Turbo finishing model. It also uses a Power Lora Loader stack with the same style-enhancement logic. This allows the final pass to keep the visual direction of the RAW stage while adding a cleaner and more finished Krea2 Turbo surface. The text encoder route uses qwen3vl_4b_fp8_scaled.safetensors with the Krea2 CLIP type. The VAE route uses qwen_image_vae.safetensors for decoding. The resolution route uses FluxResolutionNode with a 1.5 megapixel target, 9:21 Ultra Tall aspect ratio, and divisible-by-64 alignment. This makes the workflow suitable for tall vertical posters, mobile wallpapers, surreal fashion portraits, fantasy covers, and social-media visual assets. The Stage 1 sampler uses KSampler with 30 steps, CFG 4, euler sampler, simple scheduler, random seed, and denoise 1. This stage creates the full initial RAW image foundation. The Stage 2 sampler then takes the Stage 1 latent as its input and runs another 30 steps with CFG 1, euler sampler, simple scheduler, fixed seed 666, and denoise 0.42. This second pass is designed to preserve the first-stage composition while refining the final image through the Turbo model. The workflow also includes an intermediate decode and save path for the first-stage result, so the RAW

Krea 2 230 downloads
View public details
Krea2 Merged Model Implementation Workflow
Workflows 2026-07-04

Krea2 Merged Model Implementation Workflow

Watch the full video first if you want to understand how this Krea2 merged-model implementation workflow works in practice. The video demonstrates a compact Krea2 image-generation route built around a pre-merged model file, combining model-side enhancement, conditioning rebalance, and a longer DDIM sampling setup for more controlled vertical image production. This ComfyUI workflow is designed for Krea2 merged-model image generation. Unlike a standard Krea2 Turbo workflow that directly loads the original base model, this version loads merge50.safetensors as the main UNET model. That means the model fusion has already been prepared before running the workflow. The graph itself does not need to perform a live model merge every time. Instead, it uses the pre-merged model as the starting point, making the workflow simpler and more practical for repeated generation. The model route loads merge50.safetensors through UNETLoader, then passes it into ComfyUI-Krea2T-Enhancer. The enhancer is enabled, with strength set to 1. This gives the workflow an additional model-side enhancement layer before sampling. The purpose is to make the merged Krea2 model respond with stronger texture, cleaner style behavior, and better visual presence. The text encoder route uses qwen3vl_4b_fp8_scaled.safetensors with the Krea2 CLIP type. The VAE route uses qwen_image_vae.safetensors for final decoding. This keeps the workflow aligned with the normal Krea2 ecosystem while allowing the main model itself to come from a merged checkpoint. The conditioning route uses ConditioningKrea2Rebalance with a custom 12-layer weight structure: 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 2.5, 5.0, 1.1, 4.0, 1.0 The multiplier is set to 2.5, with renormalize disabled. This provides stronger prompt control than a plain conditioning route while avoiding the more extreme pressure of multiplier 4.0 setups. It is a balanced configuration for testing merged-model behavior, prompt response, color structure, composition, and stylized texture. The sampling setup uses KSampler with 20 steps, CFG 1, DDIM sampler, sgm_uniform scheduler, fixed seed, and full denoise. Compared with the common 8-step Krea2 Turbo baseline, this version gives the workflow a slower but more deliberate render path. It is suitable when the creator wants more stability, stronger structure, and more consistent testing across multiple outputs.

Krea 2 122 downloads
View public details
Krea2 Two-Stage High-Resolution Refinement Workflow
Workflows 2026-07-03

Krea2 Two-Stage High-Resolution Refinement Workflow

Watch the full video first if you want to understand how this Krea2 two-stage high-resolution refinement workflow works in practice. The video shows how a fast Krea2 Turbo image generation route can be extended into a cleaner high-definition pipeline by adding latent upscaling and a second refinement pass. This ComfyUI workflow is designed for Krea2 two-stage image generation and high-resolution polishing. Compared with a simple one-pass Krea2 workflow, this version first creates the base image, then enlarges the latent, and finally runs a second sampling pass to improve structure, detail, texture, and overall visual clarity. It is useful when a normal Krea2 result is compositionally good but still needs more sharpness, scale, and final polish. The workflow uses krea2_turbo_bf16.safetensors as the main model. The text encoder route uses qwen3vl_4b_fp8_scaled.safetensors with the Krea2 CLIP type, while qwen_image_vae.safetensors is used for final decoding. This keeps the workflow compact, fast, and suitable for RunningHub online use. The first stage starts from an EmptyLatentImage controlled by FluxResolutionNode. In the uploaded setup, the resolution route is configured for a 9:16 vertical layout, making it suitable for vertical posters, character covers, mobile-first artwork, social-media thumbnails, and short-video platform visuals. The first KSampler uses 8 steps, CFG 1, euler sampler, simple scheduler, fixed seed, and full denoise. This stage builds the main composition, subject placement, atmosphere, lighting, and overall visual direction. The prompt conditioning is strengthened through ConditioningKrea2Rebalance. It uses a custom 12-layer weight structure: 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 2.5, 5.0, 1.1, 4.0, 1.0 The multiplier is set to 2.5, which gives the image stronger prompt response without becoming as aggressive as the multiplier 4.0 workflows. This makes it more balanced for high-resolution refinement, where the goal is not just intensity, but cleaner structure and better final detail. After the first sampling stage, the latent is enlarged through LatentUpscaleBy with a 1.5 scale factor. This is the key difference from a normal Krea2 baseline workflow. Instead of decoding immediately, the workflow keeps the result in latent space and increases its scale before refinement. The second KSampler then performs a high-resolution polishing pa

Krea 2 461 downloads
View public details
Krea2 Three-Stage High-Resolution Rendering Workflow
Workflows 2026-07-03

Krea2 Three-Stage High-Resolution Rendering Workflow

Watch the full video first if you want to understand how this Krea2 three-stage rendering workflow works in practice. The video shows how a Krea2 Turbo image can be built step by step: first creating the base composition, then refining the latent image after the first upscale, and finally polishing the result through a third controlled refinement pass. This ComfyUI workflow is designed for Krea2 three-stage image rendering. Compared with a simple one-pass Krea2 workflow, this version does not stop after the first generation. It uses a staged latent pipeline to gradually improve the image structure, scale, clarity, and final visual finish. This makes it suitable for creators who want a cleaner and more production-ready result from Krea2 without building a heavy external upscaling system. The workflow uses krea2_turbo_bf16.safetensors as the main generation model. The text encoder route uses qwen3vl_4b_fp8_scaled.safetensors with the Krea2 CLIP type, and the final decoding route uses qwen_image_vae.safetensors. This keeps the workflow compact, fast, and practical for RunningHub online use. The prompt route is strengthened through ConditioningKrea2Rebalance. In this workflow, Rebalance is set to the balanced preset with a custom 12-layer weight structure: 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 2.5, 5.0, 1.1, 4.0, 1.0 The multiplier is set to 2.0, with renormalize enabled. This makes the conditioning stronger than a plain Krea2 render, but less aggressive than the heavy multiplier 4.0 workflows. It is a balanced setup for staged refinement, where the goal is not only strong prompt response, but also cleaner composition and stable detail across multiple passes. The first stage starts from an EmptyLatentImage controlled by FluxResolutionNode. The active output format is set to 3:4 Golden Ratio, making this workflow suitable for vertical posters, character images, fantasy covers, mobile thumbnails, and stylized AI artwork. Stage 1 uses 6 steps, CFG 1, euler sampler, sgm_uniform scheduler, fixed seed, and full denoise. This stage creates the base image composition. After Stage 1, the latent is enlarged through LatentUpscaleBy using nearest-exact at 1.25x scale. Stage 2 then runs another 6-step KSampler pass with denoise 0.45. This stage refines the upscaled latent while preserving the main structure. The workflow repeats the same logic once more. The Stage 2 lat

Krea 2 266 downloads
View public details
Krea2 8-Step Cinematic Portrait Benchmark Workflow
Workflows 2026-07-03

Krea2 8-Step Cinematic Portrait Benchmark Workflow

Watch the full video first if you want to understand how this Krea2 8-step cinematic portrait benchmark workflow works in practice. The video shows how a lightweight Krea2 Turbo pipeline can generate polished portrait-style images with a simple structure, fast sampling, and a tuned conditioning rebalance setup. This ComfyUI workflow is designed as a clean Krea2 Turbo benchmark for cinematic portrait and stylized image generation. Its purpose is not to be a large multi-stage production graph, but to provide a fast, stable, and repeatable baseline that creators can use to test Krea2 image quality, prompt response, composition, lighting, and cinematic style control. The workflow is built around krea2_turbo_bf16.safetensors as the main image generation model. The text encoder route uses qwen3vl_4b_fp8_scaled.safetensors with the Krea2 CLIP type, giving the workflow a compact but efficient text-conditioning setup. The VAE route uses qwen_image_vae.safetensors for final image decoding. This makes the graph lightweight and direct, suitable for quick testing and RunningHub online use. The key control module in this workflow is ConditioningKrea2Rebalance. The workflow uses a custom per-layer weight setup with a multiplier of 2.5. This rebalance node is used to strengthen the positive conditioning and improve the model’s response to cinematic prompt structure, visual hierarchy, subject placement, lighting, and style description. It is one of the main reasons this workflow can feel more controlled than a plain Krea2 Turbo render. The sampling section is intentionally simple. It uses KSampler with 8 steps, CFG 1, euler sampler, simple scheduler, and full denoise. This makes the workflow fast enough for practical iteration while still keeping enough quality for cinematic visual tests. The fixed seed also makes it easier to compare prompts, rebalance settings, model behavior, and resolution changes across different runs. The resolution section uses FluxResolutionNode to provide flexible aspect-ratio and size control. In the uploaded setup, the workflow is prepared for high-impact composition testing, including ultra-wide cinematic framing. This is useful for creators who want to test poster-like portraits, fantasy character shots, cinematic covers, key visuals, thumbnail art, and stylized AI image concepts. The negative conditioning route uses ConditioningZeroOut, k

Krea 2 286 downloads
View public details
Krea2 Civitai-Style Texture Replication Workflow
Workflows 2026-07-03

Krea2 Civitai-Style Texture Replication Workflow

Watch the full video first if you want to understand how this Krea2 Civitai-style texture replication workflow works in practice. The video shows how Krea2 Turbo can be used with a tag-driven prompt structure and strong Rebalance conditioning to create images with the kind of polished fantasy texture, character detail, and illustration-style finish often seen in community model galleries. This ComfyUI workflow is designed for Krea2 Civitai-style image generation. Its purpose is not only to generate a normal Krea2 image, but to reproduce a more tag-based, high-detail, gallery-style visual language. It is especially useful for creators who want stronger character rendering, richer fantasy texture, more detailed costume design, clearer lighting structure, and a more polished illustration finish. The workflow uses Krea2_Turbov10Fp8.safetensors as the main model. The text encoder route uses qwen3vl_4b_fp8_scaled.safetensors with the Krea2 CLIP type, while the VAE route uses qwen_image_vae.safetensors. This keeps the graph lightweight and fast, while still allowing a strong stylized image output. The main control module is ConditioningKrea2Rebalance. The workflow uses a custom 12-layer per-layer weight structure: 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 2.5, 5.0, 1.1, 4.0, 1.0 The multiplier is set to 4.0, which makes the conditioning drive much stronger than a standard baseline setup. This helps the prompt tags and visual concepts become more visible in the final image, especially for texture, contrast, costume detail, character definition, fantasy atmosphere, and overall style intensity. The prompt route uses a Civitai-style tag structure, starting with quality tags such as score_9, score_8_up, score_7_up, masterpiece, best quality, absurdres, and highly detailed. This is followed by character tags, subject description, fantasy setting, lighting, atmosphere, and style direction. The example prompt focuses on a white fox priestess, ornate ceremonial clothing, divine armor details, a ruined ancient temple city, colossal deity statues, dragon-shaped mist spirits, burning amber light, dark teal shadows, and a mythic Chinese fantasy atmosphere. The sampling setup is simple and fast. It uses KSampler with 8 steps, CFG 1, euler sampler, simple scheduler, full denoise, and a fixed seed. This makes the workflow practical for fast testing, prompt comparison, Civitai-styl

Krea 2 105 downloads
View public details
SCAIL-2 Two-Person Reference Editing Long-Video Workflow
Workflows 2026-06-18

SCAIL-2 Two-Person Reference Editing Long-Video Workflow

Watch the full video first if you want to understand how this SCAIL-2 two-person reference editing workflow works in practice. The video shows how two reference characters can replace two selected people in a driving video, while keeping left-right identity alignment, synchronized motion, original scene structure, lighting coherence, and long-video continuity more stable. This ComfyUI workflow is designed for SCAIL-2 two-person biological reference editing. Its main purpose is to take a two-person reference image and use it to replace both selected people in a two-person driving video. Unlike a simple two-person pose-driving workflow, this version is explicitly configured for character replacement. The reference image provides the two replacement identities, while the driving video provides the motion, timing, body interaction, camera rhythm, and original scene context. The workflow is built around wan2.1_14B_SCAIL_2_fp8_scaled.safetensors as the main SCAIL-2 model. It also uses WAN VAE, UMT5 WAN text encoding, CLIP Vision, SAM3 tracking, SCAIL2ColoredMask, WanSCAILToVideo, SamplerCustom, VAEDecode, ForLoop continuation, overlap-frame trimming, ColorTransfer, final video combining, and original audio restoration. A multi-LoRA chain is preserved to improve motion quality, character stability, and final visual consistency. The most important switch in this workflow is replacement_mode=true. This tells the SCAIL route to perform two-person skeleton guidance with reference character replacement. The positive prompt focuses on replacing both selected people with two reference characters, following two-person pose guidance, keeping left and right identity alignment, preserving the original scene structure, maintaining natural synchronized motion, coherent lighting, and smooth temporal consistency. The negative prompt is also built for two-person failure cases. It suppresses bad video quality, flicker, only one person being replaced, missing second person, wrong identity order, identity swap, identity drift, deformed bodies, distorted faces, extra limbs, missing hands, blur, and low-quality output. These problems are especially common in two-person editing because the model has to preserve both bodies, both identities, and their interaction at the same time. The workflow uses strict 512×896 alignment. Both the reference image and the driving video are resized

Wan Video 2.2 T2V-A14B 170 downloads
View public details
SCAIL-2 Single-Person Reference Editing Long-Video Workflow
Workflows 2026-06-12

SCAIL-2 Single-Person Reference Editing Long-Video Workflow

Watch the full video first if you want to understand how this SCAIL-2 single-person reference editing workflow works in practice. The video shows how one reference character can replace the selected person in a driving video, while the workflow preserves pose motion, original scene structure, identity consistency, lighting coherence, and long-video continuity. This ComfyUI workflow is designed for SCAIL-2 single-person biological reference editing. Its main purpose is to place one reference character onto the target person in a single-person driving video. Unlike the single-person driving workflow, this version is not only about making a reference character follow motion. It is explicitly configured for character replacement, using skeleton guidance, subject tracking, colored masks, reference identity encoding, and long-video continuation. The workflow is built around wan2.1_14B_SCAIL_2_fp8_scaled.safetensors as the main SCAIL-2 model. It also uses WAN VAE, UMT5 XXL WAN text encoding, CLIP Vision, SAM3, SCAIL2ColoredMask, WanSCAILToVideo, SamplerCustom, VAEDecode, ForLoop continuation, overlap-frame trimming, ColorTransfer, final video combining, and original audio restoration. A multi-LoRA enhancement chain is also preserved to improve motion quality, visual stability, and final rendering consistency. The most important switch in this workflow is replacement_mode=true. This tells the SCAIL route to perform single-person skeleton guidance with character replacement. The reference image provides the replacement character identity, while the driving video provides the target motion and scene structure. The positive prompt focuses on replacing the selected single target person, following one-person pose guidance, keeping the original scene structure, preserving consistent identity, natural motion, coherent lighting, and smooth temporal consistency. The negative prompt is also designed for this task. It suppresses bad video quality, flicker, wrong-area replacement, identity drift, deformed bodies, distorted faces, extra limbs, missing hands, warped hands, broken anatomy, blur, and low-quality output. This is important because single-person replacement often fails when the mask is inaccurate, the reference image is unclear, or the driving video contains heavy occlusion. The workflow uses strict 512×896 alignment. Both the reference image and the driving vide

Other 233 downloads
View public details
SCAIL-2 Single-Person Biological Long-Video Driving Workflow
Workflows 2026-06-12

SCAIL-2 Single-Person Biological Long-Video Driving Workflow

Watch the full video first if you want to understand how this SCAIL-2 single-person long-video driving workflow works in practice. The video shows how one reference character can follow a single-person driving video, while the workflow keeps identity, body structure, motion rhythm, silhouette stability, and long-video continuity more consistent. This ComfyUI workflow is designed for SCAIL-2 single-person biological long-video driving. Its main purpose is to animate one reference character by following the motion of one person in a driving video. This is not a local character replacement workflow. It is a skeleton-guided animation route where the reference character follows the full-body movement from the source video while preserving the character’s visual identity. The workflow is built around wan2.1_14B_SCAIL_2_fp8_scaled.safetensors as the main SCAIL-2 model. It also uses WAN VAE, UMT5 XXL WAN text encoding, CLIP Vision identity encoding, SAM3 subject tracking, SCAIL2ColoredMask, WanSCAILToVideo, SamplerCustom, VAEDecode, ForLoop continuation, frame trimming, ColorTransfer, final video combining, and original audio restoration. A multi-LoRA enhancement chain is also included, using modules such as LightX2V, WanAnimate relight, Wan2.2 Lightning I2V, FastWan 480p, Wan21 PusaV1, Wan2.2 Fun InP, and stage-based enhancement LoRAs. The first important rule of the workflow is strict input alignment. Both the reference image and the driving video are aligned to 512×896 before entering SAM3, CLIPVision, and SCAIL. This helps avoid mask mismatch, pose instability, identity drift, and unexpected body deformation caused by inconsistent input dimensions. The second key rule is single-subject tracking. SAM3 is configured with max_objects=1. SCAIL2ColoredMask uses object_indices=0 and sort_by=area. This tells the workflow to focus on one main character, select the dominant subject area, and use that tracked subject as the motion target. This is useful for single-person dance, character animation, creature motion transfer, digital human testing, mascot animation, anime character motion driving, and stylized biological character videos. The workflow uses replacement_mode=false. This means the goal is skeleton-guided driving rather than local replacement. The reference image provides the character identity, the driving video provides the motion structure, and the mask

Other 140 downloads
View public details
SCAIL-2 Two-Person Biological Long-Video Driving Workflow
Workflows 2026-06-12

SCAIL-2 Two-Person Biological Long-Video Driving Workflow

Watch the full video first if you want to understand how this SCAIL-2 two-person long-video driving workflow works in practice. The video shows how two characters in a reference image can follow a two-person driving video, while the workflow keeps left and right character assignments, pose structure, temporal continuity, and long-video extension more stable. This ComfyUI workflow is designed for SCAIL-2 two-person biological long-video driving. Its main purpose is to transfer global two-person motion from a driving video onto two characters in a reference image. This is not a local replacement workflow. It is a skeleton-guided animation workflow where both characters follow the full-body movement from the driving video while keeping their assigned identities and visual roles. The workflow is built around wan2.1_14B_SCAIL_2_fp8_scaled.safetensors as the main SCAIL-2 model. It also uses WAN VAE, UMT5 XXL WAN text encoding, CLIP Vision identity encoding, SAM3 subject tracking, SCAIL2ColoredMask, WanSCAILToVideo, SamplerCustom, VAEDecode, ForLoop continuation, frame trimming, color matching, video combining, and original audio restoration. The LoRA chain is also preserved for enhancement, including LightX2V, WanAnimate relight, Wan2.2 Lightning I2V, FastWan 480p, Wan21 PusaV1, Wan2.2 Fun InP, and stage-based enhancement LoRAs. The first important rule of the workflow is size alignment. Both the reference image and the driving video are aligned to 512×896 before entering SAM3, CLIPVision, and SCAIL. This avoids mismatched masks, unstable pose conditioning, and identity drift caused by inconsistent input dimensions. The second key rule is two-person tracking. SAM3 is configured with max_objects=2. SCAIL2ColoredMask uses object_indices=0,1 and sort_by=left_to_right. This means the workflow treats the left and right subjects as separate controlled identities. The goal is to reduce identity mixing, role swapping, clothing confusion, and left-right character instability during two-person motion transfer. The workflow uses replacement_mode=false, which means it focuses on two-person skeleton guidance rather than local character replacement. The reference image provides the two target characters, the driving video provides the motion, and the colored mask system links both sides together before entering WanSCAILToVideo. The long-video structure is one of the stron

Other 101 downloads
View public details
Ideogram 4 Reference Latent Reconstruction Image-to-Image Workflow
Workflows 2026-06-11

Ideogram 4 Reference Latent Reconstruction Image-to-Image Workflow

Watch the full video first if you want to understand how this Ideogram 4 reference latent reconstruction workflow works in practice. The video shows how a reference image can be analyzed, converted into a structured Ideogram 4 JSON prompt, encoded into latent space, and then regenerated through a controlled image-to-image reconstruction pipeline. This ComfyUI workflow is designed for Ideogram 4 image-to-image reference latent reconstruction. Its main purpose is to rebuild an existing image by combining two control routes: a visual-language JSON prompt route and a reference latent remix route. Compared with a pure text-to-image workflow, this graph gives Ideogram 4 both a structured description of the image and a latent-space reference starting point. Compared with ordinary image-to-image workflows, it is more experimental and more design-oriented, because the final stability comes from both the reference latent and the reconstructed JSON prompt. The workflow starts from a reference image. The image is scaled and prepared through the image scaling section, then encoded into latent space through VAEEncode. This encoded latent is sent into the sampler as the starting latent instead of using a blank EmptyFlux2LatentImage canvas. The older text-to-image empty latent path is intentionally disconnected in this version, because the workflow is focused on reference reconstruction rather than pure generation. A key technical point is the SplitSigmasDenoise stage. The scheduler output is split before entering the sampler, allowing the workflow to run a controlled denoise range over the reference latent. This makes the result behave like a forced latent remix route: it may preserve part of the reference image’s composition and structure, while still allowing Ideogram 4 to reinterpret the image according to the prompt and denoise strength. It is not a classic SD 1.5-style img2img pipeline, so the result can drift depending on denoise settings, prompt strength, and reference complexity. The second major control route is the Vision LLM prompt reconstruction chain. The workflow uses a reference image together with target width and height information, then asks the RH visual completion section to generate an Ideogram 4-compatible structured JSON prompt. This JSON can describe the subject, background, layout hierarchy, bounding boxes, readable text, color palette, lightin

Other 326 downloads
View public details
Ideogram 4 Structured JSON Image Reconstruction Workflow
Workflows 2026-06-11

Ideogram 4 Structured JSON Image Reconstruction Workflow

Watch the full video first if you want to understand how this Ideogram 4 structured JSON image reconstruction workflow works in practice. The video shows how a reference image can be analyzed, converted into an Ideogram 4 JSON prompt, and then regenerated through a structured image generation pipeline. This ComfyUI workflow is designed for Ideogram 4 image reconstruction through structured JSON prompting. Its main purpose is not ordinary text-to-image generation. Instead, the workflow starts from a reference image, analyzes its visible composition, and rebuilds the image as an Ideogram 4-compatible structured JSON prompt. This makes it especially useful for image style reconstruction, layout recovery, poster remaking, visual reference rebuilding, typography layout testing, and design-oriented “image washing” workflows. The key idea is simple: the user does not need to manually write a long prompt. The workflow uses a reference image as the main input. The visual analysis chain observes the image, extracts the subject, background, composition hierarchy, visible text, color system, lighting, medium, and layout relationship, then generates a structured JSON prompt. That JSON prompt is sent into the Ideogram 4 generation route, allowing the model to recreate the design logic instead of only guessing from a loose natural-language description. The workflow is built around Ideogram 4 image generation. It uses ideogram4_fp8_scaled.safetensors as the main model, Flux2 VAE for decoding, Ideogram4Scheduler for sampling control, DualModelGuider for guided generation, CFGOverride for guidance behavior, EmptyFlux2LatentImage for canvas creation, SamplerCustomAdvanced for final denoising, VAEDecode for image decoding, and SaveImage for final export. The workflow also includes a structured prompt encoding section where the generated JSON is passed into CLIPTextEncode. The most important part is the image-to-JSON reconstruction section. The workflow note describes the main route as LoadImage → image_scale_pixel_v2 → RHLLMChatNode image1, combined with a system prompt and target width / height. RH visual completion observes the reference image and outputs an Ideogram 4 JSON prompt. The width and height are used to help plan bounding boxes, layout proportions, and composition scale. This makes the prompt more useful for structured regeneration, especially when the original

Other 180 downloads
View public details
EverAnimate Long-Video Consistency Generation Workflow
Workflows 2026-06-09

EverAnimate Long-Video Consistency Generation Workflow

Watch the full video first if you want to understand how this EverAnimate long-video consistency workflow works in practice. The video shows how a reference character can be animated through a driving video, how face and pose information are extracted, and how the workflow extends the result into a longer continuous video while trying to keep identity, motion, and visual style stable. This ComfyUI workflow is designed for EverAnimate long-video consistency generation. Its main purpose is to solve a common problem in AI character animation: the first few seconds may look good, but as the video becomes longer, the face starts drifting, clothing changes, body proportions become unstable, and the motion loses continuity. This workflow uses a segmented generation structure with motion handoff, pose guidance, face reference, and loop-based continuation to make longer character videos more controllable. The workflow starts from a driving video. VHS video loading and video information nodes read the selected frames, FPS, width, height, and frame count. This gives the workflow a clear source timeline before generation begins. The driving video is then processed through pose and face detection. The graph includes ViTPose / YOLO-style body detection, PoseAndFaceDetection, DrawViTPose, SDPose keypoint extraction, and face image extraction. These preprocessing steps turn the original video into usable pose_video and face_video conditions. The EverAnimate generation section is the core of the workflow. ComfyEverAnimate receives the positive prompt, negative prompt, VAE, reference image, face video, pose video, width, height, length, pose strength, face strength, and motion handoff settings. The first EverAnimate pass generates the opening segment. After that, the workflow trims anchor latents and duplicate image frames, then uses continue_motion to pass motion information into the next segment. This is the key mechanism for long-video continuity. Instead of generating the entire long video in one pass, the workflow uses a ForLoop structure. The first segment establishes the character, motion, and visual identity. The loop then repeatedly generates continuation segments while receiving the previous motion context. Each continuation segment is sampled, decoded, trimmed, and batched back into the full sequence. This makes the workflow more practical for longer AI charact

Wan Video 2.2 T2V-A14B 120 downloads
View public details
EverAnimate Long-Video Consistency Editing Workflow
Workflows 2026-06-09

EverAnimate Long-Video Consistency Editing Workflow

Watch the full video first if you want to understand how this EverAnimate long-video consistency editing workflow works in practice. The video shows how an existing video can be edited with reference guidance, pose control, background preservation, character masking, and long-form continuation while keeping the edited result more stable across multiple segments. This ComfyUI workflow is designed for EverAnimate long-video consistency editing. Its main purpose is not only to generate an animated character, but to edit an existing video while preserving the important motion, body structure, background relationship, and temporal continuity. Compared with a simple one-shot video edit, this workflow is built for longer clips where character identity, pose rhythm, mask boundaries, and visual consistency often break after the first segment. The workflow starts from a driving video. VHS video loading and VHS_VideoInfo read the selected FPS, frame count, width, height, and duration. This gives the workflow a stable timeline before generation begins. The source frames are then processed through multiple pose and detection routes, including ViTPose, YOLO, SDPose, PoseAndFaceDetection, DrawViTPose, BBoxYOLO, and SDPoseKeypointExtractor. These preprocessing stages turn the original video into usable pose guidance, body structure information, and motion control signals. The editing part is where this workflow differs from the pure generation version. It includes a character mask path, mask expansion, and block-style mask processing. GrowMaskWithBlur expands and softens the mask, while BlockifyMask converts it into a more usable character-editing mask. This mask is then sent into ComfyEverAnimate as character_mask, together with background_video and pose_video. This allows the workflow to focus the edit on the character area while keeping the background or surrounding structure more stable. The core generation node is ComfyEverAnimate. In the first local editing segment, it receives the reference image, pose video, background video, character mask, positive and negative conditioning, width, height, length, pose strength, face strength, and motion handoff settings. After the first segment is generated, TrimVideoLatent and ComfyEverAnimateTrimImages remove redundant anchor latent frames and duplicate image frames. The workflow then uses continue_motion to pass motion inf

Wan Video 2.2 T2V-A14B 84 downloads
View public details
Ideogram 4 + KJ Prompt Builder Visual Composition Director Workflow
Workflows 2026-06-07

Ideogram 4 + KJ Prompt Builder Visual Composition Director Workflow

Watch the full video first if you want to understand how this Ideogram 4 + KJ Prompt Builder workflow works in practice. The video shows how structured JSON prompts can be used to control poster layout, typography, composition, color palette, object placement, and overall visual direction inside ComfyUI. This ComfyUI workflow is designed for Ideogram 4 visual composition control using KJ Prompt Builder. Its main purpose is not only to generate a beautiful image, but to make the image generation process more like a visual design system. Instead of writing one loose natural-language prompt, this workflow uses structured JSON-style prompting and a visual composition builder to define background, subject elements, text blocks, bounding boxes, color palettes, style direction, lighting, and output ratio. The workflow is built around ideogram4_fp8_scaled.safetensors as the main Ideogram 4 model. It also uses ideogram4_unconditional_fp8_scaled.safetensors as the unconditional model branch, qwen3vl_8b_fp8_scaled.safetensors as the Ideogram 4 text encoder, flux2-vae.safetensors as the VAE, Ideogram4Scheduler, DualModelGuider, CFGOverride, EmptyFlux2LatentImage, SamplerCustomAdvanced, VAEDecode, SaveImage, and multiple Ideogram4PromptBuilderKJ nodes. The most important part of this workflow is Ideogram4PromptBuilderKJ. This node allows the user to organize a visual prompt into a more controllable layout format. You can describe the global image concept, background, style, lighting, medium, color palette, and individual elements. Each element can be treated as an object or text block, with its own position, description, text content, and palette. This is especially useful for Chinese posters, thumbnails, title images, brand visuals, concept art, and design-heavy images where placement matters. The workflow also includes multiple prepared example groups, such as Chinese movie poster layouts, racing posters, realistic documentary-style scenes, Chinese fantasy worldbuilding, puzzle-adventure environments, F1 racing, MotoGP racing, gaming showcase visuals, action-horror survival scenes, and dystopian underground sci-fi compositions. These examples make the workflow more than a single generation graph; it becomes a practical visual prompt library for learning how structured Ideogram 4 prompting works. Another important part is the dual-model guidance structure. The main

Other 807 downloads
View public details
Ideogram 4 Official Image Generation Workflow
Workflows 2026-06-07

Ideogram 4 Official Image Generation Workflow

Watch the full video first if you want to understand how this Ideogram 4 official image generation workflow works in practice. The video shows how Ideogram 4 can be used inside ComfyUI for structured poster design, typography-heavy visuals, brand-style images, collage compositions, and layout-controlled image generation. This ComfyUI workflow is designed as a clean Ideogram 4 image generation route. Its main purpose is to provide a stable official-style output pipeline for creators who want to test Ideogram 4 without building a complicated multi-branch graph. The workflow focuses on model loading, prompt encoding, dual-model guidance, quality preset control, resolution handling, sampling, decoding, and final image export. The workflow is built around ideogram4_fp8_scaled.safetensors as the main generation model. It also uses ideogram4_unconditional_fp8_scaled.safetensors as the unconditional branch, qwen3vl_8b_fp8_scaled.safetensors as the Ideogram 4 text encoder, and flux2-vae.safetensors as the VAE. These are the core files required for local deployment. The graph also includes CLIPTextEncode, ConditioningZeroOut, CFGOverride, DualModelGuider, Ideogram4Scheduler, RandomNoise, KSamplerSelect, EmptyFlux2LatentImage, SamplerCustomAdvanced, VAEDecode, and SaveImage. The most important feature of this workflow is its structured prompt design. Ideogram 4 is especially strong for images that need readable text, graphic layout, poster composition, logo-like visuals, magazine-style design, and clear object placement. Instead of relying only on a loose natural-language prompt, this workflow encourages JSON-style captions. The JSON prompt can define the high-level description, background, visual style, lighting, color palette, and composition elements. This gives the model clearer design intent and makes it more useful for cover images, YouTube thumbnails, Chinese posters, advertising visuals, and social media key art. The dual-model guidance structure is another key part. The main Ideogram 4 model reads the prompt and tries to follow the target image direction. The unconditional model provides a baseline. DualModelGuider compares the two branches and pushes the output toward the prompt target. CFGOverride then controls how guidance is applied during sampling, helping the workflow stay aligned without making the image overly rigid. The workflow also includes a q

Other 323 downloads
View public details
Bernini-R Image Editing Image-to-Image Workflow
Workflows 2026-06-06

Bernini-R Image Editing Image-to-Image Workflow

Watch the full video first if you want to understand how this Bernini-R image editing workflow works in practice. The video shows how one source image can be edited through a text instruction, how the workflow expands a simple idea into a stronger Bernini prompt, and how the final result can be exported as a clean single-frame image. This ComfyUI workflow is designed for Bernini-R image-to-image editing. Its main purpose is to take an existing image, preserve the important visual identity of the original subject, and apply a controlled transformation through text. Compared with pure text-to-image generation, this workflow starts from a real source image, so it can maintain the subject’s face, clothing, composition, visual direction, and key scene structure while changing the pose, background, object interaction, lighting, or atmosphere. The workflow is built around the Bernini-R high-noise and low-noise dual-model route. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the two model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, VAEDecode, SaveImage, and PathchSageAttentionKJ. The generation chain also includes LightX2V LoRA and UnifiedReward-Flex LoRA for both high-noise and low-noise stages, helping the workflow improve speed, image quality, and final visual coherence. The source image section is the foundation of the workflow. LoadImage imports the original image, then image_scale_pixel_v2 prepares the image size and alignment before it enters BerniniConditioning. This makes the workflow suitable for controlled editing tasks such as changing a character’s pose, replacing a background, adding an object, changing the environment, converting the scene style, or creating a more cinematic version of an existing image. The prompt creation section is also important. BerniniPromptEnhancer is set to the i2i task type, meaning the workflow is optimized for image editing rather than pure generation. The user can write a short edit instruction, and the prompt enhancer builds a Bernini-specific system prompt. RHLLMChatNode then rewrites the task into a more detailed editing prompt. The output is cleaned through StringReplace nodes, removing the JSON wrapper before the final prompt is sent into CLIPTextEncode. In the uploaded example, the edit instruction cha

Wan Video 2.2 T2V-A14B 213 downloads
View public details
Bernini-R Single-Image Reference Cinematic Video Workflow
Workflows 2026-06-06

Bernini-R Single-Image Reference Cinematic Video Workflow

Watch the full video first if you want to understand how this Bernini-R single-image reference video workflow works in practice. The video shows how one reference image can be expanded into a finished cinematic video, how the prompt enhancement chain converts a rough idea into a stronger Bernini instruction, and how to run the full workflow online without rebuilding a local ComfyUI environment. This ComfyUI workflow is designed for Bernini-R single-image reference video generation. Its main purpose is to take one reference image, or an expandable batch of reference images, and generate a complete video clip from it. Unlike a pure text-to-video workflow, this graph uses the reference image as the visual anchor, so the final video can preserve subject identity, style direction, visual mood, and composition logic more effectively. The workflow is built around the Bernini-R high-noise and low-noise model structure. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, PathchSageAttentionKJ, VAEDecode, CreateVideo, and SaveVideo. The model chain also includes LightX2V-style LoRA support and UnifiedReward-Flex LoRA support for both high-noise and low-noise routes, helping the final result stay more efficient, coherent, and visually polished. The reference image side is flexible. The workflow includes multiple LoadImage nodes, image scaling nodes, and BatchImagesNode. This means the graph can be used as a single-image reference workflow, but it can also expand into multi-reference input when needed. The image is scaled and prepared before entering BerniniConditioning, where it becomes the visual condition for the generated video. The prompt side is one of the strongest parts of this workflow. BerniniPromptEnhancer is used to build a Bernini-specific prompt structure. In the uploaded graph, the task type is set around r2v / reference-to-video logic, and the example prompt describes an epic cinematic fantasy scene in a collapsing floating holy city above the clouds. The prompt is then passed into RHLLMChatNode, which rewrites the instruction into a more complete video-generation prompt. After that, StringReplace nodes clean the JSON wrapper, and the final rewritten text is automatically connected into

Wan Video 2.2 T2V-A14B 198 downloads
View public details
Bernini-R Three-Image Reference Cinematic Video Workflow
Workflows 2026-06-06

Bernini-R Three-Image Reference Cinematic Video Workflow

Watch the full video first if you want to understand how this Bernini-R three-image reference video workflow works in practice. The video shows how multiple reference images can be combined into one cinematic video generation pipeline, how the prompt enhancement system rewrites the visual concept, and how the final video can be generated online without rebuilding the full ComfyUI environment locally. This ComfyUI workflow is designed for Bernini-R three-image reference video generation. Its main purpose is to take three visual references and use them together as the foundation for a finished short video. Compared with a single-image image-to-video workflow, this graph gives the model more visual material to understand subject identity, supporting elements, atmosphere, scene structure, and cinematic direction. The workflow is built around the Bernini-R high-noise and low-noise model route. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, VAEDecode, CreateVideo, SaveVideo, and PathchSageAttentionKJ. The model chain includes both LightX2V acceleration LoRA and UnifiedReward-Flex LoRA for the high-noise and low-noise branches, helping the workflow stay more efficient while improving visual quality and coherence. The reference section is the core of this workflow. Three LoadImage nodes provide three separate image references. Each image is processed through image_scale_pixel_v2, then combined through BatchImagesNode. These batched images enter BerniniPromptEnhancer and BerniniConditioning as the multi-reference visual condition. This allows the workflow to treat the first image as the main subject, the second image as a secondary presence or object, and the third image as another visual element, environment cue, or story component. The prompt system is also important. BerniniPromptEnhancer is used to build a Bernini-specific instruction with r2v reference-to-video logic. Then RHLLMChatNode rewrites the instruction into a more complete video prompt. The output is cleaned through StringReplace nodes, removing the JSON wrapper before sending the rewritten prompt into CLIPTextEncode. This makes the workflow more practical because the user can start from a rough idea and let the system expand it in

Wan Video 2.2 T2V-A14B 202 downloads
View public details
Bernini-R Reference Background Replacement Video Editing Workflow
Workflows 2026-06-06

Bernini-R Reference Background Replacement Video Editing Workflow

Watch the full video first if you want to understand how this Bernini-R reference background replacement workflow works in practice. The video shows how a source video and a reference background image can be combined into a controlled video editing pipeline, where the original subject, motion, pose, clothing, and camera framing are preserved while the surrounding environment is replaced with a new reference-based background. This ComfyUI workflow is designed for Bernini-R reference background video replacement. Its main purpose is to take an existing video, keep the main person or subject unchanged, and replace the background with a new scene provided by a reference image. Instead of generating a completely new video or changing the entire frame, this workflow focuses on controlled background transformation: the dancer or foreground subject remains consistent, while the environment, furniture, lighting atmosphere, and spatial style are rebuilt around them. The workflow is built around the Bernini-R video editing model route. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the high-noise and low-noise model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, LoadVideo, GetVideoComponents, BatchImagesNode, KSamplerAdvanced, VAEDecode, CreateVideo, SaveVideo, and VHS_VideoCombine. The graph also includes LightX2V-style LoRA acceleration support, helping the workflow run through a more practical video editing path. The key node is BerniniConditioning. This node receives the source video, reference images, positive and negative text conditioning, VAE, width, height, video length, and reference maximum size. In this workflow, the source video provides the foreground motion and timing, while the reference image provides the new background design. This is the central logic behind background replacement: the video decides what must stay unchanged, and the reference image decides what the new environment should look like. A major advantage of this workflow is the built-in prompt creation chain. The graph uses RHLLMChatNode to analyze the source video and reference background image, then generate a detailed Bernini editing prompt. The LLM output is passed through a JSON cleanup chain using StringReplace nodes, then automatically connected into the positive prompt encoder. This help

Wan Video 2.2 T2V-A14B 115 downloads
View public details
Bernini-R Image-to-Video Source Image Animation Workflow
Workflows 2026-06-06

Bernini-R Image-to-Video Source Image Animation Workflow

Watch the full video first if you want to understand how this Bernini-R image-to-video workflow works in practice. The video shows how one source image can be turned into a dynamic video, how the prompt enhancement chain expands the motion instruction, and how the final result can be generated online without rebuilding a local ComfyUI setup. This ComfyUI workflow is designed for Bernini-R image-to-video generation. Its main purpose is to take a still image as the starting visual condition, then generate a short video based on a text instruction. Compared with pure text-to-video generation, this workflow gives the model a concrete visual anchor. The source image provides the subject, composition, framing, environment, and initial visual identity, while the prompt controls the action, reaction, camera behavior, atmosphere, and scene progression. The workflow is built around the Bernini-R high-noise and low-noise model route. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, VAEDecode, CreateVideo, SaveVideo, and PathchSageAttentionKJ. The model chain also includes LightX2V LoRA and UnifiedReward-Flex LoRA for both high-noise and low-noise stages, helping improve generation efficiency, motion coherence, and final visual quality. The source image section is the foundation of the workflow. LoadImage imports the starting image, then image_scale_pixel_v2 prepares the image size and alignment before sending it into the Bernini-R conditioning structure. This makes the workflow suitable for animating portraits, character images, trackside scenes, product photos, concept images, and cinematic still frames. The prompt creation section is also important. BerniniPromptEnhancer is set to the i2v task type. The user can write a simple instruction, and the workflow converts it into a Bernini-specific image-to-video prompt. RHLLMChatNode then rewrites the task into a more detailed cinematic instruction. The output is cleaned through StringReplace nodes, removing the JSON wrapper before sending the final prompt into CLIPTextEncode. In the uploaded example, the source image is animated into a trackside scene where a woman reacts with extreme surprise as an F1 car and a black truck race past her, creating smok

Wan Video 2.2 T2V-A14B 222 downloads
View public details
Bernini-R Video-to-Video Editing Workflow
Workflows 2026-06-06

Bernini-R Video-to-Video Editing Workflow

Watch the full video first if you want to understand how this Bernini-R V2V video editing workflow works in practice. The video shows how an existing source video can be edited through text instructions while preserving the original motion, camera structure, subject position, lighting, and scene rhythm. This ComfyUI workflow is designed for Bernini-R video-to-video editing. Its main purpose is to take a source video and apply a controlled visual edit without rebuilding the whole clip from scratch. Instead of using pure text-to-video generation, this workflow starts from real video frames, extracts the source video components, and then uses BerniniConditioning to guide the editing process around the original motion and composition. The workflow is built around the Bernini-R high-noise and low-noise model structure. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, LoadVideo, GetVideoComponents, BerniniConditioning, KSamplerAdvanced, VAEDecode, CreateVideo, SaveVideo, and PathchSageAttentionKJ. The model route includes LightX2V LoRA and UnifiedReward-Flex LoRA for both high-noise and low-noise stages, helping the workflow improve speed, stability, and final visual quality. The source video is the foundation of the workflow. LoadVideo imports the original clip, and GetVideoComponents separates the frame sequence, audio, and FPS. The extracted frames are then passed into BerniniConditioning as the source video condition. This means the original performance, body movement, camera angle, timing, and background relationship can remain stable while the edit instruction changes the visual result. The prompt system is also an important part of the graph. BerniniPromptEnhancer builds a Bernini-specific V2V instruction from a short user task. In the uploaded workflow, the example task is to make the woman in the video wear Lolita clothing. RHLLMChatNode then rewrites the task into a more detailed edit prompt. The output is cleaned through StringReplace nodes, removing the JSON wrapper before sending the final instruction into CLIPTextEncode. This allows the user to start with a simple edit request and let the workflow turn it into a more precise video editing prompt. The generation route uses BerniniConditioning with a vertical 480×848 se

Wan Video 2.2 T2V-A14B 157 downloads
View public details
Bernini-R Reference Video Conditional Editing Workflow
Workflows 2026-06-06

Bernini-R Reference Video Conditional Editing Workflow

Watch the full video first if you want to understand how this Bernini-R reference video conditioning workflow works in practice. The video shows how a source video and a reference video can be connected into one controlled editing pipeline, where the reference video becomes part of the edited scene while the original video structure remains stable. This ComfyUI workflow is designed for Bernini-R reference video conditional editing. Its main purpose is to take an existing source video and use another video as a visual condition, allowing the workflow to insert, propagate, or integrate the reference video content into the target scene. Compared with simple text-to-video generation, this is a more controlled video editing workflow because it uses both source video structure and reference video content at the same time. The workflow is built around the Bernini-R high-noise and low-noise model route. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, LoadVideo, GetVideoComponents, KSamplerAdvanced, VAEDecode, CreateVideo, and SaveVideo. The model branches are also patched through PathchSageAttentionKJ, which helps the workflow run the Bernini route with a more optimized attention setup. The key node is BerniniConditioning. In this workflow, BerniniConditioning receives both the source video and the reference video. The source video provides the base scene, motion, camera structure, timing, and audio. The reference video provides the visual content that should be inserted or used as the editing condition. This is different from a simple reference image workflow, because the reference input itself contains temporal motion and video information. The example prompt in this workflow is designed around a reference video appearing as an F1 championship video looping on a billboard in a street scene. This is a very clear use case: the source scene keeps its own environment and camera framing, while the reference video becomes embedded as a dynamic screen, billboard, or video surface inside the scene. This makes the workflow useful for screen replacement, billboard insertion, in-scene video advertising, dynamic poster replacement, and reference-video-driven scene editing. The generation route uses a two-stage KSamplerAdv

Wan Video 2.2 T2V-A14B 109 downloads
View public details
Bernini-R Text-to-Image Visual Generation Workflow
Workflows 2026-06-06

Bernini-R Text-to-Image Visual Generation Workflow

Watch the full video first if you want to understand how this Bernini-R text-to-image workflow works in practice. The video shows how a short text idea can be expanded into a detailed visual prompt, then processed through the Bernini-R dual-model route to generate a polished image-style result inside a clean ComfyUI pipeline. This ComfyUI workflow is designed for Bernini-R text-to-image generation. Its main purpose is to turn a simple creative idea into a finished visual result without requiring a source image, source video, reference image, or reference video. Unlike reference-based Bernini workflows, this version is focused on pure text input. The user writes the concept, and the workflow handles prompt enhancement, LLM rewriting, Bernini conditioning, sampling, decoding, and final output. The workflow is built around the Bernini-R high-noise and low-noise model structure. It uses Bernini_HIGH_fp8_e4m3fn_scaled.safetensors and Bernini_LOW_fp8_e4m3fn_scaled.safetensors as the dual model branches. It also uses UMT5 XXL fp8 text encoding, Wan 2.1 VAE, BerniniConditioning, KSamplerAdvanced, VAEDecode, CreateVideo, SaveVideo, and PathchSageAttentionKJ. Even though this workflow is set up as a text-to-image task, it still uses the Bernini generation pipeline and video-compatible output structure, which makes it flexible for single-frame visual generation or short output testing. The prompt section is one of the strongest parts of the workflow. BerniniPromptEnhancer is set to the t2i task type. The user can enter a rough visual idea, and the enhancer builds a Bernini-specific prompt structure. In the uploaded example, the concept is a Van Gogh Sunflowers-inspired realistic fractured cubism image. RHLLMChatNode then rewrites the short idea into a much more detailed visual instruction. The output is cleaned through StringReplace nodes, removing the JSON wrapper before sending the final prompt into CLIPTextEncode. The generation section uses BerniniConditioning with a 1280×720 setup and very short length logic, making it suitable for image-style generation and concept testing. The first KSamplerAdvanced stage handles the high-noise construction phase, where the main composition, subject, geometry, color, and structure are created. The second KSamplerAdvanced stage handles low-noise refinement, improving visual polish, texture, detail consistency, and final image

Wan Video 2.2 T2V-A14B 109 downloads
View public details