CIVITAI / Workflows

Stable Audio 3 Sound Asset Generation Workflow

Watch the full video first if you want to understand how this Stable Audio 3 workflow works in practice. The video shows how a simple text idea can be expanded into a structured audio prompt, how different sound categories affect the result, and how to launch the workflow online without building the full ComfyUI audio environment locally. This ComfyUI workflow is designed for Stable Audio 3 sound asset generation. Its main purpose is to turn text descriptions into usable audio assets, including music tracks, instrument loops, sound effects, one-shot samples, ambience, cinematic hits, UI sounds, game audio elements, and production-ready creative sound material. Instead of only generating random audio from a short prompt, this workflow adds a category-aware prompt expansion layer so the final Stable Audio prompt becomes more precise and more suitable for the target audio type. The workflow is built around the Stable Audio 3 Medium Base route. It uses stable_audio_3_medium_base.safetensors as the main checkpoint, t5gemma_b_b_ul2.safetensors as the Stable Audio text encoder, and a separate text-generation route for intelligent prompt rewriting. The generation pipeline uses CLIPTextEncode for positive and negative conditioning, EmptyLatentAudio for defining the target audio duration, KSampler for latent audio sampling, and VAEDecodeAudio to decode the generated latent into an actual audio waveform. The most important design in this workflow is the optional reprompt system. Users can input a short idea, then decide whether to enable prompt expansion. When reprompt is enabled, the workflow uses a category-aware prompt template. The available categories include Music, Instrument, SFX, and One-shot. Each category has different prompt rules. Music prompts focus on genre, instruments, layers, rhythm, mood, BPM, and track length. Instrument prompts focus on playing technique, timbre, production texture, BPM, and loop or stem length. SFX prompts focus on sound source, material texture, spatial environment, movement, attack, decay, and duration. One-shot prompts focus on short isolated audio samples such as hits, stabs, plucks, drum sounds, impacts, or short sound design elements. This structure makes the workflow more practical than a simple text-to-audio setup. In ordinary audio generation, a vague prompt like “dark cinematic sound” may produce inconsistent results.

LTXV 2.3 #character
View original on Civitai
Stable Audio 3 Sound Asset Generation Workflow

Public versions

v1.0

LTXV 2.3