CIVITAI / Workflows
Qwen3-TTS Preset Voice & Custom Voice Design Workflow
This ComfyUI workflow is designed for Qwen3-TTS voice generation, preset speaker testing, custom voice creation, and reference-based voice cloning. It combines three practical audio generation routes in one workflow: Voice Clone, Custom Voice, and Design Voice. Instead of only offering a single text-to-speech path, this workflow gives creators several ways to generate voices depending on whether they want to clone an existing reference voice, use a preset speaker, or design a new voice through descriptive character instructions. The workflow is built around Qwen3-TTS ComfyUI nodes and audio preprocessing tools. It includes FB_Qwen3TTSVoiceClone for reference-based voice cloning, FB_Qwen3TTSCustomVoice for preset speaker generation, FB_Qwen3TTSVoiceDesign for instruction-based custom voice design, MelBandRoFormer for vocal extraction, Whisper Large V3 for automatic transcription, LoadAudio for importing reference audio, PreviewAudio for quick listening, and SaveAudio for exporting final results. The first route is Voice Clone. This route is useful when you already have a reference audio sample and want Qwen3-TTS to generate new speech in a similar vocal style. The workflow loads the reference audio, separates the vocal track with MelBandRoFormer, transcribes the reference voice with Whisper, and then passes the cleaned reference audio and reference text into the Qwen3-TTS voice clone node. This makes the workflow suitable for voice imitation tests, narration style transfer, character voice reuse, AI dubbing, and digital human voice production. MelBandRoFormer is important in the cloning route because many reference audio samples are not perfectly clean. They may contain background music, room noise, ambience, or mixed sound effects. By extracting the vocal part before cloning, the workflow gives Qwen3-TTS a cleaner voice reference. This can improve speaker consistency, reduce unwanted background artifacts, and make the generated voice more stable. Whisper transcription is also important. Voice cloning works better when the reference audio and reference transcript match. The Apply Whisper node automatically transcribes the extracted vocal audio, so users do not always need to manually type the reference text. This is especially useful for longer reference clips or audio samples taken from existing videos. However, for production results, it is still recomm

公开版本
Qwen