4.3 KiB
ComfyUI Dia TTS Nodes
This node pack integrates the Nari-Labs Dia 1.6b text-to-speech model into ComfyUI using the safetensors file nari-labs provided.
Dia allows generating dialogue with speaker tags ([S1], [S2]) and non-verbal sounds ((laughs), etc.).
It requires a CUDA-enabled GPU.
Note: This version is specifically configured for the nari-labs/Dia-1.6B model architecture.
Installation
- Ensure you have a CUDA-enabled GPU and the necessary NVIDIA drivers installed.
- Download the Dia-1.6B model safetensors file from Hugging Face:
- Direct Download URL: https://huggingface.co/nari-labs/Dia-1.6B/blob/main/model.safetensors
- Place the downloaded
.safetensorsfile into thediffusion_modelsdirectory (e.g.,ComfyUI/models/diffusion_models/). - You might want to rename it to
Dia-1.6B.safetensorsfor clarity. - Navigate to your
ComfyUI/custom_nodes/directory. - Clone this repository:
Alternatively, download the ZIP and extract it into
git clone https://github.com/BobRandomNumber/ComfyUI-DiaTest.gitcustom_nodes. - Install the required dependencies:
- Activate ComfyUI's Python environment (e.g.,
source ./venv/bin/activate). - Navigate to the node directory:
cd ComfyUI/custom_nodes/ComfyUI-DiaTest - Install requirements:
pip install -r requirements.txt
- Activate ComfyUI's Python environment (e.g.,
- Restart ComfyUI.
Nodes
Dia 1.6b Loader (DiaLoader)
Loads the Dia-1.6B TTS model from a local .safetensors file located in your diffusion_models directorie. Loads the model weights and the required DAC codec onto the GPU.
Inputs:
ckpt_name: Dropdown list of found.safetensorsfiles within yourdiffusion_modelsdirectorie. Select the file corresponding to the Dia-1.6B model.
Outputs:
dia_model: A customDIA_MODELobject containing the loaded Dia model instance, ready for theDiaGeneratenode.
Dia TTS Generate (DiaGenerate)
Generates audio using a pre-loaded Dia model provided by the DiaLoader node. Displays a progress bar during generation.
Inputs:
dia_model: TheDIA_MODELoutput from theDiaLoadernode.text: The main text transcript to generate audio for. Use[S1],[S2]for speaker turns and parentheses for non-verbals like(laughs).max_tokens: Maximum number of audio tokens to generate (controls length).cfg_scale: Classifier-Free Guidance scale.temperature: Sampling temperature.top_p: Nucleus sampling probability.cfg_filter_top_k: Top-K filtering applied during CFG.speed_factor: Adjusts the speed of the generated audio (1.0 = original speed).seed: Random seed for reproducibility.
Outputs:
audio: The generated audio (AUDIOformat:{'waveform': tensor[B,C,T], 'sample_rate': sr}), ready to be saved or previewed. Sample rate is 44100 Hz.
Usage Example
-
Add the
Dia 1.6b Loadernode from theaudio/DiaTestcategory. -
Select your Dia model file (e.g.,
dia-1.6B.safetensors) from theckpt_namedropdown. -
Add the
Dia TTS Generatenode (also fromaudio/DiaTest). -
Connect the
dia_modeloutput of the Loader node to thedia_modelinput of the Generate node. -
Enter your dialogue script into the
textinput on the Generate node.Control speaker dialogue via
[S1]and[S2]ect., tags.Add tags like
(laughs),(clears throat),(sighs),(gasps),(coughs),(singing),(sings),(mumbles),(beep),(groans),(sniffs),(claps),(screams),(inhales),(exhales),(applause),(burps),(humming),(sneezes),(chuckle),(whistles)These verbal tags will be recognized, but may result in unexpected output.
-
Adjust generation parameters on the Generate node as needed.
-
Connect the
audiooutput of the Generate node to aSaveAudioorPreviewAudionode. -
Queue the prompt.
Notes
- This node pack requires a CUDA-enabled GPU.
- Only the
.safetensorsweights file is required. - The first run of the nodes descript-audio-codec may take slightly longer. Subsequent runs will be faster.
- Dependencies
descript-audio-codecmust be installed viarequirements.txt.