Rework to use available safetensor model and remove huggingface downloads
4.4 KiB
ComfyUI DiaTest TTS Node
This node pack integrates the Nari Labs Dia text-to-speech model into ComfyUI using separate nodes for loading the model from a local file and generating audio.
Dia allows generating dialogue with speaker tags ([S1], [S2]) and non-verbal sounds ((laughs), etc.). This node pack loads the model from a local .safetensors file onto the GPU and generates audio. It requires a CUDA-enabled GPU.
Note: This version is specifically configured for the nari-labs/Dia-1.6B model architecture.
Installation
- Ensure you have a CUDA-enabled GPU and the necessary NVIDIA drivers installed.
- Download the Dia-1.6B model weights file (
dia-1.6B.safetensors) from Hugging Face Hub:- Direct Download URL: https://huggingface.co/nari-labs/Dia-1.6B/resolve/main/dia-1.6B.safetensors
- Place the downloaded
.safetensorsfile into one of the directories ComfyUI recognizes fordiffusion_models(e.g.,ComfyUI/models/diffusion_models/). You might want to rename it toDia-1.6B.safetensorsfor clarity. - Navigate to your
ComfyUI/custom_nodes/directory. - Clone this repository:
Alternatively, download the ZIP and extract it into
git clone https://github.com/BobRandomNumber/ComfyUI-DiaTest.gitcustom_nodes. - Install the required dependencies:
- Activate ComfyUI's Python environment (e.g.,
source ./venv/bin/activate). - Navigate to the node directory:
cd ComfyUI/custom_nodes/ComfyUI-DiaTest - Install requirements:
pip install -r requirements.txt
- Activate ComfyUI's Python environment (e.g.,
- Restart ComfyUI.
Nodes
Dia 1.6b Loader (DiaLoader)
Loads the Dia-1.6B TTS model from a local .safetensors checkpoint file located in one of your configured diffusion_models directories. Loads the model weights and the required DAC codec onto the GPU using the embedded Dia-1.6B configuration.
Inputs:
ckpt_name: Dropdown list of found.safetensorsfiles within yourdiffusion_modelsdirectories. Select the file corresponding to the Dia-1.6B model.
Outputs:
dia_model: A customDIA_MODELobject containing the loaded Dia model instance, ready for theDiaGeneratenode.
Dia TTS Generate (DiaGenerate)
Generates audio using a pre-loaded Dia model provided by the DiaLoader node. Displays a progress bar during generation.
Inputs:
dia_model: TheDIA_MODELoutput from theDiaLoadernode.text: The main text transcript to generate audio for. Use[S1],[S2]for speaker turns and parentheses for non-verbals like(laughs).max_tokens: Maximum number of audio tokens to generate (controls length).cfg_scale: Classifier-Free Guidance scale.temperature: Sampling temperature.top_p: Nucleus sampling probability.cfg_filter_top_k: Top-K filtering applied during CFG.speed_factor: Adjusts the speed of the generated audio (1.0 = original speed).seed: Random seed for reproducibility.
Outputs:
audio: The generated audio (AUDIOformat:{'waveform': tensor[B,C,T], 'sample_rate': sr}), ready to be saved or previewed. Sample rate is 44100 Hz.
Usage Example
- Download
dia-1.6B.safetensorsand place it in adiffusion_modelsdirectory (e.g.,ComfyUI/models/diffusion_models/). - Add the
Dia 1.6b Loadernode from theaudio/DiaTestcategory. - Select your Dia model file (e.g.,
dia-1.6B.safetensors) from theckpt_namedropdown. - Add the
Dia TTS Generatenode (also fromaudio/DiaTest). - Connect the
dia_modeloutput of the Loader node to thedia_modelinput of the Generate node. - Enter your dialogue script into the
textinput on the Generate node. - Adjust generation parameters on the Generate node as needed.
- Connect the
audiooutput of the Generate node to aSaveAudioorPreviewAudionode. - Queue the prompt.
Notes
- This node pack requires a CUDA-enabled GPU.
- Only the
.safetensorsweights file is required. - The first time you load a specific model with the
DiaLoader, it loads the weights and the required DAC codec. Subsequent runs using the same loader node will use the cached model object. Switching models in the loader clears the cache. - Dependencies
descript-audio-codec,huggingface_hub, andsafetensorsmust be installed viarequirements.txt.