diff --git a/README.md b/README.md index 7abcadc..9b5b68b 100644 --- a/README.md +++ b/README.md @@ -1,38 +1,41 @@ -# ComfyUI DiaTest TTS Node +# ComfyUI Dia TTS Nodes -This node pack integrates the [Nari Labs Dia](https://github.com/nari-labs/dia) text-to-speech model into ComfyUI using separate nodes for loading the model from a local file and generating audio. +This node pack integrates the [Nari-Labs Dia](https://github.com/nari-labs/dia) 1.6b text-to-speech model into ComfyUI using the safetensors file nari-labs provided. -Dia allows generating dialogue with speaker tags (`[S1]`, `[S2]`) and non-verbal sounds (`(laughs)`, etc.). This node pack loads the model from a local `.safetensors` file onto the GPU and generates audio. It **requires a CUDA-enabled GPU**. +Dia allows generating dialogue with speaker tags (`[S1]`, `[S2]`) and non-verbal sounds (`(laughs)`, etc.). + +It **requires a CUDA-enabled GPU**. **Note:** This version is specifically configured for the `nari-labs/Dia-1.6B` model architecture. ## Installation 1. Ensure you have a CUDA-enabled GPU and the necessary NVIDIA drivers installed. -2. Download the Dia-1.6B model weights file (`dia-1.6B.safetensors`) from Hugging Face Hub: - * **Direct Download URL:** [https://huggingface.co/nari-labs/Dia-1.6B/resolve/main/dia-1.6B.safetensors](https://huggingface.co/nari-labs/Dia-1.6B/resolve/main/dia-1.6B.safetensors) -3. Place the downloaded `.safetensors` file into one of the directories ComfyUI recognizes for `diffusion_models` (e.g., `ComfyUI/models/diffusion_models/`). You might want to rename it to `Dia-1.6B.safetensors` for clarity. -4. Navigate to your `ComfyUI/custom_nodes/` directory. -5. Clone this repository: +2. Download the Dia-1.6B model safetensors file from Hugging Face: + * **Direct Download URL:** [https://huggingface.co/nari-labs/Dia-1.6B/blob/main/model.safetensors](https://huggingface.co/nari-labs/Dia-1.6B/resolve/main/model.safetensors?download=true) +3. Place the downloaded `.safetensors` file into the `diffusion_models` directory (e.g., `ComfyUI/models/diffusion_models/`). +4. You might want to rename it to `Dia-1.6B.safetensors` for clarity. +5. Navigate to your `ComfyUI/custom_nodes/` directory. +6. Clone this repository: ```bash git clone https://github.com/BobRandomNumber/ComfyUI-DiaTest.git ``` Alternatively, download the ZIP and extract it into `custom_nodes`. -6. Install the required dependencies: +7. Install the required dependencies: * Activate ComfyUI's Python environment (e.g., `source ./venv/bin/activate`). * Navigate to the node directory: `cd ComfyUI/custom_nodes/ComfyUI-DiaTest` * Install requirements: `pip install -r requirements.txt` -7. Restart ComfyUI. +8. Restart ComfyUI. ## Nodes ### Dia 1.6b Loader (`DiaLoader`) -Loads the Dia-1.6B TTS model from a local `.safetensors` checkpoint file located in one of your configured `diffusion_models` directories. Loads the model weights and the required DAC codec onto the GPU using the embedded Dia-1.6B configuration. +Loads the Dia-1.6B TTS model from a local `.safetensors` file located in your `diffusion_models` directorie. Loads the model weights and the required DAC codec onto the GPU. **Inputs:** -* `ckpt_name`: Dropdown list of found `.safetensors` files within your `diffusion_models` directories. Select the file corresponding to the Dia-1.6B model. +* `ckpt_name`: Dropdown list of found `.safetensors` files within your `diffusion_models` directorie. Select the file corresponding to the Dia-1.6B model. **Outputs:** @@ -60,12 +63,18 @@ Generates audio using a pre-loaded Dia model provided by the `DiaLoader` node. D ## Usage Example -1. Download `dia-1.6B.safetensors` and place it in a `diffusion_models` directory (e.g., `ComfyUI/models/diffusion_models/`). -2. Add the `Dia 1.6b Loader` node from the `audio/DiaTest` category. -3. Select your Dia model file (e.g., `dia-1.6B.safetensors`) from the `ckpt_name` dropdown. -4. Add the `Dia TTS Generate` node (also from `audio/DiaTest`). -5. Connect the `dia_model` output of the Loader node to the `dia_model` input of the Generate node. -6. Enter your dialogue script into the `text` input on the Generate node. +1. Add the `Dia 1.6b Loader` node from the `audio/DiaTest` category. +2. Select your Dia model file (e.g., `dia-1.6B.safetensors`) from the `ckpt_name` dropdown. +3. Add the `Dia TTS Generate` node (also from `audio/DiaTest`). +4. Connect the `dia_model` output of the Loader node to the `dia_model` input of the Generate node. +5. Enter your dialogue script into the `text` input on the Generate node. + + Control speaker dialogue via `[S1]` and `[S2]` ect., tags. + + Add tags like `(laughs)`, `(clears throat)`, `(sighs)`, `(gasps)`, `(coughs)`, `(singing)`, `(sings)`, `(mumbles)`, `(beep)`, `(groans)`, `(sniffs)`, `(claps)`, `(screams)`, `(inhales)`, `(exhales)`, `(applause)`, `(burps)`, `(humming)`, `(sneezes)`, `(chuckle)`, `(whistles)` + + These verbal tags will be recognized, but may result in unexpected output. + 7. Adjust generation parameters on the Generate node as needed. 8. Connect the `audio` output of the Generate node to a `SaveAudio` or `PreviewAudio` node. 9. Queue the prompt. @@ -74,5 +83,5 @@ Generates audio using a pre-loaded Dia model provided by the `DiaLoader` node. D * This node pack **requires a CUDA-enabled GPU**. * Only the `.safetensors` weights file is required. -* The first time you load a specific model with the `DiaLoader`, it loads the weights and the required DAC codec. Subsequent runs using the same loader node will use the cached model object. Switching models in the loader clears the cache. -* Dependencies `descript-audio-codec`, `huggingface_hub`, and `safetensors` must be installed via `requirements.txt`. \ No newline at end of file +* The first run of the nodes descript-audio-codec may take slightly longer. Subsequent runs will be faster. +* Dependencies `descript-audio-codec` must be installed via `requirements.txt`.