diff --git a/README.md b/README.md index 106c78b..12883de 100644 --- a/README.md +++ b/README.md @@ -1,7 +1,64 @@ -# ComfyUI-EdgeTTS +# ComfyUI Audio Nodes + ComfyUI-EdgeTTS is a powerful text-to-speech node for ComfyUI, leveraging Microsoft's Edge TTS capabilities. It enables seamless conversion of text into natural-sounding speech, supporting multiple languages and voices. Ideal for enhancing user interactions, this node is easy to integrate and customize, making it perfect for various applications. -## This custom node will be available soon! Follow us and leave a ⭐ to stay updated – it’s coming your way shortly. +## Features + +### Edge TTS Node +- **Edge TTS**: Convert text to speech using Microsoft Edge TTS + - Multiple languages and voices support + - Adjustable speech rate and pitch + - High-quality voice synthesis + - Configurable via config.json + +### Speech to Text Node +- **Whisper STT**: High-accuracy speech recognition + - Multiple language support with auto-detection + - Multiple model sizes (tiny to large) + - Supports ComfyUI audio format + - Language detection confidence reporting + +### Audio File Node +- **Save Audio**: Export audio files + - Supports WAV, MP3, FLAC formats + - Quality presets (high/medium/low) + - Custom file naming and paths + - Automatic file numbering + +## Installation + +### Method 1. install on ComfyUI-Manager, search `Comfyui-EdgeTTS` and install +install requirment.txt in the ComfyUI-EdgeTTS folder + ```bash + ./ComfyUI/python_embeded/python -m pip install -r requirements.txt + ``` + +### Method 2. Clone this repository to your ComfyUI custom_nodes folder: + ```bash + cd ComfyUI/custom_nodes + git clone https://github.com/1038lab/ComfyUI-EdgeTTS.git + ``` + install requirment.txt in the ComfyUI-EdgeTTS folder + ```bash + ./ComfyUI/python_embeded/python -m pip install -r requirements.txt + ``` +## Requirements +- Python packages (see requirements.txt) +- CUDA compatible GPU (optional, for faster Whisper processing) + +## Usage Examples + +### Text to Speech +1. Add Edge TTS node to workflow +2. Input text and select voice +3. Adjust speed and pitch if needed +4. Connect to Save Audio node for export + +### Speech to Text +1. Add Whisper STT node +2. Connect audio input +3. Select model size and language (or auto-detect) +4. Run to get transcription ## Supported Voices @@ -28,3 +85,7 @@ ComfyUI-EdgeTTS is a powerful text-to-speech node for ComfyUI, leveraging Micros | Indonesian | Gadis (Warm) | Ardi (Formal) | Each language provides at least one male and female voice option, allowing you to choose different voice styles based on your needs. + +## Credits +- Edge TTS: [Microsoft Edge TTS](https://github.com/rany2/edge-tts) +- Whisper: [OpenAI Whisper](https://github.com/openai/whisper)