WildAi b4fce48cff feat: report standard ComfyUI progress during TTS/ASR inference (v2.1.1)
Vendored generate() (non-streaming + streaming) gains an optional framework-agnostic progress_callback(current, total) fired once per AR loop step. modules/generation.py wraps it with comfy.utils.ProgressBar (interrupt check per step, dynamic total via update_absolute, guaranteed final 100%% in finally). ASR reports per-token progress via an HF BaseStreamer (greedy/sampling only; beam search falls back to 0->100%%). 25 new tests; full regression 560p/5f(pre-existing)/4s, zero new failures.
2026-08-16 02:24:52 +03:00
2025-08-27 14:49:07 +03:00
2026-08-14 20:26:31 +03:00
2025-08-27 14:42:07 +03:00
2026-08-14 20:26:31 +03:00
2026-08-14 20:26:31 +03:00

ComfyUI-VibeVoice

ComfyUI-VibeVoice Nodes

A custom node for ComfyUI that integrates Microsoft's VibeVoice, a frontier model for generating expressive, long-form, multi-speaker conversational audio.

Report Bug · Request Feature

Stargazers Issues Contributors Forks

About The Project

VibeVoice is a novel framework by Microsoft for generating expressive, long-form, multi-speaker conversational audio. It excels at creating natural-sounding dialogue, podcasts, and more, with consistent voices for up to 4 speakers.

ComfyUI-VibeVoice example workflow

The custom node handles everything from model downloading and memory management to audio processing, allowing you to generate high-quality speech directly from a text script and reference audio files.

✨ Key Features:

  • Multi-Speaker TTS: Generate conversations with up to 4 distinct voices in a single audio output.
  • High-Fidelity Voice Cloning: Use any audio file (.wav, .mp3) as a reference for a speaker's voice.
  • Hybrid Voice Cloning: Mix and match cloned speakers in the same script — at least one speaker_*_voice reference audio is required; other speakers are cloned from the provided reference(s).
  • Flexible Scripting: Use simple [1] tags or the classic Speaker 1: format to write your dialogue.
  • Advanced Attention Mechanisms: Choose between eager, sdpa, flash_attention_2, and the high-performance sage attention for fine-tuned control over speed and compatibility.
  • Robust 4-Bit Quantization: Run the large language model component in 4-bit mode to significantly reduce VRAM usage.
  • Automatic Model Management: Models are downloaded automatically and managed efficiently by ComfyUI to save VRAM.

(back to top)

🚀 Getting Started

The easiest way to install is through the ComfyUI Manager:

  1. Go to Manager -> Install Custom Nodes.
  2. Search for ComfyUI-VibeVoice and click "Install".
  3. Restart ComfyUI.

Alternatively, to install manually:

  1. Clone the Repository: Navigate to your ComfyUI/custom_nodes/ directory and clone this repository:

    git clone https://github.com/wildminder/ComfyUI-VibeVoice.git
    
  2. Install Dependencies: Open a terminal or command prompt, navigate into the cloned directory, and install the required Python packages. For quantization support, you must install bitsandbytes.

    cd ComfyUI-VibeVoice
    pip install -r requirements.txt
    
  3. Optional: Install SageAttention To enable the sage attention mode, you must install the sageattention library. For Windows users, a pre-compiled wheel is available at AI-windows-whl.

    Note: This is only required if you intend to use the sage attention mode.

Audio backend: Audio resampling uses torchaudio (the ComfyUI-core idiom) as the primary library; file decoding uses PyAV (av, bundled with ComfyUI) with soundfile as fallback. librosa is not required — it is only an optional last-resort fallback, and the node degrades gracefully when it (or scipy) is absent. To install the optional fallbacks: pip install scipy librosa (or pip install ComfyUI-VibeVoice[audio-extra]).

  1. Start/Restart ComfyUI: Launch ComfyUI. The "VibeVoice TTS" node will appear under the audio/tts category. The first time you use the node, it will automatically download the selected model to your ComfyUI/models/tts/VibeVoice/ folder.

Models

Model Context Length Generation Length Weight
VibeVoice-1.5B 64K ~90 min HF link
VibeVoice-Large 32K ~45 min HF link

(back to top)

🛠️ Usage

The node is designed for maximum flexibility within your ComfyUI workflow.

  1. Add Nodes: Add the VibeVoice TTS node to your graph. Use ComfyUI's built-in Load Audio node to load your reference voice files.
  2. Connect Voices (Optional): Connect the AUDIO output from each Load Audio node to the corresponding speaker_*_voice input.
  3. Write Your Script: In the text input, write your dialogue using one of the supported formats.
  4. Generate: Queue the prompt. The node will process the script and generate a single audio file containing the full conversation.

Tip: For a complete workflow, you can drag the example image from the example_workflows folder onto your ComfyUI canvas.

Scripting and Voice Modes

Speaker Tagging

You can assign lines to speakers in two ways. Both are treated identically.

  • Modern Format (Recommended): [1] This is the first speaker.
  • Classic Format: Speaker 1: This is the first speaker.

You can also add an optional colon to the modern format (e.g., [1]: ...). The node handles all variations consistently.

Hybrid Voice Generation

This is a powerful feature that lets you mix cloned voices in the same script. At least one speaker_*_voice reference audio is required — VibeVoice anchors the timbre of every speaker to a provided reference, so you cannot generate a fully reference-free ("zero-shot") voice for any speaker.

  • To Clone a Voice: Connect a Load Audio node to the speaker's input (e.g., speaker_1_voice).
  • To Reuse a Cloned Voice: Any other speaker may be left empty. When a speaker has no reference, its voice is cloned from the provided reference(s) rather than generated from scratch.

Example Hybrid Script:

[1] This line will use the audio from speaker_1_voice.
[2] This line will reuse the cloned voice from speaker 1.
[1] I'm back with my cloned voice.

In this example, you would connect an audio source to speaker_1_voice; speakers [2] are cloned from it.

Node Inputs

  • model_name: Select the VibeVoice model to use (1.5B or Large).
  • text: The conversational script. See "Scripting and Voice Modes" above for formatting.
  • quantize_llm_4bit: Enable to run the LLM component in 4-bit (NF4) mode, dramatically reducing VRAM usage.
  • attention_mode: Select the attention implementation: eager (safest), sdpa (balanced), flash_attention_2 (fastest), or sage (quantized high-performance).
  • cfg_scale: Controls how strongly the model adheres to the reference voice's timbre. Higher values are stricter. Recommended: 1.3.
  • inference_steps: Number of diffusion steps for audio generation. Recommended: 10.
  • seed: A seed for reproducibility. Set to 0 for a random seed on each run.
  • do_sample, temperature, top_p, top_k: Standard sampling parameters for controlling the creativity and determinism of the speech generation.
  • force_offload: Forces the model to be completely offloaded from VRAM after generation.

⚙️ Performance & Advanced Features

This node features a sophisticated system for managing performance, memory, and stability.

Feature Compatibility & VRAM Matrix

Quantize LLM Attention Mode Behavior / Notes Relative VRAM
OFF eager Full Precision. Most compatible baseline. High
OFF sdpa Full Precision. Recommended for balanced performance. High
OFF flash_attention_2 Full Precision. High performance on compatible GPUs. High
OFF sage Full Precision. Uses high-performance mixed-precision kernels. High
ON eager Falls back to sdpa with bfloat16 compute. Warns user. Low
ON sdpa Recommended for memory savings. Uses bfloat16 compute. Low
ON flash_attention_2 Falls back to sdpa with bfloat16 compute. Warns user. Low
ON sage Recommended for stability. Uses fp32 compute to ensure numerical stability with quantization, resulting in slightly higher VRAM usage. Medium

Changelog

v2.1.1 - Standard ComfyUI Progress Bar During Inference

✨ Highlights

  • Live progress bar: All three nodes (TTS, Realtime TTS, ASR) now drive the standard ComfyUI frontend progress bar during inference. Previously the bar sat at 0% for the whole generation and jumped to 100% only at the end.
  • Responsive cancel: the progress hook checks ComfyUI's interrupt flag on every loop step, so pressing cancel stops generation promptly instead of waiting for the current blocking call.
  • Guaranteed 100%: a final progress event is always emitted, even when generation stops early (EOS) or raises.

🔧 Changes

  • Vendored generate() (non-streaming + streaming) gained an optional, framework-agnostic progress_callback(current, total) hook fired once per AR loop step (vendored code stays comfy-free; None = disabled, fully backward compatible).
  • modules/generation.py: generate_audio() / generate_streaming_audio() wrap the hook with comfy.utils.ProgressBar (throttled WebSocket updates; the bar total self-corrects via update_absolute(value, total=...) once the loop reports its real budget).
  • modules/asr_generation.py: ASR reports per-token progress through an HF BaseStreamer (greedy/sampling only; beam search falls back to a single 0→100% bar).

🧪 Tests

  • New tests/test_generate_progress_callback.py (5 tests) and tests/test_streaming_progress_callback.py (4 tests): drive the real vendored loops with scripted mocks and lock the callback contract (monotonic, bounded, call counts, interrupt propagation, output determinism).
  • New tests/test_streaming_progress.py (4 tests); extended tests/test_generation.py (+5), tests/test_asr_generation.py (+6), tests/test_integration.py (+3 node-level tests).
v2.1.0 - torchaudio-Primary Audio Backend (librosa now optional)

✨ Highlights

  • torchaudio is now the primary audio library. Resampling uses torchaudio.functional.resample (Kaiser-windowed sinc — the same family ComfyUI core uses), matching the ComfyUI-core idiom.
  • librosa is no longer a hard dependency. It was declared but never actually imported (a phantom dependency). It is now an optional last-resort fallback; the node never crashes when librosa (or scipy) is missing or broken (e.g. the empty librosa namespace stub shipped in some embedded Pythons).
  • ComfyUI built-ins adopted: file decoding now uses PyAV (av) — ComfyUI's own audio decoder — with soundfile/torchaudio/librosa as guarded fallbacks. This also fixes .m4a/.ogg reference-audio loading, which soundfile (libsndfile) cannot decode.
  • Dependency surface minimized: librosa and scipy moved to the optional [audio-extra] group; torchaudio + soundfile remain the working default.

🔧 Changes

  • New modules/audio_backend.py: single dependency-resilient backend for resample / load / save with import-time capability detection (_HAS_* flags) and graceful fallback ordering.
  • modules/audio_utils.py: resample_audio() delegates to the backend; preprocess_comfy_audio() now resamples in tensor space (no numpy round-trip).
  • Vendored processors (_load_audio_from_path, save_audio, ASR file loading) route through the backend; no hard ffmpeg/soundfile requirement for the common wav/flac path.

🧪 Tests

  • New tests/test_audio_backend.py (41 tests): import resilience with any optional lib blocked, resample correctness/priority/fallbacks, numpy↔tensor parity, file I/O roundtrips, f32_pcm.
  • New tests/test_processor_io_backend.py (17 tests): real vendored processor I/O via the backend.
  • Extended tests/test_audio_utils.py, tests/test_pyproject.py, tests/test_imports.py.
v2.0.2 - Negative-Branch RoPE Position Fix (SDPA Shape Crash)

🐛 Fixes

  • BUG-011 — Negative-branch RoPE position_ids desync: Fixed a RuntimeError: Expected size for first two dimensions of batch2 tensor to be: [12, 3] but got: [12, 2] crash during TTS generation. In the non-streaming generate() CFG loop, the negative (unconditional) forward fed a single-token inputs_embeds (B,1,H) together with a full-length position_ids (B, step+1). In transformers 5.x the explicit position_ids drive RoPE directly, so q/k silently broadcast against full-length cos/sin and expanded to seq-len step+1 while v (never rotated) stayed length 1 — the KV cache accumulated step+1 keys but only 1 value per step, and SDPA's attn @ value crashed at AR step 1. The negative forward now passes current-only position_ids (neg_position_ids[:, -1:]), matching its single input token; the attention mask stays full-length.

🧪 Tests

  • New tests/test_generate_neg_position_ids.py (4 tests): drives the real generate() through several AR steps with a recording inner LM and locks the invariant that position_ids length always equals inputs_embeds sequence length (red/green verified against the buggy code).
v2.0.1 - CPU-First Model Loading (VRAM Round-Trip Fix)

🐛 Fixes

  • Load Device Flow — DF-001..DF-006: Fixed a GPU→RAM→GPU round-trip during model loading. Previously the checkpoint state dict was loaded directly onto CUDA (a full-model VRAM spike outside ComfyUI's arbitration), copied back to CPU-resident parameters, then moved to CUDA again. Now the loader builds the model entirely on CPU (state dict, load_state_dict, dtype cast, and 4-bit quantization all on CPU), and VibeVoicePatcher.patch_model owns the single host-to-device transfer after ComfyUI's load_models_gpu VRAM arbitration. Peak VRAM during load drops from ≈2× model size to ≈1× model size.
  • Dtype Threading — DF-004 / AUD-008: The user-selected dtype is now threaded from the node through the handler into the loader and applied on CPU before the transfer; the patcher's dtype cast is now a mismatch-only guard (no redundant GPU cast).
  • Handler No-Move — DF-003: VibeVoiceModelHandler.load_model no longer moves the model; device placement is owned solely by the patcher.

🧪 Tests

  • New tests/test_load_device_flow.py (36 tests): device-ledger doubles asserting CPU-only loading, single H2D transfer, no round-trip, dtype threading, cast guard, VRAM arbitration, and a [GPU-OPTIONAL] peak-VRAM measurement.
v2.0.0 - V3 Extension & VRAM Parity

✨ Highlights

  • V3 Extension API: Migrated the custom-node entrypoint to the ComfyUI V3 ComfyExtension / io.ComfyNode schema — type-filtered model dropdowns and declarative inputs/outputs.
  • VRAM Parity (ASR) — CRIT-001: The ASR path now runs under the same VibeVoicePatcher / model_management.load_model_gpu orchestration as TTS, clearing the dedicated ASR cache on unload.
  • Warm Re-attach — NTH-004: force_offload can retain model tensors on the intermediate device for a fast re-attach on the next run instead of reloading from disk.
  • Streaming TTS Node — NTH-001: VibeVoice-Realtime-0.5B is now reachable through a dedicated VibeVoice Realtime TTS node that shares the patcher / attention machinery.
  • Maintainability — IMP-004: TTS/ASR download, discovery, and sharded-load logic is now shared via BaseVibeVoiceLoader.
  • Device & Attention Honesty — IMP-003 / IMP-001: MPS/XPU/NPU device selection is honored when available; flash_attention_2 is only offered when flash-attn + CUDA are present.
  • Docs Consistency — CRIT-003: README zero-shot wording now matches generate_audio (at least one reference voice is required).
v1.5.0 - Stability and Prompting

✨ New Features & Improvements

  • Total Generation Stability: Fixed the bug where a speaker's voice could unintentionally change or blend with another reference voice mid-sentence.
  • Improved Voice Cloning Fidelity
  • Consistent Speaker Tagging: The node now intelligently handles multiple script formats ([1], [1]:, and Speaker 1:) to produce identical, high-quality results, removing all previous inconsistencies.
  • Hybrid Voice Cloning: Mix and match cloned speakers in the same script — at least one reference audio is required; speakers without their own reference are cloned from the provided reference(s).
v1.3.0 - SageAttention & Quantization Overhaul
  • SageAttention Support: Full integration with the sageattention library for a high-performance, mixed-precision attention option.
  • Robust 4-Bit LLM Quantization: The "Quantize LLM (4-bit)" option is now highly stable and delivers significant VRAM savings.
  • Smart Configuration & Fallbacks: The node now automatically handles incompatible settings (e.g., 4-bit with flash_attention_2) by gracefully falling back to a stable alternative (sdpa) and notifying the user.
v1.2.0 - Compatibility Update
  • Transformers Library: Includes automatic detection and compatibility for both older and newer versions of the Transformers library (pre- and post-4.56).
  • Bug Fixes: Resolved issues with Force Offload and multi-speaker generation on newer Transformers versions.

(back to top)

Tips from the Original Authors

  • Punctuation: For Chinese text, using English punctuation (commas and periods) can improve stability.
  • Model Choice: The 7B model variant (VibeVoice-Large) is generally more stable.
  • Spontaneous Sounds/Music: The model may spontaneously generate background music, especially if the reference audio contains it or if the text includes introductory phrases like "Welcome to...". This is an emergent capability and cannot be directly controlled.
  • Singing: The model was not trained on singing data, but it may attempt to sing as an emergent behavior. Results may vary.

(back to top)

License

This project is distributed under the MIT License. See LICENSE.txt for more information. The VibeVoice model and its components are subject to the licenses provided by Microsoft. Please use responsibly.

(back to top)

Acknowledgments

  • Microsoft for creating and open-sourcing the VibeVoice project.
  • The ComfyUI team for their incredible and extensible platform.

(back to top)

Star History

Star History Chart

S
Description
No description provided
Readme MIT
4.4 MiB
Languages
Python 100%