Garbled, parameter-insensitive speech came from a randomized EOS head: 5.3 re-runs _initialize_weights over acoustic_connector and tts_eos_classifier because the vendored _init_weights override had no _is_hf_initialized guard. Guard it; only checkpoint-absent weights are initialized now. - MockCacheLayer exposes both the 4.x and 5.x cache APIs, so the prefilled voice prompt is visible to 5.3's mask builder. - _ensure_cache_has_layers covers the container: offload/prefetch, batch ops, crop, and a copyable lazy prefetch stream. - max_new_tokens is a combined text+speech budget, not latents. - cfg_scale floor of 1.5 for the realtime family. - voice presets resolve against every registered TTS root. - sage excluded from the realtime path: it ignores the attention mask. - bound transformers to >=5.3.0,<5.4, the measured line.
37 lines
1.5 KiB
Python
37 lines
1.5 KiB
Python
"""Custom ComfyUI types for VibeVoice nodes.
|
|
|
|
This module defines custom ComfyUI types for passing structured data
|
|
between nodes in a modular architecture.
|
|
|
|
Types:
|
|
- VibeVoiceModel: A pre-loaded VibeVoice model bundle (state dict + config +
|
|
processor + instantiated model) produced by the "Load VibeVoice Model" node
|
|
and consumed by the canonical TTS node and the ASR node via their optional
|
|
``external_model`` input. The bundle carries both standard and realtime TTS
|
|
checkpoints as well as ASR checkpoints; the consuming node dispatches on the
|
|
``is_streaming`` / ``is_asr`` flags.
|
|
|
|
The runtime value passed through the wire is a dict containing:
|
|
|
|
{
|
|
"state_dict": dict[str, torch.Tensor], # loaded weights (CPU)
|
|
"config": VibeVoiceConfig | VibeVoiceStreamingConfig,
|
|
"processor": VibeVoiceProcessor | VibeVoiceStreamingProcessor,
|
|
"model": torch.nn.Module, # instantiated model (CPU)
|
|
"model_name": str, # display name / cache key seed
|
|
"source_path": str, # original file path (for logging)
|
|
"is_streaming": bool, # realtime (streaming TTS) flag
|
|
"is_asr": bool, # ASR flag
|
|
}
|
|
|
|
This follows the same pattern as VoxCPM's ``VoiceCloningConfig`` custom type
|
|
(via ``io.Custom()``).
|
|
"""
|
|
|
|
from comfy_api.latest import io
|
|
|
|
# Create the custom type using io.Custom
|
|
VibeVoiceModel = io.Custom("VIBEVOICE_MODEL")
|
|
|
|
__all__ = ["VibeVoiceModel"]
|