Files
WildAi d2390eb476 fix: correct realtime VibeVoice output on transformers 5.3
Garbled, parameter-insensitive speech came from a randomized EOS head:
5.3 re-runs _initialize_weights over acoustic_connector and
tts_eos_classifier because the vendored _init_weights override had no
_is_hf_initialized guard. Guard it; only checkpoint-absent weights are
initialized now.

- MockCacheLayer exposes both the 4.x and 5.x cache APIs, so the
  prefilled voice prompt is visible to 5.3's mask builder.
- _ensure_cache_has_layers covers the container: offload/prefetch,
  batch ops, crop, and a copyable lazy prefetch stream.
- max_new_tokens is a combined text+speech budget, not latents.
- cfg_scale floor of 1.5 for the realtime family.
- voice presets resolve against every registered TTS root.
- sage excluded from the realtime path: it ignores the attention mask.
- bound transformers to >=5.3.0,<5.4, the measured line.
2026-09-26 23:30:37 +03:00

37 lines
1.5 KiB
Python

"""Custom ComfyUI types for VibeVoice nodes.
This module defines custom ComfyUI types for passing structured data
between nodes in a modular architecture.
Types:
- VibeVoiceModel: A pre-loaded VibeVoice model bundle (state dict + config +
processor + instantiated model) produced by the "Load VibeVoice Model" node
and consumed by the canonical TTS node and the ASR node via their optional
``external_model`` input. The bundle carries both standard and realtime TTS
checkpoints as well as ASR checkpoints; the consuming node dispatches on the
``is_streaming`` / ``is_asr`` flags.
The runtime value passed through the wire is a dict containing:
{
"state_dict": dict[str, torch.Tensor], # loaded weights (CPU)
"config": VibeVoiceConfig | VibeVoiceStreamingConfig,
"processor": VibeVoiceProcessor | VibeVoiceStreamingProcessor,
"model": torch.nn.Module, # instantiated model (CPU)
"model_name": str, # display name / cache key seed
"source_path": str, # original file path (for logging)
"is_streaming": bool, # realtime (streaming TTS) flag
"is_asr": bool, # ASR flag
}
This follows the same pattern as VoxCPM's ``VoiceCloningConfig`` custom type
(via ``io.Custom()``).
"""
from comfy_api.latest import io
# Create the custom type using io.Custom
VibeVoiceModel = io.Custom("VIBEVOICE_MODEL")
__all__ = ["VibeVoiceModel"]