Fix 7B path (vibevoice/VibeVoice-7B), add max_length_times param to prevent early cutoff

This commit is contained in:
Hawk Lee
2025-12-29 17:52:12 +08:00
parent 121bb6c61a
commit ccfa97d03b
2 changed files with 8 additions and 5 deletions
+4 -3
View File
@@ -432,16 +432,17 @@ git clone https://github.com/havvk/ComfyUI_AIIA.git
- **当前状态**: ✅ 可用(已通过测试)
- **支持语言**: **英文 (en) 和 中文 (zh)**(官方仅在这两种语言数据集上训练)
- **可选模型**:
- `microsoft/VibeVoice-1.5B`: 轻量版,适合大多数场景(~3GB 显存)
- `microsoft/VibeVoice-7B`: 高质量版,效果更好但需要更多显存(~14GB)
- `microsoft/VibeVoice-1.5B`: 轻量版,64K 上下文(~3GB 显存)
- `vibevoice/VibeVoice-7B`: 高质量版,32K 上下文(~14GB 显存)
- **特点**:
- **即时启动**: 无需预热或编译,首次运行即可使用(CosyVoice 首次需 ~1 分钟编译)。
- **语言自动识别**: 模型会自动识别中英文文本。
- **零样本音色克隆**: 输入 `reference_audio` 即可克隆声音。
- **节点参数**:
- `cfg_scale` (默认: 3.0): CFG 引导强度。值越高,语音越忠实于文本内容。
- `cfg_scale` (默认: 3.0): CFG 引导强度。值越高,语音越忠实于文本内容,**高 CFG 音质远超 CosyVoice**。
- `ddpm_steps` (默认: 50): 扩散推理步数。越高质量越好但越慢(推荐 30-100)。
- `speed` (默认: 1.0): 播放速度。>1 = 更快, <1 = 更慢(后处理时间拉伸,保持音调)。
- `max_length_times` (默认: 5.0): 最大生成长度 = 输入长度 × 此值。如音频提前结束请增大此值。
- **环境要求**:
- **Flash Attention 2**: 强烈推荐安装(否则速度较慢)。
- **Transformers**: `>= 4.51`(重要: 旧版本不支持该模型)。
+4 -2
View File
@@ -10,7 +10,7 @@ class AIIA_VibeVoice_Loader:
def INPUT_TYPES(cls):
return {
"required": {
"model_name": (["microsoft/VibeVoice-1.5B", "microsoft/VibeVoice-7B"],),
"model_name": (["microsoft/VibeVoice-1.5B", "vibevoice/VibeVoice-7B"],),
"precision": (["fp16", "bf16", "fp32"], {"default": "fp16"}),
}
}
@@ -398,6 +398,7 @@ class AIIA_VibeVoice_TTS:
"cfg_scale": ("FLOAT", {"default": 3.0, "min": 1.0, "max": 10.0, "step": 0.5, "tooltip": "CFG scale for speech generation. Higher = more faithful to text."}),
"ddpm_steps": ("INT", {"default": 50, "min": 10, "max": 100, "step": 10, "tooltip": "Diffusion steps. Higher = better quality but slower."}),
"speed": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0, "step": 0.1, "tooltip": "Playback speed. >1 = faster, <1 = slower (post-process time-stretch)."}),
"max_length_times": ("FLOAT", {"default": 5.0, "min": 2.0, "max": 20.0, "step": 1.0, "tooltip": "Max generation length = input_length × this. Increase if audio cuts off early."}),
},
"optional": {
"reference_audio": ("AUDIO",),
@@ -409,7 +410,7 @@ class AIIA_VibeVoice_TTS:
FUNCTION = "generate"
CATEGORY = "AIIA/VibeVoice"
def generate(self, vibevoice_model, text, cfg_scale, ddpm_steps, speed, reference_audio=None):
def generate(self, vibevoice_model, text, cfg_scale, ddpm_steps, speed, max_length_times, reference_audio=None):
model = vibevoice_model["model"]
tokenizer = vibevoice_model["tokenizer"]
processor = vibevoice_model.get("processor")
@@ -484,6 +485,7 @@ class AIIA_VibeVoice_TTS:
"eos_token_id": tokenizer.eos_token_id,
"pad_token_id": tokenizer.eos_token_id,
"cfg_scale": cfg_scale, # User-controlled CFG scale for speech generation
"max_length_times": max_length_times, # Control max generation length (input_length × this)
}
# Set diffusion inference steps (crucial for quality/speed tradeoff)