diff --git a/README.md b/README.md index e86ef3c..b9510ae 100755 --- a/README.md +++ b/README.md @@ -971,11 +971,35 @@ hf download digital-avatar/ditto-talkinghead --local-dir ditto | **克隆能力 (Cloning)** | **SOTA** (Zero-Shot)`
`只需 3-10秒,对**音色质感**还原极高。 | **SOTA** (稳定性)`
`对**说话韵律/口音**的捕捉最准。 | **良好** `
`适合克隆特定语气,而非纯粹音色。 | | **多语言/方言** | **中/英** (双语优化) | **👑 霸主** (9种语言 + 18种方言) | **中/英** | | **语音转换 (VC)** (Audio-to-Audio) | ❌**不支持** `
`仅支持 TTS (Text-to-Speech)。无法改变已有音频的音色。 | ✅**支持** `
`可以将任意音频转换为任意音色 (保留语调/停顿)。 | ❌**不支持** `
`纯 TTS 模型。仅支持 Text-to-Speech。 | +| **Qwen3-TTS** (1.7B/0.6B) | ✅**支持** `
`支持 CustomVoice (内置) 和 VoiceDesign (描述)。 | ✅**支持** `
`支持 3秒极速 Zero-shot 克隆。 | ✅**支持** `
`支持 10 种语言。 | + +#### 3.13 Qwen3-TTS (New! 🔥) + +- **用途**: 阿里巴巴 Qwen 团队推出的最新旗舰级 TTS 模型,支持 10 种主要语言及多种方言,具备极高的稳定性和表现力。 +- **核心能力**: + - **CustomVoice**: 使用内置的高品质音色进行语音合成。提供 1.7B 和 0.6B 两种规格。 + - **VoiceDesign**: 通过自然语言描述(如“活泼的少女音,带点羞涩”)从零设计音色。 + - **VoiceClone**: 顶级的 3秒快速音色克隆,支持 X-Vector 模式提升稳定性。 +- **环境要求**: + - **qwen-tts**: `pip install qwen-tts` (插件会自动尝试安装)。 + - **Flash Attention 2**: 强烈推荐以获得最佳推理性能。 +- **节点**: + - `🤖 Qwen3-TTS Loader`: 加载模型。支持 `Base` (克隆)、`CustomVoice` (内置音色) 和 `VoiceDesign` (音色设计) 模型。 + - `🗣️ Qwen3-TTS Synthesis`: 执行合成。根据加载的模型类型自动切换功能。 +- **模型下载**: + - `Qwen/Qwen3-TTS-12Hz-1.7B-Base` (或 0.6B-Base) + - `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice` (或 0.6B-CustomVoice) + - `Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign` + +--- + +#### 💡 用户实测与选型指南 (Model Comparison & Selection) **选型建议**: - **追求“听起来最像真人” (音质+音色)**: 选 **VoxCPM 1.5**。它的 Tokenizer-free 架构带来了质的飞跃。 - **追求“方言/多语言/稳定性”**: 选 **CosyVoice 3.0**。目前依然是生产环境最稳的选择。 +- **追求“多样化音色设计/最新 Qwen 生态/长语音流畅度”**: 选 **Qwen3-TTS**。其 VoiceDesign 功能能让你用描述语“捏”出从未听过的声音。 - **要做“长篇广播剧/播客”**: 选 **VibeVoice**。它的长窗口上下文优势依然不可替代。 ### 4. 播客与对话生成 (Podcast & Dialogue Generation) @@ -1005,6 +1029,7 @@ https://github.com/user-attachments/assets/9a5502c5-79e3-4fc8-8a2d-2cbdbdbbc860 - **TTS Engine**: 后端引擎选择。 - **CosyVoice**: 精准控制型。 - **VibeVoice**: 自然演绎型。 + - **Qwen3-TTS**: 万能旗舰型。支持音色设计与内置高质量音色。 - **Speaker A/B/C**: - **Ref Audio**: 参考音频 (用于 Zero-Shot 克隆)。 - **ID**: 内置音色 ID (如 CosyVoice 的 `Chinese Female`)。 @@ -1026,6 +1051,7 @@ https://github.com/user-attachments/assets/9a5502c5-79e3-4fc8-8a2d-2cbdbdbbc860 - **原理**: - **CosyVoice**: 使用生成时的精确时长。 - **VibeVoice**: 使用**智能插值算法 (Smart Interpolation)**,根据字符长度自动计算长音频段内的单句时间轴。 + - **Qwen3-TTS**: 基于生成的音频振幅精准断句,支持多角色时间轴导出。 #### 4.4 AIIA Subtitle to Segments (字幕转分段) @@ -1069,13 +1095,13 @@ https://github.com/user-attachments/assets/9a5502c5-79e3-4fc8-8a2d-2cbdbdbbc860 #### 💡 引擎选型与最佳实践 (Best Practices) -| 特性 | **CosyVoice** | **VibeVoice** | -| :----------------- | :----------------------------------------- | :----------------------------------------------------------------------------------------------- | -| **核心优势** | **精准控制 (Instruction)** | **自然演绎 (Context-Aware)** | -| **情感控制** | ✅**支持** (使用 `[Happy]` 等标签) | ❌ 不支持显式标签 (依赖上下文) | -| **生成逻辑** | **逐句生成** (严格遵循每句话的指令) | **混合批处理** (Hybrid Batching) | -| **最佳场景** | 需要精确指定某句话语气、方言时 | 长篇对话、广播剧、闲聊 | -| **使用建议** | 可以在剧本中详细标注情感。 | **尽量减少 `(Pause)`**!`
`让多句对话连在一起,模型能更好地联系上下文产生自然语气。 | +| 特性 | **CosyVoice** | **VibeVoice** | **Qwen3-TTS** | +| :----------------- | :----------------------------------------- | :----------------------------------------------------------------------------------------------- | :------------------------------------------ | +| **核心优势** | **精准控制 (Instruction)** | **自然演绎 (Context-Aware)** | **万能旗舰 (Voice Design)** | +| **情感控制** | ✅**支持** (使用 `[Happy]` 等标签) | ❌ 不支持显式标签 (依赖上下文) | ✅**支持** (通过 `instruct` 或标签) | +| **生成逻辑** | **逐句生成** (严格遵循每句话的指令) | **混合批处理** (Hybrid Batching) | **动态引擎** (支持流式与批处理) | +| **最佳场景** | 需要精确指定某句话语气、方言时 | 长篇对话、广播剧、闲聊 | 音色定制、高质量配音、极速克隆 | +| **使用建议** | 可以在剧本中详细标注情感。 | **尽量减少 `(Pause)`**!`
`让多句对话连在一起,模型能更好地联系上下文产生自然语气。 | 尝试使用其 Voice Design 进行创意捏人。 | #### 📝 综合测试剧本 (Example Script) @@ -1175,6 +1201,15 @@ B: 太神奇了!那我们快去生成试试吧! ## Changelog +### [1.11.0] - 2026-02-04 + +- **Qwen3-TTS**: 新增阿里巴巴 **Qwen3-TTS** 全系列支持。 + - **🤖 Qwen3-TTS Loader**: 支持加载 Base, CustomVoice, VoiceDesign 及其 1.7B/0.6B 版本。 + - **🗣️ Qwen3-TTS Synthesis**: 实现全功能生成,包括 Zero-shot 克隆、音色设计和内置音色合成。 +- **Podcast Integration**: **AIIA Dialogue TTS** 节点现在正式集成 Qwen3-TTS 引擎。 + - 支持多角色混合场景下的 Qwen3 驱动,支持使用脚本标签触发 `instruct`。 +- **Auto-Dependency**: 首次运行 Qwen3 节点会自动检测并安装 `qwen-tts` 库。 + ### [1.10.17] - 2026-02-03 - **Subtitle**: 引入“说话人 ID 为了映射 (Speaker Mapping)”机制。 diff --git a/__init__.py b/__init__.py index f72fcf8..658b93d 100755 --- a/__init__.py +++ b/__init__.py @@ -136,6 +136,9 @@ else: # 25. 处理 aiia_debug_nodes.py (新增调试节点) _load_nodes_from_module(".aiia_debug_nodes", "aiia_debug_nodes") + # 26. 处理 aiia_qwen_nodes.py (新增 Qwen3-TTS) + _load_nodes_from_module(".aiia_qwen_nodes", "aiia_qwen_nodes") + # 告诉 ComfyUI 这个节点包有一个包含网页资源的 'js' 目录 WEB_DIRECTORY = "js" diff --git a/aiia_podcast_nodes.py b/aiia_podcast_nodes.py index bbed8ee..ba8eb26 100755 --- a/aiia_podcast_nodes.py +++ b/aiia_podcast_nodes.py @@ -176,7 +176,7 @@ class AIIA_Dialogue_TTS: return { "required": { "dialogue_json": ("STRING", {"forceInput": True}), - "tts_engine": (["CosyVoice", "VibeVoice"], {"default": "CosyVoice"}), + "tts_engine": (["CosyVoice", "VibeVoice", "Qwen3-TTS"], {"default": "CosyVoice"}), "pause_duration": ("FLOAT", {"default": 0.5, "min": 0.0, "max": 5.0, "step": 0.1}), "speed_global": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0}), "batch_mode": (["Natural (Hybrid)", "Strict (Per-Speaker)", "Whole (Single Batch)"], {"default": "Natural (Hybrid)"}), @@ -190,6 +190,7 @@ class AIIA_Dialogue_TTS: "optional": { "cosyvoice_model": ("COSYVOICE_MODEL",), "vibevoice_model": ("VIBEVOICE_MODEL",), + "qwen_model": ("QWEN_MODEL",), # Speaker A "speaker_A_ref": ("AUDIO",), @@ -248,7 +249,7 @@ class AIIA_Dialogue_TTS: return None def process_dialogue(self, dialogue_json, tts_engine, pause_duration, speed_global, - cosyvoice_model=None, vibevoice_model=None, + cosyvoice_model=None, vibevoice_model=None, qwen_model=None, cfg_scale=1.5, temperature=0.8, top_k=20, top_p=0.95, **kwargs): import json import torch @@ -260,6 +261,8 @@ class AIIA_Dialogue_TTS: raise ValueError("选择 CosyVoice 引擎时,必须连接 'cosyvoice_model'!") if tts_engine == "VibeVoice" and vibevoice_model is None: raise ValueError("选择 VibeVoice 引擎时,必须连接 'vibevoice_model'!") + if tts_engine == "Qwen3-TTS" and qwen_model is None: + raise ValueError("选择 Qwen3-TTS 引擎时,必须连接 'qwen_model'!") dialogue = json.loads(dialogue_json) full_waveform = [] @@ -267,9 +270,11 @@ class AIIA_Dialogue_TTS: from .aiia_cosyvoice_nodes import AIIA_CosyVoice_TTS from .aiia_vibevoice_nodes import AIIA_VibeVoice_TTS + from .aiia_qwen_nodes import AIIA_Qwen_TTS cosy_gen = AIIA_CosyVoice_TTS() vibe_gen = AIIA_VibeVoice_TTS() + qwen_gen = AIIA_Qwen_TTS() print(f"[AIIA Podcast] 开始处理对话,共 {len(dialogue)} 个片段。引擎: {tts_engine}") @@ -402,6 +407,74 @@ class AIIA_Dialogue_TTS: }) time_ptr[0] += 1.0 + elif tts_engine == "Qwen3-TTS": + # Qwen3-TTS (Iterative) + for i, item in enumerate(batch_items): + spk_name = item["speaker"] + spk_key = get_speaker_key(spk_name) + text = item["text"] + emotion = item.get("emotion", "None") + + # Mapping logic for Qwen + spk_id = kwargs.get(f"speaker_{spk_key}_id", "Vivian") # Default to Vivian if empty + if not spk_id.strip(): spk_id = "Vivian" + + ref_audio = get_ref_audio(spk_key) + instruct = f"{emotion}." if emotion and emotion != "None" else "" + + print(f" [Qwen Processing] {spk_name} (ID: {spk_id}): {text[:15]}...") + try: + # Call Qwen TTS + res = qwen_gen.generate( + qwen_model=qwen_model, + text=text, + language="Auto", + speaker=spk_id, + instruct=instruct, + reference_audio=ref_audio, + seed=42+i, + speed=speed_global + ) + + generated = res[0] + wav = generated["waveform"] + sr = generated["sample_rate"] + + if sr_ptr[0] != sr: + if current_full_wav: + wav = torchaudio.transforms.Resample(sr, sr_ptr[0])(wav) + else: + sr_ptr[0] = sr + + if wav.ndim == 3: wav = wav.squeeze(0) + if wav.ndim == 1: wav = wav.unsqueeze(0) + current_full_wav.append(wav) + + # --- Timestamp Tracking --- + seg_duration = wav.shape[-1] / sr + seg_start = time_ptr[0] + seg_end = seg_start + seg_duration + + segments_info.append({ + "start": round(seg_start, 3), + "end": round(seg_end, 3), + "text": text, + "speaker": spk_name, + "visual": item.get("visual") + }) + time_ptr[0] += seg_duration + + # Add a small gap between segments + gap = 0.2 + gap_samples = int(gap * sr_ptr[0]) + current_full_wav.append(torch.zeros(1, gap_samples)) + time_ptr[0] += gap + + except Exception as e: + print(f"[Error] Qwen item generation failed: {e}") + current_full_wav.append(torch.zeros(1, 24000)) + time_ptr[0] += 1.0 + else: # CosyVoice (Iterative) for i, item in enumerate(batch_items): diff --git a/aiia_qwen_nodes.py b/aiia_qwen_nodes.py new file mode 100644 index 0000000..ac12239 --- /dev/null +++ b/aiia_qwen_nodes.py @@ -0,0 +1,189 @@ +import os +import sys +import torch +import torchaudio +import numpy as np +import folder_paths +import subprocess + +# --- Robust Package Installation --- +def _install_qwen_tts_if_needed(): + try: + from qwen_tts import Qwen3TTSModel + return + except ImportError: + print("[AIIA] qwen-tts missing. Attempting installation...") + try: + subprocess.check_call([sys.executable, "-m", "pip", "install", "-U", "qwen-tts"]) + print("[AIIA] qwen-tts installed successfully.") + except Exception as e: + print(f"[AIIA] Failed to install qwen-tts: {e}") + +class AIIA_Qwen_Loader: + @classmethod + def INPUT_TYPES(s): + return { + "required": { + "model_name": ([ + "Qwen/Qwen3-TTS-12Hz-1.7B-Base", + "Qwen/Qwen3-TTS-12Hz-0.6B-Base", + "Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice", + "Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice", + "Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign" + ], {"default": "Qwen/Qwen3-TTS-12Hz-1.7B-Base"}), + "device": (["cuda", "cpu", "auto", "mps"], {"default": "auto"}), + "dtype": (["bf16", "fp16", "fp32"], {"default": "bf16"}), + }, + "optional": { + "local_path": ("STRING", {"default": ""}), + } + } + + RETURN_TYPES = ("QWEN_MODEL",) + RETURN_NAMES = ("qwen_model",) + FUNCTION = "load_model" + CATEGORY = "AIIA/Loaders" + + def load_model(self, model_name, device, dtype, local_path=""): + _install_qwen_tts_if_needed() + from qwen_tts import Qwen3TTSModel + + # Resolve device + if device == "auto": + if torch.cuda.is_available(): device = "cuda" + elif torch.backends.mps.is_available(): device = "mps" + else: device = "cpu" + + # Resolve dtype + torch_dtype = torch.bfloat16 if dtype == "bf16" else (torch.float16 if dtype == "fp16" else torch.float32) + + # Resolve path + path = local_path if local_path and os.path.exists(local_path) else model_name + + print(f"[AIIA] Loading Qwen3-TTS: {path} on {device} with {dtype}") + + # Flash Attention check + attn_impl = "flash_attention_2" if (device == "cuda" and torch.cuda.get_device_capability()[0] >= 8) else "sdpa" + + model = Qwen3TTSModel.from_pretrained( + path, + device_map=device, + torch_dtype=torch_dtype, + attn_implementation=attn_impl + ) + + model_type = "Base" + if "CustomVoice" in path: model_type = "CustomVoice" + elif "VoiceDesign" in path: model_type = "VoiceDesign" + + return ({"model": model, "type": model_type, "name": path, "device": device},) + +class AIIA_Qwen_TTS: + @classmethod + def INPUT_TYPES(s): + return { + "required": { + "qwen_model": ("QWEN_MODEL",), + "text": ("STRING", {"multiline": True, "default": "你好,这是 Qwen3-TTS 的测试。"}), + "language": (["Auto", "Chinese", "English", "Japanese", "Korean", "German", "French", "Russian", "Portuguese", "Spanish", "Italian"], {"default": "Chinese"}), + }, + "optional": { + "speaker": ("STRING", {"default": "Vivian"}), + "instruct": ("STRING", {"multiline": True, "default": ""}), + "reference_audio": ("AUDIO",), + "reference_text": ("STRING", {"multiline": True, "default": ""}), + "x_vector_only": ("BOOLEAN", {"default": False}), + "seed": ("INT", {"default": 42, "min": -1, "max": 2147483647}), + "speed": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0}), + } + } + + RETURN_TYPES = ("AUDIO",) + RETURN_NAMES = ("audio",) + FUNCTION = "generate" + CATEGORY = "AIIA/Synthesis" + + def generate(self, qwen_model, text, language, speaker="Vivian", instruct="", reference_audio=None, reference_text="", x_vector_only=False, seed=42, speed=1.0): + model = qwen_model["model"] + m_type = qwen_model["type"] + + if seed >= 0: + torch.manual_seed(seed) + if torch.cuda.is_available(): torch.cuda.manual_seed_all(seed) + + lang_param = language if language != "Auto" else "Auto" + + wavs = None + sr = 24000 # Default if unknown + + try: + if m_type == "CustomVoice": + print(f"[AIIA] Qwen3-TTS CustomVoice: {speaker} | Instruct: {instruct}") + wavs, sr = model.generate_custom_voice( + text=text, + language=lang_param, + speaker=speaker, + instruct=instruct if instruct else None + ) + elif m_type == "VoiceDesign": + print(f"[AIIA] Qwen3-TTS VoiceDesign: {instruct}") + wavs, sr = model.generate_voice_design( + text=text, + language=lang_param, + instruct=instruct + ) + else: # Base / Clone + if reference_audio is not None: + # Convert ComfyUI Audio format to (numpy, sr) tuple + device = qwen_model["device"] + ref_wav = reference_audio["waveform"] + ref_sr = reference_audio["sample_rate"] + + # Convert to mono if needed + if ref_wav.ndim == 3: ref_wav = ref_wav[0] + if ref_wav.shape[0] > 1: ref_wav = torch.mean(ref_wav, dim=0, keepdim=True) + + ref_audio_data = (ref_wav.squeeze().cpu().numpy(), ref_sr) + + print(f"[AIIA] Qwen3-TTS VoiceClone: Using provided reference.") + wavs, sr = model.generate_voice_clone( + text=text, + language=lang_param, + ref_audio=ref_audio_data, + ref_text=reference_text if reference_text else None, + x_vector_only_mode=x_vector_only + ) + else: + # Fallback if no reference provided for Base model + # Typically Base model MUST have reference. + # We might want to provide a default one or error out. + raise ValueError("Qwen3-TTS Base model requires 'reference_audio' and 'reference_text' for cloning.") + + # Process output + if wavs is not None and len(wavs) > 0: + audio_out = torch.from_numpy(wavs[0]).float() + if audio_out.ndim == 1: audio_out = audio_out.unsqueeze(0) + + # Speed adj (Qwen3-TTS might not have native speed param in generate_* yet, so we use torchaudio if needed) + if speed != 1.0: + # Simple speed change via resampling (pitch change) - matches CosyVoice fallback + resampler = torchaudio.transforms.Resample(orig_freq=int(sr*speed), new_freq=sr) + audio_out = resampler(audio_out) + + return ({"waveform": audio_out.unsqueeze(0), "sample_rate": sr},) + + except Exception as e: + print(f"[AIIA] Qwen3-TTS Generation Error: {e}") + import traceback + traceback.print_exc() + raise e + +NODE_CLASS_MAPPINGS = { + "AIIA_Qwen_Loader": AIIA_Qwen_Loader, + "AIIA_Qwen_TTS": AIIA_Qwen_TTS +} + +NODE_DISPLAY_NAME_MAPPINGS = { + "AIIA_Qwen_Loader": "🤖 Qwen3-TTS Loader", + "AIIA_Qwen_TTS": "🗣️ Qwen3-TTS Synthesis" +} diff --git a/pyproject.toml b/pyproject.toml index 9cfa867..c278482 100755 --- a/pyproject.toml +++ b/pyproject.toml @@ -1,7 +1,7 @@ [project] name = "aiia" description = "The Ultimate AI Audio/Video toolkit for ComfyUI. Features an enhanced Ditto (with optimizations that outperform official demos and other SOTA talking head models in lip-sync accuracy and natural motion), EchoMimic V3 & FLOAT, VibeVoice & CosyVoice 3.0 (Zero-Shot Voice Cloning), Multi-Role Podcast Generation, and a powerful Media Browser." -version = "1.10.20" +version = "1.11.0" license = {file = "LICENSE"} readme = "README.md" authors = [ @@ -17,6 +17,7 @@ dependencies = [ "huggingface_hub", "opencv-python", "ffmpeg-python", + "qwen-tts", ] [project.urls]