diff --git a/README.md b/README.md
index e86ef3c..b9510ae 100755
--- a/README.md
+++ b/README.md
@@ -971,11 +971,35 @@ hf download digital-avatar/ditto-talkinghead --local-dir ditto
| **克隆能力 (Cloning)** | **SOTA** (Zero-Shot)`
`只需 3-10秒,对**音色质感**还原极高。 | **SOTA** (稳定性)`
`对**说话韵律/口音**的捕捉最准。 | **良好** `
`适合克隆特定语气,而非纯粹音色。 |
| **多语言/方言** | **中/英** (双语优化) | **👑 霸主** (9种语言 + 18种方言) | **中/英** |
| **语音转换 (VC)** (Audio-to-Audio) | ❌**不支持** `
`仅支持 TTS (Text-to-Speech)。无法改变已有音频的音色。 | ✅**支持** `
`可以将任意音频转换为任意音色 (保留语调/停顿)。 | ❌**不支持** `
`纯 TTS 模型。仅支持 Text-to-Speech。 |
+| **Qwen3-TTS** (1.7B/0.6B) | ✅**支持** `
`支持 CustomVoice (内置) 和 VoiceDesign (描述)。 | ✅**支持** `
`支持 3秒极速 Zero-shot 克隆。 | ✅**支持** `
`支持 10 种语言。 |
+
+#### 3.13 Qwen3-TTS (New! 🔥)
+
+- **用途**: 阿里巴巴 Qwen 团队推出的最新旗舰级 TTS 模型,支持 10 种主要语言及多种方言,具备极高的稳定性和表现力。
+- **核心能力**:
+ - **CustomVoice**: 使用内置的高品质音色进行语音合成。提供 1.7B 和 0.6B 两种规格。
+ - **VoiceDesign**: 通过自然语言描述(如“活泼的少女音,带点羞涩”)从零设计音色。
+ - **VoiceClone**: 顶级的 3秒快速音色克隆,支持 X-Vector 模式提升稳定性。
+- **环境要求**:
+ - **qwen-tts**: `pip install qwen-tts` (插件会自动尝试安装)。
+ - **Flash Attention 2**: 强烈推荐以获得最佳推理性能。
+- **节点**:
+ - `🤖 Qwen3-TTS Loader`: 加载模型。支持 `Base` (克隆)、`CustomVoice` (内置音色) 和 `VoiceDesign` (音色设计) 模型。
+ - `🗣️ Qwen3-TTS Synthesis`: 执行合成。根据加载的模型类型自动切换功能。
+- **模型下载**:
+ - `Qwen/Qwen3-TTS-12Hz-1.7B-Base` (或 0.6B-Base)
+ - `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice` (或 0.6B-CustomVoice)
+ - `Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign`
+
+---
+
+#### 💡 用户实测与选型指南 (Model Comparison & Selection)
**选型建议**:
- **追求“听起来最像真人” (音质+音色)**: 选 **VoxCPM 1.5**。它的 Tokenizer-free 架构带来了质的飞跃。
- **追求“方言/多语言/稳定性”**: 选 **CosyVoice 3.0**。目前依然是生产环境最稳的选择。
+- **追求“多样化音色设计/最新 Qwen 生态/长语音流畅度”**: 选 **Qwen3-TTS**。其 VoiceDesign 功能能让你用描述语“捏”出从未听过的声音。
- **要做“长篇广播剧/播客”**: 选 **VibeVoice**。它的长窗口上下文优势依然不可替代。
### 4. 播客与对话生成 (Podcast & Dialogue Generation)
@@ -1005,6 +1029,7 @@ https://github.com/user-attachments/assets/9a5502c5-79e3-4fc8-8a2d-2cbdbdbbc860
- **TTS Engine**: 后端引擎选择。
- **CosyVoice**: 精准控制型。
- **VibeVoice**: 自然演绎型。
+ - **Qwen3-TTS**: 万能旗舰型。支持音色设计与内置高质量音色。
- **Speaker A/B/C**:
- **Ref Audio**: 参考音频 (用于 Zero-Shot 克隆)。
- **ID**: 内置音色 ID (如 CosyVoice 的 `Chinese Female`)。
@@ -1026,6 +1051,7 @@ https://github.com/user-attachments/assets/9a5502c5-79e3-4fc8-8a2d-2cbdbdbbc860
- **原理**:
- **CosyVoice**: 使用生成时的精确时长。
- **VibeVoice**: 使用**智能插值算法 (Smart Interpolation)**,根据字符长度自动计算长音频段内的单句时间轴。
+ - **Qwen3-TTS**: 基于生成的音频振幅精准断句,支持多角色时间轴导出。
#### 4.4 AIIA Subtitle to Segments (字幕转分段)
@@ -1069,13 +1095,13 @@ https://github.com/user-attachments/assets/9a5502c5-79e3-4fc8-8a2d-2cbdbdbbc860
#### 💡 引擎选型与最佳实践 (Best Practices)
-| 特性 | **CosyVoice** | **VibeVoice** |
-| :----------------- | :----------------------------------------- | :----------------------------------------------------------------------------------------------- |
-| **核心优势** | **精准控制 (Instruction)** | **自然演绎 (Context-Aware)** |
-| **情感控制** | ✅**支持** (使用 `[Happy]` 等标签) | ❌ 不支持显式标签 (依赖上下文) |
-| **生成逻辑** | **逐句生成** (严格遵循每句话的指令) | **混合批处理** (Hybrid Batching) |
-| **最佳场景** | 需要精确指定某句话语气、方言时 | 长篇对话、广播剧、闲聊 |
-| **使用建议** | 可以在剧本中详细标注情感。 | **尽量减少 `(Pause)`**!`
`让多句对话连在一起,模型能更好地联系上下文产生自然语气。 |
+| 特性 | **CosyVoice** | **VibeVoice** | **Qwen3-TTS** |
+| :----------------- | :----------------------------------------- | :----------------------------------------------------------------------------------------------- | :------------------------------------------ |
+| **核心优势** | **精准控制 (Instruction)** | **自然演绎 (Context-Aware)** | **万能旗舰 (Voice Design)** |
+| **情感控制** | ✅**支持** (使用 `[Happy]` 等标签) | ❌ 不支持显式标签 (依赖上下文) | ✅**支持** (通过 `instruct` 或标签) |
+| **生成逻辑** | **逐句生成** (严格遵循每句话的指令) | **混合批处理** (Hybrid Batching) | **动态引擎** (支持流式与批处理) |
+| **最佳场景** | 需要精确指定某句话语气、方言时 | 长篇对话、广播剧、闲聊 | 音色定制、高质量配音、极速克隆 |
+| **使用建议** | 可以在剧本中详细标注情感。 | **尽量减少 `(Pause)`**!`
`让多句对话连在一起,模型能更好地联系上下文产生自然语气。 | 尝试使用其 Voice Design 进行创意捏人。 |
#### 📝 综合测试剧本 (Example Script)
@@ -1175,6 +1201,15 @@ B: 太神奇了!那我们快去生成试试吧!
## Changelog
+### [1.11.0] - 2026-02-04
+
+- **Qwen3-TTS**: 新增阿里巴巴 **Qwen3-TTS** 全系列支持。
+ - **🤖 Qwen3-TTS Loader**: 支持加载 Base, CustomVoice, VoiceDesign 及其 1.7B/0.6B 版本。
+ - **🗣️ Qwen3-TTS Synthesis**: 实现全功能生成,包括 Zero-shot 克隆、音色设计和内置音色合成。
+- **Podcast Integration**: **AIIA Dialogue TTS** 节点现在正式集成 Qwen3-TTS 引擎。
+ - 支持多角色混合场景下的 Qwen3 驱动,支持使用脚本标签触发 `instruct`。
+- **Auto-Dependency**: 首次运行 Qwen3 节点会自动检测并安装 `qwen-tts` 库。
+
### [1.10.17] - 2026-02-03
- **Subtitle**: 引入“说话人 ID 为了映射 (Speaker Mapping)”机制。
diff --git a/__init__.py b/__init__.py
index f72fcf8..658b93d 100755
--- a/__init__.py
+++ b/__init__.py
@@ -136,6 +136,9 @@ else:
# 25. 处理 aiia_debug_nodes.py (新增调试节点)
_load_nodes_from_module(".aiia_debug_nodes", "aiia_debug_nodes")
+ # 26. 处理 aiia_qwen_nodes.py (新增 Qwen3-TTS)
+ _load_nodes_from_module(".aiia_qwen_nodes", "aiia_qwen_nodes")
+
# 告诉 ComfyUI 这个节点包有一个包含网页资源的 'js' 目录
WEB_DIRECTORY = "js"
diff --git a/aiia_podcast_nodes.py b/aiia_podcast_nodes.py
index bbed8ee..ba8eb26 100755
--- a/aiia_podcast_nodes.py
+++ b/aiia_podcast_nodes.py
@@ -176,7 +176,7 @@ class AIIA_Dialogue_TTS:
return {
"required": {
"dialogue_json": ("STRING", {"forceInput": True}),
- "tts_engine": (["CosyVoice", "VibeVoice"], {"default": "CosyVoice"}),
+ "tts_engine": (["CosyVoice", "VibeVoice", "Qwen3-TTS"], {"default": "CosyVoice"}),
"pause_duration": ("FLOAT", {"default": 0.5, "min": 0.0, "max": 5.0, "step": 0.1}),
"speed_global": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0}),
"batch_mode": (["Natural (Hybrid)", "Strict (Per-Speaker)", "Whole (Single Batch)"], {"default": "Natural (Hybrid)"}),
@@ -190,6 +190,7 @@ class AIIA_Dialogue_TTS:
"optional": {
"cosyvoice_model": ("COSYVOICE_MODEL",),
"vibevoice_model": ("VIBEVOICE_MODEL",),
+ "qwen_model": ("QWEN_MODEL",),
# Speaker A
"speaker_A_ref": ("AUDIO",),
@@ -248,7 +249,7 @@ class AIIA_Dialogue_TTS:
return None
def process_dialogue(self, dialogue_json, tts_engine, pause_duration, speed_global,
- cosyvoice_model=None, vibevoice_model=None,
+ cosyvoice_model=None, vibevoice_model=None, qwen_model=None,
cfg_scale=1.5, temperature=0.8, top_k=20, top_p=0.95, **kwargs):
import json
import torch
@@ -260,6 +261,8 @@ class AIIA_Dialogue_TTS:
raise ValueError("选择 CosyVoice 引擎时,必须连接 'cosyvoice_model'!")
if tts_engine == "VibeVoice" and vibevoice_model is None:
raise ValueError("选择 VibeVoice 引擎时,必须连接 'vibevoice_model'!")
+ if tts_engine == "Qwen3-TTS" and qwen_model is None:
+ raise ValueError("选择 Qwen3-TTS 引擎时,必须连接 'qwen_model'!")
dialogue = json.loads(dialogue_json)
full_waveform = []
@@ -267,9 +270,11 @@ class AIIA_Dialogue_TTS:
from .aiia_cosyvoice_nodes import AIIA_CosyVoice_TTS
from .aiia_vibevoice_nodes import AIIA_VibeVoice_TTS
+ from .aiia_qwen_nodes import AIIA_Qwen_TTS
cosy_gen = AIIA_CosyVoice_TTS()
vibe_gen = AIIA_VibeVoice_TTS()
+ qwen_gen = AIIA_Qwen_TTS()
print(f"[AIIA Podcast] 开始处理对话,共 {len(dialogue)} 个片段。引擎: {tts_engine}")
@@ -402,6 +407,74 @@ class AIIA_Dialogue_TTS:
})
time_ptr[0] += 1.0
+ elif tts_engine == "Qwen3-TTS":
+ # Qwen3-TTS (Iterative)
+ for i, item in enumerate(batch_items):
+ spk_name = item["speaker"]
+ spk_key = get_speaker_key(spk_name)
+ text = item["text"]
+ emotion = item.get("emotion", "None")
+
+ # Mapping logic for Qwen
+ spk_id = kwargs.get(f"speaker_{spk_key}_id", "Vivian") # Default to Vivian if empty
+ if not spk_id.strip(): spk_id = "Vivian"
+
+ ref_audio = get_ref_audio(spk_key)
+ instruct = f"{emotion}." if emotion and emotion != "None" else ""
+
+ print(f" [Qwen Processing] {spk_name} (ID: {spk_id}): {text[:15]}...")
+ try:
+ # Call Qwen TTS
+ res = qwen_gen.generate(
+ qwen_model=qwen_model,
+ text=text,
+ language="Auto",
+ speaker=spk_id,
+ instruct=instruct,
+ reference_audio=ref_audio,
+ seed=42+i,
+ speed=speed_global
+ )
+
+ generated = res[0]
+ wav = generated["waveform"]
+ sr = generated["sample_rate"]
+
+ if sr_ptr[0] != sr:
+ if current_full_wav:
+ wav = torchaudio.transforms.Resample(sr, sr_ptr[0])(wav)
+ else:
+ sr_ptr[0] = sr
+
+ if wav.ndim == 3: wav = wav.squeeze(0)
+ if wav.ndim == 1: wav = wav.unsqueeze(0)
+ current_full_wav.append(wav)
+
+ # --- Timestamp Tracking ---
+ seg_duration = wav.shape[-1] / sr
+ seg_start = time_ptr[0]
+ seg_end = seg_start + seg_duration
+
+ segments_info.append({
+ "start": round(seg_start, 3),
+ "end": round(seg_end, 3),
+ "text": text,
+ "speaker": spk_name,
+ "visual": item.get("visual")
+ })
+ time_ptr[0] += seg_duration
+
+ # Add a small gap between segments
+ gap = 0.2
+ gap_samples = int(gap * sr_ptr[0])
+ current_full_wav.append(torch.zeros(1, gap_samples))
+ time_ptr[0] += gap
+
+ except Exception as e:
+ print(f"[Error] Qwen item generation failed: {e}")
+ current_full_wav.append(torch.zeros(1, 24000))
+ time_ptr[0] += 1.0
+
else:
# CosyVoice (Iterative)
for i, item in enumerate(batch_items):
diff --git a/aiia_qwen_nodes.py b/aiia_qwen_nodes.py
new file mode 100644
index 0000000..ac12239
--- /dev/null
+++ b/aiia_qwen_nodes.py
@@ -0,0 +1,189 @@
+import os
+import sys
+import torch
+import torchaudio
+import numpy as np
+import folder_paths
+import subprocess
+
+# --- Robust Package Installation ---
+def _install_qwen_tts_if_needed():
+ try:
+ from qwen_tts import Qwen3TTSModel
+ return
+ except ImportError:
+ print("[AIIA] qwen-tts missing. Attempting installation...")
+ try:
+ subprocess.check_call([sys.executable, "-m", "pip", "install", "-U", "qwen-tts"])
+ print("[AIIA] qwen-tts installed successfully.")
+ except Exception as e:
+ print(f"[AIIA] Failed to install qwen-tts: {e}")
+
+class AIIA_Qwen_Loader:
+ @classmethod
+ def INPUT_TYPES(s):
+ return {
+ "required": {
+ "model_name": ([
+ "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
+ "Qwen/Qwen3-TTS-12Hz-0.6B-Base",
+ "Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
+ "Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice",
+ "Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign"
+ ], {"default": "Qwen/Qwen3-TTS-12Hz-1.7B-Base"}),
+ "device": (["cuda", "cpu", "auto", "mps"], {"default": "auto"}),
+ "dtype": (["bf16", "fp16", "fp32"], {"default": "bf16"}),
+ },
+ "optional": {
+ "local_path": ("STRING", {"default": ""}),
+ }
+ }
+
+ RETURN_TYPES = ("QWEN_MODEL",)
+ RETURN_NAMES = ("qwen_model",)
+ FUNCTION = "load_model"
+ CATEGORY = "AIIA/Loaders"
+
+ def load_model(self, model_name, device, dtype, local_path=""):
+ _install_qwen_tts_if_needed()
+ from qwen_tts import Qwen3TTSModel
+
+ # Resolve device
+ if device == "auto":
+ if torch.cuda.is_available(): device = "cuda"
+ elif torch.backends.mps.is_available(): device = "mps"
+ else: device = "cpu"
+
+ # Resolve dtype
+ torch_dtype = torch.bfloat16 if dtype == "bf16" else (torch.float16 if dtype == "fp16" else torch.float32)
+
+ # Resolve path
+ path = local_path if local_path and os.path.exists(local_path) else model_name
+
+ print(f"[AIIA] Loading Qwen3-TTS: {path} on {device} with {dtype}")
+
+ # Flash Attention check
+ attn_impl = "flash_attention_2" if (device == "cuda" and torch.cuda.get_device_capability()[0] >= 8) else "sdpa"
+
+ model = Qwen3TTSModel.from_pretrained(
+ path,
+ device_map=device,
+ torch_dtype=torch_dtype,
+ attn_implementation=attn_impl
+ )
+
+ model_type = "Base"
+ if "CustomVoice" in path: model_type = "CustomVoice"
+ elif "VoiceDesign" in path: model_type = "VoiceDesign"
+
+ return ({"model": model, "type": model_type, "name": path, "device": device},)
+
+class AIIA_Qwen_TTS:
+ @classmethod
+ def INPUT_TYPES(s):
+ return {
+ "required": {
+ "qwen_model": ("QWEN_MODEL",),
+ "text": ("STRING", {"multiline": True, "default": "你好,这是 Qwen3-TTS 的测试。"}),
+ "language": (["Auto", "Chinese", "English", "Japanese", "Korean", "German", "French", "Russian", "Portuguese", "Spanish", "Italian"], {"default": "Chinese"}),
+ },
+ "optional": {
+ "speaker": ("STRING", {"default": "Vivian"}),
+ "instruct": ("STRING", {"multiline": True, "default": ""}),
+ "reference_audio": ("AUDIO",),
+ "reference_text": ("STRING", {"multiline": True, "default": ""}),
+ "x_vector_only": ("BOOLEAN", {"default": False}),
+ "seed": ("INT", {"default": 42, "min": -1, "max": 2147483647}),
+ "speed": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0}),
+ }
+ }
+
+ RETURN_TYPES = ("AUDIO",)
+ RETURN_NAMES = ("audio",)
+ FUNCTION = "generate"
+ CATEGORY = "AIIA/Synthesis"
+
+ def generate(self, qwen_model, text, language, speaker="Vivian", instruct="", reference_audio=None, reference_text="", x_vector_only=False, seed=42, speed=1.0):
+ model = qwen_model["model"]
+ m_type = qwen_model["type"]
+
+ if seed >= 0:
+ torch.manual_seed(seed)
+ if torch.cuda.is_available(): torch.cuda.manual_seed_all(seed)
+
+ lang_param = language if language != "Auto" else "Auto"
+
+ wavs = None
+ sr = 24000 # Default if unknown
+
+ try:
+ if m_type == "CustomVoice":
+ print(f"[AIIA] Qwen3-TTS CustomVoice: {speaker} | Instruct: {instruct}")
+ wavs, sr = model.generate_custom_voice(
+ text=text,
+ language=lang_param,
+ speaker=speaker,
+ instruct=instruct if instruct else None
+ )
+ elif m_type == "VoiceDesign":
+ print(f"[AIIA] Qwen3-TTS VoiceDesign: {instruct}")
+ wavs, sr = model.generate_voice_design(
+ text=text,
+ language=lang_param,
+ instruct=instruct
+ )
+ else: # Base / Clone
+ if reference_audio is not None:
+ # Convert ComfyUI Audio format to (numpy, sr) tuple
+ device = qwen_model["device"]
+ ref_wav = reference_audio["waveform"]
+ ref_sr = reference_audio["sample_rate"]
+
+ # Convert to mono if needed
+ if ref_wav.ndim == 3: ref_wav = ref_wav[0]
+ if ref_wav.shape[0] > 1: ref_wav = torch.mean(ref_wav, dim=0, keepdim=True)
+
+ ref_audio_data = (ref_wav.squeeze().cpu().numpy(), ref_sr)
+
+ print(f"[AIIA] Qwen3-TTS VoiceClone: Using provided reference.")
+ wavs, sr = model.generate_voice_clone(
+ text=text,
+ language=lang_param,
+ ref_audio=ref_audio_data,
+ ref_text=reference_text if reference_text else None,
+ x_vector_only_mode=x_vector_only
+ )
+ else:
+ # Fallback if no reference provided for Base model
+ # Typically Base model MUST have reference.
+ # We might want to provide a default one or error out.
+ raise ValueError("Qwen3-TTS Base model requires 'reference_audio' and 'reference_text' for cloning.")
+
+ # Process output
+ if wavs is not None and len(wavs) > 0:
+ audio_out = torch.from_numpy(wavs[0]).float()
+ if audio_out.ndim == 1: audio_out = audio_out.unsqueeze(0)
+
+ # Speed adj (Qwen3-TTS might not have native speed param in generate_* yet, so we use torchaudio if needed)
+ if speed != 1.0:
+ # Simple speed change via resampling (pitch change) - matches CosyVoice fallback
+ resampler = torchaudio.transforms.Resample(orig_freq=int(sr*speed), new_freq=sr)
+ audio_out = resampler(audio_out)
+
+ return ({"waveform": audio_out.unsqueeze(0), "sample_rate": sr},)
+
+ except Exception as e:
+ print(f"[AIIA] Qwen3-TTS Generation Error: {e}")
+ import traceback
+ traceback.print_exc()
+ raise e
+
+NODE_CLASS_MAPPINGS = {
+ "AIIA_Qwen_Loader": AIIA_Qwen_Loader,
+ "AIIA_Qwen_TTS": AIIA_Qwen_TTS
+}
+
+NODE_DISPLAY_NAME_MAPPINGS = {
+ "AIIA_Qwen_Loader": "🤖 Qwen3-TTS Loader",
+ "AIIA_Qwen_TTS": "🗣️ Qwen3-TTS Synthesis"
+}
diff --git a/pyproject.toml b/pyproject.toml
index 9cfa867..c278482 100755
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -1,7 +1,7 @@
[project]
name = "aiia"
description = "The Ultimate AI Audio/Video toolkit for ComfyUI. Features an enhanced Ditto (with optimizations that outperform official demos and other SOTA talking head models in lip-sync accuracy and natural motion), EchoMimic V3 & FLOAT, VibeVoice & CosyVoice 3.0 (Zero-Shot Voice Cloning), Multi-Role Podcast Generation, and a powerful Media Browser."
-version = "1.10.20"
+version = "1.11.0"
license = {file = "LICENSE"}
readme = "README.md"
authors = [
@@ -17,6 +17,7 @@ dependencies = [
"huggingface_hub",
"opencv-python",
"ffmpeg-python",
+ "qwen-tts",
]
[project.urls]