diff --git a/.github/workflows/main.yml b/.github/workflows/main.yml
old mode 100755
new mode 100644
diff --git a/.gitignore b/.gitignore
old mode 100755
new mode 100644
index e7f0663..5e2b0cc
--- a/.gitignore
+++ b/.gitignore
@@ -2,5 +2,4 @@ GEMINI.md
.gemini/
__pycache__/
*.pyc
-.DS_Store
-.env
+.DS_Store
\ No newline at end of file
diff --git a/.vscode/extensions.json b/.vscode/extensions.json
old mode 100755
new mode 100644
diff --git a/README.md b/README.md
old mode 100755
new mode 100644
index a08c447..9914a7a
--- a/README.md
+++ b/README.md
@@ -1,6 +1,6 @@

-
AIIA Nodes for ComfyUI
+# AIIA Nodes for ComfyUI
欢迎来到 AIIA Nodes for ComfyUI 仓库!这是一个旨在为 ComfyUI 提供一系列强大、直观且高度可定制的节点的集合。这些节点专注于简化复杂的工作流,并为创意工作者提供最大的灵活性。
@@ -13,7 +13,6 @@
**还在费力地翻找 `output` 文件夹,或者对着一堆时间戳命名的文件猜内容吗?**
我们隆重推出 **AIIA 媒体浏览器**——一个完全集成在 ComfyUI 内部的、功能完备的媒体文件管理中心。它的诞生,旨在彻底改变你管理和使用生成结果的方式,让整个过程变得高效、直观且充满乐趣。
-

### ✨ 为何选择 AIIA 浏览器?
@@ -59,21 +58,14 @@
处理成百上千张高清图像帧时,轻易就会耗尽 VRAM 和系统内存,导致工作流中断。AIIA 节点通过 **增量式处理(Incremental Processing)** 的策略从根本上解决了这个问题。
-#### v1.9.21+ 内存优化
+无论是从磁盘流式读取帧进行视频合并,还是将生成结果逐帧保存到磁盘,我们的节点都**永远不需要将所有图像一次性加载到内存中**。这意味着您可以轻松生成数千帧的视频,而无需再为内存限制而烦恼。
-自 v1.9.21 起,**AIIA Body Sway** 和 **AIIA Video Combine** 节点实现了激进的内存管理策略:
-* **分批处理**:每次只处理 50 帧,处理完立即释放中间变量。
-* **逐帧释放**:每帧转换后立即 `del` 并定期 `gc.collect()`。
-* **GPU 内存同步释放**:处理前移至 CPU 并调用 `torch.cuda.empty_cache()`。
-
-这意味着即使处理 **1500+ 帧的高分辨率视频**(如 1288×1920),也能在合理的内存占用下完成,**无需磁盘中转**。
-
-#### 两种工作模式
+### ✨ 无缝与高效的平衡
我们提供了两种工作模式,以适应不同场景:
-- **内存模式 (推荐)**: 直接将上游节点的 `IMAGE` 张量输入。v1.9.21+ 的优化使其可处理数千帧而不 OOM。
-- **磁盘模式**: 对于极端长序列或内存受限环境,仍可通过 `frames_directory` 从磁盘流式读取帧。
+- **内存模式**: 对于短序列或测试,可以直接将上游节点的 `IMAGE` 张量输入,实现无缝、快速的内存内处理。
+- **磁盘模式**: 对于长序列的最终生成,节点会高效地从磁盘流式读取/写入帧,保证了稳定性和极低的内存占用。
### 🔧 强大且可扩展的预设系统
@@ -91,15 +83,7 @@
- **强烈建议**将 FFmpeg 的 `bin` 目录添加到您系统的 `PATH` 环境变量中。
- 在终端中运行 `ffmpeg -version` 和 `ffprobe -version` 来验证安装。
-### 2. 安装 SoX (VibeVoice 变速不变调必须)
-
-VibeVoice 节点的 `speed` 参数依赖系统级 `sox` 命令。
-
-- **Ubuntu/Debian**: `sudo apt-get update && sudo apt-get install -y libsox-dev sox`
-- **macOS**: `brew install sox`
-- **Windows**: 下载 [SoX 编译版](https://sourceforge.net/projects/sox/files/sox/) 并将目录添加到 `PATH`。
-
-### 3. 安装 NeMo 模型 (音频AI节点必须)
+### 2. 安装 NeMo 模型 (音频AI节点必须)
音频处理节点(如说话人日志)依赖 NeMo 模型。
@@ -124,6 +108,43 @@ hf download nvidia/nemo-models diar_sortformer_4spk-v1.nemo --local-dir nemo_mod
hf download nvidia/nemo-models diar_streaming_sortformer_4spk-v2.1.nemo --local-dir nemo_models
```
+### 3. 安装 FunASR 模型 (ASR 节点 / 播客防泄漏管线)
+
+ASR 节点使用阿里达摩院的 **FunASR** 框架进行语音识别,支持字级时间戳。模型需手动下载至 `ComfyUI/models/funasr/` 目录。
+
+```text
+ComfyUI/models/funasr/
+├── paraformer-zh/ <-- 中文 ASR (推荐,支持字级时间戳)
+│ ├── model.pt
+│ ├── configuration.json
+│ └── ...
+└── SenseVoiceSmall/ <-- 多语言 ASR (中/英/日/韩/粤,无时间戳)
+ ├── model.pt
+ └── ...
+```
+
+**下载命令**:
+
+```bash
+cd ComfyUI/models
+mkdir -p funasr
+
+# 推荐: Paraformer-zh (中文,支持字级时间戳,~950MB)
+modelscope download --model iic/speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch --local_dir funasr/paraformer-zh
+
+# 可选: SenseVoiceSmall (多语言,无时间戳,~450MB)
+modelscope download --model iic/SenseVoiceSmall --local_dir funasr/SenseVoiceSmall
+```
+
+**Python 依赖**:
+
+```bash
+pip install funasr
+```
+
+> [!NOTE]
+> **Paraformer-zh** 是播客防泄漏管线(Stitcher 节点)的**必需**模型,因为它提供精确的字级时间戳用于音频切分对齐。SenseVoiceSmall 适合通用多语言识别场景,但不输出时间戳。
+
### 4. 安装本节点套件
进入 ComfyUI 的自定义节点目录,然后克隆本仓库:
@@ -151,7 +172,7 @@ git clone https://github.com/havvk/ComfyUI_AIIA.git
### 2. 视频生成与合成 (Video Generation & Compositing)
-#### 2.1 视频合并 (AIIA, 图像或目录)
+#### 视频合并 (AIIA, 图像或目录)
这是一个功能强大且高度可定制的视频合并节点,是您工作流中处理视频生成的终极解决方案。
@@ -164,7 +185,7 @@ git clone https://github.com/havvk/ComfyUI_AIIA.git
- **全面的音频控制**: 支持 `AUDIO` 张量和外部文件,并提供对编解码器和码率的精细控制。
- **智能自动配置**: `auto` 模式能自动应用格式预设中的音频参数,并能自动检测源文件的码率。
-#### 2.2 FLOAT 影片生成 (内存与磁盘模式)
+#### FLOAT 影片生成 (内存与磁盘模式)
这组节点封装了先进的 **FLOAT** 模型,能够根据参考图像和音频生成高质量的口型同步影片。我们提供了两种模式,以应对不同长度的生成需求。
@@ -182,7 +203,7 @@ git clone https://github.com/havvk/ComfyUI_AIIA.git
- **优势**: 在解码过程中,节点以小批量方式处理帧并**逐帧保存到磁盘**,内存占用极低,可以处理任意长度的音频。
- **工作流**: 此节点的输出目录可以直接作为 **视频合并节点** 的 `frames_directory` 输入,构建一个完整的、内存高效的 talking head 视频生成管线。
-#### 2.3 PersonaLive 视频驱动 (AIIA Integrated)
+#### PersonaLive 视频驱动 (AIIA Integrated)
这组节点基于强大的 [PersonaLive](https://github.com/GVCLab/PersonaLive) 模型,专为生成高质量的 Talking Head 视频而设计。我们将原版代码完全重构并集成到 ComfyUI 中,通过特有的分块处理和磁盘流式技术,**彻底解决了长视频生成时的显存和内存溢出 (OOM) 问题**。
@@ -203,262 +224,7 @@ git clone https://github.com/havvk/ComfyUI_AIIA.git
- **输出**: `STRING` (包含生成帧的目录路径) 和 `INT` (帧数)。
- **最佳实践**: 将此节点的输出目录直接连接到 **AIIA Video Combine** 节点,即可实现从生成到合成的全流程 OOM-Safe。
-#### 2.4 EchoMimic V3 (AIIA Integrated)
-
-这组节点集成了最新的 **EchoMimic V3** (1.3B Parameters) 模型,它是目前开源界效果最惊艳的 Talking Head 解决方案之一。
-
-**特点**:
-
-- **多模态驱动**: 支持 **Audio Only** (仅音频驱动) 和 **Audio + Reference Pose** (音频+参考姿态) 驱动。
-- **自然度极高**: 相比 float 等早期模型,V3 在头部运动、表情微表情的自然度上有巨大提升。
-- **ComfyUI 原生**: 我们将其封装为标准的 Loader 和 Sampler 节点,支持流式生成和内存优化。
-
-**1. EchoMimic V3 Loader**
-
-- **用途**: 加载模型权重 (Transformer, VAE, Wav2Vec, etc.)。
-- **参数**:
- - `model_subfolder`: 模型子目录名 (默认 `Wan2.1-Fun-V1.1-1.3B-InP`)。
- - `device`: 指定运行设备 (CUDA)。
-
-**2. EchoMimic V3 Sampler**
-
-- **用途**: 执行推理生成。
-- **输入**:
- - `ref_image`: 参考人物图片 (建议 1:1 比例,如 768x768)。
- - `ref_audio`: 驱动音频。
-- **参数**:
- - `cfg`: 视觉引导系数 (默认 4.0)。
- - `audio_cfg`: 音频引导系数 (默认 2.9)。
- - `enable_teacache`: **True** (默认)。开启后生成速度提升 1.5 倍以上,且质量无损。
- - `keep_model_loaded`: **True** (默认)。即使显存占用增加,也强制将模型保留在 GPU 上,显著减少多段视频生成时的加载时间。
- - `negative_prompt`: 已内置优化过的 **眼部修复 (Eye Correction)** 提示词,有效防止翻白眼和眼神飘忽。
-
-**🚀 性能优化 (Performance)**:
-
-- **Flash Attention 2**: 强烈推荐安装。检测到时会自动启用,大幅提升推理速度。
-- **TeaCache**: 默认启用。通过缓存部分 Transformer 层计算,大幅加速生成。
-- **Full GPU Mode**: 默认启用。适合显存充足 (24GB+) 用户,享受极致流畅的生成体验。
-
-**🛠️ 模型下载指南 (Manual Download Guide)**
-
-由于 EchoMimic V3 模型较大且组件较多,目前**不支持自动下载**,请按以下步骤手动准备模型。
-
-目标目录: `ComfyUI/models/EchoMimicV3/`
-
-**目录结构**:
-
-```text
-ComfyUI/models/EchoMimicV3/
-├── Wan2.1-Fun-V1.1-1.3B-InP/ <-- 主模型目录
-│ ├── transformer/
-│ │ ├── config.json
-│ │ └── diffusion_pytorch_model.safetensors
-│ ├── vae/
-│ │ ├── config.json
-│ │ └── diffusion_pytorch_model.safetensors
-│ ├── text_encoder/
-│ ├── tokenizer/
-│ ├── image_encoder/
-│ └── scheduler/
-└── wav2vec2-base-960h/ <-- 音频编码器 (必需)
- ├── config.json
- ├── pytorch_model.bin
- └── ...
-```
-
-**下载地址**:
-
-1. **主模型 (EchoMimicV3)**:
-
- - HuggingFace: [BadToBest/EchoMimicV3](https://huggingface.co/BadToBest/EchoMimicV3)
- - **下载命令 (推荐)**:
- ```bash
- hf download BadToBest/EchoMimicV3 --local-dir models/EchoMimicV3/EchoMimicV3
- ```
- - *注意:此模型包含 EchoMimic 特有的 Transformer 权重,是生成嘴型的核心。*
-2. **底模 (Wan2.1-Fun-V1.1-1.3B-InP)**:
-
- - HuggingFace: [alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP)
- - **下载命令**:
- ```bash
- hf download alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP --local-dir models/EchoMimicV3/Wan2.1-Fun-V1.1-1.3B-InP
- ```
- - *注意:作为 fallback 来源,提供 VAE、Text Encoder 和 Image Encoder权重。*
-3. **音频编码器 (wav2vec2-base-960h)**:
-
- - HuggingFace: [facebook/wav2vec2-base-960h](https://huggingface.co/facebook/wav2vec2-base-960h)
- - **下载命令**:
- ```bash
- hf download facebook/wav2vec2-base-960h --local-dir models/EchoMimicV3/wav2vec2-base-960h
- ```
-
-**环境依赖**:
-
-- 请确保安装了 `requirements.txt` 中的依赖,如 `diffusers>=0.30.1`。节点加载时会尝试自动引用,但如果报错缺包,请手动安装。
-
-#### 2.5 Ditto Talking Head (AIIA Integrated)
-
-这组节点集成了 [Ditto](https://github.com/antgroup/ditto-talkinghead) 数字人模型。我们采用了 **PyTorch** 原生实现,避免了复杂的 TensorRT 编译过程,让用户能够“开箱即用”地生成高质量的 Talking Head 视频。
-
-**特点**:
-
-- **PyTorch Native**: 无需安装 TensorRT,兼容性更好。
-- **In-Memory Pipeline**: 针对 ComfyUI 优化的内存内处理流程,无需生成中间视频文件。
-- **自动模型管理**: 支持自动下载模型权重。
-
-**1. AIIA Ditto Loader**
-
-- **用途**: 下载并加载 Ditto 模型 (约 1.2GB)。
-- **参数**:
- - `model_name`: 模型名称 (默认 `ditto-talkinghead`)。
- - `device`: 运行设备 (CUDA/CPU)。
-
-**2. AIIA Ditto Sampler**
-
-- **用途**: 执行推理生成。
-- **输入**:
- - `pipe`: 来自 Loader 的模型管道。
- - `ref_image`: 参考人物图片 (建议正方形,人脸居中)。
- - `audio`: 驱动音频。
- - `fps`: 建议 **25** (Ditto 针对 25FPS 训练)。即使输入其他值,目前内部逻辑也会优先保证 25FPS 的同步率。
-- **输出**: `IMAGE` (视频帧), `AUDIO`。
-- **高级参数 (Advanced Parameters)**:
- - `seed`: **随机种子 (Random Seed)**。
- - 控制扩散模型的噪声生成。
- - **关键作用**: 在长语音生成中,当触发“静音重置”时,种子会被重置,从而确保每一句话的生成条件都与第一句话完全一致,彻底消除“长语音嘴型漂移”和“对口型不准”的问题。
- - `crop_scale`: (默认 2.3) **面部工作区视野 (Face Context Scale)**。
- - **注意**: 此参数**不会改变输出视频的分辨率**,它决定了模型“看”到了多少人脸周围的内容。
- - **数值越大 (如 2.5)**: **广角视野**。模型能覆盖更多头发、脖子和背景。
- - ✅ 优点:适合头部运动幅度大的场景,不容易出框。
- - ❌ 缺点:在固定的推理画布中,人脸占比变小,生成的五官细节(如牙齿、眼神)可能会变糊。
- - **数值越小 (如 2.0)**: **特写视野**。模型聚焦于面部核心区域。
- - ✅ 优点:人脸占比大,五官细节极其清晰锐利。
- - ❌ 缺点:容易裁掉下巴或额头,头部大幅运动时可能会出现“断头”或伪影。
- - **推荐值**:
- - 标准场景: **2.3**
- - 大动态/全身/半身: **2.5** (牺牲细节换稳定性)
- - 大头照/证件照: **2.0** (追求极致细节)
- - `emo`: (默认 Neutral) 表情控制。可选 Angry, Happy, Sad 等。
- - `drive_eye`: (默认 True) 是否驱动眼睛。关闭后眼睛将保持参考图状态(或微动),适合原图眼神较好的情况。
- - `chk_eye_blink`: (已废弃,请使用 `blink_mode`)。
- - `blink_mode`: (默认 Natural) **眨眼模式控制**。
- - `Natural`: **拟人化随机眨眼**。
- - 基础频率:90-150帧/次 (约 3.6s - 6.0s)。
- - `Slow`: 慢速沉稳眨眼 (120-200帧/次)。
- - `Fast`: 快速频繁眨眼 (10-40帧/次)。
- - `None`: **彻底关闭眨眼**。
- - `blink_amp`: (v1.9.1 New) **眨眼幅度控制**。
- - **1.0 (默认)**: 标准闭眼幅度。
- - **< 1.0 (推荐 0.8)**: 适合**男性角色**或眼睛较小的人物,避免“用力挤眼”的感觉。
- - **> 1.0**: 加深闭眼力度。
- - `mouth_amp`: (v1.9.1 New) **嘴型幅度控制**。
- - **1.0 (默认)**: 标准嘴部开合幅度。
- - **> 1.0 (推荐 1.1-1.2)**: 适合**大声说话**或需要更夸张表情的场景,增强口型辨识度。
- - **< 1.0**: 适合轻声细语。
- - `relax_on_silence`: (默认 True) **静音归位 (Relax Face on Silence)**。
- - 结合下方的 `silence_release` 参数,针对静音片段进行**平滑过渡**(慢速闭合),避免“紧绷抿嘴”。
- - **智能预测 (Predictive Logic)**: 自动识别短停顿(如逗号)与长静音(如句号)。短停顿不触发闭嘴动画,由模型自由发挥;长静音则触发优雅的慢速闭合。
- - **防止累积误差 (Drift Correction)**:
- - **精密网格对齐 (Exact Onset Alignment)**: 将模型的时间步长精确对齐到语音的开始(Onset),而不是固定的处理块。
- - **状态重置 (State & RNG Reset)**: 在长静音后,强制重置模型状态和**随机种子**,确保每一句话的生成质量一致,彻底消除“长语音嘴型漂移”。
- - `silence_release`: (v1.9.2 New) **静音闭嘴速度 (Adsr Release Control)**。
- - **Natural (0.8s) [默认]**:
- - 自然平衡模式。适合大多数常速对话。
- - 触发阈值: >0.88s。
- - **Fast (0.5s)**:
- - 快速响应模式。适合语速极快、充满激情的演讲。
- - 触发阈值: >0.56s。
- - **Deep (1.3s)**:
- - 深沉模式。适合朗诵、讲故事或情感类内容。超长尾韵,极度平滑。
- - 触发阈值: >1.4s。
- - `ref_threshold`: (默认 0.005) **静音检测相对阈值 (Relative Silence Threshold)**。
- - 现在的阈值是基于**全段音频峰值音量**的比例 (例如 0.005 = 0.5% 的峰值音量)。
- - 这意味着无论音频整体是大声还是小声,系统都能自动适应,准确捕捉微弱的语音片段。
- - 只有低于此比例的音量才会被视为静音。
- - `smo_k_d`: (默认 3) 运动平滑系数。数值越大动作越柔和,可抑制面部抖动。
- - `hd_rot_p` / `y` / `r`: 头部旋转微调 (Pitch/Yaw/Roll)。
- - `speech_pitch`: (v1.10.0 New) **说话时俯仰角补偿 (Speech Pitch Offset)**。
- - 用于修正"说话时头抬得太高"或"需要低头说话"的场景。此偏移量仅在说话期间生效,并随语音强度平滑切入切出。
- - **正值 (+) = 低头 (Look Down)**。例如 `5.0` 表示说话时微微低头。
- - **负值 (-) = 抬头 (Look Up)**。
- - `mouth_smoothing`: (v1.9.5 New) **嘴型惯性平滑 (Mouth Motion Inertia)**。
- - 防止模型输出的嘴型瞬间开合(如爆破音时),增加物理惯性感。
- - **`None (Raw)`**: 无平滑,模型原始输出。追求极致对口型,容忍偶尔快速开合。
- - **`Light`** (0.3): 轻微平滑,**推荐快语速使用**。
- - **`Normal`** (0.5) [默认]: 适中平滑,常规对话推荐。
- - **`Heavy`** (0.7): 强力平滑,适合低质量音频或模型输出抖动严重的情况。
- - `save_to_disk`: (v1.9.24 New) **OOM 安全模式 (OOM-Safe Mode)**。
- - **`Memory (Default)`**: 传统模式,所有帧保存在内存中。适合短视频(<1000帧)。
- - **`Disk (OOM-Safe)`**: **长视频推荐**。边生成边保存到磁盘,无 OOM 风险。
- - 选择 Disk 模式时,`frames_dir` 输出会包含帧保存路径,可直接连接 **AIIA Video Combine** 节点的 `frames_directory` 输入。
- - ⚠️ Disk 模式下 `images` 输出为占位符,请使用 `frames_dir` 连接后续节点。
-- **输出**:
- - `images`: 生成的视频帧序列(Memory 模式)或占位符(Disk 模式)。
- - `audio`: 透传的音频。
- - `frames_dir`: (v1.9.24 New) Disk 模式下的帧保存路径。Memory 模式下为空字符串。
-
-**🛠️ 模型下载指南 (Manual Download Guide)**
-
-如果自动下载失败,请手动下载模型并放入 `ComfyUI/models/ditto/` 目录。
-
-**目标目录结构**:
-
-```text
-ComfyUI/models/ditto/
-├── ditto_pytorch/
-│ ├── audio2motion.pth
-│ ├── ...
-└── ditto_cfg/
- ├── v0.4_hubert_cfg_pytorch.pkl
- ├── ...
-```
-
-**下载地址**:
-
-- HuggingFace: [digital-avatar/ditto-talkinghead](https://huggingface.co/digital-avatar/ditto-talkinghead)
-
-**下载命令**:
-
-```bash
-# 进入 models 目录
-cd ComfyUI/models
-
-# 下载模型 (直接下载到 ditto 目录,避免多层嵌套)
-hf download digital-avatar/ditto-talkinghead --local-dir ditto
-```
-
-#### 2.6 身体微动后处理 (Body Sway Post-Processing)
-
-这个轻量级后处理节点可以为 Ditto 等 Talking Head 模型的输出添加**模拟的身体晃动效果**,让人物看起来更加自然、有呼吸感。
-
-**工作原理**:
-* 通过**裁切平移 + 轻微旋转**模拟人体站立或坐着时的自然重心漂移和呼吸起伏。
-* 使用多频正弦波叠加(基于黄金比例)生成平滑、有机的运动轨迹。
-* **纯裁切**方式,不放大图像,保持原始画质。
-
-**AIIA Body Sway 节点**
-
-- **输入**:
- - `images` (可选): 来自 Ditto 等节点的视频帧张量 (Memory 模式)
- - `frames_directory` (可选, v1.9.25 New): 帧目录路径 (Disk 模式,连接 Ditto 的 `frames_dir` 输出)
-- **参数**:
- - `crop_ratio`: (默认 0.99) 输出尺寸占输入的比例。
- - 0.99 = 保留 99%,晃动幅度较小 (推荐)
- - 0.98 = 保留 98%,晃动幅度中等
- - 支持三位小数 (如 0.995)
- - `rotation_amplitude`: (默认 0.1) 最大旋转角度 (度)。
- - `smoothness`: (默认 0.02) Perlin 噪声平滑度。数值越小,运动越缓慢。
- - `seed`: 控制随机轨迹。
-- **输出**:
- - `images`: 应用了微动效果的帧 (Memory 模式) 或占位符 (Disk 模式)。
- - `output_frames_dir` (v1.9.25 New): Disk 模式下处理后的帧保存路径。
-
-> [!NOTE]
-> v1.9.17 改进:使用 **Perlin 噪声** 替代正弦波,运动更有机自然。**已移除垂直方向位移**,减少叠加 Ditto 头部运动时的"晕船"感。
-
-> [!TIP]
-> **OOM-Safe 工作流** (v1.9.24+): Ditto (`Disk`) → BodySway (`frames_directory`) → VideoCombine (`frames_directory`),全流程无 OOM 风险。
-> **性能无损** (v1.9.28+): 采用并行 I/O 和零压缩策略,**Disk 模式生成速度与 Memory 模式完全一致** (~30fps+),且极大降低 RAM 占用。强烈推荐长视频生成使用!
+---
### 3. 音频智能处理 (Intelligent Audio Processing)
@@ -839,7 +605,7 @@ hf download digital-avatar/ditto-talkinghead --local-dir ditto
#### 1. 🗣️ VibeVoice TTS (Standard)
- **适用模型**: `VibeVoice-1.5B`, `VibeVoice-7B`
-- **参考音频 (Reference Audio)**: 可选 (`optional`)。如果不连接,将自动使用内置的高品质女声种子 (Fallback Seed) 进行生成。
+- **必选参数**: `reference_audio` (参考音频) - **必须连接**。
- **功能**: 支持零样本音色克隆 (Zero-shot Cloning)。输入任何音频,它都会模仿该音色。
- **不支持**: `voice_preset` (预设)。
@@ -964,136 +730,18 @@ hf download digital-avatar/ditto-talkinghead --local-dir ditto
经过深度测试,我们在三个主流模型中整理了以下对比,助您选择最适合的引擎:
-| 维度 | **VoxCPM 1.5** (800M) | **CosyVoice 3.0** (0.5B/1.5B) | **VibeVoice** (1.5B/7B) |
-| :--------------------------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------------------------------------------------------------------------- | :-------------------------------------------------------------- |
-| **音质 (Fidelity)** | **44.1kHz 格式** `
`虽然物理格式为 44.1k,但因采用 **Neural Upsampling** (神经升频) 技术,听感上会有**含混 (Muffled)** 或**金属感**,且伴有底噪。 | **优秀** `
`听感最自然,但采样率稍低 (22/24kHz),有时需 AI 增强。 | **良好** `
`主要强在语气自然度,纯音质略逊。 |
-| **推理速度 (Speed)** | **🚀 冠军 (RTF ~0.17)**`
`得益于 Tokenizer-free,极其高效。 | **极快** `
`流式响应仅 150ms,且支持 TensorRT 加速。 | **一般/较慢** `
`7B 版本较重,更适合离线生成。 |
-| **克隆能力 (Cloning)** | **SOTA** (Zero-Shot)`
`只需 3-10秒,对**音色质感**还原极高。 | **SOTA** (稳定性)`
`对**说话韵律/口音**的捕捉最准。 | **良好** `
`适合克隆特定语气,而非纯粹音色。 |
-| **多语言/方言** | **中/英** (双语优化) | **👑 霸主** (9种语言 + 18种方言) | **中/英** |
-| **语音转换 (VC)** (Audio-to-Audio) | ❌**不支持** `
`仅支持 TTS (Text-to-Speech)。无法改变已有音频的音色。 | ✅**支持** `
`可以将任意音频转换为任意音色 (保留语调/停顿)。 | ❌**不支持** `
`纯 TTS 模型。仅支持 Text-to-Speech。 |
-| **Qwen3-TTS** (1.7B/0.6B) | ✅**支持** `
`支持 Presets (内置音色) 和 VoiceDesign (描述)。 | ✅**支持** `
`支持 3秒极速 Clone (克隆) 模型。 | ✅**支持** `
`支持 10 种语言。 |
-
-#### 3.13 Qwen3-TTS (New! 🔥)
-
-- **用途**: 阿里巴巴 Qwen 团队推出的最新旗舰级 TTS 模型,支持 10 种主要语言及多种方言,具备极高的稳定性和表现力。
-- **核心能力**:
- - **Base (Clone)**: 核心能力为 **3秒极速音色克隆**。支持 X-Vector 模式提升稳定性。
- - **CustomVoice (Presets)**: 阿里巴巴官方提供的 **9 种高品质内置音色** (如 Vivian, Zack 等),支持极强的情感和方言控制。
- - **VoiceDesign**: 通过自然语言描述(如“活泼的少女音,带点羞涩”)从零设计音色。
-- **环境要求**:
- - **qwen-tts**: `pip install qwen-tts` (插件会自动尝试安装)。
- - **Flash Attention 2**: 强烈推荐以获得最佳推理性能。
-- **节点**:
- - `🤖 Qwen3-TTS Loader`: 加载模型。支持 `Base (Clone)`、`CustomVoice (Presets)` 和 `VoiceDesign` 模型。
- - `🗣️ Qwen3-TTS Synthesis`: 执行合成。支持单模型连接或通过 Router 连接的 Bundle。
- - `🔌 Qwen3-Model Router (Bundle)`: **[新]** 路由节点。将多个分立的 Qwen 模型捆绑成一个,供对话节点自动调用。
- - `🎙️ Qwen3-TTS Dialogue (Specialist)`: **[旗帜级]** 专为 Qwen3 设计的对话节点。单输入设计,支持通过 Router 实现混合克隆/捏人。
-- **模型列表**:
- - `Qwen/Qwen3-TTS-12Hz-1.7B-Base` (或 0.6B-Base)
- - `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice` (或 0.6B-CustomVoice)
- - `Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign`
-
-#### 📊 模型功能映射表 (Model Capability Mapping)
-
-| 模型版本 | **音色克隆 (Clone)** | **情感控制 (Emotion)** | **文字捏人 (Design)** | **方言支持 (Dialect)** |
-| :--- | :---: | :---: | :---: | :---: |
-| **Base** (1.7B/0.6B) | **👑 最强** | ❌ 仅限录音自带 | ❌ 不支持 | ⚠️ 仅限录音自带 |
-| **CustomVoice** (1.7B) | ⚠️ 效果极差 | ✅ 支持 | ⚠️ 指令干扰严重 | ⚠️ 效果一般 |
-| **VoiceDesign** (1.7B) | ❌ 不支持 | **👑 专家** | **👑 专家** | **👑 完美支持** |
-| **CustomVoice** (0.6B) | ⚠️ 效果极差 | ✅ 支持 | ✅ 表现优异 | ✅ 表现优异 |
-
-> [!TIP]
-> **关于 UI 简化**:
-> 现在的对话节点只有一个 `qwen_model` 输入槽。
-> - 如果你只需要一种模型,直接连上即可。
-> - 如果你想实现“Speaker A 克隆,Speaker B 捏人”的混合效果,请使用 `🔌 Qwen3-Model Router` 节点进行打包连接。
-
-> [!IMPORTANT]
-> **结论**:
-> 1. 做 **3秒音色克隆**:必须连 `Base` 模型。
-> 2. 说 **方言** 或 **文字定制音色**:优先连 `VoiceDesign`(1.7B)或 `CustomVoice`(0.6B)。
-> 3. 使用 **Vivian/Zack 内置音色**:连接 `CustomVoice` 模型。
-
-**🛠️ 手工下载指南 (Manual Download Guide)**:
-
-如果节点无法自动下载,或您需要在离线环境使用,请手动从 HuggingFace 或 ModelScope 下载模型文件夹,并放入以下目录(文件夹建议保留原名):
-
-```text
-ComfyUI/models/qwen_tts/Qwen/
-├── Qwen3-TTS-12Hz-1.7B-Base/ <-- 对应 1.7B Base (Clone)
-├── Qwen3-TTS-12Hz-1.7B-CustomVoice/ <-- 对应 1.7B CustomVoice
-├── Qwen3-TTS-12Hz-1.7B-VoiceDesign/ <-- 对应 1.7B VoiceDesign
-├── Qwen3-TTS-12Hz-0.6B-Base/ <-- 对应 0.6B Base (Clone)
-└── Qwen3-TTS-12Hz-0.6B-CustomVoice/ <-- 对应 0.6B CustomVoice/VoiceDesign
-```
-
-**下载命令 (HuggingFace CLI)**:
-
-```bash
-mkdir -p models/qwen_tts/Qwen
-
-# 1.7B 系列
-hf download Qwen/Qwen3-TTS-12Hz-1.7B-Base --local-dir models/qwen_tts/Qwen/Qwen3-TTS-12Hz-1.7B-Base
-hf download Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --local-dir models/qwen_tts/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
-hf download Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign --local-dir models/qwen_tts/Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
-
-# 0.6B 系列
-hf download Qwen/Qwen3-TTS-12Hz-0.6B-Base --local-dir models/qwen_tts/Qwen/Qwen3-TTS-12Hz-0.6B-Base
-hf download Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice --local-dir models/qwen_tts/Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice
-```
-
-**ModelScope 下载 (国内推荐)**:
-
-```bash
-# 0.6B 示例
-pip install modelscope
-modelscope download --model qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice --local_dir models/qwen_tts/Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice
-```
-
-> [!NOTE]
-> 对于 **0.6B** 系列,官方目前将 `CustomVoice` 和 `VoiceDesign`(文字设计)能力集成在同一个模型中。因此在设计模式下,加载 `0.6B-CustomVoice` 即可获得极佳效果。
-
-#### 🎭 掌握指令控制 (Instruct Control)
-
-Qwen3-TTS 最强大的特性之一是其**自然语言指令驱动**的能力。与传统的“固定标签”不同,你可以直接在 `instruct` 输入框中用一段描述来控制声音的表现。
-
-**1. 情感与语气控制 (Emotion & Tone)**
-虽然官方没有强制的固定标签列表,但以下描述词被证明效果极佳(支持中文或英文):
-- **基础情感**: "开心" (Happy), "悲伤" (Sad), "生气" (Angry), "兴奋" (Excited), "温柔" (Gentle), "严肃" (Serious)。
-- **微表情控制 (New!)**: 在 Specialist 节点中,你可以叠加更细腻的语气,如 "带点羞涩的" (With a hint of shyness), "语气充满诱惑力" (Seductive tone), "语气带着哭腔" (Crying tone), "语气充满笑意" (Cheerful tone) 等。
-- **提示**: 这些指令可以组合,例如 `生气且激动的。` 或 `Very happy and excited.`
-
-**2. 语速与节奏 (Prosody)**
-虽然节点有专门的 `speed` 滑块,但通过 `instruct` 可以实现更自然的节奏控制:
-- "语速极快" (Very fast speaking rate), "缓慢且深情地" (Slow and soulful), "中间有明显的停顿" (Dramatic pauses)。
-
-**3. 音色设计 (Voice Design)**
-在加载 **VoiceDesign** 模型时,指令框即为你的“捏人”引擎:
-- **特征描述**: "沙哑的男低音" (Raspy deep male voice), "甜美的少女音" (Sweet young girl's voice), "充满磁性的中年女性" (Magnetic middle-aged female)。
-- **示例**: `A young woman with a clear, bright voice, speaking with great confidence.`
-
-- **示例**: `A young woman with a clear, bright voice, speaking with great confidence.`
-- **方言与口音 (Dialect & Accent)**:
- - 虽然官方称全系列支持,但实测发现不同模型遵循度不同:
- - **VoiceDesign (1.7B/0.6B-Custom)**: **👑 效果最强**。因为没有固定身份限制,能完美呈现粤语、上海话等方言的韵律。
- - **CustomVoice (Presets)**: 效果一般。由于 Vivian 等音色有固定的标准语设定,方言指令常会被弱化以维持音色一致性。
- - **Base (Clone)**: 效果最弱。主要取决于你的参考音频本身是什么口音。
-
-**4. 使用技巧**:
-- **句尾符号**: 指令末尾建议加一个句号(如 `开心地。`),这有助于模型更稳定地理解指令边界。
-- **对话剧本**: 在 `🎙️ Qwen3-TTS Dialogue (Specialist)` 节点中,如果某位 Speaker 处于 `Preset` 或 `Design` 模式,系统会自动将剧本中的情感标签(如 `[开心]`)转换为对应的 `instruct` 指令。
-
----
-
----
-
-#### 💡 用户实测与选型指南 (Model Comparison & Selection)
+| 维度 | **VoxCPM 1.5** (800M) | **CosyVoice 3.0** (0.5B/1.5B) | **VibeVoice** (1.5B/7B) |
+| :--------------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :-------------------------------------------------------------------------- | :------------------------------------------------------------- |
+| **音质 (Fidelity)** | **44.1kHz 格式**`
`虽然物理格式为 44.1k,但因采用 **Neural Upsampling** (神经升频) 技术,听感上会有**含混 (Muffled)** 或**金属感**,且伴有底噪。 | **优秀**`
`听感最自然,但采样率稍低 (22/24kHz),有时需 AI 增强。 | **良好**`
`主要强在语气自然度,纯音质略逊。 |
+| **推理速度 (Speed)** | **🚀 冠军 (RTF ~0.17)**`
`得益于 Tokenizer-free,极其高效。 | **极快**`
`流式响应仅 150ms,且支持 TensorRT 加速。 | **一般/较慢**`
`7B 版本较重,更适合离线生成。 |
+| **克隆能力 (Cloning)** | **SOTA** (Zero-Shot)`
`只需 3-10秒,对**音色质感**还原极高。 | **SOTA** (稳定性)`
`对**说话韵律/口音**的捕捉最准。 | **良好**`
`适合克隆特定语气,而非纯粹音色。 |
+| **多语言/方言** | **中/英** (双语优化) | **👑 霸主** (9种语言 + 18种方言) | **中/英** |
+| **语音转换 (VC)** (Audio-to-Audio) | ❌**不支持**`
`仅支持 TTS (Text-to-Speech)。无法改变已有音频的音色。 | ✅**支持**`
`可以将任意音频转换为任意音色 (保留语调/停顿)。 | ❌**不支持**`
`纯 TTS 模型。仅支持 Text-to-Speech。 |
**选型建议**:
- **追求“听起来最像真人” (音质+音色)**: 选 **VoxCPM 1.5**。它的 Tokenizer-free 架构带来了质的飞跃。
- **追求“方言/多语言/稳定性”**: 选 **CosyVoice 3.0**。目前依然是生产环境最稳的选择。
-- **追求“多样化音色设计/最新 Qwen 生态/长语音流畅度”**: 选 **Qwen3-TTS**。其 VoiceDesign 功能能让你用描述语“捏”出从未听过的声音。
- **要做“长篇广播剧/播客”**: 选 **VibeVoice**。它的长窗口上下文优势依然不可替代。
### 4. 播客与对话生成 (Podcast & Dialogue Generation)
@@ -1122,71 +770,29 @@ https://github.com/user-attachments/assets/9a5502c5-79e3-4fc8-8a2d-2cbdbdbbc860
- **TTS Engine**: 后端引擎选择。
- **CosyVoice**: 精准控制型。
- - **Qwen3-TTS**: 万能旗舰型。支持混合模式:通过连接多个 Qwen 模型,可实现在一个对话中同时使用克隆和内置音色。
-- **Qwen Model Pins (Multi-Routing)**:
- - `qwen_model`: 默认主模型。
- - `qwen_base_model` (可选): 连接 `Base` 模型,专门处理有参考音频 (Clone) 的角色。
- - `qwen_custom_model` (可选): 连接 `Custom` 模型,专门处理使用内置 ID (Presets) 的角色。
- - `qwen_design_model` (可选): 连接 `VoiceDesign` 模型,专门处理复杂描述的角色。
+ - **VibeVoice**: 自然演绎型。
+- **Speaker A/B/C**:
- **Ref Audio**: 参考音频 (用于 Zero-Shot 克隆)。
- **ID**: 内置音色 ID (如 CosyVoice 的 `Chinese Female`)。
- **Batch Mode**: 生成模式控制。
- `Natural (Hybrid)`: 混合批处理。仅在 `(Pause)` 处断开。语流最自然,但可能发生音色泄漏。
- `Strict (Per-Speaker)`: 严格模式。每句话都会强制断开重置。彻底杜绝音色泄漏,但对话流畅度略低。
- `Whole (Single Batch)`: 全量模式。无视所有暂停,一次性生成整本剧本。连贯性最强,但无法控制停顿时间。
-- **Batching Parameters**:
- - `max_batch_char` (Default 1500): 单次批处理的最大字符上限。增加此值可大幅提升 Qwen3 的对话连贯性和情感一致性。最高支持模型上限 **32,768**。
-- **Emotion Safeguard (New!)**:
- - **智能检测**: 系统会自动嗅探加载模型的元数据。如果你使用 CosyVoice SFT/Base 或 VibeVoice 等不支持 `Instruct` 功能的模型,系统将自动跳过 `[Emotion]` 标签插入,防止模型读出方括号。
-#### 4.6 AIIA Qwen Dialogue TTS (Qwen 旗舰对话节点)
-
-**[v1.11.0 New]** 深度集成 Qwen3-TTS 的多模式特性,支持复杂的混合角色场景。
-
-- **Parameters**:
- - `seed`: 随机种子。
- - `speed`: 语速调节。
- - `cfg_scale`: 指令遵循强度 (Classifier-Free Guidance)。建议值 1.5 - 7.0。
- - `emotion`: **[v1.11.1 New]** 选中预设情感(开心、悲伤、幽默、愤怒等系统预置微调)。
- - `dialect`: **[v1.11.1 New]** 选中预设方言(粤语、上海话、东北话、四川话等)。
- - `temperature`: 采样温度。
- - `max_batch_char`: 单次批处理上限(最高 32,768)。
-- **Speaker A/B/C Configuration**:
- - **Mode**: 选择 `Clone` (音色克隆)、`Preset` (官方预设) 或 `Design` (文字设计)。
- - **ID**: 当模式为 Preset 时,输入预设音色名 (如 `Vivian`, `Serena`, `Uncle_Fu`, `Dylan`, `Eric`, `Ryan`, `Aiden`, `Ono_Anna`, `Sohee`)。
- - **Expression**: (New!) 为当前角色选择专属微表情描述。
- - **Dialect**: **[v1.11.1 New]** 为当前角色选择方言/口音(支持粤语、上海话、四川话、东北话等)。
- - **Design Description**: 当模式为 Design 时,输入对音色的详细自然语言描述。
- - **Ref Audio**: 当模式为 Clone 时,连接参考音频。
-- **特点**: 相对于通用对话节点,此节点能根据每个人的模式自动路由到最合适的 Qwen 引擎,且支持在 UI 直接输入设计描述。
-
-#### 4.7 AIIA Subtitle Gen (字幕生成器)
+#### 4.3 AIIA Subtitle Gen (字幕生成器)
**[v1.7.0 New]** 无需 STT,直接从生成过程中提取精准时间轴。
- **Input**:
- - `segments_info`: 来自 `AIIA Dialogue TTS` 或 `AIIA Generate Segments` 的输出。
- - `calibration_info` (可选): **[v1.10.2 新增]** 接入 `AIIA Generate Speaker Segments` 的输出。用于将估算的时间轴自动“吸附”到真实的 VAD 语音活动区间,解决 VibeVoice 等批处理引擎的时间轴偏移问题。
+ - `segments_info`: 来自 `AIIA Dialogue TTS` 的输出。
- **Output**:
- `SRT`: 通用字幕格式。
- `ASS`: 高级排版字幕格式 (自动区分角色颜色)。
- **原理**:
- **CosyVoice**: 使用生成时的精确时长。
- **VibeVoice**: 使用**智能插值算法 (Smart Interpolation)**,根据字符长度自动计算长音频段内的单句时间轴。
- - **Qwen3-TTS**: 基于生成的音频振幅精准断句,支持多角色时间轴导出。
-#### 4.4 AIIA Subtitle to Segments (字幕转分段)
-
-**[v1.10.3 New]** 将现有的 SRT/ASS 字幕文件转换为 `segments_info` 格式,以便进行时间轴重新校准。
-
-- **Input**:
- - `subtitle_text`: SRT 或 ASS 格式的文本内容。
- - `subtitle_path` (可选): 字幕文件的本地路径(如果提供,将优先读取文件)。
-- **Output**:
- - `segments_info`: 标准化的 JSON 字符串,可直接输入到 `AIIA Subtitle Gen`。
-- **用途**: 结合 `Subtitle Gen` 的 `calibration_info` 输入,可以将**旧的、不准的字幕**自动对齐到**新的、精准的音轨**上。
-
-#### 4.5 AIIA Subtitle Preview (字幕预览)
+#### 4.4 AIIA Subtitle Preview (字幕预览)
**[v1.7.1 New]** 实时校验音画同步效果。
@@ -1199,31 +805,100 @@ https://github.com/user-attachments/assets/9a5502c5-79e3-4fc8-8a2d-2cbdbdbbc860
#### 4.5 Interactive Teaching (Web Export) (互动式教学导出)
+> [!TIP]
+> **音色泄漏问题?** 如果 VibeVoice 在多角色对话中出现音色混串(Speaker Leakage),请使用下方的 **4.6-4.8 防泄漏管线** 代替直接使用 Dialogue TTS 节点。
+
**[v1.8.1 New]** 将播客升级为视听同步的互动网页。支持“读写分离”的缓存优化,修改 Visual 标签无需重跑 TTS。
- **工作流 (Workflow)**:
- 1. `Script Parser` 输出 `tts_data` (连接到 TTS) 和 `full_script` (连接到 Merge)。
- 2. `AIIA Dialogue TTS` 生成音频和 `segments_info`。
- 3. `AIIA Segment Merge` 将 `full_script` 中的 Visual 标签重新贴回到 `segments_info` 时间轴上。
- 4. `AIIA Web Export` 生成最终 HTML。
+ 1. `Script Parser` 输出 `tts_data` (连接到 TTS) 和 `full_script` (连接到 Merge)。
+ 2. `AIIA Dialogue TTS` 生成音频和 `segments_info`。
+ 3. `AIIA Segment Merge` 将 `full_script` 中的 Visual 标签重新贴回到 `segments_info` 时间轴上。
+ 4. `AIIA Web Export` 生成最终 HTML。
- **Input**:
- - `audio`: 音频信号。
- - `segments_info`: 来自 Merge 节点的包含 Visual 信息的 JSON。
- - `template`: `Split Screen` (适合宽屏) 或 `Presentation` (适合演示)。
+ - `audio`: 音频信号。
+ - `segments_info`: 来自 Merge 节点的包含 Visual 信息的 JSON。
+ - `template`: `Split Screen` (适合宽屏) 或 `Presentation` (适合演示)。
- **Visual Tag 语法**:
- - 在剧本中插入 `(Visual: url)`。
- - 支持绝对 URL: `(Visual: https://example.com)`
- - 支持相对路径: `(Visual: ./slides/01.jpg)` (相对于导出 HTML 的位置)
+ - 在剧本中插入 `(Visual: url)`。
+ - 支持绝对 URL: `(Visual: https://example.com)`
+ - 支持相对路径: `(Visual: ./slides/01.jpg)` (相对于导出 HTML 的位置)
+
+#### 4.6 🎙️ AIIA ASR (通用语音识别)
+
+**[v1.9.0 New]** 基于 FunASR 的通用语音识别节点,提供**字级时间戳**,是防泄漏管线的核心组件。
+
+- **Input**:
+ - `audio`: 待识别的音频信号。
+ - `model`: 选择 ASR 模型(自动扫描 `ComfyUI/models/funasr/` 目录)。
+ - `device`: `cuda` 或 `cpu`。
+ - `batch_size_s`: 动态 batch 大小(秒),越大越快但越占显存。
+ - `hotword` (可选): 热词列表,提高特定词汇识别率。
+- **Output**:
+ - `asr_result` (`ASR_RESULT`): 包含 `text`(完整文本)和 `words`(词级时间戳列表 `[{word, start, end}]`,单位:秒)。
+ - `text` (`STRING`): 识别出的纯文本。
+- **支持模型**:
+ | 模型 | 语言 | 时间戳 | 大小 |
+ |------|------|--------|------|
+ | **paraformer-zh** | 中文(含中英混合) | ✅ 字级 | ~950MB |
+ | SenseVoiceSmall | 中/英/日/韩/粤 | ❌ | ~450MB |
+- **特性**:
+ - 🔄 模型缓存:加载一次,后续调用直接复用。
+ - 🎵 自动重采样:非 16kHz 音频自动转换。
+ - 📝 热词增强:提高专业术语识别率(仅 Paraformer)。
+
+#### 4.7 ✂️ AIIA Podcast Splitter (文本拆分)
+
+**[v1.9.0 New]** 将多角色对话脚本按说话人拆分为独立文本段,用于分轨 TTS。
+
+- **Input**:
+ - `dialogue_json`: 来自 `Script Parser` 的 `dialogue_json` 输出。
+- **Output**:
+ - `speaker_A_text`: Speaker A 的所有台词拼接(换行分隔)。
+ - `speaker_B_text`: Speaker B 的所有台词拼接(换行分隔)。
+ - `split_map`: 原始对话顺序映射 JSON(记录每句话的说话人、索引、文本)。
+- **原理**: 解析对话 JSON,按出场顺序将前两个说话人分别归为 A/B,保留 `(Pause)` 暂停信息。
+
+#### 4.8 🧵 AIIA Podcast Stitcher (音频拼接)
+
+**[v1.9.0 New]** 利用 ASR 时间戳精确切分分轨音频,按原始对话顺序重组为最终播客。**彻底消除 VibeVoice 的音色泄漏问题。**
+
+- **Input**:
+ - `split_map`: 来自 Splitter 的顺序映射。
+ - `audio_A` / `audio_B`: 分别为两个说话人独立生成的 TTS 音频。
+ - `asr_A` / `asr_B`: 分别对应的 ASR 识别结果(含字级时间戳)。
+ - `gap_duration`: 说话人交替时插入的静音时长(默认 0.3s)。
+ - `padding`: 每个切片前后保留的余量,保护呼吸声和尾音(默认 0.05s)。
+- **Output**:
+ - `audio`: 最终拼接好的完整播客音频。
+ - `segments_info`: 包含每个语音段时间轴的 JSON(可直接用于字幕生成)。
+- **核心算法**:
+ 1. **文本-ASR 字符对齐**: 将 ASR 词级文本与原始句子逐字匹配,精确定位每句话在音频中的起止时间。
+ 2. **边界扩展到中点**: 切割点扩展到相邻句间隙的中点,避免截断尾音。
+ 3. **模糊匹配回退**: ASR 识别与原文不完全一致时,使用前缀模糊匹配。
+ 4. **等分回退**: ASR 完全失败时,按字符数等比例分配时间。
+
+#### 🔗 防泄漏管线连线方式 (Anti-Leakage Pipeline)
+
+```text
+ ┌─ speaker_A_text → VibeVoice TTS (A) → audio_A → ASR → asr_A ─┐
+Script Parser → Splitter ─────┤ ├→ Stitcher → Final Audio
+ ├─ speaker_B_text → VibeVoice TTS (B) → audio_B → ASR → asr_B ─┘
+ └─ split_map ──────────────────────────────────────────────────────→
+```
+
+> [!IMPORTANT]
+> **关键原理**: 每个说话人的音频由独立的 TTS 节点生成(各自使用不同的参考音频),从根本上杜绝了音色泄漏。Stitcher 节点再利用 ASR 时间戳精确地将各段重新交错拼接,还原原始对话节奏。
#### 💡 引擎选型与最佳实践 (Best Practices)
-| 特性 | **CosyVoice** | **VibeVoice** | **Qwen3-TTS** |
-| :----------------- | :----------------------------------------- | :----------------------------------------------------------------------------------------------- | :------------------------------------------ |
-| **核心优势** | **精准控制 (Instruction)** | **自然演绎 (Context-Aware)** | **万能旗舰 (Voice Design)** |
-| **情感控制** | ✅**支持** (使用 `[Happy]` 等标签) | ❌ 不支持显式标签 (依赖上下文) | ✅**支持** (通过 `instruct` 或标签) |
-| **生成逻辑** | **逐句生成** (严格遵循每句话的指令) | **混合批处理** (Hybrid Batching) | **动态引擎** (支持流式与批处理) |
-| **最佳场景** | 需要精确指定某句话语气、方言时 | 长篇对话、广播剧、闲聊 | 音色定制、高质量配音、极速克隆 |
-| **使用建议** | 可以在剧本中详细标注情感。 | **尽量减少 `(Pause)`**!`
`让多句对话连在一起,模型能更好地联系上下文产生自然语气。 | 尝试使用其 Voice Design 进行创意捏人。 |
+| 特性 | **CosyVoice** | **VibeVoice** |
+| :----------------- | :----------------------------------------- | :----------------------------------------------------------------------------------------------- |
+| **核心优势** | **精准控制 (Instruction)** | **自然演绎 (Context-Aware)** |
+| **情感控制** | ✅**支持** (使用 `[Happy]` 等标签) | ❌ 不支持显式标签 (依赖上下文) |
+| **生成逻辑** | **逐句生成** (严格遵循每句话的指令) | **混合批处理** (Hybrid Batching) |
+| **最佳场景** | 需要精确指定某句话语气、方言时 | 长篇对话、广播剧、闲聊 |
+| **使用建议** | 可以在剧本中详细标注情感。 | **尽量减少 `(Pause)`**!`
`让多句对话连在一起,模型能更好地联系上下文产生自然语气。 |
#### 📝 综合测试剧本 (Example Script)
@@ -1273,38 +948,7 @@ B: 太神奇了!那我们快去生成试试吧!
- 可自动调整其中一个图像序列的尺寸以匹配另一个,并保持宽高比。
- 可自定义背景填充颜色。
- **输出**: `STRING` (包含所有拼接后帧的新目录路径)。
-
-#### AIIA Image Smart Crop (智能图像裁切)
-
-- **用途**: 一个功能全面的智能裁切节点,专为解决人脸比例、视频构图等问题设计。
-- **场景**: 强烈建议在 **Ditto Sampler** 或其他视频生成节点之前使用,以确保输入图像(特别是人脸)处于最佳位置和比例,避免“嘴巴太大”或“五官漂移”等问题。
-- **参数**:
- - `crop_basis`: 裁切基准。
- - `fixed_width` / `fixed_height`: 锁定一条边 (使用 width/height 参数),另一条边自适应。
- - `fixed_long_side`: **匹配原图长边**。裁切出的长边长度等于原图长边长度 (忽略 width/height 参数)。
- - `fixed_short_side`: **匹配原图短边**。裁切出的短边长度等于原图短边长度 (忽略 width/height 参数)。适合“最大化裁切”。
- - `custom_size`: 强制裁切为指定的 `width` x `height`。
- - `aspect_ratio`: 裁切比例。
- - 默认为 `original` (保持原图比例或使用 custom_size 的宽高)。
- - 可选 `1:1`, `16:9`, `custom` 等。
- - 选择非 original 时,会根据 `crop_basis` 自动计算另一条边的长度。
- - `custom_aspect_ratio`: 自定义比例值 (例如 2.35)。仅在 `aspect_ratio` 选 `custom` 时生效。
- - `position`: 锚点位置 (九宫格)。支持 `center`, `top`, `bottom_left` 等。
- - `offset_x` / `offset_y`: 相对偏移量。用于在自动定位的基础上进行微调 (范围 -1.0 到 1.0)。
-- **输出**: `IMAGE` (裁切后的图像)。
-
----
-
-### 6. 调试与实用工具 (Debug & Utilities)
-
-#### 6.1 文本调试拼接 (Text Debug Splicer)
-
-- **用途**: 方便地将多段文本拼接为一个字符串,支持自定义标题和分隔符,常用于构建和调试复杂的 Prompt 或记录中间结果。
-- **功能**:
- - **多路输入**: 支持最多 3 路文本输入 (`text_1`~`3`) 和自定义标题 (`title`)。
- - **灵活分隔**: 内置多种常用分隔符 (换行、分割线等)。
- - **自动归档**: 支持将拼接结果自动保存为 `.txt` 文件,文件名支持**自定义前缀** (save_prefix),方便回溯。
-- **输出**: `STRING` (拼接后的文本)。
+- **输出**: `STRING` (包含所有拼接后帧的新目录路径)。
---
@@ -1323,63 +967,14 @@ B: 太神奇了!那我们快去生成试试吧!
## Changelog
-### [1.11.0] - 2026-02-04
+### [1.9.0] - 2026-02-15
-- **Qwen3-TTS**: 新增阿里巴巴 **Qwen3-TTS** 全系列支持。
- - **🤖 Qwen3-TTS Loader**: 支持加载 Base, CustomVoice, VoiceDesign 及其 1.7B/0.6B 版本。
- - **🗣️ Qwen3-TTS Synthesis**: 实现全功能生成,包括 Zero-shot 克隆、音色设计和内置音色合成。
-- **Podcast Integration**: **AIIA Dialogue TTS** 节点现在正式集成 Qwen3-TTS 引擎。
- - 支持多角色混合场景下的 Qwen3 驱动,支持使用脚本标签触发 `instruct`。
-- **Auto-Dependency**: 首次运行 Qwen3 节点会自动检测并安装 `qwen-tts` 库。
-
-### [1.10.17] - 2026-02-03
-
-- **Subtitle**: 引入“说话人 ID 为了映射 (Speaker Mapping)”机制。
- - 在字幕校准过程中,系统会建立脚本角色与 VAD 检测角色的对应关系。这确保了即使在短句重叠(Spillover)的情况下,字幕也能强制匹配到正确的说话人音频,避免被相邻的音量大/时长长的角色“抢走”。
-
-### [1.10.16] - 2026-02-02
-
-- **Subtitle**: 优化了多片段合并逻辑,引入“贪婪说话人占用”原则。
- - 对于同一说话人的连续音频片段,只要中间没有被其他说话人占用且停顿小于 3s,都会自动合并到当前行字幕中。解决多句/长句被意外截断的问题。
-
-### [1.10.15] - 2026-02-02
-
-- **Subtitle**: 修复了字幕时间校准逻辑中的 Bug。
- - 增加了说话人一致性检查,防止上一句音频片段(VAD Chunk)被错误地共享给下一个不同说话人的句子,从而导致当前句字幕被截断。
-
-### [1.10.14] - 2026-02-02
-
-- **Ditto Sampler**: 修复了由于 `comfy.model_management` 接口版本差异导致的 `AttributeError`。
-
-### [1.10.13] - 2026-02-02
-
-- **Ditto Sampler**: 修复了采样过程中无法正常响应 ComfyUI 中断/取消信号的问题。
- - 为所有工作线程增加了超时检测和状态轮询,支持在长任务执行期间即时退出。
-
-### [1.10.12] - 2026-02-02
-
-- **Debug & Utilities**: 新增 **Text Debug Splicer** 节点。
- - 支持多路文本拼接、自定义分隔符和自动文件归档。
-
-### [1.9.0] - 2026-01-20
-
-- **Ditto Talking Head**: 新增 Ditto 模型支持 (PyTorch 版)。
- - **AIIA Ditto Loader**: 支持自动下载与加载。
- - **AIIA Ditto Sampler**: 支持内存内流式生成。
-- **EchoMimic V3**: 优化了音频同步逻辑,修复了唇形漂移问题。
-
-### [1.8.4] - 2026-01-19
-
-- **VibeVoice Speed Control**: 实现了基于系统 `sox` 命令的**变速不变调**(Time Stretching)。
-- **稳定性修复**:
- - 修复了 VibeVoice 在调整速度时由于张量类型不匹配(Half vs Float)导致的崩溃。
- - 强制所有音频输出为 `float32`,解决了在 `speed=1.0` 且有参考音频时,下游节点(如 Resemble Enhance)报错的问题。
-- **依赖更新**: 新增 `sox` 依赖。Linux 服务器用户请确保安装系统库:`sudo apt-get install libsox-dev sox`。
-
-### [1.8.3] - 2026-01-07
-
-- **VibeVoice TTS (Standard)**: `reference_audio` 变为可选参数。如果不输入,节点会自动加载内置的高品质女声种子,方便快速测试。
-- **Fix**: 修复 GitHub Actions 发布的子模块错误。
+- **防音色泄漏管线 (Anti-Leakage Pipeline)**: 新增 3 个节点,彻底解决 VibeVoice 多角色对话中的音色混串问题。
+ - **🎙️ AIIA ASR**: 通用语音识别节点,基于 FunASR Paraformer,提供字级时间戳。
+ - **✂️ AIIA Podcast Splitter**: 按说话人拆分对话脚本,输出分轨文本和顺序映射。
+ - **🧵 AIIA Podcast Stitcher**: 利用 ASR 时间戳精确切分分轨音频,按原始对话顺序重组。
+- **Speaker Tag 优化**: VibeVoice 的说话人标签从 `Speaker N:` 改为 `[N]:` 格式,减少注意力泄漏。
+- **FunASR 模型支持**: 支持 Paraformer-zh(中文/字级时间戳)和 SenseVoiceSmall(多语言)。
### [1.8.1] - 2026-01-05
diff --git a/__init__.py b/__init__.py
old mode 100755
new mode 100644
index 658b93d..0ee5ce7
--- a/__init__.py
+++ b/__init__.py
@@ -53,9 +53,13 @@ else:
NODE_DISPLAY_NAME_MAPPINGS.update(module_object.NODE_DISPLAY_NAME_MAPPINGS)
except ImportError as e_import:
+ import traceback
print(f"错误: 导入 {module_alias_for_log} ({module_name_relative}) 失败: {e_import}")
+ traceback.print_exc()
except Exception as e_generic:
+ import traceback
print(f"错误: 在 {module_alias_for_log} ({module_name_relative}) 导入或处理过程中发生错误: {e_generic}")
+ traceback.print_exc()
# 1. 处理 aiia_float_nodes.py
@@ -124,21 +128,30 @@ else:
# 21. 处理 aiia_web_export_nodes.py (新增网页导出)
_load_nodes_from_module(".aiia_web_export_nodes", "aiia_web_export_nodes")
- # 22. 处理 aiia_echomimic_nodes.py (新增 EchoMimic V3)
- _load_nodes_from_module(".aiia_echomimic_nodes", "aiia_echomimic_nodes")
-
- # 23. 处理 aiia_ditto_nodes.py (新增 Ditto)
- _load_nodes_from_module(".aiia_ditto_nodes", "aiia_ditto_nodes")
-
- # 24. 处理 aiia_image_nodes.py (新增 Smart Crop)
- _load_nodes_from_module(".aiia_image_nodes", "aiia_image_nodes")
-
- # 25. 处理 aiia_debug_nodes.py (新增调试节点)
+ # 22. 处理 aiia_debug_nodes.py
_load_nodes_from_module(".aiia_debug_nodes", "aiia_debug_nodes")
- # 26. 处理 aiia_qwen_nodes.py (新增 Qwen3-TTS)
+ # 23. 处理 aiia_ditto_nodes.py (Ditto TTS)
+ _load_nodes_from_module(".aiia_ditto_nodes", "aiia_ditto_nodes")
+
+ # 24. 处理 aiia_echomimic_nodes.py (EchoMimic)
+ _load_nodes_from_module(".aiia_echomimic_nodes", "aiia_echomimic_nodes")
+
+ # 25. 处理 aiia_image_nodes.py (图像工具)
+ _load_nodes_from_module(".aiia_image_nodes", "aiia_image_nodes")
+
+ # 26. 处理 aiia_qwen_nodes.py (Qwen 模型)
_load_nodes_from_module(".aiia_qwen_nodes", "aiia_qwen_nodes")
+ # 27. 处理 aiia_asr_nodes.py (ASR 语音识别)
+ _load_nodes_from_module(".aiia_asr_nodes", "aiia_asr_nodes")
+
+ # 23. 处理 aiia_podcast_splitter.py (播客文本拆分)
+ _load_nodes_from_module(".aiia_podcast_splitter", "aiia_podcast_splitter")
+
+ # 24. 处理 aiia_podcast_stitcher.py (播客音频拼接)
+ _load_nodes_from_module(".aiia_podcast_stitcher", "aiia_podcast_stitcher")
+
# 告诉 ComfyUI 这个节点包有一个包含网页资源的 'js' 目录
WEB_DIRECTORY = "js"
diff --git a/aiia_asr_nodes.py b/aiia_asr_nodes.py
new file mode 100644
index 0000000..c453c33
--- /dev/null
+++ b/aiia_asr_nodes.py
@@ -0,0 +1,219 @@
+import torch
+import os
+import json
+import tempfile
+import numpy as np
+import soundfile as sf
+import folder_paths
+
+# --- 模型路径初始化 ---
+_FUNASR_MODELS_DIR = None
+_AVAILABLE_MODELS = {}
+
+try:
+ _models_base = os.path.join(folder_paths.base_path, "models", "funasr")
+ if os.path.isdir(_models_base):
+ _FUNASR_MODELS_DIR = _models_base
+ for entry in os.listdir(_models_base):
+ full_path = os.path.join(_models_base, entry)
+ if os.path.isdir(full_path):
+ _AVAILABLE_MODELS[entry] = full_path
+ print(f"[AIIA ASR] 发现模型: {entry} -> {full_path}")
+ else:
+ print(f"[AIIA ASR] 警告: funasr 模型目录不存在: {_models_base}")
+except Exception as e:
+ print(f"[AIIA ASR] 模型路径初始化错误: {e}")
+
+
+class AIIA_ASR:
+ """通用 ASR 语音识别节点,基于 FunASR,支持字级时间戳输出。"""
+
+ NODE_NAME = "AIIA ASR"
+ _model_cache = {} # 类级别模型缓存: {model_key: model_instance}
+
+ @classmethod
+ def INPUT_TYPES(cls):
+ model_choices = list(_AVAILABLE_MODELS.keys()) if _AVAILABLE_MODELS else ["NO_MODELS_FOUND"]
+ default_model = "paraformer-zh" if "paraformer-zh" in _AVAILABLE_MODELS else model_choices[0]
+
+ return {
+ "required": {
+ "audio": ("AUDIO",),
+ "model": (model_choices, {"default": default_model}),
+ },
+ "optional": {
+ "device": (["cuda", "cpu"], {"default": "cuda"}),
+ "batch_size_s": ("INT", {
+ "default": 300, "min": 1, "max": 3600, "step": 10,
+ "tooltip": "以秒为单位的动态 batch 大小。越大越快但占用更多显存。"
+ }),
+ "hotword": ("STRING", {
+ "default": "",
+ "tooltip": "热词列表,每行一个词。提高这些词的识别准确率。"
+ }),
+ }
+ }
+
+ RETURN_TYPES = ("ASR_RESULT", "STRING",)
+ RETURN_NAMES = ("asr_result", "text",)
+ FUNCTION = "recognize"
+ CATEGORY = "AIIA/Audio"
+
+ def _ensure_model(self, model_name: str, device: str):
+ """加载或从缓存获取模型实例。"""
+ cache_key = f"{model_name}_{device}"
+ if cache_key in self._model_cache:
+ print(f"[{self.NODE_NAME}] 使用缓存模型: {model_name} on {device}")
+ return self._model_cache[cache_key]
+
+ model_path = _AVAILABLE_MODELS.get(model_name)
+ if not model_path:
+ raise RuntimeError(f"模型 '{model_name}' 未找到。可用模型: {list(_AVAILABLE_MODELS.keys())}")
+
+ from funasr import AutoModel
+
+ # 检测是否为 SenseVoice 系列(需要 trust_remote_code)
+ is_sensevoice = "sensevoice" in model_name.lower()
+
+ print(f"[{self.NODE_NAME}] 加载模型: {model_path} on {device}...")
+ model = AutoModel(
+ model=model_path,
+ device=device,
+ disable_update=True,
+ trust_remote_code=is_sensevoice,
+ )
+ print(f"[{self.NODE_NAME}] 模型加载完成。")
+
+ self._model_cache[cache_key] = model
+ return model
+
+ def _audio_to_numpy(self, audio: dict) -> tuple:
+ """将 ComfyUI AUDIO 格式转换为 16kHz mono numpy 数组。"""
+ waveform = audio["waveform"] # (batch, channels, samples)
+ sample_rate = audio["sample_rate"]
+
+ # 取第一个 batch
+ if waveform.ndim == 3:
+ wav = waveform[0]
+ else:
+ wav = waveform
+
+ # 转 mono
+ if wav.ndim == 2 and wav.shape[0] > 1:
+ wav = wav.mean(dim=0)
+ elif wav.ndim == 2:
+ wav = wav.squeeze(0)
+
+ wav_np = wav.cpu().numpy().astype(np.float32)
+
+ # 重采样到 16kHz(FunASR 要求)
+ if sample_rate != 16000:
+ try:
+ import torchaudio.functional as F
+ wav_tensor = torch.from_numpy(wav_np).unsqueeze(0)
+ wav_resampled = F.resample(wav_tensor, sample_rate, 16000)
+ wav_np = wav_resampled.squeeze(0).numpy()
+ print(f"[{self.NODE_NAME}] 重采样: {sample_rate}Hz -> 16000Hz")
+ except ImportError:
+ # 如果 torchaudio 不可用,写临时文件让 FunASR 自行处理
+ print(f"[{self.NODE_NAME}] 警告: torchaudio 不可用,尝试直接传入音频")
+ sample_rate = 16000
+
+ return wav_np, sample_rate
+
+ def recognize(self, audio, model, device="cuda", batch_size_s=300, hotword=""):
+ log = f"[{self.NODE_NAME}]"
+
+ if model == "NO_MODELS_FOUND":
+ error_result = {
+ "text": "",
+ "words": [],
+ "error": "未找到 FunASR 模型。请将模型放在 ComfyUI/models/funasr/ 目录下。"
+ }
+ return (error_result, "")
+
+ # 验证音频
+ if audio is None or "waveform" not in audio:
+ error_result = {"text": "", "words": [], "error": "输入音频无效"}
+ return (error_result, "")
+
+ wav_np, sr = self._audio_to_numpy(audio)
+ duration = len(wav_np) / sr
+ print(f"{log} 音频时长: {duration:.2f}s, 采样率: {sr}Hz")
+
+ if duration < 0.1:
+ print(f"{log} 音频太短 ({duration:.3f}s),跳过识别")
+ return ({"text": "", "words": []}, "")
+
+ # 加载模型
+ asr_model = self._ensure_model(model, device)
+
+ # 构建生成参数
+ generate_kwargs = {
+ "input": wav_np,
+ "batch_size_s": batch_size_s,
+ }
+
+ # 热词支持(仅 Paraformer 支持)
+ if hotword and hotword.strip() and "paraformer" in model.lower():
+ generate_kwargs["hotword"] = hotword.strip()
+ print(f"{log} 使用热词: {hotword.strip()[:50]}...")
+
+ # SenseVoice 特殊参数
+ if "sensevoice" in model.lower():
+ generate_kwargs["language"] = "auto"
+ generate_kwargs["use_itn"] = True
+
+ print(f"{log} 开始识别...")
+ results = asr_model.generate(**generate_kwargs)
+
+ if not results or len(results) == 0:
+ print(f"{log} 识别结果为空")
+ return ({"text": "", "words": []}, "")
+
+ result = results[0]
+ raw_text = result.get("text", "")
+ raw_timestamps = result.get("timestamp", [])
+
+ # 构建 words 列表
+ words = []
+ if raw_timestamps and raw_text:
+ # FunASR paraformer: text 是空格分隔的词, timestamp 是 [[start_ms, end_ms], ...]
+ text_tokens = raw_text.split()
+ if len(text_tokens) == len(raw_timestamps):
+ for token, ts in zip(text_tokens, raw_timestamps):
+ words.append({
+ "word": token,
+ "start": round(ts[0] / 1000.0, 3), # ms -> s
+ "end": round(ts[1] / 1000.0, 3),
+ })
+ else:
+ print(f"{log} 警告: 词数 ({len(text_tokens)}) 与时间戳数 ({len(raw_timestamps)}) 不匹配")
+ # 尽力匹配
+ for i, ts in enumerate(raw_timestamps):
+ token = text_tokens[i] if i < len(text_tokens) else "?"
+ words.append({
+ "word": token,
+ "start": round(ts[0] / 1000.0, 3),
+ "end": round(ts[1] / 1000.0, 3),
+ })
+
+ # 去掉空格,构建完整文本
+ clean_text = raw_text.replace(" ", "") if raw_text else ""
+
+ asr_result = {
+ "text": clean_text,
+ "words": words,
+ }
+
+ print(f"{log} 识别完成: {len(words)} 个词, 文本: {clean_text[:80]}...")
+ return (asr_result, clean_text)
+
+
+# --- ComfyUI 节点注册 ---
+NODE_CLASS_MAPPINGS = {
+ "AIIA_ASR": AIIA_ASR,
+}
+NODE_DISPLAY_NAME_MAPPINGS = {
+ "AIIA_ASR": "🎙️ AIIA ASR (Word Timestamps)",
+}
diff --git a/aiia_audio_debug.py b/aiia_audio_debug.py
old mode 100755
new mode 100644
diff --git a/aiia_audio_denoise.py b/aiia_audio_denoise.py
old mode 100755
new mode 100644
diff --git a/aiia_audio_enhance.py b/aiia_audio_enhance.py
old mode 100755
new mode 100644
index a74fc99..f07fc6d
--- a/aiia_audio_enhance.py
+++ b/aiia_audio_enhance.py
@@ -624,13 +624,6 @@ class AIIA_Audio_Enhance:
new_splice_info["sample_rate"] = result_sr
new_splice_info["scale_factor"] = scale
- # Cleanup: Move global enhancer to CPU
- try:
- if _cached_enhancer is not None:
- _cached_enhancer.to("cpu")
- torch.cuda.empty_cache()
- except: pass
-
return ({"waveform": processed_wav, "sample_rate": result_sr}, new_splice_info)
NODE_CLASS_MAPPINGS = {
diff --git a/aiia_audio_info.py b/aiia_audio_info.py
old mode 100755
new mode 100644
diff --git a/aiia_audio_isolator.py b/aiia_audio_isolator.py
old mode 100755
new mode 100644
diff --git a/aiia_audio_merger.py b/aiia_audio_merger.py
old mode 100755
new mode 100644
diff --git a/aiia_audio_processor.py b/aiia_audio_processor.py
old mode 100755
new mode 100644
diff --git a/aiia_browser_node.py b/aiia_browser_node.py
old mode 100755
new mode 100644
index ead0b64..daacbb2
--- a/aiia_browser_node.py
+++ b/aiia_browser_node.py
@@ -248,35 +248,6 @@ async def get_video_poster(request):
traceback.print_exc()
return web.Response(status=500, text=str(e))
-async def delete_item(request):
- data = await request.json()
- relative_path_str = data.get("path", "")
- filename = data.get("filename", "")
- if not filename:
- return web.Response(status=400, text="Filename is required.")
-
- try:
- file_path = get_safe_path(output_dir, os.path.join(relative_path_str, filename))
-
- if file_path.is_file():
- os.remove(file_path)
- # Also cleanup potential cache files
- cache_thumb = get_safe_path(image_thumb_dir, os.path.join(relative_path_str, f"{file_path.stem}.jpg"))
- if cache_thumb.exists(): os.remove(cache_thumb)
- cache_poster = get_safe_path(video_poster_dir, os.path.join(relative_path_str, f"{file_path.stem}.jpg"))
- if cache_poster.exists(): os.remove(cache_poster)
-
- return web.json_response({"status": "success", "message": f"File {filename} deleted."})
- elif file_path.is_dir():
- shutil.rmtree(file_path)
- return web.json_response({"status": "success", "message": f"Directory {filename} deleted."})
- else:
- return web.Response(status=404, text="Item not found")
-
- except Exception as e:
- traceback.print_exc()
- return web.Response(status=500, text=str(e))
-
async def get_batch_metadata(request):
data = await request.json()
path = data.get("path", "")
@@ -365,7 +336,6 @@ server.PromptServer.instance.app.router.add_post('/api/aiia/v1/browser/get_batch
server.PromptServer.instance.app.router.add_get('/api/aiia/v1/browser/thumbnail', get_thumbnail)
server.PromptServer.instance.app.router.add_get('/api/aiia/v1/browser/poster', get_video_poster)
server.PromptServer.instance.app.router.add_get('/api/aiia/v1/browser/get_workflow', get_workflow)
-server.PromptServer.instance.app.router.add_post('/api/aiia/v1/browser/delete_item', delete_item)
NODE_CLASS_MAPPINGS = {}
NODE_DISPLAY_NAME_MAPPINGS = {}
\ No newline at end of file
diff --git a/aiia_cosyvoice_nodes.py b/aiia_cosyvoice_nodes.py
old mode 100755
new mode 100644
index f8a742e..ae3061f
--- a/aiia_cosyvoice_nodes.py
+++ b/aiia_cosyvoice_nodes.py
@@ -543,14 +543,14 @@ class AIIA_CosyVoice_TTS:
return {
"required": {
"model": ("COSYVOICE_MODEL",),
- "prompt_label_1": ("STRING", {"default": "Step 1: Enter TTS Text here.", "is_label": True}),
- "tts_text": ("STRING", {"multiline": True, "default": "Hello, this is a test of CosyVoice 3.0."}),
- "prompt_label_2": ("STRING", {"default": "Step 2: Enter Style Description here.", "is_label": True}),
- "instruct_text": ("STRING", {"multiline": True, "default": "Slow speed, magnetic tone, full of emotion.", "tooltip": "Text description to control style/emotion."}),
- "base_gender": (["Female", "Male"], {"default": "Female", "tooltip": "Base gender for description-based synthesis."}),
- "dialect": (["None (Auto)", "Cantonese", "Northeastern", "Sichuan", "Henan", "Tianjin", "Shanghai", "Shandong", "Hubei", "Hunan", "Shaanxi", "Shanxi", "Gansu", "Ningxia", "Hokkien", "Guizhou", "Yunnan", "Jiangxi"], {"default": "None (Auto)", "tooltip": "Preset dialect instruction."}),
- "emotion": (["None (Neutral)", "Happy", "Sad", "Angry", "Robotic", "Peppa Pig"], {"default": "None (Neutral)", "tooltip": "Preset emotion instruction."}),
- "spk_id": ("STRING", {"default": "", "tooltip": "Fixed Speaker ID (e.g. pure_1). Leave empty for Zero-Shot models."}),
+ "提示1_说的内容": ("STRING", {"default": "📖 第一步:在此输入您想让 AI 说的话 (TTS Text)", "is_label": True}),
+ "tts_text": ("STRING", {"multiline": True, "default": "你好,这是 CosyVoice 3.0 的全能模式测试。"}),
+ "提示2_音色描述": ("STRING", {"default": "🎨 第二步:在此输入对表现力/情感的文字描述 (Style Description)", "is_label": True}),
+ "instruct_text": ("STRING", {"multiline": True, "default": "语速非常慢,语气充满磁性,情感饱满。", "tooltip": "文字描述:在 0.5B 中主要控制情感、方言、语速等‘表现风格’,而非从零生成音色身份。"}),
+ "base_gender": (["Female", "Male"], {"default": "Female", "tooltip": "基础性别底色。在“描述生成”模式下,这提供初始的声音身份(性别/音感底色)。"}),
+ "dialect": (["None (Auto)", "广东话 (Cantonese)", "东北话 (Northeastern)", "四川话 (Sichuan)", "河南话 (Henan)", "天津话 (Tianjin)", "上海话 (Shanghai)", "山东话 (Shandong)", "湖北话 (Hubei)", "湖南话 (Hunan)", "陕西话 (Shaanxi)", "山西话 (Shanxi)", "甘肃话 (Gansu)", "宁夏话 (Ningxia)", "闽南话 (Hokkien)", "贵州话 (Guizhou)", "云南话 (Yunnan)", "江西话 (Jiangxi)"], {"default": "None (Auto)", "tooltip": "预设方言指令。会自动添加在自定义描述之前。若与自定义文字描述冲突,模型表现将不可预测。"}),
+ "emotion": (["None (Neutral)", "开心 (Happy)", "伤心 (Sad)", "生气 (Angry)", "机器人的方式 (Robotic)", "小猪佩奇风格 (Peppa Pig)"], {"default": "None (Neutral)", "tooltip": "预设情感指令。会自动添加在自定义描述之前。"}),
+ "spk_id": ("STRING", {"default": "", "tooltip": "固定音色 ID (如 pure_1)。对于 0.5B/V3 等 Zero-Shot 模型,此项通常为空,需配合参考音频使用。"}),
"speed": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0, "step": 0.1}),
"seed": ("INT", {"default": 42, "min": -1, "max": 2147483647}),
},
diff --git a/aiia_cosyvoice_nodes.py.bak b/aiia_cosyvoice_nodes.py.bak
new file mode 100644
index 0000000..b5a7e8b
--- /dev/null
+++ b/aiia_cosyvoice_nodes.py.bak
@@ -0,0 +1,212 @@
+import torch
+import numpy as np
+import os
+import random
+import tempfile
+import soundfile as sf
+import warnings
+import sys
+import subprocess
+import folder_paths
+from huggingface_hub import snapshot_download
+
+# Suppress annoying warnings
+warnings.filterwarnings("ignore", category=FutureWarning)
+warnings.filterwarnings("ignore", category=UserWarning, module="onnxruntime")
+os.environ["KMP_DUPLICATE_LIB_OK"] = "TRUE"
+os.environ["ONNXRUNTIME_QUIET"] = "1"
+
+# Lazy-loaded global variable
+CosyVoice = None
+
+def _install_cosyvoice_if_needed():
+ global CosyVoice
+ if CosyVoice is not None: return
+ try:
+ from cosyvoice.cli.cosyvoice import CosyVoice as CV
+ CosyVoice = CV
+ return
+ except ImportError: pass
+
+ try:
+ libs_dir = os.path.join(os.path.dirname(__file__), "libs")
+ cosyvoice_dir = os.path.join(libs_dir, "CosyVoice")
+ matcha_dir = os.path.join(cosyvoice_dir, "third_party", "Matcha-TTS")
+ if not os.path.exists(libs_dir): os.makedirs(libs_dir, exist_ok=True)
+ if not os.path.exists(cosyvoice_dir):
+ subprocess.check_call(["git", "clone", "--recursive", "https://github.com/FunAudioLLM/CosyVoice.git", cosyvoice_dir])
+ if cosyvoice_dir not in sys.path: sys.path.insert(0, cosyvoice_dir)
+ if matcha_dir not in sys.path: sys.path.insert(0, matcha_dir)
+ from cosyvoice.cli.cosyvoice import CosyVoice as CV
+ CosyVoice = CV
+ except Exception as e:
+ print(f"[AIIA] Failed to install/import CosyVoice: {e}")
+
+class AIIA_CosyVoice_ModelLoader:
+ @classmethod
+ def INPUT_TYPES(cls):
+ return {
+ "required": {
+ "model_name": ([
+ "FunAudioLLM/Fun-CosyVoice3-0.5B-2512",
+ "FunAudioLLM/CosyVoice2-0.5B",
+ "CosyVoice-300M",
+ "CosyVoice-300M-SFT",
+ "CosyVoice-300M-Instruct"
+ ],),
+ "use_fp16": ("BOOLEAN", {"default": True}),
+ }
+ }
+ RETURN_TYPES = ("COSYVOICE_MODEL",)
+ RETURN_NAMES = ("model",)
+ FUNCTION = "load_model"
+ CATEGORY = "AIIA/Loaders"
+
+ def load_model(self, model_name, use_fp16):
+ _install_cosyvoice_if_needed()
+ if model_name.startswith("FunAudioLLM/"):
+ model_dir = os.path.join(folder_paths.models_dir, "cosyvoice", model_name.split("/")[-1])
+ if not os.path.exists(model_dir):
+ snapshot_download(repo_id=model_name, local_dir=model_dir)
+ else:
+ model_dir = os.path.join(folder_paths.models_dir, "cosyvoice", model_name)
+
+ from cosyvoice.cli.cosyvoice import AutoModel
+ is_v3 = os.path.exists(os.path.join(model_dir, "cosyvoice3.yaml"))
+ is_v2 = os.path.exists(os.path.join(model_dir, "cosyvoice2.yaml")) or (not is_v3 and os.path.exists(os.path.join(model_dir, "flow.pt")))
+
+ print(f"[AIIA] Loading {'V3' if is_v3 else ('V2' if is_v2 else 'V1')} model from {model_dir}")
+ model_instance = AutoModel(model_dir=model_dir, fp16=use_fp16)
+
+ # Identity detection
+ available_spks = []
+ spk2info_path = os.path.join(model_dir, "spk2info.pt")
+ if os.path.exists(spk2info_path):
+ try: available_spks = list(torch.load(spk2info_path, map_location='cpu').keys())
+ except: pass
+ if "instruct" in model_dir.lower() and not is_v2 and not is_v3:
+ available_spks = sorted(list(set(available_spks + ["中文男", "中文女", "英文男", "英文女", "日语男", "粤语女", "韩语女"])))
+
+ return ({"model": model_instance, "model_dir": model_dir, "is_v3": is_v3, "is_v2": is_v2, "available_spks": available_spks},)
+
+class AIIA_CosyVoice_V1_TTS:
+ """Specialized node for 300M (V1) models with Surgical Fix for Male voices."""
+ @classmethod
+ def INPUT_TYPES(cls):
+ return {
+ "required": {
+ "model": ("COSYVOICE_MODEL",),
+ "tts_text": ("STRING", {"multiline": True, "default": "你好,这是V1专号节点的测试。"}),
+ "instruct_text": ("STRING", {"multiline": True, "default": "Theo 'Crimson', is a fiery, passionate rebel leader."}),
+ "spk_id": ("STRING", {"default": "中文男"}),
+ "speed": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0, "step": 0.1}),
+ "seed": ("INT", {"default": 42, "min": -1, "max": 2147483647}),
+ },
+ "optional": {
+ "reference_audio": ("AUDIO",),
+ "prompt_text": ("STRING", {"multiline": True, "default": ""}),
+ }
+ }
+ RETURN_TYPES = ("AUDIO",)
+ FUNCTION = "generate"
+ CATEGORY = "AIIA/Synthesis"
+
+ def generate(self, model, tts_text, instruct_text, spk_id, speed, seed, reference_audio=None, prompt_text=""):
+ cosyvoice_model = model["model"]
+ if seed >= 0:
+ torch.manual_seed(seed)
+ if torch.cuda.is_available(): torch.cuda.manual_seed_all(seed)
+
+ # 1. Surgical Fix Logic for V1
+ # Check if it's actually V1
+ if model.get("is_v2") or model.get("is_v3"):
+ print("[AIIA] Warning: V1 node used with V2/V3 model. Falling back to native wrapper.")
+ output = cosyvoice_model.inference_instruct(tts_text, instruct_text, None, speed=speed)
+ else:
+ # PURE V1 SURGICAL PATH
+ if instruct_text:
+ print(f"[AIIA] V1 Surgical Instruct | Spk: {spk_id}")
+ clean_inst = instruct_text.strip().split("<|")[0].strip() + "<|endofprompt|>"
+ def gen():
+ chunks = cosyvoice_model.frontend.text_normalize(tts_text, split=True)
+ for c in chunks:
+ mi = cosyvoice_model.frontend.frontend_instruct(c, spk_id, clean_inst)
+ if 'llm_embedding' in mi: del mi['llm_embedding']
+ for o in cosyvoice_model.model.tts(**mi, stream=False, speed=speed): yield o
+ output = gen()
+ elif reference_audio is not None and prompt_text:
+ print("[AIIA] V1 Zero-shot")
+ with tempfile.NamedTemporaryFile(delete=False, suffix=".wav") as tmp:
+ wav = reference_audio["waveform"].squeeze().cpu().numpy()
+ if wav.ndim == 2: wav = wav.T
+ sf.write(tmp.name, wav, cosyvoice_model.sample_rate)
+ output = cosyvoice_model.inference_zero_shot(tts_text, prompt_text, tmp.name, speed=speed)
+ os.unlink(tmp.name)
+ else:
+ print(f"[AIIA] V1 SFT | Spk: {spk_id}")
+ output = cosyvoice_model.inference_sft(tts_text, spk_id, speed=speed)
+
+ all_speech = [c['tts_speech'] for c in output]
+ final_wav = torch.cat(all_speech, dim=-1)
+ return ({"waveform": final_wav.unsqueeze(0).cpu(), "sample_rate": cosyvoice_model.sample_rate},)
+
+class AIIA_CosyVoice_V2V3_TTS:
+ """Native node for 0.5B (V2/V3) models using official APIs."""
+ @classmethod
+ def INPUT_TYPES(cls):
+ return {
+ "required": {
+ "model": ("COSYVOICE_MODEL",),
+ "tts_text": ("STRING", {"multiline": True, "default": "你好,这是V2/V3专用节点的测试。"}),
+ "instruct_text": ("STRING", {"multiline": True, "default": ""}),
+ "spk_id": ("STRING", {"default": ""}),
+ "speed": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0, "step": 0.1}),
+ "seed": ("INT", {"default": 42, "min": -1, "max": 2147483647}),
+ },
+ "optional": {
+ "reference_audio": ("AUDIO",),
+ }
+ }
+ RETURN_TYPES = ("AUDIO",)
+ FUNCTION = "generate"
+ CATEGORY = "AIIA/Synthesis"
+
+ def generate(self, model, tts_text, instruct_text, spk_id, speed, seed, reference_audio=None):
+ cosyvoice_model = model["model"]
+ if seed >= 0:
+ torch.manual_seed(seed)
+ if torch.cuda.is_available(): torch.cuda.manual_seed_all(seed)
+
+ ref_path = None
+ if reference_audio:
+ with tempfile.NamedTemporaryFile(delete=False, suffix=".wav") as tmp:
+ ref_path = tmp.name
+ wav = reference_audio["waveform"].squeeze().cpu().numpy()
+ if wav.ndim == 2: wav = wav.T
+ sf.write(ref_path, wav, cosyvoice_model.sample_rate)
+
+ try:
+ if model["is_v3"]:
+ print(f"[AIIA] V3 Native | Spk: {spk_id}")
+ output = cosyvoice_model.inference_instruct2(tts_text, instruct_text, ref_path, zero_shot_spk_id=spk_id, speed=speed)
+ else:
+ print(f"[AIIA] V2 Native | Spk: {spk_id}")
+ output = cosyvoice_model.inference_instruct(tts_text, instruct_text, ref_path, zero_shot_spk_id=spk_id, speed=speed)
+
+ all_speech = [c['tts_speech'] for c in output]
+ final_wav = torch.cat(all_speech, dim=-1)
+ finally:
+ if ref_path and os.path.exists(ref_path): os.unlink(ref_path)
+
+ return ({"waveform": final_wav.unsqueeze(0).cpu(), "sample_rate": cosyvoice_model.sample_rate},)
+
+NODE_CLASS_MAPPINGS = {
+ "AIIA_CosyVoice_ModelLoader": AIIA_CosyVoice_ModelLoader,
+ "AIIA_CosyVoice_V1_TTS": AIIA_CosyVoice_V1_TTS,
+ "AIIA_CosyVoice_V2V3_TTS": AIIA_CosyVoice_V2V3_TTS
+}
+NODE_DISPLAY_NAME_MAPPINGS = {
+ "AIIA_CosyVoice_ModelLoader": "CosyVoice Model Loader (AIIA)",
+ "AIIA_CosyVoice_V1_TTS": "CosyVoice V1 (300M) TTS",
+ "AIIA_CosyVoice_V2V3_TTS": "CosyVoice V2/V3 (0.5B+) TTS"
+}
diff --git a/aiia_e2e_diarizer.py b/aiia_e2e_diarizer.py
old mode 100755
new mode 100644
index 9ab6f34..9742cb2
--- a/aiia_e2e_diarizer.py
+++ b/aiia_e2e_diarizer.py
@@ -181,36 +181,11 @@ class AIIA_E2E_Speaker_Diarization:
if not model_path:
return (self._assign_speakers_to_chunks(whisper_chunks, [{"start":0, "end":0, "speaker":f"error_model_not_found_{backend_model}"}]),)
- if audio is None:
- print("错误: [AIIA E2E Diarization] 音频数据为 None")
- return (self._assign_speakers_to_chunks(whisper_chunks, [{"start":0, "end":0, "speaker":"error_no_audio"}]),)
-
- # Handle cases where audio might be passed as a single-item list
- if isinstance(audio, list) and len(audio) > 0:
- audio = audio[0]
-
- # Try to treat as a dictionary or object with waveform/sample_rate
- try:
- waveform = audio["waveform"]
- sample_rate = audio["sample_rate"]
- except (KeyError, TypeError):
- try:
- waveform = getattr(audio, "waveform", None)
- sample_rate = getattr(audio, "sample_rate", None)
- except:
- waveform, sample_rate = None, None
-
- if waveform is None or sample_rate is None:
- print(f"错误: [AIIA E2E Diarization] 音频数据格式错误: 无法获取 waveform 或 sample_rate (输入类型: {type(audio)})")
- return (self._assign_speakers_to_chunks(whisper_chunks, [{"start":0, "end":0, "speaker":"error_no_audio"}]),)
-
- # Ensure waveform is a tensor and sample_rate is a number
- if not isinstance(waveform, torch.Tensor) or not isinstance(sample_rate, (int, float)):
- print(f"错误: [AIIA E2E Diarization] 音频数据类型错误: waveform={type(waveform)}, sample_rate={type(sample_rate)}")
- return (self._assign_speakers_to_chunks(whisper_chunks, [{"start":0, "end":0, "speaker":"error_no_audio"}]),)
-
- if waveform.ndim < 1:
- print("错误: [AIIA E2E Diarization] 音频波形维度不足")
+ if audio is None or not isinstance(audio, dict) or \
+ "waveform" not in audio or not isinstance(audio["waveform"], torch.Tensor) or \
+ "sample_rate" not in audio or not isinstance(audio["sample_rate"], int) or \
+ audio["waveform"].ndim < 1:
+ print("错误: [AIIA E2E Diarization] 音频数据缺失、格式不正确或无效。")
return (self._assign_speakers_to_chunks(whisper_chunks, [{"start":0, "end":0, "speaker":"error_no_audio"}]),)
if not isinstance(whisper_chunks, dict) or not isinstance(whisper_chunks.get("chunks"), list) :
diff --git a/aiia_float_nodes.py b/aiia_float_nodes.py
old mode 100755
new mode 100644
index 14b9930..15717f1
--- a/aiia_float_nodes.py
+++ b/aiia_float_nodes.py
@@ -51,14 +51,6 @@ def _patched_decode_for_in_memory_stack(
img_t_gpu_raw, _ = self_float_model.motion_autoencoder.dec(s_r_plus_motion, alpha=None, feats=s_r_feats)
img_t_gpu_clamped = torch.clamp(img_t_gpu_raw, -1, 1) # 值域 [-1, 1]
- # --- AIIA FIX: Top Edge Cropping ---
- mask_top_edge = getattr(self_float_model, '_aiia_mask_top_edge', 0)
- if mask_top_edge > 0:
- # Crop the top N rows to remove artifacts
- if img_t_gpu_clamped.shape[-2] > mask_top_edge:
- img_t_gpu_clamped = img_t_gpu_clamped[..., mask_top_edge:, :]
- # ----------------------------------
-
gpu_frame_buffer.append(img_t_gpu_clamped.squeeze(0) if B == 1 else img_t_gpu_clamped[0])
if len(gpu_frame_buffer) >= FRAMES_PER_GPU_CHUNK or \
@@ -113,14 +105,6 @@ def _patched_decode_and_save_to_disk(
s_r_plus_motion = s_r + current_motion_vector
img_t_gpu_raw, _ = self_float_model.motion_autoencoder.dec(s_r_plus_motion, alpha=None, feats=s_r_feats)
img_t_gpu_clamped = torch.clamp(img_t_gpu_raw, -1, 1)
-
- # --- AIIA FIX: Top Edge Cropping ---
- mask_top_edge = getattr(self_float_model, '_aiia_mask_top_edge', 0)
- if mask_top_edge > 0:
- if img_t_gpu_clamped.shape[-2] > mask_top_edge:
- img_t_gpu_clamped = img_t_gpu_clamped[..., mask_top_edge:, :]
- # ----------------------------------
-
gpu_frame_buffer.append(img_t_gpu_clamped.squeeze(0) if B == 1 else img_t_gpu_clamped[0])
if len(gpu_frame_buffer) >= FRAMES_PER_GPU_CHUNK_FOR_PROCESSING or \
@@ -167,7 +151,7 @@ class AIIA_FloatProcess_InMemory:
@classmethod
def INPUT_TYPES(cls):
- return {"required": {"float_pipe": ("FLOAT_PIPE",),"ref_image": ("IMAGE",),"ref_audio": ("AUDIO",),"a_cfg_scale": ("FLOAT", {"default": 2.0,"min": 0.0, "max": 10.0, "step": 0.1}),"r_cfg_scale": ("FLOAT", {"default": 1.0,"min": 0.0, "max": 10.0, "step": 0.1}),"e_cfg_scale": ("FLOAT", {"default": 1.0,"min": 0.0, "max": 10.0, "step": 0.1}),"fps": ("FLOAT", {"default": 25.0, "min":1.0, "max": 60.0, "step": 0.5}),"emotion": (['none', 'angry', 'disgust', 'fear', 'happy', 'neutral', 'sad', 'surprise'], {"default": "none"}),"crop_input_image": ("BOOLEAN",{"default":False},),"seed": ("INT", {"default": 0, "min": 0, "max": 0xffffffffffffffff}),"nfe": ("INT", {"default": 10, "min": 1, "max": 100, "step": 1}), },"optional": {"device_override": (["default", "cuda", "cpu"], {"default": "default"}), "decode_gpu_chunk_size": ("INT", {"default": 32, "min":1, "max":128, "step":1, "tooltip":"(In-Memory) GPU解码后一次转移多少帧到CPU。影响显存和速度。"}), "mask_top_edge_pixels": ("INT", {"default": 0, "min": 0, "max": 64, "step": 1, "tooltip": "CROPS the top N rows of pixels to remove artifacts. Output height will be smaller."})}}
+ return {"required": {"float_pipe": ("FLOAT_PIPE",),"ref_image": ("IMAGE",),"ref_audio": ("AUDIO",),"a_cfg_scale": ("FLOAT", {"default": 2.0,"min": 0.0, "max": 10.0, "step": 0.1}),"r_cfg_scale": ("FLOAT", {"default": 1.0,"min": 0.0, "max": 10.0, "step": 0.1}),"e_cfg_scale": ("FLOAT", {"default": 1.0,"min": 0.0, "max": 10.0, "step": 0.1}),"fps": ("FLOAT", {"default": 25.0, "min":1.0, "max": 60.0, "step": 0.5}),"emotion": (['none', 'angry', 'disgust', 'fear', 'happy', 'neutral', 'sad', 'surprise'], {"default": "none"}),"crop_input_image": ("BOOLEAN",{"default":False},),"seed": ("INT", {"default": 0, "min": 0, "max": 0xffffffffffffffff}),"nfe": ("INT", {"default": 10, "min": 1, "max": 100, "step": 1}), },"optional": {"device_override": (["default", "cuda", "cpu"], {"default": "default"}), "decode_gpu_chunk_size": ("INT", {"default": 32, "min":1, "max":128, "step":1, "tooltip":"(In-Memory) GPU解码后一次转移多少帧到CPU。影响显存和速度。"}),}}
def _create_error_image(self, error_message_text: str, log_message: bool = True) -> tuple:
if log_message:
@@ -178,8 +162,7 @@ class AIIA_FloatProcess_InMemory:
a_cfg_scale, r_cfg_scale, e_cfg_scale,
fps, emotion, crop_input_image, seed, nfe,
device_override: str = "default",
- decode_gpu_chunk_size: int = 32,
- mask_top_edge_pixels: int = 0):
+ decode_gpu_chunk_size: int = 32):
node_name_log = f"[{self.__class__.NODE_NAME}]"
print(f"{node_name_log} 流程开始 (内存输出模式)。")
start_time_process = time.time()
@@ -225,8 +208,7 @@ class AIIA_FloatProcess_InMemory:
float_pipe.opt.rank = processing_device.index if processing_device.type == 'cuda' and processing_device.index is not None else (0 if processing_device.type == 'cuda' else -1)
if hasattr(float_pipe.opt, 'fps'): float_pipe.opt.fps = float(fps)
float_pipe.opt.decode_gpu_chunk_size = decode_gpu_chunk_size
- float_pipe.G._aiia_mask_top_edge = mask_top_edge_pixels # Inject param for patch
- print(f"{node_name_log} opt 更新: rank={getattr(float_pipe.opt, 'rank', 'N/A')}, fps={getattr(float_pipe.opt, 'fps', 'N/A')}, decode_chunk={getattr(float_pipe.opt, 'decode_gpu_chunk_size', 'N/A')}, mask_top={mask_top_edge_pixels}")
+ print(f"{node_name_log} opt 更新: rank={getattr(float_pipe.opt, 'rank', 'N/A')}, fps={getattr(float_pipe.opt, 'fps', 'N/A')}, decode_chunk={getattr(float_pipe.opt, 'decode_gpu_chunk_size', 'N/A')}")
model_current_device_before_move = next(float_pipe.G.parameters()).device
if model_current_device_before_move != processing_device: float_pipe.G.to(processing_device)
@@ -269,10 +251,6 @@ class AIIA_FloatProcess_InMemory:
del float_pipe.opt.decode_gpu_chunk_size
except AttributeError:
pass
-
- if hasattr(float_pipe.G, '_aiia_mask_top_edge'):
- try: del float_pipe.G._aiia_mask_top_edge
- except: pass
current_g_device_after_proc = next(float_pipe.G.parameters()).device
if current_g_device_after_proc.type == 'cuda':
@@ -315,8 +293,7 @@ class AIIA_FloatProcess_ToDisk:
fps, emotion, crop_input_image, seed, nfe,
device_override: str = "default",
output_subdir_name: str = "float_frames_AIIA",
- decode_gpu_chunk_size: int = 16,
- mask_top_edge_pixels: int = 0):
+ decode_gpu_chunk_size: int = 16):
node_name_log = f"[{self.__class__.NODE_NAME}]"
print(f"{node_name_log} 流程开始 (输出到磁盘模式)。")
@@ -373,8 +350,7 @@ class AIIA_FloatProcess_ToDisk:
float_pipe.opt.rank = processing_device.index if processing_device.type == 'cuda' and processing_device.index is not None else (0 if processing_device.type == 'cuda' else -1)
if hasattr(float_pipe.opt, 'fps'): float_pipe.opt.fps = float(fps)
float_pipe.opt.frames_per_gpu_chunk_for_processing = decode_gpu_chunk_size
- float_pipe.G._aiia_mask_top_edge = mask_top_edge_pixels # Inject param
- print(f"{node_name_log} opt 更新: rank={getattr(float_pipe.opt, 'rank', 'N/A')}, fps={getattr(float_pipe.opt, 'fps', 'N/A')}, frames_chunk_for_processing={getattr(float_pipe.opt, 'frames_per_gpu_chunk_for_processing', 'N/A')}, mask_top={mask_top_edge_pixels}")
+ print(f"{node_name_log} opt 更新: rank={getattr(float_pipe.opt, 'rank', 'N/A')}, fps={getattr(float_pipe.opt, 'fps', 'N/A')}, frames_chunk_for_processing={getattr(float_pipe.opt, 'frames_per_gpu_chunk_for_processing', 'N/A')}")
model_current_device_before_move = next(float_pipe.G.parameters()).device
if model_current_device_before_move != processing_device: float_pipe.G.to(processing_device)
@@ -439,10 +415,6 @@ class AIIA_FloatProcess_ToDisk:
del float_pipe.opt.frames_per_gpu_chunk_for_processing
except AttributeError:
pass
-
- if hasattr(float_pipe.G, '_aiia_mask_top_edge'):
- try: del float_pipe.G._aiia_mask_top_edge
- except: pass
current_g_device_after_proc = next(float_pipe.G.parameters()).device
if current_g_device_after_proc.type == 'cuda':
diff --git a/aiia_generate_segments.py b/aiia_generate_segments.py
old mode 100755
new mode 100644
index aa99956..2032321
--- a/aiia_generate_segments.py
+++ b/aiia_generate_segments.py
@@ -165,33 +165,11 @@ class AIIA_GenerateSpeakerSegments:
if not model_path:
return self._create_error_output(f"模型 '{e2e_backend_model}' 文件路径无效")
- if audio is None:
- return self._create_error_output("音频数据为 None")
-
- # Handle cases where audio might be passed as a single-item list
- if isinstance(audio, list) and len(audio) > 0:
- audio = audio[0]
-
- # Try to treat as a dictionary or object with waveform/sample_rate
- try:
- waveform = audio["waveform"]
- sample_rate = audio["sample_rate"]
- except (KeyError, TypeError):
- try:
- waveform = getattr(audio, "waveform", None)
- sample_rate = getattr(audio, "sample_rate", None)
- except:
- waveform, sample_rate = None, None
-
- if waveform is None or sample_rate is None:
- return self._create_error_output(f"音频数据格式错误: 无法获取 waveform 或 sample_rate (输入类型: {type(audio)})")
-
- # Ensure waveform is a tensor and sample_rate is a number
- if not isinstance(waveform, torch.Tensor) or not isinstance(sample_rate, (int, float)):
- return self._create_error_output(f"音频数据类型错误: waveform={type(waveform)}, sample_rate={type(sample_rate)}")
-
- if waveform.ndim < 1:
- return self._create_error_output("音频波形维度不足")
+ if audio is None or not isinstance(audio, dict) or \
+ "waveform" not in audio or not isinstance(audio["waveform"], torch.Tensor) or \
+ "sample_rate" not in audio or not isinstance(audio["sample_rate"], int) or \
+ audio["waveform"].ndim < 1:
+ return self._create_error_output("音频数据缺失或无效")
# 检查音频长度
if audio["waveform"].shape[-1] == 0:
@@ -333,19 +311,6 @@ class AIIA_GenerateSpeakerSegments:
print(f"警告: [{self.NODE_NAME}] 最终未能获取任何说话人分段。")
output_data_structure = {"text": "", "chunks": speaker_segments_for_json_chunks, "language": ""}
-
- # Cleanup: Move model to CPU and delete
- try:
- if 'diar_model' in locals() and diar_model is not None:
- print(f"{node_name_log} Cleaning up NeMo model (Moving to CPU)...")
- diar_model.to("cpu")
- if hasattr(diar_model, 'encoder'): diar_model.encoder.to("cpu")
- if hasattr(diar_model, 'decoder'): diar_model.decoder.to("cpu")
- del diar_model
- torch.cuda.empty_cache()
- except Exception as cleanup_err:
- print(f"Warning: Cleanup failed: {cleanup_err}")
-
print(f"{node_name_log} 流程结束。")
return (output_data_structure,)
diff --git a/aiia_personalive_nodes.py b/aiia_personalive_nodes.py
old mode 100755
new mode 100644
diff --git a/aiia_podcast_nodes.py b/aiia_podcast_nodes.py
old mode 100755
new mode 100644
index 7afba54..b8e6a09
--- a/aiia_podcast_nodes.py
+++ b/aiia_podcast_nodes.py
@@ -2,50 +2,6 @@
import json
import re
-AIIA_EMOTION_LIST = [
- "None",
- "Happy (开心)", "Sad (悲伤)", "Angry (愤怒)", "Excited (兴奋)",
- "Gentle (温柔)", "Fearful (恐惧)", "Surprised (惊讶)", "Disappointed (失望)",
- "Proud (骄傲)", "Anxious (焦虑)", "Calm (冷静)", "Neutral (中性)",
- "Affectionate (深情)", "Awkward (尴尬)", "Determined (坚定)", "Hesitant (犹豫)",
- "With a hint of shyness (带点羞涩)",
- "With a hint of a smile (带有一丝笑意)",
- "Seductive tone (充满诱惑力)",
- "Crying tone (带着哭腔)",
- "Cheerful tone (充满笑意)",
- "Serious tone (语气严肃)",
- "Sarcastic tone (冷嘲热讽)",
- "Arrogant tone (语气傲慢)",
- "Cold tone (语气冷淡)",
- "Affectionate tone (充满爱意)",
- "Whispering (轻声耳语)",
- "Shouting (大声叫喊)",
- "Rapid fire (语速较快)",
- "Slow and deliberate (语速较慢)",
- "Tired (疲惫不堪)",
- "Sleepy tone (睡意朦胧)",
- "Drunken tone (醉意微醺)",
- "Professional tone (专业播音)",
- "Magnetic tone (磁性嗓音)",
- "Breathless (气喘吁吁)",
- "Terrified (惊恐万分)",
- "Nervous (紧张不安)",
- "Mysterious (语气神秘)",
- "Enthusiastic (热情高涨)",
- "Lazy tone (语气慵懒)",
- "Gossip tone (八卦语气)",
- "Innocent (语气天真)"
-]
-
-AIIA_DIALECT_LIST = [
- "None",
- "Mandarin (普通话)", "Cantonese (粤语)", "Shanghainese (上海话)",
- "Sichuanese (四川话)", "Northeastern (东北话)", "Hokkien (闽南话)",
- "Hakka (客家话)", "Tianjinese (天津话)", "Shandongnese (山东话)",
- "Henan (河南话)", "Shaanxi (陕西话)", "Hunan (湖南话)", "Jiangxi (江西话)",
- "Hubei (湖北话)", "Guizhou (贵州话)", "Yunnan (云南话)", "Gansu (甘肃话)", "Ningxia (宁夏话)"
-]
-
class AIIA_Podcast_Script_Parser:
def __init__(self):
pass
@@ -219,40 +175,33 @@ class AIIA_Dialogue_TTS:
def INPUT_TYPES(s):
return {
"required": {
- "dialogue_json": ("STRING", {"multiline": True}),
- "tts_engine": (["CosyVoice", "VibeVoice", "Qwen3-TTS"], {"default": "CosyVoice"}),
+ "dialogue_json": ("STRING", {"forceInput": True}),
+ "tts_engine": (["CosyVoice", "VibeVoice"], {"default": "CosyVoice"}),
"pause_duration": ("FLOAT", {"default": 0.5, "min": 0.0, "max": 5.0, "step": 0.1}),
"speed_global": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0}),
"batch_mode": (["Natural (Hybrid)", "Strict (Per-Speaker)", "Whole (Single Batch)"], {"default": "Natural (Hybrid)"}),
- },
- "optional": {
- # Speaker A
- "speaker_A_ref": ("AUDIO",),
- "speaker_A_id": ("STRING", {"default": "", "placeholder": "CosyVoice Internal ID (Optional)"}),
- "speaker_A_emotion": (AIIA_EMOTION_LIST, {"default": "None"}),
- "speaker_A_dialect": (AIIA_DIALECT_LIST, {"default": "None"}),
- # Speaker B
- "speaker_B_ref": ("AUDIO",),
- "speaker_B_id": ("STRING", {"default": "", "placeholder": "CosyVoice Internal ID (Optional)"}),
- "speaker_B_emotion": (AIIA_EMOTION_LIST, {"default": "None"}),
- "speaker_B_dialect": (AIIA_DIALECT_LIST, {"default": "None"}),
-
- # Speaker C
- "speaker_C_ref": ("AUDIO",),
- "speaker_C_id": ("STRING", {"default": "", "placeholder": "CosyVoice Internal ID (Optional)"}),
- "speaker_C_emotion": (AIIA_EMOTION_LIST, {"default": "None"}),
- "speaker_C_dialect": (AIIA_DIALECT_LIST, {"default": "None"}),
-
- # Model Slots and Params (Appended to prevent shift)
- "cosyvoice_model": ("COSYVOICE_MODEL",),
- "vibevoice_model": ("VIBEVOICE_MODEL",),
- "qwen_model": ("QWEN_MODEL",),
- "max_batch_char": ("INT", {"default": 1000, "min": 100, "max": 32768}),
+ # VibeVoice Specific Params
"cfg_scale": ("FLOAT", {"default": 1.5, "min": 1.0, "max": 10.0, "step": 0.1}),
"temperature": ("FLOAT", {"default": 0.8, "min": 0.1, "max": 2.0}),
"top_k": ("INT", {"default": 20, "min": 0, "max": 100}),
- "top_p": ("FLOAT", {"default": 0.95, "min": 0.0, "max": 1.0, "step": 0.05}),
+ "top_p": ("FLOAT", {"default": 0.95, "min": 0.0, "max": 1.0}),
+ },
+ "optional": {
+ "cosyvoice_model": ("COSYVOICE_MODEL",),
+ "vibevoice_model": ("VIBEVOICE_MODEL",),
+
+ # Speaker A
+ "speaker_A_ref": ("AUDIO",),
+ "speaker_A_id": ("STRING", {"default": "", "placeholder": "CosyVoice 内部音色ID (可选)"}),
+
+ # Speaker B
+ "speaker_B_ref": ("AUDIO",),
+ "speaker_B_id": ("STRING", {"default": "", "placeholder": "CosyVoice 内部音色ID (可选)"}),
+
+ # Speaker C
+ "speaker_C_ref": ("AUDIO",),
+ "speaker_C_id": ("STRING", {"default": "", "placeholder": "CosyVoice 内部音色ID (可选)"}),
}
}
@@ -298,103 +247,9 @@ class AIIA_Dialogue_TTS:
print(f"[AIIA Error] Failed to load fallback audio: {e}")
return None
- def _generate_qwen_batch(self, batch_data, qwen_gen, current_full_wav, sr_ptr, segments_info, time_ptr, speed_global, cfg_scale, temperature, top_k, top_p):
- # This helper processes a batch of Qwen items that are compatible (same routed model, etc.)
- # Qwen's `generate` method takes a single text, so we iterate through the batch.
- for i, item_params in enumerate(batch_data):
- target_model = item_params["tm"]
- text = item_params["tx"]
- spk_id = item_params["sid"]
- ref_audio = item_params["ref"]
- instruct = item_params["ins"]
- spk_name = item_params["original_speaker"] # Added this to item_params in get_qwen_params
- original_item = item_params["original_item"] # Added this to item_params in get_qwen_params
-
- print(f" [Qwen Batch] {spk_name}: {text[:30]}... ({target_model['type']})")
- if instruct:
- print(f" [Qwen Instruct] {instruct}")
-
- try:
- # Call Qwen TTS with routed model
- res = qwen_gen.generate(
- qwen_model=target_model,
- text=text,
- language="Auto",
- speaker=spk_id,
- instruct=instruct,
- reference_audio=ref_audio,
- dialect=item_params.get("dialect", "None"),
- seed=42+i, # Use a seed for reproducibility within the batch
- speed=speed_global,
- cfg_scale=cfg_scale,
- temperature=temperature,
- top_k=top_k,
- top_p=top_p
- )
-
- generated = res[0]
- wav = generated["waveform"]
- sr = generated["sample_rate"]
-
- if sr_ptr[0] != sr:
- if current_full_wav:
- wav = torchaudio.transforms.Resample(sr, sr_ptr[0])(wav)
- else:
- sr_ptr[0] = sr
-
- if wav.ndim == 3: wav = wav.squeeze(0)
- if wav.ndim == 1: wav = wav.unsqueeze(0)
-
- # AIIA Fix: Apply tiny fade-in/out to prevent clicks at boundaries
- fade_len = int(sr * 0.05) # 50ms fade
- if wav.shape[-1] > fade_len * 2:
- fade_in = torch.linspace(0, 1, fade_len, device=wav.device)
- fade_out = torch.linspace(1, 0, fade_len, device=wav.device)
- wav[..., :fade_len] *= fade_in
- wav[..., -fade_len:] *= fade_out
-
- current_full_wav.append(wav)
-
- # --- Timestamp Tracking ---
- seg_duration = wav.shape[-1] / sr
- seg_start = time_ptr[0]
- seg_end = seg_start + seg_duration
-
- segments_info.append({
- "start": round(seg_start, 3),
- "end": round(seg_end, 3),
- "text": text,
- "speaker": spk_name,
- "visual": original_item.get("visual")
- })
- time_ptr[0] += seg_duration
-
- # Add a small gap between segments within a Qwen batch
- gap = 0.2
- gap_samples = int(gap * sr_ptr[0])
- current_full_wav.append(torch.zeros(1, gap_samples))
- time_ptr[0] += gap
-
- except Exception as e:
- print(f"[Error] Qwen item generation failed: {e}")
- current_full_wav.append(torch.zeros(1, 24000))
- time_ptr[0] += 1.0
-
-
- def process_dialogue(self, dialogue_json, tts_engine, pause_duration, speed_global, batch_mode, **kwargs):
- # Extract optional and model-specific params from kwargs
- max_batch_char = kwargs.get("max_batch_char", 1000)
- cfg_scale = kwargs.get("cfg_scale", 1.5)
- temperature = kwargs.get("temperature", 0.8)
- top_k = kwargs.get("top_k", 20)
- top_p = kwargs.get("top_p", 0.95)
-
- cosyvoice_model = kwargs.get("cosyvoice_model")
- vibevoice_model = kwargs.get("vibevoice_model")
- qwen_model = kwargs.get("qwen_model")
-
- # Robustness: ensure max_batch_char is correctly picked up even if shifted or provided as kwarg
- max_batch_char = kwargs.get("max_batch_char", max_batch_char)
+ def process_dialogue(self, dialogue_json, tts_engine, pause_duration, speed_global,
+ cosyvoice_model=None, vibevoice_model=None,
+ cfg_scale=1.5, temperature=0.8, top_k=20, top_p=0.95, **kwargs):
import json
import torch
import os
@@ -405,9 +260,6 @@ class AIIA_Dialogue_TTS:
raise ValueError("选择 CosyVoice 引擎时,必须连接 'cosyvoice_model'!")
if tts_engine == "VibeVoice" and vibevoice_model is None:
raise ValueError("选择 VibeVoice 引擎时,必须连接 'vibevoice_model'!")
- if tts_engine == "Qwen3-TTS":
- if qwen_model is None:
- raise ValueError("选择 Qwen3-TTS 引擎时,必须连接 'qwen_model'!(如果需要多个模型,请使用 Router 节点打包)")
dialogue = json.loads(dialogue_json)
full_waveform = []
@@ -415,11 +267,9 @@ class AIIA_Dialogue_TTS:
from .aiia_cosyvoice_nodes import AIIA_CosyVoice_TTS
from .aiia_vibevoice_nodes import AIIA_VibeVoice_TTS
- from .aiia_qwen_nodes import AIIA_Qwen_TTS
cosy_gen = AIIA_CosyVoice_TTS()
vibe_gen = AIIA_VibeVoice_TTS()
- qwen_gen = AIIA_Qwen_TTS()
print(f"[AIIA Podcast] 开始处理对话,共 {len(dialogue)} 个片段。引擎: {tts_engine}")
@@ -476,13 +326,13 @@ class AIIA_Dialogue_TTS:
internal_id = unique_speakers[spk_key]
text = item["text"]
- # VibeVoice does not support emotion macro text tags.
- # We send only pure text to prevent the model from reading tags aloud.
- char_len = len(text) if text else 1
+ # Clean text for length calc (approx)
+ clean_text = re.sub(r'\[.*?\]', '', text).strip()
+ char_len = len(clean_text) if clean_text else 1
total_char_len += char_len
item_lengths.append(char_len)
- final_text_lines.append(f"Speaker {internal_id}: {text}")
+ final_text_lines.append(f"[{internal_id}]: {text}")
full_text = "\n".join(final_text_lines)
print(f" [Batch Process] Processing {len(batch_items)} segments using {len(unique_speakers)} speakers.")
@@ -552,114 +402,20 @@ class AIIA_Dialogue_TTS:
})
time_ptr[0] += 1.0
- elif tts_engine == "Qwen3-TTS":
- # --- Qwen3-TTS Batch Maximization ---
- current_batch = []
- current_batch_char = 0
- current_hash = None
-
- # Batching items by "compatibility"
- # Compatibility = Same routed model, speaker_id, and reference_audio
- def get_qwen_params(it):
- sk = get_speaker_key(it["speaker"])
- tx = it["text"]
- em = it.get("emotion", "None")
- sid = kwargs.get(f"speaker_{sk}_id", "Vivian") # Default to Vivian if empty
- if not sid.strip(): sid = "Vivian"
- ref = get_ref_audio(sk)
- pemf = kwargs.get(f"speaker_{sk}_emotion", "None")
- dia = kwargs.get(f"speaker_{sk}_dialect", "None")
-
- me = em if em and em != "None" else ""
- if pemf and pemf != "None":
- el = pemf.split(" (")[0] if " (" in pemf else pemf
- me = f"{me},{el}" if me else el
- ins = f"{me}。" if me else ""
-
- # Routing: Use bundle if available, else check direct slots
- tm = qwen_model
- if qwen_model and qwen_model.get("is_bundle"):
- if ref is not None: tm = qwen_model.get("base") or qwen_model.get("default")
- elif ins: tm = qwen_model.get("design") or qwen_model.get("default")
- else: tm = qwen_model.get("custom") or qwen_model.get("default")
- elif tm is None:
- # Fallback for deprecated single-slot inputs
- if ref is not None: tm = qwen_base_model or qwen_custom_model
- elif ins: tm = qwen_design_model or qwen_custom_model
- else: tm = qwen_custom_model or qwen_base_model or qwen_design_model
-
- # Dialect is part of compatibility
- return {
- "tm": tm, "tx": tx, "sid": sid, "ref": ref, "ins": ins, "me": me, "sk": sk,
- "dialect": dia,
- "h": (id(tm), dia), # Gouping key
- "original_speaker": it["speaker"],
- "original_item": it # Keep original item for visual tag
- }
-
- for it in batch_items:
- p = get_qwen_params(it)
-
- # Check if the current item is compatible with the current batch
- # Compatibility: same routed model (via hash), and total char count within limit
- can_m = (current_hash is not None and p["h"] == current_hash and (current_batch_char + len(p["tx"]) < max_batch_char))
-
- if not can_m:
- # If not compatible, or if it's the first item, flush the previous batch (if any)
- if current_batch:
- self._generate_qwen_batch(current_batch, qwen_gen, current_full_wav, sr_ptr, segments_info, time_ptr, speed_global, cfg_scale, temperature, top_k, top_p)
-
- # Start a new batch
- current_batch = [p]
- current_batch_char = len(p["tx"])
- current_hash = p["h"]
- else:
- # Add to current batch
- current_batch.append(p)
- current_batch_char += len(p["tx"])
-
- # Flush any remaining items in the last batch
- if current_batch:
- self._generate_qwen_batch(current_batch, qwen_gen, current_full_wav, sr_ptr, segments_info, time_ptr, speed_global, cfg_scale, temperature, top_k, top_p)
-
else:
# CosyVoice (Iterative)
for i, item in enumerate(batch_items):
spk_name = item["speaker"]
spk_key = get_speaker_key(spk_name)
text = item["text"]
- emotion = item.get("emotion") # In CosyVoice, we put it in [] in text
+ emotion = item.get("emotion", "None")
- # Emotion compatibility check
- is_expressive = False
- if cosyvoice_model:
- is_expressive = cosyvoice_model.get("is_instruct") or cosyvoice_model.get("is_v2") or cosyvoice_model.get("is_v3")
-
- if is_expressive:
- # Merge preset emotion
- preset_emo_full = kwargs.get(f"speaker_{spk_key}_emotion", "None")
- if preset_emo_full and preset_emo_full != "None":
- emo_label = preset_emo_full.split(" (")[0] if " (" in preset_emo_full else preset_emo_full
- if emotion: text = f"[{emotion}, {emo_label}] {text}"
- else: text = f"[{emo_label}] {text}"
- elif emotion:
- text = f"[{emotion}] {text}"
-
+ spk_id = kwargs.get(f"speaker_{spk_key}_id", "")
ref_audio = get_ref_audio(spk_key)
- # CosyVoice uses instruct_text for emotion, so we use the merged emotion for it
- merged_emo_for_instruct = ""
- if is_expressive:
- if preset_emo_full and preset_emo_full != "None":
- merged_emo_for_instruct = preset_emo_full.split(" (")[0] if " (" in preset_emo_full else preset_emo_full
- elif item.get("emotion") and item.get("emotion") != "None": # Use script emotion if no preset
- merged_emo_for_instruct = item.get("emotion")
-
- instruct = f"{merged_emo_for_instruct}." if merged_emo_for_instruct else ""
+ instruct = f"{emotion}." if emotion and emotion != "None" else ""
- print(f" [CosyVoice Text] {spk_name}: {text}")
- if instruct:
- print(f" [CosyVoice Instruct] {instruct}")
+ print(f" [Processing] {spk_name}: {text[:15]}...")
try:
res = cosy_gen.generate(
model=cosyvoice_model,
@@ -668,7 +424,7 @@ class AIIA_Dialogue_TTS:
spk_id=spk_id,
speed=speed_global,
seed=42+i,
- dialect=kwargs.get(f"speaker_{spk_key}_dialect", "None (Auto)"),
+ dialect="None (Auto)",
emotion="None (Neutral)",
reference_audio=ref_audio
)
@@ -779,7 +535,7 @@ NODE_CLASS_MAPPINGS = {
}
NODE_DISPLAY_NAME_MAPPINGS = {
- "AIIA_Podcast_Script_Parser": "AIIA Podcast Script Parser",
- "AIIA_Dialogue_TTS": "AIIA Dialogue TTS (Multi-Role)",
- "AIIA_Segment_Merge": "AIIA Segment Merge (Visual)"
+ "AIIA_Podcast_Script_Parser": "📜 AIIA Podcast Script Parser",
+ "AIIA_Dialogue_TTS": "🎧 AIIA Dialogue TTS (Multi-Role)",
+ "AIIA_Segment_Merge": "🔗 AIIA Segment Merge (Visual)"
}
diff --git a/aiia_podcast_splitter.py b/aiia_podcast_splitter.py
new file mode 100644
index 0000000..996b2c9
--- /dev/null
+++ b/aiia_podcast_splitter.py
@@ -0,0 +1,124 @@
+import json
+
+
+class AIIA_Podcast_Splitter:
+ """
+ 将多角色对话脚本按说话人拆分为独立文本。
+
+ 输入: Script Parser 输出的 dialogue_json
+ 输出: 每个说话人的拼接文本(用于独立 TTS)+ 重组映射表
+ """
+
+ NODE_NAME = "AIIA Podcast Splitter"
+
+ @classmethod
+ def INPUT_TYPES(cls):
+ return {
+ "required": {
+ "dialogue_json": ("STRING", {"forceInput": True}),
+ },
+ }
+
+ RETURN_TYPES = ("STRING", "STRING", "STRING",)
+ RETURN_NAMES = ("speaker_A_text", "speaker_B_text", "split_map",)
+ FUNCTION = "split_dialogue"
+ CATEGORY = "AIIA/Podcast"
+
+ def split_dialogue(self, dialogue_json):
+ log = f"[{self.NODE_NAME}]"
+
+ # 解析 dialogue_json
+ try:
+ dialogue = json.loads(dialogue_json)
+ except json.JSONDecodeError as e:
+ print(f"{log} JSON 解析失败: {e}")
+ empty_map = json.dumps([], ensure_ascii=False)
+ return ("", "", empty_map)
+
+ if not isinstance(dialogue, list):
+ print(f"{log} 错误: dialogue_json 不是列表")
+ empty_map = json.dumps([], ensure_ascii=False)
+ return ("", "", empty_map)
+
+ # 收集所有说话人
+ speakers_seen = []
+ for item in dialogue:
+ if item.get("type") == "speech":
+ spk = item["speaker"]
+ if spk not in speakers_seen:
+ speakers_seen.append(spk)
+
+ if len(speakers_seen) == 0:
+ print(f"{log} 警告: 没有找到任何说话人")
+ empty_map = json.dumps([], ensure_ascii=False)
+ return ("", "", empty_map)
+
+ if len(speakers_seen) > 2:
+ print(f"{log} 警告: 发现 {len(speakers_seen)} 个说话人 ({speakers_seen}),仅使用前两个")
+
+ speaker_A = speakers_seen[0] if len(speakers_seen) > 0 else None
+ speaker_B = speakers_seen[1] if len(speakers_seen) > 1 else None
+
+ print(f"{log} Speaker A: {speaker_A}, Speaker B: {speaker_B}")
+
+ # 按说话人分组,同时记录顺序映射
+ texts_A = [] # Speaker A 的所有台词
+ texts_B = [] # Speaker B 的所有台词
+ split_map = [] # 原始顺序映射
+
+ for item in dialogue:
+ if item.get("type") != "speech":
+ # 暂停等非语音条目也记录到 split_map
+ if item.get("type") == "pause":
+ split_map.append({
+ "type": "pause",
+ "duration": item.get("duration", 0.3),
+ })
+ continue
+
+ text = item["text"]
+ speaker = item["speaker"]
+
+ if speaker == speaker_A:
+ split_map.append({
+ "type": "speech",
+ "speaker": "A",
+ "index": len(texts_A),
+ "text": text,
+ "original_speaker": speaker,
+ })
+ texts_A.append(text)
+ elif speaker == speaker_B:
+ split_map.append({
+ "type": "speech",
+ "speaker": "B",
+ "index": len(texts_B),
+ "text": text,
+ "original_speaker": speaker,
+ })
+ texts_B.append(text)
+ else:
+ print(f"{log} 跳过第三个说话人 '{speaker}' 的台词: {text[:30]}...")
+
+ # 拼接每个说话人的文本
+ # 每句之间用换行分隔(TTS 会在换行处产生自然停顿)
+ speaker_A_text = "\n".join(texts_A)
+ speaker_B_text = "\n".join(texts_B)
+
+ split_map_json = json.dumps(split_map, ensure_ascii=False, indent=2)
+
+ print(f"{log} 拆分完成:")
+ print(f" Speaker A ({speaker_A}): {len(texts_A)} 句, {len(speaker_A_text)} 字符")
+ print(f" Speaker B ({speaker_B}): {len(texts_B)} 句, {len(speaker_B_text)} 字符")
+ print(f" split_map: {len(split_map)} 条目")
+
+ return (speaker_A_text, speaker_B_text, split_map_json)
+
+
+# --- ComfyUI 节点注册 ---
+NODE_CLASS_MAPPINGS = {
+ "AIIA_Podcast_Splitter": AIIA_Podcast_Splitter,
+}
+NODE_DISPLAY_NAME_MAPPINGS = {
+ "AIIA_Podcast_Splitter": "✂️ AIIA Podcast Splitter",
+}
diff --git a/aiia_podcast_stitcher.py b/aiia_podcast_stitcher.py
new file mode 100644
index 0000000..14aa231
--- /dev/null
+++ b/aiia_podcast_stitcher.py
@@ -0,0 +1,481 @@
+import json
+import torch
+import numpy as np
+
+
+class AIIA_Podcast_Stitcher:
+ """
+ 将分轨生成的多角色音频按原始对话顺序精确拼接。
+
+ 利用 ASR 词级时间戳找到每句话在音频中的边界,切分后交错拼接。
+ """
+
+ NODE_NAME = "AIIA Podcast Stitcher"
+
+ @classmethod
+ def INPUT_TYPES(cls):
+ return {
+ "required": {
+ "split_map": ("STRING", {"forceInput": True}),
+ "audio_A": ("AUDIO",),
+ "audio_B": ("AUDIO",),
+ "asr_A": ("ASR_RESULT",),
+ "asr_B": ("ASR_RESULT",),
+ },
+ "optional": {
+ "gap_duration": ("FLOAT", {
+ "default": 0.3, "min": 0.0, "max": 2.0, "step": 0.05,
+ "tooltip": "说话人交替时插入的静音时长(秒)"
+ }),
+ "padding": ("FLOAT", {
+ "default": 0.05, "min": 0.0, "max": 0.5, "step": 0.01,
+ "tooltip": "每个切片前后保留的呼吸/尾音余量(秒)"
+ }),
+ }
+ }
+
+ RETURN_TYPES = ("AUDIO", "STRING",)
+ RETURN_NAMES = ("audio", "segments_info",)
+ FUNCTION = "stitch"
+ CATEGORY = "AIIA/Podcast"
+
+ def _audio_to_numpy(self, audio: dict) -> tuple:
+ """将 ComfyUI AUDIO 转为 numpy 数组和采样率。"""
+ waveform = audio["waveform"]
+ sr = audio["sample_rate"]
+
+ if waveform.ndim == 3:
+ wav = waveform[0]
+ else:
+ wav = waveform
+
+ if wav.ndim == 2 and wav.shape[0] > 1:
+ wav = wav.mean(dim=0)
+ elif wav.ndim == 2:
+ wav = wav.squeeze(0)
+
+ return wav.cpu().numpy().astype(np.float32), sr
+
+ def _find_sentence_boundaries(self, asr_words: list, sentences: list, total_duration: float) -> list:
+ """
+ 将 ASR 词级时间戳与原始句子列表对齐,找到每句话在音频中的时间范围。
+
+ 三层匹配策略:
+ 1. 精确子串匹配(去标点后)
+ 2. 编辑距离模糊匹配(滑动窗口,容忍 ASR 错字/漏字)
+ 3. 间隙填补 / 等分回退
+ """
+ log = f"[{self.NODE_NAME}]"
+
+ if not asr_words:
+ print(f"{log} ASR 结果为空,使用等分策略")
+ return self._fallback_equal_split(sentences, total_duration)
+
+ if not sentences:
+ return []
+
+ # 构建 ASR 文本和字符到词索引的映射
+ asr_full_text = ""
+ char_to_word_idx = [] # char_to_word_idx[i] = 该字符属于哪个 word
+ for word_idx, w in enumerate(asr_words):
+ word_text = w["word"]
+ for ch in word_text:
+ char_to_word_idx.append(word_idx)
+ asr_full_text += word_text
+
+ print(f"{log} ASR 全文 ({len(asr_full_text)} 字): {asr_full_text[:100]}...")
+
+ # 为每句话找到在 ASR 文本中的匹配位置
+ boundaries = []
+ search_start = 0 # 保证顺序匹配
+
+ for sent_idx, sentence in enumerate(sentences):
+ # 清理句子文本(去除标点符号和空格,与 ASR 输出对齐)
+ clean_sent = self._clean_text_for_matching(sentence)
+
+ if not clean_sent:
+ print(f"{log} 句子 {sent_idx} 清理后为空: '{sentence}'")
+ boundaries.append(None)
+ continue
+
+ # === 第 1 层:精确子串匹配 ===
+ match_pos = asr_full_text.find(clean_sent, search_start)
+
+ if match_pos != -1:
+ match_end = match_pos + len(clean_sent) - 1
+ match_quality = "精确"
+ else:
+ # === 第 2 层:编辑距离模糊匹配 ===
+ match_pos, match_end, edit_dist = self._fuzzy_find(
+ asr_full_text, clean_sent, search_start
+ )
+
+ if match_pos != -1:
+ match_quality = f"模糊(ed={edit_dist})"
+ else:
+ print(f"{log} 句子 {sent_idx} 无法匹配: '{clean_sent[:30]}...'")
+ boundaries.append(None)
+ continue
+
+ # 映射字符位置到词索引
+ start_word_idx = char_to_word_idx[match_pos] if match_pos < len(char_to_word_idx) else len(asr_words) - 1
+ end_word_idx = char_to_word_idx[min(match_end, len(char_to_word_idx) - 1)]
+
+ start_time = asr_words[start_word_idx]["start"]
+ end_time = asr_words[end_word_idx]["end"]
+
+ print(f"{log} 句子 {sent_idx} [{match_quality}]: "
+ f"'{clean_sent[:15]}' → pos={match_pos}-{match_end}, "
+ f"time={start_time:.2f}-{end_time:.2f}s")
+
+ boundaries.append({
+ "start": start_time,
+ "end": end_time,
+ "start_word_idx": start_word_idx,
+ "end_word_idx": end_word_idx,
+ })
+
+ # 更新搜索起点
+ search_start = match_end + 1
+
+ # 填补未匹配的句子(使用前后句子的时间插值)
+ boundaries = self._fill_missing_boundaries(boundaries, asr_words, total_duration)
+
+ # 扩展边界到句间间隙的中点(避免截断尾音)
+ boundaries = self._expand_to_midpoints(boundaries, total_duration)
+
+ return boundaries
+
+ @staticmethod
+ def _edit_distance(s1: str, s2: str) -> int:
+ """计算两个字符串的编辑距离(Levenshtein distance),使用空间优化的 DP。"""
+ m, n = len(s1), len(s2)
+ if m == 0:
+ return n
+ if n == 0:
+ return m
+
+ # 只需两行
+ prev = list(range(n + 1))
+ curr = [0] * (n + 1)
+
+ for i in range(1, m + 1):
+ curr[0] = i
+ for j in range(1, n + 1):
+ if s1[i - 1] == s2[j - 1]:
+ curr[j] = prev[j - 1]
+ else:
+ curr[j] = 1 + min(prev[j], curr[j - 1], prev[j - 1])
+ prev, curr = curr, prev
+
+ return prev[n]
+
+ def _fuzzy_find(self, haystack: str, needle: str, search_start: int = 0,
+ max_error_ratio: float = 0.4) -> tuple:
+ """
+ 在 haystack 中从 search_start 开始,用滑动窗口+编辑距离找到与 needle 最相似的子串。
+
+ 参数:
+ haystack: ASR 全文
+ needle: 待匹配的原始句子(已去标点)
+ search_start: 搜索起始位置
+ max_error_ratio: 允许的最大错误率(编辑距离 / needle 长度)
+
+ 返回:
+ (match_pos, match_end, edit_distance) 或 (-1, -1, -1) 表示失败
+ """
+ needle_len = len(needle)
+ if needle_len == 0:
+ return (-1, -1, -1)
+
+ max_errors = int(needle_len * max_error_ratio)
+ remaining = haystack[search_start:]
+ remaining_len = len(remaining)
+
+ if remaining_len == 0:
+ return (-1, -1, -1)
+
+ best_pos = -1
+ best_end = -1
+ best_dist = needle_len + 1 # 初始化为一个大值
+
+ # 尝试多种窗口大小(needle 长度的 ±30%),处理 ASR 漏字/多字的情况
+ window_sizes = set()
+ for ratio in [1.0, 0.85, 0.9, 0.95, 1.05, 1.1, 1.15, 1.2]:
+ ws = max(1, int(needle_len * ratio))
+ if ws <= remaining_len:
+ window_sizes.add(ws)
+
+ # 限制搜索范围以避免 O(n²) 爆炸
+ # 在合理的搜索范围内:从 search_start 开始,最多搜到 needle 长度的 3 倍
+ max_search_len = min(remaining_len, needle_len * 3 + 20)
+
+ for window_size in sorted(window_sizes):
+ for i in range(0, max_search_len - window_size + 1):
+ candidate = remaining[i:i + window_size]
+ dist = self._edit_distance(needle, candidate)
+
+ if dist < best_dist:
+ best_dist = dist
+ best_pos = search_start + i
+ best_end = search_start + i + window_size - 1
+
+ # 如果编辑距离为 0 或 1,可以提前退出
+ if dist <= 1:
+ break
+
+ if best_dist <= 1:
+ break
+
+ # 只接受错误率在阈值内的匹配
+ if best_dist <= max_errors:
+ return (best_pos, best_end, best_dist)
+ else:
+ return (-1, -1, -1)
+
+ def _clean_text_for_matching(self, text: str) -> str:
+ """清理文本用于与 ASR 输出匹配:去除标点、空格、英文转小写。"""
+ import re
+ # 去除常见中英文标点和空格
+ cleaned = re.sub(r'[,。!?、;:""''「」【】()《》\s,\.!?\-\;\:\"\'\(\)\[\]\{\}…—~~·]', '', text)
+ # 英文转小写(ASR 可能输出不同大小写)
+ cleaned = cleaned.lower()
+ return cleaned
+
+ def _fallback_equal_split(self, sentences: list, total_duration: float) -> list:
+ """回退策略:按句子字符数等比例分配时间。"""
+ if not sentences:
+ return []
+
+ total_chars = sum(len(s) for s in sentences)
+ if total_chars == 0:
+ segment_duration = total_duration / len(sentences)
+ return [{"start": i * segment_duration, "end": (i + 1) * segment_duration}
+ for i in range(len(sentences))]
+
+ boundaries = []
+ current_time = 0.0
+ for sent in sentences:
+ ratio = len(sent) / total_chars
+ duration = ratio * total_duration
+ boundaries.append({
+ "start": round(current_time, 3),
+ "end": round(current_time + duration, 3),
+ })
+ current_time += duration
+
+ return boundaries
+
+ def _fill_missing_boundaries(self, boundaries: list, asr_words: list, total_duration: float) -> list:
+ """填补未能匹配的句子边界。"""
+ filled = list(boundaries)
+
+ for i in range(len(filled)):
+ if filled[i] is not None:
+ continue
+
+ # 找前一个已知边界
+ prev_end = 0.0
+ for j in range(i - 1, -1, -1):
+ if filled[j] is not None:
+ prev_end = filled[j]["end"]
+ break
+
+ # 找后一个已知边界
+ next_start = total_duration
+ for j in range(i + 1, len(filled)):
+ if filled[j] is not None:
+ next_start = filled[j]["start"]
+ break
+
+ # 在空隙中均匀分配
+ gap_count = 0
+ gap_start_idx = i
+ for j in range(i, len(filled)):
+ if filled[j] is None:
+ gap_count += 1
+ else:
+ break
+
+ gap_duration = (next_start - prev_end) / gap_count
+ for k in range(gap_count):
+ filled[gap_start_idx + k] = {
+ "start": round(prev_end + k * gap_duration, 3),
+ "end": round(prev_end + (k + 1) * gap_duration, 3),
+ }
+
+ return filled
+
+ def _expand_to_midpoints(self, boundaries: list, total_duration: float) -> list:
+ """将切割点扩展到相邻句子间隙的中点,避免截断尾音/吸气声。"""
+ if len(boundaries) <= 1:
+ if boundaries:
+ boundaries[0]["cut_start"] = 0.0
+ boundaries[0]["cut_end"] = total_duration
+ return boundaries
+
+ for i in range(len(boundaries)):
+ if i == 0:
+ boundaries[i]["cut_start"] = 0.0
+ else:
+ # 与前一句的间隙中点
+ gap_mid = (boundaries[i - 1]["end"] + boundaries[i]["start"]) / 2
+ boundaries[i]["cut_start"] = round(gap_mid, 3)
+
+ if i == len(boundaries) - 1:
+ boundaries[i]["cut_end"] = total_duration
+ else:
+ # 与后一句的间隙中点
+ gap_mid = (boundaries[i]["end"] + boundaries[i + 1]["start"]) / 2
+ boundaries[i]["cut_end"] = round(gap_mid, 3)
+
+ return boundaries
+
+ def stitch(self, split_map, audio_A, audio_B, asr_A, asr_B,
+ gap_duration=0.3, padding=0.05):
+ log = f"[{self.NODE_NAME}]"
+
+ # 解析 split_map
+ try:
+ map_items = json.loads(split_map)
+ except json.JSONDecodeError as e:
+ print(f"{log} split_map JSON 解析失败: {e}")
+ return (audio_A, "[]")
+
+ # 提取音频数据
+ wav_A, sr_A = self._audio_to_numpy(audio_A)
+ wav_B, sr_B = self._audio_to_numpy(audio_B)
+ duration_A = len(wav_A) / sr_A
+ duration_B = len(wav_B) / sr_B
+
+ # 使用统一采样率
+ sr = sr_A
+ if sr_A != sr_B:
+ print(f"{log} 警告: sr_A={sr_A} != sr_B={sr_B}, 使用 sr_A")
+
+ print(f"{log} Audio A: {duration_A:.2f}s, Audio B: {duration_B:.2f}s, SR: {sr}")
+
+ # 收集每个说话人的句子列表
+ sentences_A = [item["text"] for item in map_items if item.get("type") == "speech" and item.get("speaker") == "A"]
+ sentences_B = [item["text"] for item in map_items if item.get("type") == "speech" and item.get("speaker") == "B"]
+
+ print(f"{log} 句子数 - A: {len(sentences_A)}, B: {len(sentences_B)}")
+
+ # ASR 对齐切分
+ words_A = asr_A.get("words", []) if isinstance(asr_A, dict) else []
+ words_B = asr_B.get("words", []) if isinstance(asr_B, dict) else []
+
+ print(f"{log} ASR 词数 - A: {len(words_A)}, B: {len(words_B)}")
+
+ boundaries_A = self._find_sentence_boundaries(words_A, sentences_A, duration_A)
+ boundaries_B = self._find_sentence_boundaries(words_B, sentences_B, duration_B)
+
+ print(f"{log} 边界数 - A: {len(boundaries_A)}, B: {len(boundaries_B)}")
+
+ # 按 split_map 顺序拼接
+ audio_segments = []
+ segments_info = []
+ current_time = 0.0
+ idx_A = 0
+ idx_B = 0
+ prev_speaker = None
+
+ for item in map_items:
+ if item.get("type") == "pause":
+ # 显式暂停
+ pause_dur = item.get("duration", 0.3)
+ pause_samples = int(pause_dur * sr)
+ audio_segments.append(np.zeros(pause_samples, dtype=np.float32))
+ current_time += pause_dur
+ continue
+
+ if item.get("type") != "speech":
+ continue
+
+ speaker = item["speaker"]
+
+ # 说话人切换时插入间隙
+ if prev_speaker is not None and speaker != prev_speaker:
+ gap_samples = int(gap_duration * sr)
+ audio_segments.append(np.zeros(gap_samples, dtype=np.float32))
+ current_time += gap_duration
+
+ # 获取对应的边界和音频
+ if speaker == "A":
+ if idx_A >= len(boundaries_A):
+ print(f"{log} 警告: A 的句子索引 {idx_A} 超出边界数 {len(boundaries_A)}")
+ idx_A += 1
+ continue
+ boundary = boundaries_A[idx_A]
+ wav = wav_A
+ idx_A += 1
+ elif speaker == "B":
+ if idx_B >= len(boundaries_B):
+ print(f"{log} 警告: B 的句子索引 {idx_B} 超出边界数 {len(boundaries_B)}")
+ idx_B += 1
+ continue
+ boundary = boundaries_B[idx_B]
+ wav = wav_B
+ idx_B += 1
+ else:
+ continue
+
+ # 切割音频片段(使用 cut_start/cut_end,带 padding)
+ cut_start = boundary.get("cut_start", boundary["start"])
+ cut_end = boundary.get("cut_end", boundary["end"])
+
+ # 应用 padding
+ cut_start = max(0, cut_start - padding)
+ cut_end = min(len(wav) / sr, cut_end + padding)
+
+ start_sample = int(cut_start * sr)
+ end_sample = int(cut_end * sr)
+ end_sample = min(end_sample, len(wav))
+
+ segment = wav[start_sample:end_sample]
+
+ if len(segment) == 0:
+ print(f"{log} 警告: 空片段 at {cut_start:.3f}-{cut_end:.3f}s")
+ continue
+
+ seg_duration = len(segment) / sr
+
+ audio_segments.append(segment)
+
+ # 记录 segment info
+ original_speaker = item.get("original_speaker", speaker)
+ segments_info.append({
+ "start": round(current_time, 3),
+ "end": round(current_time + seg_duration, 3),
+ "text": item["text"],
+ "speaker": original_speaker,
+ })
+
+ current_time += seg_duration
+ prev_speaker = speaker
+
+ # 拼接所有片段
+ if not audio_segments:
+ print(f"{log} 错误: 没有任何音频片段")
+ return (audio_A, "[]")
+
+ final_audio = np.concatenate(audio_segments)
+ total_duration = len(final_audio) / sr
+ print(f"{log} 拼接完成: {total_duration:.2f}s, {len(segments_info)} 个语音段")
+
+ # 转为 ComfyUI AUDIO 格式
+ audio_tensor = torch.from_numpy(final_audio).unsqueeze(0).unsqueeze(0) # (1, 1, samples)
+ audio_output = {"waveform": audio_tensor, "sample_rate": sr}
+
+ segments_info_json = json.dumps(segments_info, ensure_ascii=False, indent=2)
+
+ return (audio_output, segments_info_json)
+
+
+# --- ComfyUI 节点注册 ---
+NODE_CLASS_MAPPINGS = {
+ "AIIA_Podcast_Stitcher": AIIA_Podcast_Stitcher,
+}
+NODE_DISPLAY_NAME_MAPPINGS = {
+ "AIIA_Podcast_Stitcher": "🧵 AIIA Podcast Stitcher",
+}
diff --git a/aiia_subtitle_nodes.py b/aiia_subtitle_nodes.py
old mode 100755
new mode 100644
index 835aff3..0428ab9
--- a/aiia_subtitle_nodes.py
+++ b/aiia_subtitle_nodes.py
@@ -2,7 +2,6 @@ import json
import datetime
import os
import random
-import re
import torchaudio
import folder_paths
@@ -16,12 +15,9 @@ class AIIA_Subtitle_Gen:
"required": {
"segments_info": ("STRING", {"forceInput": True}),
"format": (["SRT", "ASS", "Both"], {"default": "SRT"}),
- "save_file": ("BOOLEAN", {"default": False, "label_on": "Save to Disk", "label_off": "Memory Only"}),
},
"optional": {
- "calibration_info": ("WHISPER_CHUNKS",),
"ass_style": ("STRING", {"default": "Default", "multiline": False}),
- "filename_prefix": ("STRING", {"default": "aiia_subtitle"}),
}
}
@@ -29,9 +25,8 @@ class AIIA_Subtitle_Gen:
RETURN_NAMES = ("srt_content", "ass_content")
FUNCTION = "generate_subtitle"
CATEGORY = "AIIA/Subtitle"
- OUTPUT_NODE = True
- def generate_subtitle(self, segments_info, format="SRT", save_file=False, ass_style="Default", filename_prefix="aiia_subtitle", calibration_info=None):
+ def generate_subtitle(self, segments_info, format="SRT", ass_style="Default"):
try:
segments = json.loads(segments_info)
except Exception as e:
@@ -42,15 +37,6 @@ class AIIA_Subtitle_Gen:
print("[AIIA Subtitle] Segments info must be a list of dicts.")
return ("", "")
- # --- Subtitle Calibration (v1.10.2) ---
- if calibration_info and "chunks" in calibration_info:
- print(f"[AIIA Subtitle] Calibrating {len(segments)} segments using {len(calibration_info['chunks'])} high-precision chunks.")
- segments = self._calibrate_segments(segments, calibration_info["chunks"])
-
- if not isinstance(segments, list):
- print("[AIIA Subtitle] Segments info must be a list of dicts.")
- return ("", "")
-
srt_out = ""
ass_out = ""
@@ -59,36 +45,12 @@ class AIIA_Subtitle_Gen:
if format in ["ASS", "Both"]:
ass_out = self._generate_ass(segments, ass_style)
-
- # File Saving Logic
- if save_file:
- output_dir = folder_paths.get_output_directory()
-
- # Timestamp (uniques)
- timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
-
- if format in ["SRT", "Both"]:
- srt_name = f"{filename_prefix}_{timestamp}.srt"
- srt_path = os.path.join(output_dir, srt_name)
- with open(srt_path, "w", encoding="utf-8") as f:
- f.write(srt_out)
- print(f"[AIIA Subtitle] Saved SRT to: {srt_path}")
-
- if format in ["ASS", "Both"]:
- ass_name = f"{filename_prefix}_{timestamp}.ass"
- ass_path = os.path.join(output_dir, ass_name)
- with open(ass_path, "w", encoding="utf-8") as f:
- f.write(ass_out)
- print(f"[AIIA Subtitle] Saved ASS to: {ass_path}")
return (srt_out, ass_out)
def _generate_srt(self, segments):
output = []
- # Filter out zero-duration (silenced) segments
- valid_segments = [s for s in segments if s["end"] - s["start"] >= 0.01]
-
- for i, seg in enumerate(valid_segments):
+ for i, seg in enumerate(segments):
start = self._format_srt_time(seg["start"])
end = self._format_srt_time(seg["end"])
text = seg["text"]
@@ -100,21 +62,18 @@ class AIIA_Subtitle_Gen:
return "\n".join(output)
def _generate_ass(self, segments, style_name="Default"):
- # Filter out zero-duration (silenced) segments
- valid_segments = [s for s in segments if s["end"] - s["start"] >= 0.01]
-
# 1. Collect unique speakers
speakers = set()
- for seg in valid_segments:
+ for seg in segments:
speakers.add(seg.get("speaker", "Unknown"))
# 2. Assign colors to speakers
# Simple palette: White, Yellow, Cyan, Green, Orange, Pink, LightBlue
palette = [
- "&H00FFFFFF", # White (Pure)
- "&H0000D7FF", # Gold/Yellow (Cinematic)
- "&H00FFFF00", # Cyan (Standard)
- "&H0000FF00", # Green (Lime)
+ "&H00FFFFFF", # White
+ "&H0000FFFF", # Yellow (BGR)
+ "&H00FFFF00", # Cyan
+ "&H0000FF00", # Green
"&H000080FF", # Orange
"&H00FF80FF", # Pink
"&H00FFC0C0" # LightBlue
@@ -125,10 +84,7 @@ class AIIA_Subtitle_Gen:
# Base Style String Template
# Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, ...
- # Base Style String Template
- # Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, ...
- # Changes: Bold=1, Outline=2, Shadow=1 for professional look
- base_style = "Arial,40,{primary_color},&H000000FF,&H00000000,&H00000000,0,0,0,0,100,100,0,0,1,2,1,2,10,10,10,1"
+ base_style = "Arial,20,{primary_color},&H000000FF,&H00000000,&H00000000,0,0,0,0,100,100,0,0,1,1,0,2,10,10,10,1"
sorted_speakers = sorted(list(speakers))
for i, spk in enumerate(sorted_speakers):
@@ -161,12 +117,10 @@ class AIIA_Subtitle_Gen:
]
events = []
- for seg in valid_segments:
+ for seg in segments:
start = self._format_ass_time(seg["start"])
end = self._format_ass_time(seg["end"])
text = seg["text"].replace("\n", "\\N")
- # [v1.10.8] Support Markdown Bold (**text**) -> ASS Bold ({\b1}text{\b0})
- text = re.sub(r'\*\*(.*?)\*\*', r'{\\b1}\1{\\b0}', text)
speaker = seg.get("speaker", "Unknown")
# Use the mapped style name
style_for_event = speaker_map.get(speaker, "Default")
@@ -188,260 +142,6 @@ class AIIA_Subtitle_Gen:
return f"{hours:02}:{minutes:02}:{secs:02},{millis:03}"
- def _calibrate_segments(self, segments, chunks):
- """
- Calibrate estimated segments using high-precision VAD chunks.
- Algorithm: Iterative sequence matching with speaker-centric isolation (v1.10.5).
- """
- if not chunks:
- return segments
-
- # 1. Ensure chunks are sorted chronologically
- sorted_chunks = sorted(chunks, key=lambda x: x["timestamp"][0])
-
- calibrated = []
- chunk_idx = 0
- num_chunks = len(sorted_chunks)
-
- def normalize_spk(s):
- if not s: return ""
- return str(s).lower().replace("speaker_", "").replace("speaker ", "").strip()
-
- # [v1.10.17] Speaker Mapping: Script Identity -> VAD Identity
- speaker_map = {} # e.g. {"speaker_a": "speaker_00", "speaker_b": "speaker_01"}
-
- for i, seg in enumerate(segments):
- seg_start = seg["start"]
- seg_end = seg["end"]
- seg_dur = seg_end - seg_start
- seg_spk = normalize_spk(seg.get("speaker"))
-
- # --- Speaker-Centric Magic (v1.10.5) ---
- # 1. Find the "Winner Speaker" for this segment based on maximum overlap duration
- speaker_overlaps = {}
- # Large window for initial scan to be robust
- scan_idx = chunk_idx
- while scan_idx < num_chunks:
- c = sorted_chunks[scan_idx]
- c_start, c_end = c["timestamp"]
- # Hard break if the chunk is way past our segment (relaxed to 60s for sequential matching)
- if c_start > seg_end + 60.0: break
-
- # Calculate overlap duration
- overlap = min(seg_end, c_end) - max(seg_start, c_start)
- if overlap > 0:
- spk = normalize_spk(c.get("speaker", "unknown"))
- speaker_overlaps[spk] = speaker_overlaps.get(spk, 0.0) + overlap
- scan_idx += 1
-
- winner_spk = None
-
- # [v1.10.17] Consistency Logic
- # If we already know who this Script Speaker maps to, try to find THAT VAD Speaker first.
- if seg_spk in speaker_map:
- mapped_vad_spk = speaker_map[seg_spk]
- # Trust the map (Sequence Priority). Even if no overlap, we search for this speaker.
- winner_spk = mapped_vad_spk
-
- # Fallback (First time seeing this speaker OR mapped speaker missing): Max Overlap
- if winner_spk is None and speaker_overlaps:
- # Get speaker with most accumulated overlap duration
- winner_spk = max(speaker_overlaps, key=speaker_overlaps.get)
-
- # Update Map
- if winner_spk and seg_spk:
- speaker_map[seg_spk] = winner_spk
-
- # 2. Find chunks belonging to the winner spk to use for snapping
- # [v1.10.16] Greedy Speaker Turn Collection: collect all consecutive chunks for winner_spk
- matched_chunks = []
- find_idx = chunk_idx
- found_start = False
-
- while find_idx < num_chunks:
- chunk = sorted_chunks[find_idx]
- c_start, c_end = chunk["timestamp"]
- c_spk = normalize_spk(chunk.get("speaker", "unknown"))
-
- overlap = min(seg_end, c_end) - max(seg_start, c_start)
- is_overlap = overlap > 0.05
-
- # Case: First segment special snapping
- if not is_overlap and i == 0 and find_idx == 0:
- if abs(c_start - seg_start) < 0.5 and c_end > seg_start - 0.1:
- is_overlap = True
-
- if not found_start:
- # Looking for the first chunk that belongs to our speaker
- # Trust Sequence: If we have a winner_spk, find their next chunk regardless of overlap
- match_condition = is_overlap
- if winner_spk is not None:
- match_condition = is_overlap or (c_spk == winner_spk)
-
- if match_condition and (winner_spk is None or c_spk == winner_spk):
- matched_chunks.append(find_idx)
- found_start = True
- if winner_spk is None: winner_spk = c_spk
- elif c_start > seg_end + 60.0:
- # Way past the estimated end, give up
- break
- else:
- # Already started collecting.
- if c_spk == winner_spk:
- # Ensure no massive gap (e.g. 8s) within a single turn, unless it's within expected segment time
- last_matched_end = sorted_chunks[matched_chunks[-1]]["timestamp"][1]
- if (c_start - last_matched_end < 8.0) or (c_start < seg_end + 1.0):
- matched_chunks.append(find_idx)
- else:
- break
- else:
- # [v1.10.18 Fix] Encountered unrelated speaker.
- # Check if this is just a brief interruption (noise/other speaker)
- # and if our speaker resumes shortly.
- # Mismatch!
- # [v1.10.19] Check if this is a real turn of another speaker.
- # If a different speaker talks for more than 0.5s, we must yield the turn.
- interruption_dur = c["timestamp"][1] - c["timestamp"][0]
- if interruption_dur > 0.5:
- break
-
- # Lookahead mechanism
- resume_idx = -1
- lookahead_limit = 5 # Check next 5 chunks
-
- for k in range(1, lookahead_limit + 1):
- next_idx = find_idx + k
- if next_idx >= num_chunks: break
-
- nc = sorted_chunks[next_idx]
- nc_spk = normalize_spk(nc.get("speaker", "unknown"))
- nc_start = nc["timestamp"][0]
-
-
- # If we find our speaker again
- if nc_spk == winner_spk:
- # Check if the gap is acceptable:
- # 1. Short interruption (< 0.8s)
- # 2. OR the resume chunk starts reasonably close to the EXPECTED end of the segment.
- # (This handles cases where VAD has a large gap but Script says it should be one segment)
- last_matched_end = sorted_chunks[matched_chunks[-1]]["timestamp"][1]
- if (nc_start - last_matched_end < 0.8) or (nc_start < seg_end + 1.0):
- resume_idx = next_idx
- break
-
- # If we hit a substantial chunk of another speaker, stop looking
- if nc["timestamp"][1] - nc["timestamp"][0] > 0.5:
- break
-
- if resume_idx != -1:
- # Resume found! Skip intermediate chunks.
- # We do NOT add the intermediate chunks to matched_chunks (they belong to noise/others)
- # But we continue the loop from the resume point.
-
- # Note: The intermediate chunks are effectively "skipped" by this segment.
- # If they were important for another segment, that segment logic needs to handle them.
- # But typically, if they are "interruptions" inside a sentence, they are noise.
-
- # Advance find_idx to just before resume_idx (loop will increment)
- find_idx = resume_idx - 1
- # We will pick up the resumed chunk in next iteration
- else:
- # comprehensive stop
- break
-
- find_idx += 1
- # Limit lookahead for safety (total span)
- if found_start and (find_idx - matched_chunks[0] > 50): break
- if not found_start and (find_idx - chunk_idx > 20): break
-
- if matched_chunks:
- # Use min/max over all matched chunks
- actual_starts = [sorted_chunks[idx]["timestamp"][0] for idx in matched_chunks]
- actual_ends = [sorted_chunks[idx]["timestamp"][1] for idx in matched_chunks]
-
- min_s = min(actual_starts)
- max_e = max(actual_ends)
-
- # [v1.10.7 Fix] Handle Multi-Segment Chunks (Shared Chunk Logic)
- # If we are reusing a chunk from previous segment, we must start AFTER previous segment
- new_start = min_s
- if i > 0:
- prev_end = calibrated[-1]["end"]
- if prev_end > new_start and prev_end < max_e:
- new_start = prev_end
-
- # Determine if we should consume the chunk or share it
- # Check if next segment also wants this chunk (overlaps with the tail of this chunk)
- is_shared = False
- last_matched_idx = max(matched_chunks)
- chunk_end_time = sorted_chunks[last_matched_idx]["timestamp"][1]
-
- # Predicted end for this segment
- predicted_end = new_start + seg_dur
-
- # Only check for sharing if there is significant leftover time in the chunk
- # AND if the next segment belongs to the same speaker (critical fix v1.10.15)
- if chunk_end_time - predicted_end > 0.5 and i + 1 < len(segments):
- next_seg = segments[i+1]
- # If next segment effectively overlaps the remainder of this chunk
- if next_seg["start"] < chunk_end_time:
- # [v1.10.15] Ensure speaker match before sharing
- current_spk = normalize_spk(seg.get("speaker"))
- next_spk = normalize_spk(next_seg.get("speaker"))
- if current_spk == next_spk:
- is_shared = True
-
- if is_shared:
- # If shared, we limit our end to our duration (trust TTS relative duration)
- new_end = predicted_end
- # And we DO NOT advance past this chunk, so next segment can pick it up
- chunk_idx = last_matched_idx
- else:
- # If not shared, we consume the full VAD chunk (snap to VAD end)
- new_end = max_e
- chunk_idx = last_matched_idx + 1
-
- seg["start"] = round(new_start, 3)
- seg["end"] = round(new_end, 3)
- else:
- # No match found for the specific speaker.
- # [v1.10.19] Silence-Aware Fallback:
- # Check if there is ANY speech (any speaker) in the estimated range.
- has_speech = False
- fallback_scan_idx = chunk_idx
- while fallback_scan_idx < num_chunks:
- fc = sorted_chunks[fallback_scan_idx]
- fc_start, fc_end = fc["timestamp"]
- if fc_start > seg["end"]: break # Past our range
-
- # Check overlap
- overlap = min(seg["end"], fc_end) - max(seg["start"], fc_start)
- if overlap > 0.1: # Threshold for "speech exists"
- has_speech = True
- break
-
- fallback_scan_idx += 1
-
- start_point = calibrated[-1]["end"] if i > 0 else 0.0
-
- if has_speech:
- # Someone is speaking, so we keep the text but shift it to avoid overlap
- duration = seg["end"] - seg["start"]
- if seg["start"] < start_point:
- seg["start"] = start_point
- seg["end"] = start_point + duration
- else:
- # [Silence Detected]
- # User feedback: segments_info is the "Plan". Even if silent, the dialogue must be shown to preserve order.
- # We place the subtitle in the gap (after previous segment) with its estimated duration.
- duration = seg["end"] - seg["start"]
- seg["start"] = start_point
- seg["end"] = start_point + duration
-
- calibrated.append(seg)
-
- return calibrated
-
def _format_ass_time(self, seconds):
# H:MM:SS.cs (centiseconds)
td = datetime.timedelta(seconds=seconds)
@@ -497,128 +197,12 @@ class AIIA_Subtitle_Preview:
return {"ui": {"text": [subtitle_content], "audio": [audio_info] if audio_info else []}}
-class AIIA_Subtitle_To_Segments:
- """Convert SRT/ASS text or files into segments_info format."""
- @classmethod
- def INPUT_TYPES(cls):
- return {
- "required": {
- "subtitle_text": ("STRING", {"multiline": True, "default": ""}),
- },
- "optional": {
- "subtitle_path": ("STRING", {"default": ""}),
- }
- }
-
- RETURN_TYPES = ("STRING",)
- RETURN_NAMES = ("segments_info",)
- FUNCTION = "convert"
- CATEGORY = "AIIA/Subtitle"
-
- def convert(self, subtitle_text, subtitle_path=""):
- import re
- content = subtitle_text.strip()
-
- # If path provided and exists, read it
- if subtitle_path and os.path.exists(subtitle_path):
- try:
- with open(subtitle_path, 'r', encoding='utf-8', errors='ignore') as f:
- content = f.read().strip()
- except Exception as e:
- print(f"[AIIA Subtitle Convert] Error reading file: {e}")
-
- if not content:
- return (json.dumps([]),)
-
- segments = []
-
- # Detect Format
- if "Dialogue:" in content:
- segments = self._parse_ass(content)
- elif " --> " in content:
- segments = self._parse_srt(content)
- else:
- print("[AIIA Subtitle Convert] Unknown format or empty content.")
-
- return (json.dumps(segments, ensure_ascii=False, indent=2),)
-
- def _parse_srt(self, text):
- import re
- segments = []
- # Pattern: Index, Time, Text
- # Handles \n and \r\n
- blocks = re.split(r'\n\s*\n', text.strip())
- for block in blocks:
- lines = [l.strip() for l in block.split('\n') if l.strip()]
- if len(lines) < 2: continue
-
- # Find time line
- time_match = re.search(r'(\d+:\d+:\d+,\d+) --> (\d+:\d+:\d+,\d+)', lines[0] if "-->" in lines[0] else lines[1])
- if not time_match: continue
-
- start_s = self._time_to_seconds(time_match.group(1), "srt")
- end_s = self._time_to_seconds(time_match.group(2), "srt")
-
- # Content is everything after the time line
- idx = 1 if "-->" in lines[0] else 2
- content = " ".join(lines[idx:])
-
- segments.append({
- "start": round(start_s, 3),
- "end": round(end_s, 3),
- "text": content,
- "speaker": "Unknown"
- })
- return segments
-
- def _parse_ass(self, text):
- import re
- segments = []
- # Look for Dialogue: lines
- for line in text.split('\n'):
- if line.startswith("Dialogue:"):
- # Dialogue: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text
- parts = line.split(',', 9)
- if len(parts) < 10: continue
-
- start_s = self._time_to_seconds(parts[1].strip(), "ass")
- end_s = self._time_to_seconds(parts[2].strip(), "ass")
- speaker = parts[4].strip() or "Unknown"
- content = parts[9].strip().replace('\\N', ' ').replace('\\n', ' ')
- # Clean ASS tags like {\pos(x,y)}
- content = re.sub(r'\{.*?\}', '', content)
-
- segments.append({
- "start": round(start_s, 3),
- "end": round(end_s, 3),
- "text": content,
- "speaker": speaker
- })
- return segments
-
- def _time_to_seconds(self, t_str, fmt):
- try:
- if fmt == "srt":
- # HH:MM:SS,mmm
- h, m, s_ms = t_str.split(':')
- s, ms = s_ms.split(',')
- return int(h)*3600 + int(m)*60 + int(s) + int(ms)/1000.0
- else:
- # H:MM:SS.cc
- h, m, s_cs = t_str.split(':')
- s, cs = s_cs.split('.')
- return int(h)*3600 + int(m)*60 + int(s) + int(cs)/100.0
- except:
- return 0.0
-
NODE_CLASS_MAPPINGS = {
"AIIA_Subtitle_Gen": AIIA_Subtitle_Gen,
- "AIIA_Subtitle_Preview": AIIA_Subtitle_Preview,
- "AIIA_Subtitle_To_Segments": AIIA_Subtitle_To_Segments
+ "AIIA_Subtitle_Preview": AIIA_Subtitle_Preview
}
NODE_DISPLAY_NAME_MAPPINGS = {
- "AIIA_Subtitle_Gen": "AIIA Subtitle Generation",
- "AIIA_Subtitle_Preview": "AIIA Subtitle Preview",
- "AIIA_Subtitle_To_Segments": "AIIA Subtitle to Segments"
+ "AIIA_Subtitle_Gen": "📝 AIIA Subtitle Generation",
+ "AIIA_Subtitle_Preview": "🎬 AIIA Subtitle Preview"
}
diff --git a/aiia_utils_nodes.py b/aiia_utils_nodes.py
old mode 100755
new mode 100644
diff --git a/aiia_vibevoice_nodes.py b/aiia_vibevoice_nodes.py
old mode 100755
new mode 100644
index 06de81a..7997927
--- a/aiia_vibevoice_nodes.py
+++ b/aiia_vibevoice_nodes.py
@@ -4,9 +4,6 @@ import torch
import numpy as np
import torchaudio
import folder_paths
-import subprocess
-import tempfile
-import soundfile as sf
# print(f"\n[AIIA DEBUG] Loaded aiia_vibevoice_nodes.py from: {os.path.abspath(__file__)}\n")
from tqdm import tqdm
from transformers import AutoConfig, AutoModel, AutoTokenizer, Qwen2TokenizerFast
@@ -205,7 +202,7 @@ class AIIA_VibeVoice_TTS:
"required": {
"vibevoice_model": ("VIBEVOICE_MODEL",),
"text": ("STRING", {"multiline": True, "default": "Hello, this is a test of VibeVoice."}),
- # "reference_audio": ("AUDIO",), <-- Moved to optional
+ "reference_audio": ("AUDIO",),
"cfg_scale": ("FLOAT", {"default": 1.3, "min": 1.0, "max": 10.0, "step": 0.1}),
"ddpm_steps": ("INT", {"default": 20, "min": 10, "max": 100, "step": 1}),
"speed": ("FLOAT", {"default": 1.0, "min": 0.5, "max": 2.0, "step": 0.1}),
@@ -214,9 +211,6 @@ class AIIA_VibeVoice_TTS:
"temperature": ("FLOAT", {"default": 0.8, "min": 0.1, "max": 2.0}),
"top_k": ("INT", {"default": 20, "min": 0, "max": 100}),
"top_p": ("FLOAT", {"default": 0.95, "min": 0.0, "max": 1.0}),
- },
- "optional": {
- "reference_audio": ("AUDIO",),
}
}
@@ -225,95 +219,12 @@ class AIIA_VibeVoice_TTS:
FUNCTION = "generate"
CATEGORY = "AIIA/VibeVoice"
- def _load_fallback_audio(self, target_name="Female_HQ"):
- import torchaudio
- # 定位 assets 目录 (Shared with Podcast nodes)
- nodes_path = os.path.dirname(os.path.abspath(__file__))
- assets_dir = os.path.join(nodes_path, "assets")
-
- # Consistent mapping with Podcast node
- filename_map = {
- "Female_HQ": "seed_female_hq.wav",
- "Male_HQ": "seed_male_hq.wav",
- "Female": "seed_female.wav",
- "Male": "seed_male.wav"
- }
-
- filename = filename_map.get(target_name, "seed_female_hq.wav")
- path = os.path.join(assets_dir, filename)
-
- if not os.path.exists(path):
- print(f"[AIIA Warning] Fallback seed not found at {path}")
- return None
-
- try:
- waveform, sample_rate = torchaudio.load(path)
- if waveform.shape[0] > 1:
- waveform = torch.mean(waveform, dim=0, keepdim=True)
- waveform = waveform * 0.8 # Attenuate
- return {"waveform": waveform, "sample_rate": sample_rate}
- except Exception as e:
- print(f"[AIIA Error] Failed to load fallback audio: {e}")
- return None
-
- def _normalize_roles(self, text):
- """
- Detects custom roles (e.g. 'Host A:', 'User:') and normalizes them to 'Speaker N:'.
- Returns: (normalized_text, role_mapping)
- """
- import re
- lines = text.split('\n')
- # Matches "Role Name:" at start of line.
- # Excludes "Speaker N:" which is already valid.
- # Limit role name to 30 chars to avoid matching long sentences.
- role_pattern = re.compile(r'^([^\n:]{1,30}):\s+')
- speaker_pattern = re.compile(r'^Speaker\s*\d+', re.IGNORECASE)
-
- roles_map = {}
- next_id = 1
- normalized_lines = []
-
- for line in lines:
- stripped = line.strip()
- if not stripped:
- normalized_lines.append(line)
- continue
-
- match = role_pattern.match(stripped)
- if match:
- role_name = match.group(1).strip()
-
- # If already standard format, keep it
- if speaker_pattern.match(role_name):
- normalized_lines.append(line)
- continue
-
- # Map custom role
- if role_name not in roles_map:
- roles_map[role_name] = next_id
- next_id += 1
-
- spk_id = roles_map[role_name]
- # Replace prefix with Speaker N
- # We reconstruct the line to ensure standard formatting
- content = stripped[match.end():]
- normalized_lines.append(f"Speaker {spk_id}: {content}")
- else:
- normalized_lines.append(line)
-
- return "\n".join(normalized_lines), roles_map
-
- def generate(self, vibevoice_model, text, cfg_scale, ddpm_steps, speed, normalize_text,
- do_sample, temperature, top_k, top_p, reference_audio=None):
+ def generate(self, vibevoice_model, text, reference_audio, cfg_scale, ddpm_steps, speed, normalize_text,
+ do_sample, temperature, top_k, top_p):
model = vibevoice_model["model"]
tokenizer = vibevoice_model["tokenizer"]
processor = vibevoice_model.get("processor")
is_streaming = vibevoice_model.get("is_streaming", False)
-
- # AIIA Fix: Ensure model is on GPU (it might have been offloaded to CPU by previous run)
- if torch.cuda.is_available():
- model.to("cuda")
-
device = model.device
if processor is None: raise RuntimeError("Processor is missing.")
@@ -330,39 +241,12 @@ class AIIA_VibeVoice_TTS:
text = re.sub(r'(\d+年)\s*[-—–]\s*(\d+年)', r'\1至\2', text)
text = text.replace('"', '').replace("'", '')
- # [AIIA v1.10.8] Auto-Normalize Roles (e.g. "Host A:" -> "Speaker 1:")
- text, role_map = self._normalize_roles(text)
- num_roles = len(role_map) if role_map else 0
-
- # Determine unique speakers count from text if no role map (e.g. manual Speaker 1, Speaker 2)
- if not role_map:
- # Basic regex count of unique "Speaker N"
- spk_ids = set(re.findall(r'^Speaker\s+(\d+):', text, re.MULTILINE))
- num_roles = len(spk_ids) if spk_ids else 1
+ # Default Speaker Tag
+ if not re.search(r'^Speaker\s+\d+\s*:', text, re.IGNORECASE | re.MULTILINE):
+ lines = text.split('\n')
+ text = "\n".join([f"Speaker 1: {line.strip()}" for line in lines if line.strip()])
# Process Reference Audio
- if reference_audio is None:
- print(f"[AIIA INFO] No reference audio provided. Auto-loading fallbacks for {num_roles} speakers...")
- reference_audio = []
-
- # Simple alternating strategy
- # Speaker 1 (or Host A) -> Female HQ
- # Speaker 2 (or Host B) -> Male HQ
- # Speaker 3 -> Female
- # Speaker 4 -> Male
- patterns = ["Female_HQ", "Male_HQ", "Female", "Male"]
-
- for i in range(max(num_roles, 1)):
- target = patterns[i % len(patterns)]
- fb = self._load_fallback_audio(target)
- if fb: reference_audio.append(fb)
- else:
- # Should not happen if assets exist, but fallback to anything
- if reference_audio: reference_audio.append(reference_audio[0])
-
- if not reference_audio:
- raise ValueError("Could not load any fallback audio!")
-
voice_samples = []
# Determine if input is list or single item
@@ -370,19 +254,6 @@ class AIIA_VibeVoice_TTS:
raw_refs = reference_audio
else:
raw_refs = [reference_audio]
-
- # [AIIA v1.10.9] Smart Pad: If we have more roles than refs, recycle or pad?
- # If user provided 1 ref but text has 2 speakers, previous behavior: Speaker 2 gets nothing?
- # VibeVoice processor slices refs[:num_speakers].
- # If we pad, we can give Speaker 2 the SAME voice, or a fallback?
- # Usually if user gives 1 ref, they might want cloning for Speaker 1, but what for Speaker 2?
- # Safest is to Repeat, or maybe Fallback?
- # Let's Repeat the last ref to avoid errors, assuming 'Cloning' context.
- # But for distinct roles, users SHOULD provide distinct audios.
- if len(raw_refs) < num_roles:
- print(f"[AIIA Warning] Text has {num_roles} roles but only {len(raw_refs)} reference audios. Recycling last audio for remaining speakers.")
- while len(raw_refs) < num_roles:
- raw_refs.append(raw_refs[-1])
for ref_item in raw_refs:
if ref_item is None:
@@ -453,67 +324,9 @@ class AIIA_VibeVoice_TTS:
if audio_out.ndim == 1: audio_out = audio_out.unsqueeze(0)
if audio_out.ndim == 3: audio_out = audio_out.squeeze(0)
- # Ensure float32 (Vital for downstream nodes like Resemble Enhance which fail on fp16)
- audio_out = audio_out.float()
-
# Speed adj
if speed != 1.0:
- original_device = audio_out.device
-
- # Try System 'sox' command for time stretching (pitch preservation)
- # This is more robust than torchaudio.sox_effects which may be missing in some builds
- try:
- # Prepare input
- audio_cpu = audio_out.cpu().numpy()
- # Ensure [C, T]
- if audio_cpu.ndim == 1: audio_cpu = audio_cpu[np.newaxis, :] # [1, T]
- elif audio_cpu.ndim == 3: audio_cpu = audio_cpu.squeeze(0) # [C, T]
-
- # soundfile writes [T, C]
- audio_cpu_t = audio_cpu.T
-
- # Normalize to -1dB (approx 0.9) to prevent clipping during sox processing
- # Sox 'tempo' effect can increase peak amplitude
- max_val = np.abs(audio_cpu_t).max()
- if max_val > 0.9:
- audio_cpu_t = audio_cpu_t * (0.9 / max_val)
-
- with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as in_f, \
- tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as out_f:
- in_path = in_f.name
- out_path = out_f.name
-
- try:
- # Write temp file
- sf.write(in_path, audio_cpu_t, 24000)
-
- # Call sox
- # tempo command: changes speed without pitch
- # -q: quiet
- # -s: use sequence search (better quality for speech)
- cmd = ["sox", "-q", in_path, out_path, "tempo", "-s", str(speed)]
- subprocess.run(cmd, check=True)
- print(f"[AIIA] Applied time stretch: speed={speed}x (pitch preserved)")
-
- # Read back
- out_wav, out_sr = sf.read(out_path)
- # sf reads as [T, C] or [T] if mono
- if out_wav.ndim == 1:
- out_wav = out_wav[np.newaxis, :] # [1, T]
- else:
- out_wav = out_wav.T # [C, T]
-
- audio_out = torch.from_numpy(out_wav).float().to(original_device)
-
- finally:
- if os.path.exists(in_path): os.remove(in_path)
- if os.path.exists(out_path): os.remove(out_path)
-
- except Exception as e:
- # Fallback to Resample (Pitch Shift)
- print(f"[AIIA WARNING] System 'sox' failed ({e}), using Resample (Pitch Shift).")
- resampler = torchaudio.transforms.Resample(orig_freq=int(24000*speed), new_freq=24000).to(original_device)
- audio_out = resampler(audio_out)
+ audio_out = torchaudio.transforms.Resample(orig_freq=int(24000*speed), new_freq=24000)(audio_out)
if audio_out.ndim == 2: audio_out = audio_out.unsqueeze(0)
@@ -553,14 +366,6 @@ class AIIA_VibeVoice_TTS:
except Exception as e:
print(f"[AIIA WARNING] Failed to trim audio: {e}")
- # Cleanup: Move model back to CPU to release VRAM
- try:
- model.to("cpu")
- if hasattr(model, "model") and hasattr(model.model, "to"):
- model.model.to("cpu")
- torch.cuda.empty_cache()
- except: pass
-
return ({"waveform": audio_out.cpu(), "sample_rate": 24000},)
except Exception as e:
@@ -575,7 +380,7 @@ NODE_CLASS_MAPPINGS = {
}
NODE_DISPLAY_NAME_MAPPINGS = {
- "AIIA_VibeVoice_Loader": "VibeVoice Loader",
- "AIIA_VibeVoice_TTS": "VibeVoice TTS (Standard)"
+ "AIIA_VibeVoice_Loader": "🎤 VibeVoice Loader",
+ "AIIA_VibeVoice_TTS": "🗣️ VibeVoice TTS (Standard)"
}
diff --git a/aiia_vibevoice_preset_maker.py b/aiia_vibevoice_preset_maker.py
old mode 100755
new mode 100644
diff --git a/aiia_vibevoice_realtime_tts.py b/aiia_vibevoice_realtime_tts.py
old mode 100755
new mode 100644
diff --git a/aiia_video_nodes.py b/aiia_video_nodes.py
old mode 100755
new mode 100644
index dd617ca..11b8953
--- a/aiia_video_nodes.py
+++ b/aiia_video_nodes.py
@@ -177,7 +177,7 @@ def aiia_apply_video_format_config(format_ui_name: str, user_inputs: dict) -> di
return processed
class AIIA_VideoCombine:
- NODE_NAME = "AIIA Video Combine (Images or Directory)"
+ NODE_NAME = "AIIA 视频合并 (图像或目录)"
CATEGORY = "AIIA/视频"
FUNCTION = "combine_video"
RETURN_TYPES = ("STRING",)
@@ -209,7 +209,7 @@ class AIIA_VideoCombine:
},
"optional": {
"images": ("IMAGE",), "frames_directory": ("STRING", {"default": ""}),
- "filename_pattern": ("STRING", {"default": "frame_%08d.png"}),
+ "filename_pattern": ("STRING", {"default": "frame_%06d.png"}),
"audio_tensor": ("AUDIO",), "audio_file_path": ("STRING", {"default": ""}),
# 【逻辑修改】将 'auto' 设为默认值
"audio_codec": (audio_codec_options, {"default": "auto"}),
@@ -233,88 +233,17 @@ class AIIA_VideoCombine:
try:
effective_frames_dir, effective_filename_pattern = None, filename_pattern
if images is not None:
- num_frames = images.shape[0]
- logger.info(f"检测到 {num_frames} 帧的图像张量输入...")
- from tqdm import tqdm
- import gc
+ logger.info(f"检测到 {images.shape[0]} 帧的图像张量输入...")
temp_image_dir_to_delete = tempfile.mkdtemp(prefix="aiia_frames_")
effective_frames_dir, effective_filename_pattern = temp_image_dir_to_delete, "frame_%08d.png"
- pbar = ProgressBar(num_frames)
- # Direct tqdm to stdout for console logs
- console_pbar = tqdm(total=num_frames, desc="[AIIA Video] Saving Frames", unit="frame", file=sys.stdout)
-
- # Move to CPU if on GPU to free GPU memory
- if images.device.type == 'cuda':
- images = images.cpu()
- torch.cuda.empty_cache()
-
- # Convert to list so we can delete individual frames to free memory
- # This is crucial for large frame counts
- frame_list = [images[i] for i in range(num_frames)]
- del images # Release the original tensor immediately
- gc.collect()
-
- for i in range(num_frames):
- # Get frame, convert, save, and release immediately
- frame_np = (frame_list[i].numpy() * 255).astype(np.uint8)
- frame_list[i] = None # Release this frame from the list
- img = Image.fromarray(frame_np)
- img.save(os.path.join(effective_frames_dir, effective_filename_pattern % (i + 1)))
- del frame_np, img
+ pbar = ProgressBar(images.shape[0])
+ for i, frame_tensor in enumerate(images):
+ Image.fromarray((frame_tensor.cpu().numpy() * 255).astype(np.uint8)).save(os.path.join(effective_frames_dir, effective_filename_pattern % (i + 1)))
pbar.update(1)
- console_pbar.update(1)
-
- # Periodic garbage collection every 50 frames
- if i > 0 and i % 50 == 0:
- gc.collect()
- if torch.cuda.is_available():
- torch.cuda.empty_cache()
-
- console_pbar.close()
-
- # Final cleanup
- del frame_list
- gc.collect()
- if torch.cuda.is_available():
- torch.cuda.empty_cache()
elif frames_directory:
effective_frames_dir = strip_path_aiia(frames_directory)
if not validate_path_aiia(effective_frames_dir, check_is_dir=True):
raise ValueError(f"帧目录验证失败: {effective_frames_dir}")
-
- # Intelligent Pattern Detection (Fix for v1.9.27 compatibility)
- # If user workflow has old default (%06d) but files are new (%08d), auto-correct it.
- try:
- # Look for frame_*.png files
- search_glob = os.path.join(effective_frames_dir, "frame_*.png")
- sample_files = sorted(glob.glob(search_glob))
-
- # Log detection status
- if not sample_files:
- logger.warning(f"[AIIA_VideoCombine] No 'frame_*.png' files found in {effective_frames_dir}. Detection skipped.")
- else:
- first_file = os.path.basename(sample_files[0])
- logger.info(f"[AIIA_VideoCombine] Auto-detecting pattern from first file: {first_file}")
-
- # Extract the numeric part: frame_00000123.png -> 00000123
- match = re.search(r"frame_(\d+)\.png", first_file)
- if match:
- digit_count = len(match.group(1))
- auto_pattern = f"frame_%0{digit_count}d.png"
-
- if auto_pattern != effective_filename_pattern:
- logger.warning(f"[AIIA_VideoCombine] Pattern Mismatch! User: '{effective_filename_pattern}', Actual: '{auto_pattern}'. Auto-correcting.")
- effective_filename_pattern = auto_pattern
-
- # Verification: Check if ffmpeg will find it
- test_glob = os.path.join(effective_frames_dir, f"frame_{'0'*digit_count}.png")
- if not os.path.exists(test_glob) and len(sample_files) > 0:
- # Sometimes start index is not 0?
- logger.warning(f"[AIIA_VideoCombine] Verify: frame_{'0'*digit_count}.png not found, but we found {len(sample_files)} files.")
- else:
- logger.warning(f"[AIIA_VideoCombine] Regex failed on {first_file}")
- except Exception as e:
- logger.warning(f"[AIIA_VideoCombine] Auto-pattern detection failed: {e}")
else: raise ValueError("错误: 必须提供 'images' 或 'frames_directory' 输入。")
output_dir = folder_paths.get_output_directory() if save_output else folder_paths.get_temp_directory()
@@ -394,312 +323,5 @@ class AIIA_VideoCombine:
try: shutil.rmtree(temp_image_dir_to_delete); logger.info(f"已清理临时图像帧目录: {temp_image_dir_to_delete}")
except Exception as e_del: logger.error(f"清理临时图像帧目录失败: {e_del}")
-
-class AIIA_BodySway:
- """Simulate subtle body movement through crop-based pan and rotation."""
-
- NODE_NAME = "AIIA Body Sway"
- CATEGORY = "AIIA/视频"
- FUNCTION = "apply_sway"
- RETURN_TYPES = ("IMAGE", "STRING")
- RETURN_NAMES = ("images", "output_frames_dir")
-
- @classmethod
- def INPUT_TYPES(cls):
- return {
- "required": {
- "crop_ratio": ("FLOAT", {"default": 0.99, "min": 0.90, "max": 1.0, "step": 0.001,
- "tooltip": "Output size as ratio of input (0.99 = keep 99%, crop 1%)"}),
- "rotation_amplitude": ("FLOAT", {"default": 0.1, "min": 0.0, "max": 2.0, "step": 0.1,
- "tooltip": "Max rotation in degrees"}),
- "smoothness": ("FLOAT", {"default": 0.02, "min": 0.005, "max": 0.1, "step": 0.005,
- "tooltip": "Perlin noise smoothness (smaller = slower drift)"}),
- "seed": ("INT", {"default": 0, "min": 0, "max": 0xffffffffffffffff}),
- },
- "optional": {
- "images": ("IMAGE",), # [B, H, W, C] tensor
- "frames_directory": ("STRING", {"default": "", "tooltip": "Path to directory containing frames (from Ditto disk mode)"}),
- }
- }
-
- def apply_sway(self, crop_ratio: float, rotation_amplitude: float,
- smoothness: float, seed: int, images: torch.Tensor = None, frames_directory: str = ""):
- """
- Apply body sway effect using crop + rotation with Perlin noise.
-
- images: [B, H, W, C] tensor (float32, 0-1 range) - optional
- frames_directory: Path to directory containing frames - optional (for OOM-safe mode)
- crop_ratio: Output size as ratio of input (e.g., 0.99)
- smoothness: Perlin noise smoothness (smaller = slower drift)
- """
- import math
- from PIL import Image
- import random
- import os
- import glob
- import tempfile
- import gc
-
- # Validate inputs
- frames_directory = frames_directory.strip() if frames_directory else ""
- use_disk_mode = bool(frames_directory and os.path.isdir(frames_directory))
-
- if images is None and not use_disk_mode:
- raise ValueError("必须提供 images 或 frames_directory 输入")
-
- # Limit seed to 32-bit for numpy compatibility
- seed_32 = seed % (2**32)
- random.seed(seed_32)
- np.random.seed(seed_32)
-
- # Determine input dimensions and frame count
- if use_disk_mode:
- # Get frame list from directory
- frame_files = sorted(glob.glob(os.path.join(frames_directory, "*.png")))
- if not frame_files:
- frame_files = sorted(glob.glob(os.path.join(frames_directory, "*.jpg")))
- if not frame_files:
- raise ValueError(f"目录中未找到帧文件: {frames_directory}")
-
- batch_size = len(frame_files)
- first_frame = Image.open(frame_files[0])
- in_w, in_h = first_frame.size
- channels = 3
- first_frame.close()
- logger.info(f"[BodySway] Disk mode: {batch_size} frames from {frames_directory}")
- else:
- batch_size, in_h, in_w, channels = images.shape
-
- # Auto-calculate target size and sway amplitude from crop_ratio
- target_width = int(in_w * crop_ratio)
- target_height = int(in_h * crop_ratio)
-
- # Ensure target is even (for compatibility with video encoders)
- target_width = target_width - (target_width % 2)
- target_height = target_height - (target_height % 2)
-
- # Calculate total available margin
- margin_x = (in_w - target_width) / 2
- margin_y = (in_h - target_height) / 2
- total_margin = min(margin_x, margin_y)
-
- # Reserve margin for rotation (rotation introduces black corners)
- diagonal = math.sqrt(in_w**2 + in_h**2)
- rotation_margin = (diagonal / 2) * math.sin(math.radians(rotation_amplitude)) if rotation_amplitude > 0 else 0
-
- # Safe margin for sway = total margin - rotation margin - 20% safety buffer
- safe_margin = max(0, total_margin - rotation_margin)
- sway_amplitude = safe_margin * 0.8 # Use 80% of remaining safe margin
-
- # Clamp rotation to prevent exceeding available margin
- max_safe_rotation = math.degrees(math.asin(min(1.0, total_margin * 0.5 / (diagonal / 2 + 0.001))))
- actual_rotation = min(rotation_amplitude, max_safe_rotation)
- if actual_rotation < rotation_amplitude:
- logger.warning(f"[BodySway] Rotation clamped from {rotation_amplitude}° to {actual_rotation:.2f}° to prevent black corners")
-
- logger.info(f"[BodySway] Input: {in_w}x{in_h}, Output: {target_width}x{target_height}, "
- f"Margin: {total_margin:.1f}px, RotMargin: {rotation_margin:.1f}px, Sway: {sway_amplitude:.1f}px")
-
- # Generate Perlin-like noise trajectory (1D)
- def generate_perlin_trajectory(n_frames: int, amplitude: float, scale: float):
- """Generate smooth organic trajectory using 1D Perlin-like noise."""
- # Use cumulative random walk with smoothing
- trajectory = np.zeros(n_frames)
-
- # Generate multi-octave noise
- octaves = 4
- persistence = 0.5
-
- for octave in range(octaves):
- freq = scale * (2 ** octave)
- amp = amplitude * (persistence ** octave)
- phase = random.random() * 1000
-
- for t in range(n_frames):
- # Smooth interpolated noise using sine-based sampling
- x = t * freq + phase
- # Interpolate between random values
- i = int(x)
- f = x - i
- # Smooth interpolation (cosine)
- f = (1 - math.cos(f * math.pi)) / 2
-
- # Get random values seeded by position
- random.seed(seed_32 + i + octave * 10000)
- v0 = random.random() * 2 - 1
- random.seed(seed_32 + i + 1 + octave * 10000)
- v1 = random.random() * 2 - 1
-
- trajectory[t] += (v0 * (1 - f) + v1 * f) * amp
-
- # Normalize to amplitude range
- if np.max(np.abs(trajectory)) > 0:
- trajectory = trajectory / np.max(np.abs(trajectory)) * amplitude
-
- return trajectory
-
- # Generate X and Rotation trajectories (NO Y translation to reduce dizziness)
- traj_x = generate_perlin_trajectory(batch_size, sway_amplitude, smoothness)
- traj_y = np.zeros(batch_size) # No vertical movement
- traj_rot = generate_perlin_trajectory(batch_size, actual_rotation, smoothness * 0.7)
-
- # Calculate GLOBAL valid region based on MAX rotation angle
- if actual_rotation > 0:
- max_rot_rad = math.radians(actual_rotation)
- global_rot_margin_x = int(math.ceil((in_h * math.sin(max_rot_rad) + in_w * (1 - math.cos(max_rot_rad))) / 2))
- global_rot_margin_y = int(math.ceil((in_w * math.sin(max_rot_rad) + in_h * (1 - math.cos(max_rot_rad))) / 2))
- valid_left = global_rot_margin_x
- valid_top = global_rot_margin_y
- valid_right = in_w - global_rot_margin_x
- valid_bottom = in_h - global_rot_margin_y
- logger.info(f"[BodySway] Global valid region: ({valid_left}, {valid_top}) to ({valid_right}, {valid_bottom})")
- else:
- valid_left, valid_top = 0, 0
- valid_right, valid_bottom = in_w, in_h
-
- # Check if target fits within valid region
- if target_width > (valid_right - valid_left) or target_height > (valid_bottom - valid_top):
- logger.warning(f"[BodySway] Target size ({target_width}x{target_height}) exceeds valid region. Consider reducing rotation_amplitude or crop_ratio.")
-
- # Create output directory for disk mode
- output_frames_dir = ""
- if use_disk_mode:
- import folder_paths
- output_frames_dir = tempfile.mkdtemp(prefix="bodysway_frames_", dir=folder_paths.get_temp_directory())
- logger.info(f"[BodySway] Output frames will be saved to: {output_frames_dir}")
-
- # Process frames: disk mode or memory mode
- if use_disk_mode:
- from concurrent.futures import ThreadPoolExecutor
- import threading
-
- # Disk mode: parallel processing and saving
- # Use 4 workers to balance CPU usage (cropping/rotating is fast, saving is slow)
- process_executor = ThreadPoolExecutor(max_workers=4)
- write_sem = threading.Semaphore(50) # Limit pending writes
-
- def process_and_save_task(idx, in_path, out_path, sem):
- try:
- img = Image.open(in_path)
-
- # Apply rotation
- if actual_rotation > 0:
- rotated = img.rotate(traj_rot[idx], resample=Image.BILINEAR, expand=False)
- img.close()
- img = rotated
-
- # Calculate crop box
- center_x = in_w / 2 + traj_x[idx]
- center_y = in_h / 2 + traj_y[idx] # traj_y is actually 0 but for consistency
-
- left = int(center_x - target_width / 2)
- top = int(center_y - target_height / 2)
- right = left + target_width
- bottom = top + target_height
-
- # Clamp
- left = max(valid_left, min(left, valid_right - target_width))
- top = max(valid_top, min(top, valid_bottom - target_height))
- right = left + target_width
- bottom = top + target_height
-
- # Crop and save
- cropped = img.crop((left, top, right, bottom))
- img.close()
-
- # Fast save
- cropped.save(out_path, format="PNG", compress_level=0)
- cropped.close()
- except Exception as e:
- logger.error(f"[BodySway] Error processing frame {idx}: {e}")
- finally:
- sem.release()
-
- logger.info("[BodySway] Starting parallel disk processing...")
-
- for i, frame_path in enumerate(frame_files):
- output_path = os.path.join(output_frames_dir, f"frame_{i:08d}.png")
-
- # Wait for slot
- write_sem.acquire()
-
- # Submit task
- process_executor.submit(process_and_save_task, i, frame_path, output_path, write_sem)
-
- # Periodic logging
- if i > 0 and i % 100 == 0:
- if batch_size > 200:
- logger.info(f"[BodySway] Submitted: {i}/{batch_size} frames")
-
- # Wait for all tasks to complete
- process_executor.shutdown(wait=True)
-
- # Return placeholder tensor for disk mode
- output_tensor = torch.zeros((1, 1, 1, 3), dtype=torch.float32)
- logger.info(f"[BodySway] Disk mode: Processed {batch_size} frames, saved to {output_frames_dir}")
-
- else:
- # Memory mode: original logic with tensor input/output
- output_tensor = torch.zeros((batch_size, target_height, target_width, channels),
- dtype=torch.float32, device='cpu')
-
- if images.device.type == 'cuda':
- images = images.cpu()
- torch.cuda.empty_cache()
-
- batch_chunk_size = 50
-
- for batch_start in range(0, batch_size, batch_chunk_size):
- batch_end = min(batch_start + batch_chunk_size, batch_size)
-
- for i in range(batch_start, batch_end):
- frame_np = (images[i].numpy() * 255).astype(np.uint8)
- img = Image.fromarray(frame_np)
- del frame_np
-
- if actual_rotation > 0:
- rotated = img.rotate(traj_rot[i], resample=Image.BILINEAR, expand=False)
- del img
- img = rotated
-
- center_x = in_w / 2 + traj_x[i]
- center_y = in_h / 2 + traj_y[i]
-
- left = int(center_x - target_width / 2)
- top = int(center_y - target_height / 2)
- right = left + target_width
- bottom = top + target_height
-
- left = max(valid_left, min(left, valid_right - target_width))
- top = max(valid_top, min(top, valid_bottom - target_height))
- right = left + target_width
- bottom = top + target_height
-
- cropped = img.crop((left, top, right, bottom))
- del img
-
- cropped_arr = np.array(cropped, dtype=np.float32)
- del cropped
- output_tensor[i] = torch.from_numpy(cropped_arr / 255.0)
- del cropped_arr
-
- gc.collect()
- if torch.cuda.is_available():
- torch.cuda.empty_cache()
-
- if batch_size > 200:
- logger.info(f"[BodySway] Progress: {batch_end}/{batch_size} frames processed")
-
- logger.info(f"[BodySway] Processed {batch_size} frames with Perlin noise (smoothness={smoothness})")
-
- return (output_tensor, output_frames_dir)
-
-
-NODE_CLASS_MAPPINGS = {
- "AIIA_VideoCombine": AIIA_VideoCombine,
- "AIIA_BodySway": AIIA_BodySway,
-}
-NODE_DISPLAY_NAME_MAPPINGS = {
- "AIIA_VideoCombine": "AIIA Video Combine (Images or Dir)",
- "AIIA_BodySway": "AIIA Body Sway",
-}
\ No newline at end of file
+NODE_CLASS_MAPPINGS = { "AIIA_VideoCombine": AIIA_VideoCombine }
+NODE_DISPLAY_NAME_MAPPINGS = { "AIIA_VideoCombine": "视频合并 (AIIA, 图像或目录)" }
\ No newline at end of file
diff --git a/aiia_voxcpm_nodes.py b/aiia_voxcpm_nodes.py
old mode 100755
new mode 100644
index e0314ab..643432a
--- a/aiia_voxcpm_nodes.py
+++ b/aiia_voxcpm_nodes.py
@@ -254,6 +254,6 @@ NODE_CLASS_MAPPINGS = {
}
NODE_DISPLAY_NAME_MAPPINGS = {
- "AIIA_VoxCPM_Loader": "VoxCPM Loader",
- "AIIA_VoxCPM_TTS": "VoxCPM 1.5 TTS"
+ "AIIA_VoxCPM_Loader": "🎤 VoxCPM Loader",
+ "AIIA_VoxCPM_TTS": "🗣️ VoxCPM 1.5 TTS"
}
diff --git a/aiia_web_export_nodes.py b/aiia_web_export_nodes.py
old mode 100755
new mode 100644
diff --git a/assets/seed_female.wav b/assets/seed_female.wav
old mode 100755
new mode 100644
diff --git a/assets/seed_female_hq.wav b/assets/seed_female_hq.wav
old mode 100755
new mode 100644
diff --git a/assets/seed_male.wav b/assets/seed_male.wav
old mode 100755
new mode 100644
diff --git a/assets/seed_male_hq.wav b/assets/seed_male_hq.wav
old mode 100755
new mode 100644
diff --git a/convert_to_bf16.py b/convert_to_bf16.py
old mode 100755
new mode 100644
diff --git a/debug_config.py b/debug_config.py
old mode 100755
new mode 100644
diff --git a/debug_preset.py b/debug_preset.py
old mode 100755
new mode 100644
diff --git a/debug_sr.py b/debug_sr.py
old mode 100755
new mode 100644
diff --git a/experiment_instruct.py b/experiment_instruct.py
old mode 100755
new mode 100644
diff --git a/identity_grid.py b/identity_grid.py
old mode 100755
new mode 100644
diff --git a/inspect_resemble.py b/inspect_resemble.py
old mode 100755
new mode 100644
diff --git a/js/aiia_browser.js b/js/aiia_browser.js
old mode 100755
new mode 100644
index 1285750..e54f71e
--- a/js/aiia_browser.js
+++ b/js/aiia_browser.js
@@ -77,8 +77,6 @@ class AIIABrowserDialog extends ComfyDialog {
this.applyFocusRaf = null;
this.iconLoadTimers = new Map();
- this.outsideClickListener = null;
-
this.tooltipImage = $el("img.aiia-tooltip-image");
this.tooltipVideo = $el("video.aiia-tooltip-video", { autoplay: true, muted: true, loop: true, controls: false, volume: 0.8 });
this.tooltipAudio = $el("audio", { autoplay: true });
@@ -93,10 +91,7 @@ class AIIABrowserDialog extends ComfyDialog {
const splitViewContainer = $el("div.aiia-split-view-container", [this.directoryPanel, this.contentPanel, this.tooltipElement]);
- this.iconViewObserver = new IntersectionObserver(this.handleIconIntersection.bind(this), {
- root: null, // Use root: null to observe relative to browser viewport
- rootMargin: '500px 0px 500px 0px'
- });
+ this.iconViewObserver = new IntersectionObserver(this.handleIconIntersection.bind(this), { root: this.contentArea, rootMargin: '300px 0px 300px 0px' });
this.titleElement = $el("span", { textContent: "AIIA Media Browser" });
this.closeButton = $el("button.close", { textContent: "✖", title: "Close" });
@@ -303,25 +298,6 @@ class AIIABrowserDialog extends ComfyDialog {
}
}
}
- async deleteItem(itemName) {
- if (!itemName) return;
- if (!confirm(`Are you sure you want to delete "${itemName}"? This action cannot be undone.`)) return;
- try {
- const res = await api.fetchApi('/aiia/v1/browser/delete_item', {
- method: 'POST',
- body: JSON.stringify({ path: this.currentPath, filename: itemName })
- });
- if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
- const result = await res.json();
- if (result.status === 'success') {
- this.setFocus(null, null);
- this.refreshCurrentDirectory();
- }
- } catch (e) {
- alert(`Error deleting item: ${e.message}`);
- console.error(e);
- }
- }
updateTooltipSizeLimits() { const rect = this.contentPanel.getBoundingClientRect(); const maxSize = Math.min(rect.width, rect.height) * 0.30; this.tooltipElement.style.setProperty('--aiia-tooltip-media-max-size', `${maxSize}px`); }
updateSortControls() { this.sortKeySelect.value = this.sortKey; this.sortDirButton.textContent = this.sortDir === 'asc' ? '▲' : '▼'; this.sortDirButton.dataset.tooltipText = `Sort ${this.sortDir === 'asc' ? 'Descending' : 'Ascending'}`; }
showPathInput() { this.breadcrumbs.style.display = 'none'; this.pathInputContainer.style.display = 'flex'; this.pathInput.value = this.currentPath; this.pathInput.focus(); this.pathInput.select(); }
@@ -1335,15 +1311,11 @@ class AIIABrowserDialog extends ComfyDialog {
if ((target.tagName === 'INPUT' && target !== this.pathInput) || target.tagName === 'SELECT') return;
- const keyMap = { 'ArrowUp': 'moveFocus', 'ArrowDown': 'moveFocus', 'ArrowLeft': 'moveFocus', 'ArrowRight': 'moveFocus', 'Enter': 'activateFocusedItem', ' ': 'toggleTooltipForFocusedItem', 'Escape': 'close', 'Delete': 'deleteFocusedItem' };
+ const keyMap = { 'ArrowUp': 'moveFocus', 'ArrowDown': 'moveFocus', 'ArrowLeft': 'moveFocus', 'ArrowRight': 'moveFocus', 'Enter': 'activateFocusedItem', ' ': 'toggleTooltipForFocusedItem' };
if (keyMap[e.key]) {
e.preventDefault(); e.stopPropagation();
this.isKeyboardNavigating = true;
if (keyMap[e.key] === 'moveFocus') this.moveFocus(e.key);
- else if (keyMap[e.key] === 'deleteFocusedItem') {
- const item = this.getAllNavigableItems()[this.currentFocusIndex];
- if (item) this.deleteItem(item.name);
- }
else this[keyMap[e.key]]();
}
}
@@ -1355,35 +1327,6 @@ class AIIABrowserDialog extends ComfyDialog {
super.show();
this.updateTooltipSizeLimits();
this.element.focus({ preventScroll: true });
-
- // [v1.9.184] Click-outside to close - targeting the native ComfyUI modal backdrop
- if (!this.outsideClickListener) {
- this.outsideClickListener = (e) => {
- if (this.element.style.display !== "none") {
- // If the click is on the parent container (the backdrop) and NOT on the element itself
- if (e.target === this.element.parentElement || e.target.classList.contains('comfy-modal')) {
- // Safety check: Don't close if we are interacting with another aiia modal
- if (this.fullscreenViewer && this.fullscreenViewer.element.style.display !== 'none') return;
- this.close();
- }
- }
- };
- setTimeout(() => document.addEventListener("mousedown", this.outsideClickListener), 10);
- }
-
- // Force a resize/scroll event to trigger initial icon rendering
- setTimeout(() => { this.render(); }, 100);
- }
-
- close() {
- if (this.outsideClickListener) {
- document.removeEventListener("mousedown", this.outsideClickListener);
- this.outsideClickListener = null;
- }
- this.element.style.display = "none";
- this.hideTooltip();
- if (typeof super.close === "function") super.close();
- else if (typeof super.hide === "function") super.hide();
}
}
@@ -1418,16 +1361,8 @@ app.registerExtension({
let browserDialog = null;
const createDialog = () => { if (!browserDialog) browserDialog = new AIIABrowserDialog(); browserDialog.show(); };
document.addEventListener('keydown', (e) => {
- if (!browserDialog || (browserDialog.element && browserDialog.element.style.display === 'none')) return;
- if (browserDialog.fullscreenViewer && browserDialog.fullscreenViewer.element && browserDialog.fullscreenViewer.element.style.display !== 'none') return;
-
- if (e.key === 'Escape') {
- e.preventDefault();
- e.stopPropagation();
- browserDialog.close();
- return;
- }
-
+ if (!browserDialog || browserDialog.element.style.display === 'none') return;
+ if (browserDialog.fullscreenViewer && browserDialog.fullscreenViewer.element.style.display !== 'none') return;
if (!browserDialog.element.contains(document.activeElement)) {
const isNavKey = ['ArrowUp', 'ArrowDown', 'ArrowLeft', 'ArrowRight', 'Enter', ' '].includes(e.key);
if (isNavKey) {
diff --git a/js/aiia_browser_styles.js b/js/aiia_browser_styles.js
old mode 100755
new mode 100644
index 43f0e29..6ec0a35
--- a/js/aiia_browser_styles.js
+++ b/js/aiia_browser_styles.js
@@ -2,7 +2,7 @@
export const browserStyles = `
/* Main Browser Styles */
#aiia-browser-menu-button { margin-left: 10px; }
- .comfy-modal.aiia-browser-dialog-root { top: calc(50% + 40px); left: 50%; transform: translate(-50%, -50%); width: 80vw; height: 85vh; max-width: 1400px; max-height: 1000px; min-width: 800px; min-height: 500px; display: flex; flex-direction: column; padding: 0; border-radius: 8px; box-shadow: 0 10px 30px rgba(0,0,0,0.2); resize: both; overflow: hidden; }
+ .comfy-modal.aiia-browser-dialog-root { top: 50%; left: 50%; transform: translate(-50%, -50%); width: 80vw; height: 85vh; max-width: 1400px; max-height: 1000px; min-width: 800px; min-height: 500px; display: flex; flex-direction: column; padding: 0; border-radius: 8px; box-shadow: 0 10px 30px rgba(0,0,0,0.2); resize: both; overflow: hidden; }
.aiia-browser-titlebar { background: var(--comfy-box-bg); padding: 4px 8px; font-weight: bold; display: flex; justify-content: space-between; align-items: center; border-bottom: 1px solid var(--border-color); flex-shrink: 0; color: #F9FAFB; cursor: default; }
.aiia-browser-main-container { display: flex; flex-direction: column; flex-grow: 1; padding: 8px; overflow: hidden; }
.aiia-browser-header-controls { display: flex; justify-content: space-between; align-items: center; margin-bottom: 8px; flex-shrink: 0; gap: 8px; flex-wrap: wrap; }
diff --git a/js/aiia_fullscreen_viewer.js b/js/aiia_fullscreen_viewer.js
old mode 100755
new mode 100644
diff --git a/js/aiia_labels.js b/js/aiia_labels.js
new file mode 100644
index 0000000..6d403b5
--- /dev/null
+++ b/js/aiia_labels.js
@@ -0,0 +1,55 @@
+
+import { app } from "../../../scripts/app.js";
+
+// Extension to handle static labels (identifying by 'is_label' flag in metadata)
+app.registerExtension({
+ name: "AIIA.Labels",
+ async beforeRegisterNodeDef(nodeType, nodeData, app) {
+ // Iterate through required inputs to find ones marked as is_label
+ const requiredInputs = nodeData.input?.required || {};
+
+ let labelWidgetNames = [];
+ for (const [name, inputDef] of Object.entries(requiredInputs)) {
+ // inputDef is [type, metadata_dict]
+ if (inputDef[1] && inputDef[1].is_label === true) {
+ labelWidgetNames.push(name);
+ }
+ }
+
+ if (labelWidgetNames.length > 0) {
+ const onNodeCreated = nodeType.prototype.onNodeCreated;
+ nodeType.prototype.onNodeCreated = function () {
+ const r = onNodeCreated ? onNodeCreated.apply(this, arguments) : undefined;
+
+ for (const w of this.widgets) {
+ if (labelWidgetNames.includes(w.name)) {
+ // Change type to avoid standard text box rendering
+ w.type = "AIIA_STATIC_TEXT";
+
+ w.draw = function (ctx, node, widget_width, y, widget_height) {
+ ctx.save();
+ // Background or subtle underline if needed
+ // ctx.fillStyle = "#222222";
+ // ctx.fillRect(0, y, widget_width, widget_height);
+
+ ctx.fillStyle = "#AAAAAA"; // Label color
+ ctx.font = "italic 12px Arial";
+ // Draw the text
+ ctx.fillText(this.value, 15, y + widget_height * 0.7);
+ ctx.restore();
+ };
+
+ // Disable interaction
+ w.mouse = () => { };
+ w.computeSize = () => [200, 20];
+
+ // Prevent this widget from being converted to a socket (though it's already a Primitive STRING)
+ w.inputKey = null;
+ w.serializeValue = async () => ""; // Don't send label text back to Python
+ }
+ }
+ return r;
+ };
+ }
+ }
+});
diff --git a/js/aiia_subtitle_preview.js b/js/aiia_subtitle_preview.js
old mode 100755
new mode 100644
diff --git a/js/aiia_video_nodes.js b/js/aiia_video_nodes.js
old mode 100755
new mode 100644
index 9fee6f8..1c86639
--- a/js/aiia_video_nodes.js
+++ b/js/aiia_video_nodes.js
@@ -12,15 +12,15 @@ function toggleWidget(node, widget, show = false) {
// console.log(`AIIA Debug (toggleWidget): Toggling '${widget.name}'. Should show: ${show}`);
// --- Debug End ---
- if (!widget) return;
- if (!origProps[widget.name]) {
- origProps[widget.name] = {
- origType: widget.type,
- origComputeSize: widget.computeSize
- };
- }
- widget.type = show ? origProps[widget.name].origType : "AIIA_HIDDEN";
- widget.computeSize = show ? origProps[widget.name].origComputeSize : () => [0, -4];
+ if (!widget) return;
+ if (!origProps[widget.name]) {
+ origProps[widget.name] = {
+ origType: widget.type,
+ origComputeSize: widget.computeSize
+ };
+ }
+ widget.type = show ? origProps[widget.name].origType : "AIIA_HIDDEN";
+ widget.computeSize = show ? origProps[widget.name].origComputeSize : () => [0, -4];
node.setDirtyCanvas(true);
}
@@ -28,7 +28,7 @@ function toggleWidget(node, widget, show = false) {
function chainCallback(object, property, callback) {
if (object[property]) {
const original = object[property];
- object[property] = function () {
+ object[property] = function() {
original.apply(this, arguments);
callback.apply(this, arguments);
};
@@ -41,15 +41,15 @@ app.registerExtension({
name: "AIIA.VideoNodes.DynamicWidgets.Final",
async beforeRegisterNodeDef(nodeType, nodeData, app) {
if (nodeData.name === "AIIA_VideoCombine") {
-
+
const widgetsByFormat = nodeData.input.required.format[1].formats;
if (!widgetsByFormat) return;
- chainCallback(nodeType.prototype, "onNodeCreated", function () {
+ chainCallback(nodeType.prototype, "onNodeCreated", function() {
const node = this;
const formatWidget = findWidgetByName(node, "format");
if (!formatWidget) return;
-
+
const allDynamicWidgetNames = new Set(Object.values(widgetsByFormat).flat().map(p => p[0]));
const updateWidgetsVisibility = (formatValue) => {
@@ -57,37 +57,23 @@ app.registerExtension({
const visibleWidgetNames = new Set(
(widgetsByFormat[formatValue] || []).map(p => p[0])
);
-
+
for (const widgetName of allDynamicWidgetNames) {
const widget = findWidgetByName(node, widgetName);
if (widget) {
toggleWidget(node, widget, visibleWidgetNames.has(widgetName));
}
}
-
+
node.setSize([node.size[0], node.computeSize()[1]]);
};
// 为format widget的callback链接上更新函数
chainCallback(formatWidget, "callback", updateWidgetsVisibility);
-
- // Expose function for onConfigure
- node.aiiaUpdateVideoWidgets = updateWidgetsVisibility;
-
+
// 初始加载时触发
updateWidgetsVisibility(formatWidget.value);
});
-
- chainCallback(nodeType.prototype, "onConfigure", function () {
- const node = this;
- // 使用 requestAnimationFrame 确保在所有widget值被写入后再执行更新
- requestAnimationFrame(() => {
- const formatWidget = findWidgetByName(node, "format");
- if (formatWidget && node.aiiaUpdateVideoWidgets) {
- node.aiiaUpdateVideoWidgets(formatWidget.value);
- }
- });
- });
}
}
});
\ No newline at end of file
diff --git a/libs/CosyVoice/.github/ISSUE_TEMPLATE/bug_report.md b/libs/CosyVoice/.github/ISSUE_TEMPLATE/bug_report.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/.github/ISSUE_TEMPLATE/feature_request.md b/libs/CosyVoice/.github/ISSUE_TEMPLATE/feature_request.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/.github/workflows/lint.yml b/libs/CosyVoice/.github/workflows/lint.yml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/.github/workflows/stale-issues.yml b/libs/CosyVoice/.github/workflows/stale-issues.yml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/.gitignore b/libs/CosyVoice/.gitignore
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/CODE_OF_CONDUCT.md b/libs/CosyVoice/CODE_OF_CONDUCT.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/FAQ.md b/libs/CosyVoice/FAQ.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/LICENSE b/libs/CosyVoice/LICENSE
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/README.md b/libs/CosyVoice/README.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/asset/dingding.png b/libs/CosyVoice/asset/dingding.png
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/__init__.py b/libs/CosyVoice/cosyvoice/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/bin/average_model.py b/libs/CosyVoice/cosyvoice/bin/average_model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/bin/export_jit.py b/libs/CosyVoice/cosyvoice/bin/export_jit.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/bin/export_onnx.py b/libs/CosyVoice/cosyvoice/bin/export_onnx.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/bin/train.py b/libs/CosyVoice/cosyvoice/bin/train.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/cli/__init__.py b/libs/CosyVoice/cosyvoice/cli/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/cli/cosyvoice.py b/libs/CosyVoice/cosyvoice/cli/cosyvoice.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/cli/frontend.py b/libs/CosyVoice/cosyvoice/cli/frontend.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/cli/model.py b/libs/CosyVoice/cosyvoice/cli/model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/dataset/__init__.py b/libs/CosyVoice/cosyvoice/dataset/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/dataset/dataset.py b/libs/CosyVoice/cosyvoice/dataset/dataset.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/dataset/processor.py b/libs/CosyVoice/cosyvoice/dataset/processor.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/flow/DiT/dit.py b/libs/CosyVoice/cosyvoice/flow/DiT/dit.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/flow/DiT/modules.py b/libs/CosyVoice/cosyvoice/flow/DiT/modules.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/flow/decoder.py b/libs/CosyVoice/cosyvoice/flow/decoder.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/flow/flow.py b/libs/CosyVoice/cosyvoice/flow/flow.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/flow/flow_matching.py b/libs/CosyVoice/cosyvoice/flow/flow_matching.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/flow/length_regulator.py b/libs/CosyVoice/cosyvoice/flow/length_regulator.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/hifigan/discriminator.py b/libs/CosyVoice/cosyvoice/hifigan/discriminator.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/hifigan/f0_predictor.py b/libs/CosyVoice/cosyvoice/hifigan/f0_predictor.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/hifigan/generator.py b/libs/CosyVoice/cosyvoice/hifigan/generator.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/hifigan/hifigan.py b/libs/CosyVoice/cosyvoice/hifigan/hifigan.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/llm/llm.py b/libs/CosyVoice/cosyvoice/llm/llm.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/tokenizer/assets/multilingual_zh_ja_yue_char_del.tiktoken b/libs/CosyVoice/cosyvoice/tokenizer/assets/multilingual_zh_ja_yue_char_del.tiktoken
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/tokenizer/tokenizer.py b/libs/CosyVoice/cosyvoice/tokenizer/tokenizer.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/__init__.py b/libs/CosyVoice/cosyvoice/transformer/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/activation.py b/libs/CosyVoice/cosyvoice/transformer/activation.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/attention.py b/libs/CosyVoice/cosyvoice/transformer/attention.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/convolution.py b/libs/CosyVoice/cosyvoice/transformer/convolution.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/decoder.py b/libs/CosyVoice/cosyvoice/transformer/decoder.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/decoder_layer.py b/libs/CosyVoice/cosyvoice/transformer/decoder_layer.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/embedding.py b/libs/CosyVoice/cosyvoice/transformer/embedding.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/encoder.py b/libs/CosyVoice/cosyvoice/transformer/encoder.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/encoder_layer.py b/libs/CosyVoice/cosyvoice/transformer/encoder_layer.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/label_smoothing_loss.py b/libs/CosyVoice/cosyvoice/transformer/label_smoothing_loss.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/positionwise_feed_forward.py b/libs/CosyVoice/cosyvoice/transformer/positionwise_feed_forward.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/subsampling.py b/libs/CosyVoice/cosyvoice/transformer/subsampling.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/transformer/upsample_encoder.py b/libs/CosyVoice/cosyvoice/transformer/upsample_encoder.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/__init__.py b/libs/CosyVoice/cosyvoice/utils/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/class_utils.py b/libs/CosyVoice/cosyvoice/utils/class_utils.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/common.py b/libs/CosyVoice/cosyvoice/utils/common.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/executor.py b/libs/CosyVoice/cosyvoice/utils/executor.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/file_utils.py b/libs/CosyVoice/cosyvoice/utils/file_utils.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/frontend_utils.py b/libs/CosyVoice/cosyvoice/utils/frontend_utils.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/losses.py b/libs/CosyVoice/cosyvoice/utils/losses.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/mask.py b/libs/CosyVoice/cosyvoice/utils/mask.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/scheduler.py b/libs/CosyVoice/cosyvoice/utils/scheduler.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/utils/train_utils.py b/libs/CosyVoice/cosyvoice/utils/train_utils.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/cosyvoice/vllm/cosyvoice2.py b/libs/CosyVoice/cosyvoice/vllm/cosyvoice2.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/docker/Dockerfile b/libs/CosyVoice/docker/Dockerfile
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/example.py b/libs/CosyVoice/example.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/Dockerfile b/libs/CosyVoice/examples/grpo/cosyvoice2/Dockerfile
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/README.md b/libs/CosyVoice/examples/grpo/cosyvoice2/README.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/huggingface_to_pretrained.py b/libs/CosyVoice/examples/grpo/cosyvoice2/huggingface_to_pretrained.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/infer_dataset.py b/libs/CosyVoice/examples/grpo/cosyvoice2/infer_dataset.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/prepare_data.py b/libs/CosyVoice/examples/grpo/cosyvoice2/prepare_data.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/pretrained_to_huggingface.py b/libs/CosyVoice/examples/grpo/cosyvoice2/pretrained_to_huggingface.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/requirements.txt b/libs/CosyVoice/examples/grpo/cosyvoice2/requirements.txt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/reward_tts.py b/libs/CosyVoice/examples/grpo/cosyvoice2/reward_tts.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/run.sh b/libs/CosyVoice/examples/grpo/cosyvoice2/run.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/scripts/compute_wer.sh b/libs/CosyVoice/examples/grpo/cosyvoice2/scripts/compute_wer.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/scripts/offline-decode-files.py b/libs/CosyVoice/examples/grpo/cosyvoice2/scripts/offline-decode-files.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/grpo/cosyvoice2/token2wav_asr_server.py b/libs/CosyVoice/examples/grpo/cosyvoice2/token2wav_asr_server.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/conf/cosyvoice.yaml b/libs/CosyVoice/examples/libritts/cosyvoice/conf/cosyvoice.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/conf/ds_stage2.json b/libs/CosyVoice/examples/libritts/cosyvoice/conf/ds_stage2.json
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/cosyvoice b/libs/CosyVoice/examples/libritts/cosyvoice/cosyvoice
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/local/download_and_untar.sh b/libs/CosyVoice/examples/libritts/cosyvoice/local/download_and_untar.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/local/prepare_data.py b/libs/CosyVoice/examples/libritts/cosyvoice/local/prepare_data.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/local/prepare_reject_sample.py b/libs/CosyVoice/examples/libritts/cosyvoice/local/prepare_reject_sample.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/path.sh b/libs/CosyVoice/examples/libritts/cosyvoice/path.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/run.sh b/libs/CosyVoice/examples/libritts/cosyvoice/run.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/tools b/libs/CosyVoice/examples/libritts/cosyvoice/tools
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice/tts_text.json b/libs/CosyVoice/examples/libritts/cosyvoice/tts_text.json
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/conf/cosyvoice2.yaml b/libs/CosyVoice/examples/libritts/cosyvoice2/conf/cosyvoice2.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/conf/ds_stage2.json b/libs/CosyVoice/examples/libritts/cosyvoice2/conf/ds_stage2.json
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/cosyvoice b/libs/CosyVoice/examples/libritts/cosyvoice2/cosyvoice
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/local b/libs/CosyVoice/examples/libritts/cosyvoice2/local
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/path.sh b/libs/CosyVoice/examples/libritts/cosyvoice2/path.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/run.sh b/libs/CosyVoice/examples/libritts/cosyvoice2/run.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/run_dpo.sh b/libs/CosyVoice/examples/libritts/cosyvoice2/run_dpo.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/tools b/libs/CosyVoice/examples/libritts/cosyvoice2/tools
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice2/tts_text.json b/libs/CosyVoice/examples/libritts/cosyvoice2/tts_text.json
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice3/conf/cosyvoice3.yaml b/libs/CosyVoice/examples/libritts/cosyvoice3/conf/cosyvoice3.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice3/conf/ds_stage2.json b/libs/CosyVoice/examples/libritts/cosyvoice3/conf/ds_stage2.json
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice3/cosyvoice b/libs/CosyVoice/examples/libritts/cosyvoice3/cosyvoice
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice3/local b/libs/CosyVoice/examples/libritts/cosyvoice3/local
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice3/path.sh b/libs/CosyVoice/examples/libritts/cosyvoice3/path.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice3/run.sh b/libs/CosyVoice/examples/libritts/cosyvoice3/run.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/libritts/cosyvoice3/tools b/libs/CosyVoice/examples/libritts/cosyvoice3/tools
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/magicdata-read/cosyvoice/conf b/libs/CosyVoice/examples/magicdata-read/cosyvoice/conf
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/magicdata-read/cosyvoice/cosyvoice b/libs/CosyVoice/examples/magicdata-read/cosyvoice/cosyvoice
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/magicdata-read/cosyvoice/local/download_and_untar.sh b/libs/CosyVoice/examples/magicdata-read/cosyvoice/local/download_and_untar.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/magicdata-read/cosyvoice/local/prepare_data.py b/libs/CosyVoice/examples/magicdata-read/cosyvoice/local/prepare_data.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/magicdata-read/cosyvoice/path.sh b/libs/CosyVoice/examples/magicdata-read/cosyvoice/path.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/magicdata-read/cosyvoice/run.sh b/libs/CosyVoice/examples/magicdata-read/cosyvoice/run.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/magicdata-read/cosyvoice/tools b/libs/CosyVoice/examples/magicdata-read/cosyvoice/tools
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/examples/magicdata-read/cosyvoice/tts_text.json b/libs/CosyVoice/examples/magicdata-read/cosyvoice/tts_text.json
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/requirements.txt b/libs/CosyVoice/requirements.txt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/python/Dockerfile b/libs/CosyVoice/runtime/python/Dockerfile
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/python/fastapi/client.py b/libs/CosyVoice/runtime/python/fastapi/client.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/python/fastapi/server.py b/libs/CosyVoice/runtime/python/fastapi/server.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/python/grpc/client.py b/libs/CosyVoice/runtime/python/grpc/client.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/python/grpc/cosyvoice.proto b/libs/CosyVoice/runtime/python/grpc/cosyvoice.proto
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/python/grpc/server.py b/libs/CosyVoice/runtime/python/grpc/server.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/Dockerfile.server b/libs/CosyVoice/runtime/triton_trtllm/Dockerfile.server
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/README.DIT.md b/libs/CosyVoice/runtime/triton_trtllm/README.DIT.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/README.md b/libs/CosyVoice/runtime/triton_trtllm/README.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/client_grpc.py b/libs/CosyVoice/runtime/triton_trtllm/client_grpc.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/client_http.py b/libs/CosyVoice/runtime/triton_trtllm/client_http.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/docker-compose.dit.yml b/libs/CosyVoice/runtime/triton_trtllm/docker-compose.dit.yml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/docker-compose.yml b/libs/CosyVoice/runtime/triton_trtllm/docker-compose.yml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/audio_tokenizer/1/model.py b/libs/CosyVoice/runtime/triton_trtllm/model_repo/audio_tokenizer/1/model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/audio_tokenizer/config.pbtxt b/libs/CosyVoice/runtime/triton_trtllm/model_repo/audio_tokenizer/config.pbtxt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/cosyvoice2/1/model.py b/libs/CosyVoice/runtime/triton_trtllm/model_repo/cosyvoice2/1/model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/cosyvoice2/config.pbtxt b/libs/CosyVoice/runtime/triton_trtllm/model_repo/cosyvoice2/config.pbtxt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/cosyvoice2_dit/1/model.py b/libs/CosyVoice/runtime/triton_trtllm/model_repo/cosyvoice2_dit/1/model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/cosyvoice2_dit/config.pbtxt b/libs/CosyVoice/runtime/triton_trtllm/model_repo/cosyvoice2_dit/config.pbtxt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/speaker_embedding/1/model.py b/libs/CosyVoice/runtime/triton_trtllm/model_repo/speaker_embedding/1/model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/speaker_embedding/config.pbtxt b/libs/CosyVoice/runtime/triton_trtllm/model_repo/speaker_embedding/config.pbtxt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/tensorrt_llm/1/.gitkeep b/libs/CosyVoice/runtime/triton_trtllm/model_repo/tensorrt_llm/1/.gitkeep
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/tensorrt_llm/config.pbtxt b/libs/CosyVoice/runtime/triton_trtllm/model_repo/tensorrt_llm/config.pbtxt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav/1/model.py b/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav/1/model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav/config.pbtxt b/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav/config.pbtxt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav_dit/1/model.py b/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav_dit/1/model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav_dit/1/token2wav_dit.py b/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav_dit/1/token2wav_dit.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav_dit/config.pbtxt b/libs/CosyVoice/runtime/triton_trtllm/model_repo/token2wav_dit/config.pbtxt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/offline_inference.py b/libs/CosyVoice/runtime/triton_trtllm/offline_inference.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/requirements.txt b/libs/CosyVoice/runtime/triton_trtllm/requirements.txt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/run.sh b/libs/CosyVoice/runtime/triton_trtllm/run.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/run_stepaudio2_dit_token2wav.sh b/libs/CosyVoice/runtime/triton_trtllm/run_stepaudio2_dit_token2wav.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/scripts/convert_checkpoint.py b/libs/CosyVoice/runtime/triton_trtllm/scripts/convert_checkpoint.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/scripts/fill_template.py b/libs/CosyVoice/runtime/triton_trtllm/scripts/fill_template.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/scripts/test_llm.py b/libs/CosyVoice/runtime/triton_trtllm/scripts/test_llm.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/streaming_inference.py b/libs/CosyVoice/runtime/triton_trtllm/streaming_inference.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/token2wav.py b/libs/CosyVoice/runtime/triton_trtllm/token2wav.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/runtime/triton_trtllm/token2wav_dit.py b/libs/CosyVoice/runtime/triton_trtllm/token2wav_dit.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.env.example b/libs/CosyVoice/third_party/Matcha-TTS/.env.example
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.github/PULL_REQUEST_TEMPLATE.md b/libs/CosyVoice/third_party/Matcha-TTS/.github/PULL_REQUEST_TEMPLATE.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.github/codecov.yml b/libs/CosyVoice/third_party/Matcha-TTS/.github/codecov.yml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.github/dependabot.yml b/libs/CosyVoice/third_party/Matcha-TTS/.github/dependabot.yml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.github/release-drafter.yml b/libs/CosyVoice/third_party/Matcha-TTS/.github/release-drafter.yml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.gitignore b/libs/CosyVoice/third_party/Matcha-TTS/.gitignore
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.pre-commit-config.yaml b/libs/CosyVoice/third_party/Matcha-TTS/.pre-commit-config.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.project-root b/libs/CosyVoice/third_party/Matcha-TTS/.project-root
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/.pylintrc b/libs/CosyVoice/third_party/Matcha-TTS/.pylintrc
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/LICENSE b/libs/CosyVoice/third_party/Matcha-TTS/LICENSE
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/MANIFEST.in b/libs/CosyVoice/third_party/Matcha-TTS/MANIFEST.in
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/Makefile b/libs/CosyVoice/third_party/Matcha-TTS/Makefile
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/README.md b/libs/CosyVoice/third_party/Matcha-TTS/README.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/configs/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/model_checkpoint.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/model_checkpoint.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/model_summary.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/model_summary.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/none.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/none.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/rich_progress_bar.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/callbacks/rich_progress_bar.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/fdr.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/fdr.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/limit.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/limit.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/overfit.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/overfit.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/profiler.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/debug/profiler.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/eval.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/eval.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/experiment/hifi_dataset_piper_phonemizer.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/experiment/hifi_dataset_piper_phonemizer.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/experiment/ljspeech.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/experiment/ljspeech.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/experiment/ljspeech_min_memory.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/experiment/ljspeech_min_memory.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/experiment/multispeaker.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/experiment/multispeaker.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/extras/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/extras/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/hparams_search/mnist_optuna.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/hparams_search/mnist_optuna.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/hydra/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/hydra/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/local/.gitkeep b/libs/CosyVoice/third_party/Matcha-TTS/configs/local/.gitkeep
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/aim.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/aim.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/comet.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/comet.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/csv.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/csv.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/many_loggers.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/many_loggers.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/mlflow.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/mlflow.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/neptune.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/neptune.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/tensorboard.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/tensorboard.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/wandb.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/logger/wandb.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/model/cfm/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/model/cfm/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/model/decoder/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/model/decoder/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/model/encoder/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/model/encoder/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/model/matcha.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/model/matcha.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/model/optimizer/adam.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/model/optimizer/adam.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/paths/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/paths/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/train.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/train.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/cpu.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/cpu.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/ddp.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/ddp.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/ddp_sim.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/ddp_sim.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/default.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/default.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/gpu.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/gpu.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/mps.yaml b/libs/CosyVoice/third_party/Matcha-TTS/configs/trainer/mps.yaml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/VERSION b/libs/CosyVoice/third_party/Matcha-TTS/matcha/VERSION
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/app.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/app.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/cli.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/cli.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/LICENSE b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/LICENSE
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/README.md b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/README.md
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/config.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/config.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/denoiser.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/denoiser.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/env.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/env.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/meldataset.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/meldataset.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/models.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/models.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/xutils.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/hifigan/xutils.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/baselightningmodule.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/baselightningmodule.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/decoder.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/decoder.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/flow_matching.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/flow_matching.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/text_encoder.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/text_encoder.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/transformer.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/components/transformer.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/matcha_tts.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/models/matcha_tts.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/onnx/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/onnx/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/onnx/export.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/onnx/export.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/onnx/infer.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/onnx/infer.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/text/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/text/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/text/cleaners.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/text/cleaners.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/text/numbers.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/text/numbers.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/text/symbols.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/text/symbols.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/train.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/train.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/audio.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/audio.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/generate_data_statistics.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/generate_data_statistics.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/instantiators.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/instantiators.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/logging_utils.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/logging_utils.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/model.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/model.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/monotonic_align/__init__.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/monotonic_align/__init__.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/monotonic_align/core.pyx b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/monotonic_align/core.pyx
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/monotonic_align/setup.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/monotonic_align/setup.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/pylogger.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/pylogger.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/rich_utils.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/rich_utils.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/utils.py b/libs/CosyVoice/third_party/Matcha-TTS/matcha/utils/utils.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/notebooks/.gitkeep b/libs/CosyVoice/third_party/Matcha-TTS/notebooks/.gitkeep
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/pyproject.toml b/libs/CosyVoice/third_party/Matcha-TTS/pyproject.toml
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/requirements.txt b/libs/CosyVoice/third_party/Matcha-TTS/requirements.txt
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/scripts/schedule.sh b/libs/CosyVoice/third_party/Matcha-TTS/scripts/schedule.sh
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/setup.py b/libs/CosyVoice/third_party/Matcha-TTS/setup.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/third_party/Matcha-TTS/synthesis.ipynb b/libs/CosyVoice/third_party/Matcha-TTS/synthesis.ipynb
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/tools/extract_embedding.py b/libs/CosyVoice/tools/extract_embedding.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/tools/extract_speech_token.py b/libs/CosyVoice/tools/extract_speech_token.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/tools/make_parquet_list.py b/libs/CosyVoice/tools/make_parquet_list.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/vllm_example.py b/libs/CosyVoice/vllm_example.py
old mode 100755
new mode 100644
diff --git a/libs/CosyVoice/webui.py b/libs/CosyVoice/webui.py
old mode 100755
new mode 100644
diff --git a/libs/Ditto/.gitignore b/libs/Ditto/.gitignore
deleted file mode 100644
index 3f46987..0000000
--- a/libs/Ditto/.gitignore
+++ /dev/null
@@ -1,40 +0,0 @@
-# Byte-compiled / optimized / DLL files
-__pycache__/
-*__pycache__
-**/__pycache__/
-*.py[cod]
-**/*.py[cod]
-*$py.class
-
-# Model weights
-checkpoints
-**/*.pth
-**/*.onnx
-**/*.pt
-**/*.pth.tar
-
-.idea
-.vscode
-.DS_Store
-*.DS_Store
-
-*.swp
-tmp*
-
-*build
-*.egg-info/
-*.mp4
-
-log/*
-*.mp4
-*.png
-*.jpg
-*.wav
-*.pth
-*.pyc
-*.jpeg
-
-prepare_data
-
-!example/audio.wav
-!example/image.png
diff --git a/libs/Ditto_v1.9.513_backup/.gitignore b/libs/Ditto_v1.9.513_backup/.gitignore
deleted file mode 100644
index 3f46987..0000000
--- a/libs/Ditto_v1.9.513_backup/.gitignore
+++ /dev/null
@@ -1,40 +0,0 @@
-# Byte-compiled / optimized / DLL files
-__pycache__/
-*__pycache__
-**/__pycache__/
-*.py[cod]
-**/*.py[cod]
-*$py.class
-
-# Model weights
-checkpoints
-**/*.pth
-**/*.onnx
-**/*.pt
-**/*.pth.tar
-
-.idea
-.vscode
-.DS_Store
-*.DS_Store
-
-*.swp
-tmp*
-
-*build
-*.egg-info/
-*.mp4
-
-log/*
-*.mp4
-*.png
-*.jpg
-*.wav
-*.pth
-*.pyc
-*.jpeg
-
-prepare_data
-
-!example/audio.wav
-!example/image.png
diff --git a/libs/EchoMimicV3/.gitignore b/libs/EchoMimicV3/.gitignore
deleted file mode 100644
index 98fd9ed..0000000
--- a/libs/EchoMimicV3/.gitignore
+++ /dev/null
@@ -1,168 +0,0 @@
-# Byte-compiled / optimized / DLL files
-output*
-logs*
-taming*
-samples*
-datasets*
-_*
-logs*
-__pycache__/
-*.py[cod]
-*$py.class
-scripts_demo*
-
-# C extensions
-*.so
-
-# Distribution / packaging
-.Python
-build/
-develop-eggs/
-dist/
-downloads/
-eggs/
-.eggs/
-lib/
-lib64/
-parts/
-sdist/
-var/
-wheels/
-share/python-wheels/
-*.egg-info/
-.installed.cfg
-*.egg
-MANIFEST
-
-# PyInstaller
-# Usually these files are written by a python script from a template
-# before PyInstaller builds the exe, so as to inject date/other infos into it.
-*.manifest
-*.spec
-
-# Installer logs
-pip-log.txt
-pip-delete-this-directory.txt
-
-# Unit test / coverage reports
-htmlcov/
-.tox/
-.nox/
-.coverage
-.coverage.*
-.cache
-nosetests.xml
-coverage.xml
-*.cover
-*.py,cover
-.hypothesis/
-.pytest_cache/
-cover/
-
-# Translations
-*.mo
-*.pot
-
-# Django stuff:
-*.log
-local_settings.py
-db.sqlite3
-db.sqlite3-journal
-
-# Flask stuff:
-instance/
-.webassets-cache
-
-# Scrapy stuff:
-.scrapy
-
-# Sphinx documentation
-docs/_build/
-
-# PyBuilder
-.pybuilder/
-target/
-
-# Jupyter Notebook
-.ipynb_checkpoints
-
-# IPython
-profile_default/
-ipython_config.py
-
-# pyenv
-# For a library or package, you might want to ignore these files since the code is
-# intended to run in multiple environments; otherwise, check them in:
-# .python-version
-
-# pipenv
-# According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
-# However, in case of collaboration, if having platform-specific dependencies or dependencies
-# having no cross-platform support, pipenv may install dependencies that don't work, or not
-# install all needed dependencies.
-#Pipfile.lock
-
-# poetry
-# Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
-# This is especially recommended for binary packages to ensure reproducibility, and is more
-# commonly ignored for libraries.
-# https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
-#poetry.lock
-
-# pdm
-# Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
-#pdm.lock
-# pdm stores project-wide configurations in .pdm.toml, but it is recommended to not include it
-# in version control.
-# https://pdm.fming.dev/#use-with-ide
-.pdm.toml
-
-# PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
-__pypackages__/
-
-# Celery stuff
-celerybeat-schedule
-celerybeat.pid
-
-# SageMath parsed files
-*.sage.py
-
-# Environments
-.env
-.venv
-env/
-venv/
-ENV/
-env.bak/
-venv.bak/
-
-# Spyder project settings
-.spyderproject
-.spyproject
-
-# Rope project settings
-.ropeproject
-
-# mkdocs documentation
-/site
-
-# mypy
-.mypy_cache/
-.dmypy.json
-dmypy.json
-
-# Pyre type checker
-.pyre/
-
-# pytype static type analyzer
-.pytype/
-
-# Cython debug symbols
-cython_debug/
-
-# PyCharm
-# JetBrains specific template is maintained in a separate JetBrains.gitignore that can
-# be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
-# and can be added to the global gitignore or merged into this file. For a more nuclear
-# option (not recommended) you can uncomment the following to ignore the entire idea folder.
-#.idea/
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/01.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/01.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/01.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/02.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/02.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/02.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/03.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/03.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/03.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/04.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/04.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/04.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/05.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/05.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/05.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/06.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/06.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/06.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/07.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/07.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/07.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/08.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/08.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/08.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/09.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/09.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/09.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/10.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/10.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/10.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/11.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/11.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/11.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/12.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/12.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/12.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/13.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/13.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/13.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/14.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/14.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/14.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/15.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/15.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/15.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/16.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/16.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/16.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/17.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/17.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/17.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/18.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/18.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/18.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/19.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/19.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/19.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/20.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/20.WAV
new file mode 100644
index 0000000..5a0c96a
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/20.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-1036.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-1036.WAV
new file mode 100644
index 0000000..2f4a0a7
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-1036.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-1942.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-1942.WAV
new file mode 100644
index 0000000..9913766
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-1942.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-2371.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-2371.WAV
new file mode 100644
index 0000000..0f781b4
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-2371.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-3927.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-3927.WAV
new file mode 100644
index 0000000..04ff423
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-3927.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-3945.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-3945.WAV
new file mode 100644
index 0000000..58e5cff
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-3945.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-4513.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-4513.WAV
new file mode 100644
index 0000000..7384c0d
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-4513.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-6032.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-6032.WAV
new file mode 100644
index 0000000..2f4a0a7
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-6032.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7028.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7028.WAV
new file mode 100644
index 0000000..40cceeb
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7028.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7113.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7113.WAV
new file mode 100644
index 0000000..31b392c
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7113.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7335.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7335.WAV
new file mode 100644
index 0000000..69c7de3
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7335.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7825.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7825.WAV
new file mode 100644
index 0000000..04ff423
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-7825.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-8644.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-8644.WAV
new file mode 100644
index 0000000..85d6876
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/2025-07-14-8644.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/21.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/21.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/21.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/22.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/22.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/22.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/23.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/23.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/23.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/24.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/24.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/24.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/25.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/25.WAV
new file mode 100644
index 0000000..ab41cfe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/25.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_02.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_02.png
new file mode 100644
index 0000000..25e2a90
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_02.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_03.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_03.png
new file mode 100644
index 0000000..428bc93
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_03.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_04.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_04.png
new file mode 100644
index 0000000..010ec09
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_04.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_05.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_05.jpg
new file mode 100644
index 0000000..4c26cc9
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_cartoon_05.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_ch_man_01.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_ch_man_01.WAV
new file mode 100644
index 0000000..5a0c96a
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_ch_man_01.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_ch_woman_04.WAV b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_ch_woman_04.WAV
new file mode 100644
index 0000000..6d78eec
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/demo_ch_woman_04.WAV differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/guitar_man_01.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/guitar_man_01.png
new file mode 100644
index 0000000..d3345b5
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/guitar_man_01.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/guitar_woman_01.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/guitar_woman_01.png
new file mode 100644
index 0000000..c4a4ddd
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/guitar_woman_01.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/music_woman_01.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/music_woman_01.png
new file mode 100644
index 0000000..3a7cfde
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/audios/music_woman_01.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/01.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/01.jpg
new file mode 100644
index 0000000..030a327
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/01.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/02.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/02.jpg
new file mode 100644
index 0000000..d596435
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/02.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/03.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/03.jpg
new file mode 100644
index 0000000..ae620c2
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/03.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/04.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/04.jpg
new file mode 100644
index 0000000..5519015
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/04.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/05.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/05.jpg
new file mode 100644
index 0000000..3cebfe9
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/05.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/06.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/06.jpg
new file mode 100644
index 0000000..d92121c
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/06.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/07.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/07.jpg
new file mode 100644
index 0000000..271efc5
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/07.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/08.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/08.jpg
new file mode 100644
index 0000000..4406ef9
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/08.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/09.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/09.jpg
new file mode 100644
index 0000000..8748181
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/09.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/10.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/10.jpg
new file mode 100644
index 0000000..75297c9
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/10.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/11.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/11.jpg
new file mode 100644
index 0000000..736be62
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/11.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/12.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/12.jpg
new file mode 100644
index 0000000..819022b
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/12.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/13.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/13.jpg
new file mode 100644
index 0000000..851179e
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/13.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/14.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/14.jpg
new file mode 100644
index 0000000..dde57fe
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/14.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/15.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/15.jpg
new file mode 100644
index 0000000..0676229
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/15.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/16.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/16.jpg
new file mode 100644
index 0000000..6757618
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/16.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/17.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/17.jpg
new file mode 100644
index 0000000..6c2adf5
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/17.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/18.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/18.jpg
new file mode 100644
index 0000000..3df9789
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/18.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/19.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/19.jpg
new file mode 100644
index 0000000..0b8674e
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/19.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/20.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/20.jpg
new file mode 100644
index 0000000..6199a2e
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/20.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-1036.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-1036.jpg
new file mode 100644
index 0000000..9a4f233
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-1036.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-1942.jpeg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-1942.jpeg
new file mode 100644
index 0000000..60f2506
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-1942.jpeg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-2371.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-2371.jpg
new file mode 100644
index 0000000..a0c8e8a
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-2371.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-3927.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-3927.jpg
new file mode 100644
index 0000000..7a3e858
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-3927.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-3945.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-3945.jpg
new file mode 100644
index 0000000..580f1aa
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-3945.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-4513.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-4513.jpg
new file mode 100644
index 0000000..9f176d9
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-4513.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-6032.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-6032.jpg
new file mode 100644
index 0000000..e047fe9
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-6032.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7028.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7028.png
new file mode 100644
index 0000000..a7160a1
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7028.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7113.jpeg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7113.jpeg
new file mode 100644
index 0000000..e6a648b
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7113.jpeg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7335.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7335.jpg
new file mode 100644
index 0000000..2ba7939
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7335.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7825.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7825.jpg
new file mode 100644
index 0000000..c0d9939
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-7825.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-8644.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-8644.png
new file mode 100644
index 0000000..9ff8bda
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/2025-07-14-8644.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/21.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/21.jpg
new file mode 100644
index 0000000..d67c24d
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/21.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/22.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/22.jpg
new file mode 100644
index 0000000..b206bb4
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/22.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/23.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/23.jpg
new file mode 100644
index 0000000..b1ae4d5
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/23.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/24.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/24.jpg
new file mode 100644
index 0000000..2eca07d
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/24.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/25.jpg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/25.jpg
new file mode 100644
index 0000000..f6c911a
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/25.jpg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_02.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_02.txt
new file mode 100644
index 0000000..501c86c
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_02.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The character's hands are not visible in the image, so there are no specific hand or finger movements to describe. Body Positions and Posture: The character is standing upright with a relaxed posture. Body Coverage in Frame: The character is fully visible from the shoulders up. Face Expressions Change: The character has a joyful expression with wide eyes and a big smile. Eyes Movement: The eyes are wide open, indicating excitement or happiness. Head Movement: The head is slightly tilted to one side, adding to the cheerful demeanor. Overall Description: The character appears to be a young boy with a joyful and excited expression. He is wearing a yellow sweater with a red heart on it. The background is plain and neutral, which keeps the focus entirely on the character. The overall impression is that of a happy and lively individual, possibly reacting to something amusing or delightful. There is no indication of any jewelry or accessories being worn by the character.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_03.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_03.txt
new file mode 100644
index 0000000..75e3eb7
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_03.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The character's hands are not visible in the image. Body Positions and Posture: The character is standing upright with a neutral posture. Body Coverage in Frame: The character is fully visible from the shoulders up. Face Expressions Change: The character has a cheerful expression with a slight smile and bright eyes. Eyes Movement: The character's eyes are looking directly at the viewer. Head Movement: The character's head is slightly tilted to one side. Overall Description: The character appears to be a young girl with long dark hair, wearing a light-colored sleeveless top. She is standing against a plain pink background. Her facial expression is friendly and engaging, suggesting she might be interacting with someone or something off-camera. The overall tone of the image is warm and inviting. There is no visible jewelry on the character.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_04.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_04.txt
new file mode 100644
index 0000000..9b9cf43
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_04.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The character is holding a book with both hands. The fingers are positioned around the edges of the book, providing a steady grip. There are no significant changes in the hand or finger movements throughout the video. Body Positions and Posture: The character is standing upright, facing forward. The body remains stationary with minimal movement. Body Coverage in Frame: The character occupies a central position within the frame, with the upper body and face clearly visible. Front-Facing Shot: The character is in a front-facing shot, with the face directly facing the viewer. Face Expressions Change: The character has a neutral expression, with a slight smile. The eyes are open and looking straight ahead. Eyes Movement: The eyes remain fixed on the viewer, with no noticeable movement. Head Movement: The head is slightly tilted downward, focusing on the book being held. Overall Description: The character appears to be engaged in reading or studying, as indicated by the book being held in both hands. The neutral facial expression suggests concentration or calmness. The overall posture and stance convey a sense of stillness and focus. Jewelry: There is no visible jewelry on the character. Background: The background is a plain light blue color, providing a simple and unobtrusive backdrop that keeps the focus on the character. Clothing: The character is wearing a pink cardigan over a white sweater with a subtle pattern. The outfit gives a casual and comfortable appearance.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_05.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_05.txt
new file mode 100644
index 0000000..f1901ad
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_cartoon_05.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The character starts with one hand raised in a thumbs-up gesture. As the video progresses, the character extends both arms outward, palms facing forward, as if embracing or welcoming someone. Body Positions and Posture: The character maintains a standing position throughout the video, with slight adjustments to the stance to accommodate the arm movements. Face Expressions Change: The character's facial expression remains consistently cheerful and friendly, with a wide smile and bright eyes. Eyes Movement: The eyes remain fixed forward, maintaining direct engagement with the viewer. Head Movement: The head slightly tilts from side to side, adding a dynamic element to the otherwise static pose. Overall Description: The character appears to be engaging with the viewer in a friendly and welcoming manner. The thumbs-up gesture initially conveys approval or positivity, while the subsequent extended arms suggest openness and inclusivity. The consistent smile and direct gaze maintain a positive and approachable demeanor throughout the video. The character's attire, a light blue hoodie, complements its cheerful appearance, reinforcing the overall friendly and inviting tone of the video. There is no jewelry visible on the character. Background: The background is plain white, ensuring that all attention remains on the character itself. Clothing: The character is wearing a light blue hoodie with a small logo on the chest area.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_ch_man_01.jpeg b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_ch_man_01.jpeg
new file mode 100644
index 0000000..bab8b36
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_ch_man_01.jpeg differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_ch_woman_04.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_ch_woman_04.png
new file mode 100644
index 0000000..6631af4
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/demo_ch_woman_04.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/guitar_man_01.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/guitar_man_01.png
new file mode 100644
index 0000000..d3345b5
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/guitar_man_01.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/guitar_woman_01.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/guitar_woman_01.png
new file mode 100644
index 0000000..c4a4ddd
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/guitar_woman_01.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/music_woman_01.png b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/music_woman_01.png
new file mode 100644
index 0000000..3a7cfde
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/imgs/music_woman_01.png differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/01.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/01.npy
new file mode 100644
index 0000000..df1a358
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/01.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/02.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/02.npy
new file mode 100644
index 0000000..707ccf3
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/02.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/03.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/03.npy
new file mode 100644
index 0000000..6105b70
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/03.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/04.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/04.npy
new file mode 100644
index 0000000..5bb0c11
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/04.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/05.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/05.npy
new file mode 100644
index 0000000..a4bfc1b
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/05.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/06.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/06.npy
new file mode 100644
index 0000000..c1a60ce
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/06.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/07.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/07.npy
new file mode 100644
index 0000000..8ebc76f
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/07.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/08.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/08.npy
new file mode 100644
index 0000000..272782d
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/08.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/09.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/09.npy
new file mode 100644
index 0000000..24f3346
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/09.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/10.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/10.npy
new file mode 100644
index 0000000..06a9f9c
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/10.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/11.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/11.npy
new file mode 100644
index 0000000..23f7966
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/11.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/12.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/12.npy
new file mode 100644
index 0000000..a78c765
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/12.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/13.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/13.npy
new file mode 100644
index 0000000..497ac24
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/13.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/14.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/14.npy
new file mode 100644
index 0000000..69d72b0
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/14.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/15.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/15.npy
new file mode 100644
index 0000000..9dfb7f6
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/15.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/16.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/16.npy
new file mode 100644
index 0000000..de25f5b
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/16.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/17.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/17.npy
new file mode 100644
index 0000000..3712994
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/17.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/18.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/18.npy
new file mode 100644
index 0000000..1c02738
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/18.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/19.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/19.npy
new file mode 100644
index 0000000..265d7f6
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/19.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/20.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/20.npy
new file mode 100644
index 0000000..5375b51
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/20.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-1036.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-1036.npy
new file mode 100644
index 0000000..1eadd74
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-1036.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-1942.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-1942.npy
new file mode 100644
index 0000000..ef3c5fc
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-1942.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-2371.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-2371.npy
new file mode 100644
index 0000000..1ac4bac
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-2371.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-3927.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-3927.npy
new file mode 100644
index 0000000..e3c1fa5
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-3927.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-3945.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-3945.npy
new file mode 100644
index 0000000..5514040
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-3945.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-4513.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-4513.npy
new file mode 100644
index 0000000..763d584
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-4513.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-6032.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-6032.npy
new file mode 100644
index 0000000..0b0a237
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-6032.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7028.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7028.npy
new file mode 100644
index 0000000..d95fea2
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7028.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7113.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7113.npy
new file mode 100644
index 0000000..f8c19f5
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7113.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7335.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7335.npy
new file mode 100644
index 0000000..4a68ce1
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7335.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7825.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7825.npy
new file mode 100644
index 0000000..1ed13d4
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-7825.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-8644.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-8644.npy
new file mode 100644
index 0000000..0f479da
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/2025-07-14-8644.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/21.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/21.npy
new file mode 100644
index 0000000..ee8373c
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/21.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/22.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/22.npy
new file mode 100644
index 0000000..a69fc8e
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/22.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/23.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/23.npy
new file mode 100644
index 0000000..c568f01
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/23.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/24.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/24.npy
new file mode 100644
index 0000000..4f0b2bf
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/24.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/25.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/25.npy
new file mode 100644
index 0000000..281412e
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/25.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_02.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_02.npy
new file mode 100644
index 0000000..ad23ca0
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_02.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_03.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_03.npy
new file mode 100644
index 0000000..162047d
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_03.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_04.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_04.npy
new file mode 100644
index 0000000..85873a1
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_04.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_05.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_05.npy
new file mode 100644
index 0000000..162047d
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_cartoon_05.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_ch_man_01.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_ch_man_01.npy
new file mode 100644
index 0000000..a5dac93
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_ch_man_01.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_ch_woman_04.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_ch_woman_04.npy
new file mode 100644
index 0000000..7355d92
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/demo_ch_woman_04.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/guitar_man_01.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/guitar_man_01.npy
new file mode 100644
index 0000000..ad89625
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/guitar_man_01.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/guitar_woman_01.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/guitar_woman_01.npy
new file mode 100644
index 0000000..e5aad02
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/guitar_woman_01.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/music_woman_01.npy b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/music_woman_01.npy
new file mode 100644
index 0000000..f4f6888
Binary files /dev/null and b/libs/EchoMimicV3/datasets/echomimicv3_demos/masks/music_woman_01.npy differ
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/01.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/01.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/01.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/02.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/02.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/02.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/03.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/03.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/03.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/04.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/04.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/04.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/05.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/05.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/05.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/06.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/06.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/06.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/07.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/07.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/07.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/08.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/08.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/08.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/09.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/09.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/09.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/10.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/10.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/10.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/11.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/11.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/11.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/12.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/12.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/12.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/13.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/13.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/13.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/14.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/14.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/14.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/15.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/15.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/15.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/16.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/16.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/16.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/17.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/17.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/17.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/18.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/18.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/18.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/19.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/19.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/19.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/20.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/20.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/20.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-1036.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-1036.txt
new file mode 100644
index 0000000..3d6ca60
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-1036.txt
@@ -0,0 +1 @@
+Body Positions and Posture: The individual is standing upright with their hands in the pockets of their suit. The body is positioned slightly angled towards the camera, giving a confident stance. Hands and Fingers Movements: The hands are relaxed and remain in the pockets throughout the video. There are no noticeable finger movements or gestures. Gestures Changes Trajectory: The hands do not move from the pockets, indicating a static position. Fingers Movements: The fingers are not visible as they are inside the pockets. Object in Hand: No objects are held in the hands. Body Coverage in Frame: The individual occupies most of the frame, with the upper body and head clearly visible. Face Expressions Change: The face appears neutral, with a composed and professional demeanor. There are no significant changes in facial expressions throughout the video. Eyes Movement: The eyes are looking directly at the camera, maintaining steady eye contact. Head Movements: The head remains relatively still, with minimal movement. Facial Expression and Emotion: The overall expression is calm and professional, suggesting confidence and composure. Jewelry: No jewelry is visible on the individual. Background: The background is minimalistic, featuring a plain wall and a window to the left side of the frame. The lighting is soft and even, highlighting the subject without harsh shadows. Clothing: The individual is dressed in a dark blue suit with a white dress shirt. The suit jacket is unbuttoned, revealing the shirt collar. The pants are neatly tailored, and the overall attire conveys a formal and professional appearance. Overall Description: The individual is standing confidently with hands in pockets, wearing a formal suit. The setting is simple and professional, with soft lighting that emphasizes the subject. The person maintains a neutral and composed expression, suggesting a formal or business context.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-1942.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-1942.txt
new file mode 100644
index 0000000..ad2a138
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-1942.txt
@@ -0,0 +1 @@
+Hands and Fingers Movements: The character's hands are positioned in front of their body, with the right hand slightly raised and the left hand resting lower. The fingers on both hands appear relaxed, with minimal movement observed. Body Positions and Posture: The character is seated, leaning slightly forward. The body is covered from the shoulders down to the waist, with the upper body prominently displayed. Face Expressions and Emotion: The character has a neutral expression, with a slight smile. The eyes are open, looking directly at the viewer. There is no significant change in facial expressions or emotions throughout the video. Head Movements: The head remains relatively still, with minor adjustments to maintain eye contact with the viewer. There is no noticeable head tilting or turning. Jewelry: The character wears a silver headband with a crescent moon design on top of their head. Background: The background appears to be a forest setting with green foliage and trees, suggesting an outdoor environment. The lighting is natural, indicating daytime. Clothing: The character is dressed in a yellow outfit with a blue scarf tied around their neck. The outfit seems to be traditional or fantasy-inspired, fitting the character's appearance. Overall Description: The character is depicted in a seated position, wearing a yellow outfit with a blue scarf and a silver headband with a crescent moon design. The character's hands are relaxed and positioned in front of them, with minimal movement. The background suggests an outdoor forest setting, and the character's neutral expression and direct gaze indicate a calm and composed demeanor. The overall scene conveys a sense of tranquility and focus.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-2371.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-2371.txt
new file mode 100644
index 0000000..edd2f98
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-2371.txt
@@ -0,0 +1 @@
+Body Positions and Postures: The person is standing upright with their body facing forward. Their shoulders are relaxed, and their arms are positioned at their sides. Hand Movements and Gestures: Throughout the video, the person's hands are mostly out of frame or partially visible. There are no significant hand movements or gestures that can be observed clearly. Finger Movements: No specific finger movements can be discerned from the provided frames. Object in Hand: There is no object visible in the person's hands. Clothing: The person is wearing a plain white T-shirt. Jewelry: The person is wearing small earrings on both ears. Background: The background appears to be an indoor setting, possibly a room with warm lighting. There are blurred elements such as furniture and possibly a television or monitor in the distance. Face Expressions and Emotion: The person's facial expression seems neutral or slightly engaged. There are no significant changes in facial expressions or emotions observable across the frames. Head Movements: The person's head remains relatively still, with minimal movement detected. Overall Description: The person is standing in a neutral pose, facing the camera directly. The video captures the upper body and face, focusing on the person's facial expressions and slight head movements. The environment suggests a casual indoor setting, and the person appears to be speaking or presenting, given the direct gaze towards the camera.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-3927.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-3927.txt
new file mode 100644
index 0000000..f44cc5e
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-3927.txt
@@ -0,0 +1 @@
+Hands and Fingers Movements: The individual's hands are not visible in the image, as the focus is on the upper body and face. Body Positions and Posture: The person is standing upright with a straight posture. The shoulders are slightly squared, indicating a formal stance. Face Expressions Change: The individual appears to have a neutral or slightly smiling expression. The mouth is slightly open, suggesting that they might be speaking or about to speak. Eyes Movement: The eyes are looking directly at the camera, maintaining steady contact. Head Movements: The head remains relatively still, with minimal movement. There is a slight tilt of the head to the right side, which could indicate a momentary shift in attention or a gesture. Overall Description: The person is dressed in a formal suit, consisting of a dark blue blazer, a white dress shirt, and a black tie. A white pocket square is visible in the breast pocket of the blazer. The background is dark, which helps to highlight the subject. The lighting is focused on the individual, creating a professional and polished appearance. Jewelry: No jewelry is visible on the individual. Clothing: The outfit is formal and well-tailored, suitable for a professional setting or a formal event. Conclusion: The individual appears to be in a professional or formal setting, possibly giving a speech or participating in a formal interview. The neutral to slightly smiling expression suggests a composed demeanor, while the direct eye contact indicates engagement with the audience or interviewer. The overall impression is one of confidence and professionalism.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-3945.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-3945.txt
new file mode 100644
index 0000000..b823aec
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-3945.txt
@@ -0,0 +1 @@
+Hands and Fingers Movements: The person's hands are not visible in the frames provided. Body Positions and Posture: The person appears to be standing upright. There are slight shifts in posture as the person moves their head and shoulders slightly. Face Expressions Change: The person's facial expression changes from a smile to a neutral or slightly open-mouthed expression. The eyes appear to look directly at the camera initially and then shift slightly to the side. Eyes Movement: The eyes move from looking straight ahead to looking slightly to the side. Head Movements: The head tilts slightly from one side to another. There are minor movements of the neck and shoulders. Overall Description: The person seems to be standing on a street with blurred background activity, suggesting a casual outdoor setting. The individual is wearing a gray t-shirt and a black jacket, indicating a relaxed yet put-together appearance. The background shows a busy street scene with pedestrians and vehicles, giving the impression of a lively urban environment. The lighting suggests it might be daytime, possibly late afternoon given the softness of the light. Clothing: Gray t-shirt Black jacket Jewelry: No jewelry is visible on the person. Background: A busy street with blurred figures of pedestrians and vehicles. The presence of traffic lights and buildings indicates an urban setting. Relationship Between Object and Person’s Hands: Since the hands are not visible, there is no interaction with objects or gestures involving the hands. Relationship Between Object and Person’s Body: The person is positioned centrally in the frame, facing forward. The body coverage is from the shoulders up, allowing for clear visibility of the upper body and face. Role: The person appears to be posing for the camera, possibly for a portrait or a casual video recording.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-4513.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-4513.txt
new file mode 100644
index 0000000..53cf0ef
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-4513.txt
@@ -0,0 +1 @@
+The image provided does not contain any visible hands or fingers, nor does it show any objects being held. The person in the image appears to be standing in a front-facing position, wearing a white knitted sweater. The background is a plain light blue color, which suggests a simple and uncluttered setting. Body Positions and Posture: The person is standing upright with their body facing forward. There are no significant body movements observed in this still image. Face Expressions Change: The person has a neutral expression with slightly parted lips. There are no noticeable changes in facial expressions as this is a single static image. Eyes Movement: The eyes appear to be looking directly at the camera, suggesting a direct engagement with the viewer. Head Movements: The head is positioned straight, with no tilting or turning observed. Overall Description: The character in the image is standing still, facing the camera, and wearing a white knitted sweater. The background is plain and light blue, providing a clean and minimalistic backdrop. There are no visible hands, fingers, or objects in the image, and no jewelry is present on the person. The overall impression is one of simplicity and focus on the individual's attire and posture.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-6032.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-6032.txt
new file mode 100644
index 0000000..4b48a62
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-6032.txt
@@ -0,0 +1 @@
+Hands and Fingers Movements: The individual's hands are not visible in the image, as the focus is on their upper body and face. Body Positions and Posture: The person is standing upright with a relaxed posture. Their shoulders are slightly squared, suggesting a confident stance. Body Coverage in Frame: The person is fully visible from the chest up, with the background showing a clear sky and water, indicating they are outdoors. Face Expressions Change: The person has a neutral expression, with their lips closed and a calm demeanor. There are no significant changes in facial expressions throughout the image. Eyes Movement: The person’s eyes are looking directly at the camera, maintaining steady eye contact. Head Movements: The head is held straight, with minimal movement, reinforcing the calm and composed expression. Jewelry: No jewelry is visible on the person in this image. Background: The background features a clear blue sky and a body of water, possibly an ocean or sea, suggesting a serene outdoor setting. Clothing: The person is wearing a white, open-collared shirt with a pocket on the left side. The shirt appears to be made of a lightweight fabric, suitable for warm weather. Overall Description: The individual in the image is standing outdoors, likely near a body of water, dressed in a casual yet stylish white shirt. They appear relaxed and composed, with a neutral expression and direct eye contact with the camera. The background suggests a peaceful and natural environment, enhancing the overall serene atmosphere of the scene.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7028.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7028.txt
new file mode 100644
index 0000000..07158b2
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7028.txt
@@ -0,0 +1 @@
+The image provided does not contain any visible hands or fingers, nor does it show any actions or movements of the person depicted. The individual appears to be in a front-facing shot, with their face clearly visible. The person has a serious expression, with furrowed brows and a slightly open mouth, suggesting a moment of intense focus or determination. Background: The background is abstract and predominantly green with some darker tones, giving it a somewhat dramatic and historical feel. This could suggest that the setting is meant to evoke a sense of ancient or legendary times. Clothing: The person is wearing a green headscarf and a green garment with gold accents, which could indicate a historical or fantasy context. The attire suggests a warrior or noble character from a period where such clothing was common. Jewelry: There is no visible jewelry on the person in this image. Overall Description: The character appears to be a warrior or noble figure, possibly from a historical or fantasy setting. The serious expression and the detailed attire suggest a moment of contemplation or readiness for action. The background adds to the dramatic effect, enhancing the impression of a significant or pivotal moment in the narrative.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7113.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7113.txt
new file mode 100644
index 0000000..5c45a09
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7113.txt
@@ -0,0 +1 @@
+Hands and Fingers Movements: The character's hands are positioned in front of their body, with palms facing slightly outward. There is minimal movement in the hands throughout the frames, suggesting a static or contemplative pose. The fingers remain relaxed and do not show any specific gestures or movements. Object in Hand: No object is held in the hands. The hands appear empty, emphasizing a gesture of readiness or anticipation. Body Positions and Posture: The character stands upright, maintaining a straight posture. The body is fully visible within the frame, indicating a frontal shot. Face Expressions and Emotion: The character's facial expression appears calm and composed. The eyes are open, looking directly at the camera, which suggests attentiveness or engagement. There are no significant changes in facial expressions or emotions across the frames. Head Movements: The head remains steady, with minimal movement. The gaze is fixed forward, reinforcing the sense of focus or contemplation. Overall Description: The character is depicted in a frontal shot, wearing elaborate traditional attire that includes a golden robe and a ornate headdress. The setting appears to be a mystical or historical environment, possibly a temple or palace, given the architectural elements in the background. The character’s posture and expression suggest a moment of pause or reflection, with no dynamic action taking place. The overall mood conveyed is one of serenity and reverence. Jewelry: The character is adorned with a large, intricate headdress featuring gold and red accents, which is a prominent piece of jewelry. The golden robe also has detailed embroidery and embellishments, adding to the regal appearance. Background: The background depicts a traditional Chinese architectural style, with elements such as tiled roofs and stone structures. The setting appears to be indoors, possibly within a temple or palace, with a serene and somewhat mystical atmosphere enhanced by soft lighting and shadows. Clothing: The character wears a richly decorated golden robe with intricate patterns and a deep red undergarment. The headdress is elaborate, with gold and red details, and features decorative elements that suggest a high status or divine nature. Overall, the character exudes a sense of authority and spirituality, fitting the context of a traditional or mythological setting.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7335.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7335.txt
new file mode 100644
index 0000000..4ef7933
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7335.txt
@@ -0,0 +1 @@
+Body Positions and Postures: The individual is standing with their back slightly turned to the camera. Their body is angled slightly to the side, facing towards the right of the frame. Body Movements: There is minimal movement observed in the video. The individual appears to be stationary for the duration of the frames shown. Hands and Fingers Movements: The hands are not visible in the frame, suggesting that they may be behind the individual or obscured by clothing. Clothing: The individual is wearing a dark brown cowboy hat. They have on a black leather jacket with a high collar. A brown belt with a silver buckle is visible around their waist. The lower part of the outfit includes brown chaps or pants, which are partially visible at the bottom of the frame. Background: The background features a desert-like landscape with dry grass and sparse vegetation. In the distance, there are mountains under a clear blue sky, indicating a sunny day. Face Expressions and Emotion: The individual’s face is partially visible, showing a neutral expression. The mouth is slightly open, possibly indicating that they are speaking or about to speak. The eyes are looking slightly to the left, away from the camera. Jewelry: No jewelry is visible on the individual in the provided frames. Overall Description: The individual appears to be dressed in a rugged, Western-style outfit, suitable for a desert environment. They are standing in a calm and composed manner, with minimal movement. The setting suggests a scene set in a remote, arid location, possibly during a conversation or a moment of contemplation.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7825.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7825.txt
new file mode 100644
index 0000000..f64d563
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-7825.txt
@@ -0,0 +1 @@
+Hands and Fingers Movements: The individual's right hand is holding a piece of paper or document, which appears to be slightly crumpled. The fingers are positioned around the edges of the paper, suggesting that the person might be referencing it while speaking. The left hand is partially visible, with the thumb and fingers gripping the edge of the podium, indicating stability and support. Body Positions and Posture: The person is standing upright, leaning slightly forward, which suggests engagement and focus on the audience or the task at hand. The body is angled slightly to the side, but the overall posture remains straight and composed. Body Coverage in Frame: The individual occupies a significant portion of the frame, from the waist up, allowing for clear visibility of their attire and gestures. Face Expressions Change: The face shows a serious and focused expression, with the mouth slightly open as if mid-speech. The eyes appear to be looking directly ahead, possibly at an audience or a point of reference. Eyes Movement: The eyes remain fixed forward, maintaining a steady gaze, which conveys attentiveness and seriousness. Head Movements: The head is mostly stationary, with minimal movement, reinforcing the impression of a formal and deliberate presentation. Facial Expression and Emotion: The overall expression is one of concentration and authority, typical of someone delivering an important speech or presentation. Jewelry: No visible jewelry is seen on the individual. Background: The background is simple and uncluttered, featuring neutral-colored curtains and a wooden podium. This setting suggests a formal environment, such as a courtroom, lecture hall, or conference room. Clothing: The individual is dressed in a formal three-piece suit with a patterned jacket, a white shirt, and a tie. A pocket square is visible in the breast pocket, adding a touch of elegance to the outfit. The suit appears well-tailored, fitting the professional context of the scene. Overall Description: The character is delivering a speech or presentation in a formal setting. The individual's posture, hand gestures, and facial expressions all contribute to a sense of authority and seriousness. The setting and attire further emphasize the importance of the occasion.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-8644.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-8644.txt
new file mode 100644
index 0000000..a9aa48d
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/2025-07-14-8644.txt
@@ -0,0 +1 @@
+The image provided does not contain any visible hands or fingers, nor does it show any clear actions or movements of the person depicted. The individual appears to be in a front-facing shot, with their face clearly visible. The person has a stern expression, with furrowed brows and a slightly open mouth, suggesting a moment of intense focus or determination. Jewelry: The person is wearing a golden crown with intricate designs, which suggests a high-ranking or noble status. There are also ornate decorations on the green headpiece, adding to the regal appearance. Background: The background is somewhat blurred but appears to depict a natural setting, possibly a forest or mountainous area, with hints of greenery and earthy tones. This setting might suggest a historical or fantasy context. Clothing: The individual is dressed in elaborate armor and attire, predominantly green with gold accents. The armor includes detailed patterns and embellishments, indicating a warrior or noble character from a historical or fictional setting. The clothing and armor suggest a significant role, possibly a leader or a warrior of importance. Overall Description: The character appears to be a warrior or noble figure, likely from a historical or fantasy context. The stern expression and detailed attire suggest a moment of contemplation or readiness for action. The background and jewelry further emphasize the character's elevated status and the significance of the scene.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/21.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/21.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/21.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/22.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/22.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/22.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/23.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/23.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/23.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/24.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/24.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/24.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/25.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/25.txt
new file mode 100644
index 0000000..fb9f7b6
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/25.txt
@@ -0,0 +1 @@
+A person is holding an object in a relaxed pose. As the video progresses, the character speaks while arm and body movements are minimal and consistent with a natural speaking posture. Hand movements remain minimal. Don't blink too often. Preserve background integrity matching the reference image's spatial configuration, lighting conditions, and color temperature.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_02.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_02.txt
new file mode 100644
index 0000000..501c86c
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_02.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The character's hands are not visible in the image, so there are no specific hand or finger movements to describe. Body Positions and Posture: The character is standing upright with a relaxed posture. Body Coverage in Frame: The character is fully visible from the shoulders up. Face Expressions Change: The character has a joyful expression with wide eyes and a big smile. Eyes Movement: The eyes are wide open, indicating excitement or happiness. Head Movement: The head is slightly tilted to one side, adding to the cheerful demeanor. Overall Description: The character appears to be a young boy with a joyful and excited expression. He is wearing a yellow sweater with a red heart on it. The background is plain and neutral, which keeps the focus entirely on the character. The overall impression is that of a happy and lively individual, possibly reacting to something amusing or delightful. There is no indication of any jewelry or accessories being worn by the character.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_03.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_03.txt
new file mode 100644
index 0000000..75e3eb7
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_03.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The character's hands are not visible in the image. Body Positions and Posture: The character is standing upright with a neutral posture. Body Coverage in Frame: The character is fully visible from the shoulders up. Face Expressions Change: The character has a cheerful expression with a slight smile and bright eyes. Eyes Movement: The character's eyes are looking directly at the viewer. Head Movement: The character's head is slightly tilted to one side. Overall Description: The character appears to be a young girl with long dark hair, wearing a light-colored sleeveless top. She is standing against a plain pink background. Her facial expression is friendly and engaging, suggesting she might be interacting with someone or something off-camera. The overall tone of the image is warm and inviting. There is no visible jewelry on the character.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_04.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_04.txt
new file mode 100644
index 0000000..9b9cf43
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_04.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The character is holding a book with both hands. The fingers are positioned around the edges of the book, providing a steady grip. There are no significant changes in the hand or finger movements throughout the video. Body Positions and Posture: The character is standing upright, facing forward. The body remains stationary with minimal movement. Body Coverage in Frame: The character occupies a central position within the frame, with the upper body and face clearly visible. Front-Facing Shot: The character is in a front-facing shot, with the face directly facing the viewer. Face Expressions Change: The character has a neutral expression, with a slight smile. The eyes are open and looking straight ahead. Eyes Movement: The eyes remain fixed on the viewer, with no noticeable movement. Head Movement: The head is slightly tilted downward, focusing on the book being held. Overall Description: The character appears to be engaged in reading or studying, as indicated by the book being held in both hands. The neutral facial expression suggests concentration or calmness. The overall posture and stance convey a sense of stillness and focus. Jewelry: There is no visible jewelry on the character. Background: The background is a plain light blue color, providing a simple and unobtrusive backdrop that keeps the focus on the character. Clothing: The character is wearing a pink cardigan over a white sweater with a subtle pattern. The outfit gives a casual and comfortable appearance.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_05.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_05.txt
new file mode 100644
index 0000000..f1901ad
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_cartoon_05.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The character starts with one hand raised in a thumbs-up gesture. As the video progresses, the character extends both arms outward, palms facing forward, as if embracing or welcoming someone. Body Positions and Posture: The character maintains a standing position throughout the video, with slight adjustments to the stance to accommodate the arm movements. Face Expressions Change: The character's facial expression remains consistently cheerful and friendly, with a wide smile and bright eyes. Eyes Movement: The eyes remain fixed forward, maintaining direct engagement with the viewer. Head Movement: The head slightly tilts from side to side, adding a dynamic element to the otherwise static pose. Overall Description: The character appears to be engaging with the viewer in a friendly and welcoming manner. The thumbs-up gesture initially conveys approval or positivity, while the subsequent extended arms suggest openness and inclusivity. The consistent smile and direct gaze maintain a positive and approachable demeanor throughout the video. The character's attire, a light blue hoodie, complements its cheerful appearance, reinforcing the overall friendly and inviting tone of the video. There is no jewelry visible on the character. Background: The background is plain white, ensuring that all attention remains on the character itself. Clothing: The character is wearing a light blue hoodie with a small logo on the chest area.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_ch_man_01.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_ch_man_01.txt
new file mode 100644
index 0000000..0363c0d
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_ch_man_01.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The individual is actively gesturing with both hands. Initially, the hands are open with palms facing up, fingers slightly spread apart. As the video progresses, the hands move closer together, fingers interlocking briefly before spreading out again. The hands then return to the initial position with palms facing up. Body Positions and Posture: The person remains seated at a table throughout the video. There is minimal movement of the torso, maintaining a relatively stable posture. Body Coverage in Frame: The person is fully visible from the waist up, covering most of the frame. Face Expressions Change: The individual appears to be speaking or explaining something, as indicated by their hand gestures and facial expressions. The mouth is open, suggesting speech, and the eyes are focused forward, indicating engagement with the audience. Eyes Movement: The eyes remain mostly fixed forward, occasionally shifting slightly but generally maintaining focus on the camera. Head Movement: The head is mostly stationary, with slight movements corresponding to the person’s gestures. Jewelry: Wristwatch: The individual is wearing a wristwatch on the left wrist. Background: The background features a modern, tech-oriented setting with shelves holding various items, possibly tech gadgets or equipment. The lighting is bright, with a blue hue dominating the scene. Clothing: The individual is wearing a dark-colored t-shirt with a small logo on the left chest area. The shirt is short-sleeved, and the person has a wristwatch on the left wrist. Overall Description: The character appears to be delivering a presentation or explanation, using hand gestures to emphasize points. The setting suggests a tech-related context, possibly a review or tutorial video. The individual is actively engaging with the audience through their gestures and facial expressions.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_ch_woman_04.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_ch_woman_04.txt
new file mode 100644
index 0000000..e4e02d3
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/demo_ch_woman_04.txt
@@ -0,0 +1 @@
+ Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The person's hands are not visible in the image, so there are no specific hand or finger movements to describe. Body Positions and Posture: The person is standing upright with a straight posture. Body Coverage in Frame: The person is fully visible from the waist up. Face Expressions Change: The person appears to have a neutral expression with a slight smile. Eyes Movement: The eyes are looking directly at the camera. Head Movement: The head is slightly tilted forward. Overall Description: The character is standing in a front-facing shot, wearing a pink knitted vest over a white collared shirt and a white pleated skirt. The background appears to be a studio setting with soft lighting and some blurred elements that suggest a modern, clean environment. The person is adorned with pearl earrings, adding a touch of elegance to their appearance. The overall impression is one of a professional or casual presentation, possibly for a broadcast or a photoshoot.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/guitar_man_01.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/guitar_man_01.txt
new file mode 100644
index 0000000..18a0d04
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/guitar_man_01.txt
@@ -0,0 +1 @@
+Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The person is playing an acoustic guitar. Their left hand is positioned on the fretboard, pressing down on the strings to form chords. The right hand is strumming the strings with a pick. The fingers on the left hand move up and down the fretboard as they change chords. The right hand moves rhythmically, indicating a steady strumming motion. Body Coverage in Frame: The person occupies a significant portion of the frame, with their upper body and part of their lower body visible. Face Expressions Change: The person appears focused and concentrated on playing the guitar. There are subtle movements of the mouth and eyes that suggest they might be singing along or reacting to the music. The head remains mostly stationary.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/guitar_woman_01.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/guitar_woman_01.txt
new file mode 100644
index 0000000..18a0d04
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/guitar_woman_01.txt
@@ -0,0 +1 @@
+Gesture, Body, Face Expressions and Movements: Hands and Fingers Movements: The person is playing an acoustic guitar. Their left hand is positioned on the fretboard, pressing down on the strings to form chords. The right hand is strumming the strings with a pick. The fingers on the left hand move up and down the fretboard as they change chords. The right hand moves rhythmically, indicating a steady strumming motion. Body Coverage in Frame: The person occupies a significant portion of the frame, with their upper body and part of their lower body visible. Face Expressions Change: The person appears focused and concentrated on playing the guitar. There are subtle movements of the mouth and eyes that suggest they might be singing along or reacting to the music. The head remains mostly stationary.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/music_woman_01.txt b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/music_woman_01.txt
new file mode 100644
index 0000000..aa96a62
--- /dev/null
+++ b/libs/EchoMimicV3/datasets/echomimicv3_demos/prompts/music_woman_01.txt
@@ -0,0 +1 @@
+Precise Description: Hands and Fingers Movements: The person is holding a vintage-style microphone with both hands. Throughout the video, the hands remain steady, gripping the microphone firmly. There are no significant changes in the position or movement of the hands. Body Positions and Posture: The person is standing upright, facing forward. The body remains relatively still, with minimal movement. Facial Expression and Emotion: The person appears to be singing or speaking into the microphone, as indicated by their open mouth and expressive eyes. The facial expression suggests engagement and passion. Clothing: The individual is wearing a sleeveless, sparkly dress that covers the upper torso. The dress has thin straps and a fitted design. Jewelry: The person is wearing long, dangling earrings and a pearl necklace. Brief Description: The person is standing in front of a dark background, holding a vintage-style microphone with both hands. They are dressed in a sleeveless, sparkly dress and accessorized with long earrings and a pearl necklace. The person appears to be singing or speaking into the microphone, with an engaged and passionate facial expression. The overall scene suggests a performance or recording setting.
\ No newline at end of file
diff --git a/libs/EchoMimicV3/echomimic_v3_src/__init__.py b/libs/EchoMimicV3/echomimic_v3_src/__init__.py
new file mode 100644
index 0000000..e69de29
diff --git a/libs/EchoMimicV3/echomimic_v3_src/dist/__init__.py b/libs/EchoMimicV3/echomimic_v3_src/dist/__init__.py
deleted file mode 100644
index a641a7b..0000000
--- a/libs/EchoMimicV3/echomimic_v3_src/dist/__init__.py
+++ /dev/null
@@ -1,62 +0,0 @@
-import torch
-import torch.distributed as dist
-
-from .fsdp import shard_model
-
-try:
- try:
- import pai_fuser
- from pai_fuser.core.distributed import (
- get_sequence_parallel_rank, get_sequence_parallel_world_size,
- get_sp_group, get_world_group, init_distributed_environment,
- initialize_model_parallel)
- from pai_fuser.core.long_ctx_attention import \
- xFuserLongContextAttention
- except Exception as ex:
- import xfuser
- from xfuser.core.distributed import (get_sequence_parallel_rank,
- get_sequence_parallel_world_size,
- get_sp_group, get_world_group,
- init_distributed_environment,
- initialize_model_parallel)
- from xfuser.core.long_ctx_attention import xFuserLongContextAttention
-except Exception as ex:
- get_sequence_parallel_world_size = None
- get_sequence_parallel_rank = None
- xFuserLongContextAttention = None
- get_sp_group = None
- get_world_group = None
- init_distributed_environment = None
- initialize_model_parallel = None
-
-try:
- from pai_fuser.core import parallel_magvit_vae
-except:
- def parallel_magvit_vae(multi_gpus_overlap_scale, spatial_compression_ratio):
- def decorator(func):
- def wrapper(self, z, *args, **kwargs):
- decoded = func(self, z, *args, **kwargs)
- return decoded
- return wrapper
- return decorator
-
-def set_multi_gpus_devices(ulysses_degree, ring_degree):
- if ulysses_degree > 1 or ring_degree > 1:
- if get_sp_group is None:
- raise RuntimeError("xfuser is not installed.")
- dist.init_process_group("nccl")
- print('parallel inference enabled: ulysses_degree=%d ring_degree=%d rank=%d world_size=%d' % (
- ulysses_degree, ring_degree, dist.get_rank(),
- dist.get_world_size()))
- assert dist.get_world_size() == ring_degree * ulysses_degree, \
- "number of GPUs(%d) should be equal to ring_degree * ulysses_degree." % dist.get_world_size()
- init_distributed_environment(rank=dist.get_rank(), world_size=dist.get_world_size())
- initialize_model_parallel(sequence_parallel_degree=dist.get_world_size(),
- ring_degree=ring_degree,
- ulysses_degree=ulysses_degree)
- # device = torch.device("cuda:%d" % dist.get_rank())
- device = torch.device(f"cuda:{get_world_group().local_rank}")
- print('rank=%d device=%s' % (get_world_group().rank, str(device)))
- else:
- device = "cuda"
- return device
diff --git a/libs/EchoMimicV3/echomimic_v3_src/dist/fsdp.py b/libs/EchoMimicV3/echomimic_v3_src/dist/fsdp.py
deleted file mode 100644
index 569621a..0000000
--- a/libs/EchoMimicV3/echomimic_v3_src/dist/fsdp.py
+++ /dev/null
@@ -1,42 +0,0 @@
-# Copyied from https://github.com/Wan-Video/Wan2.1/blob/main/wan/distributed/fsdp.py
-# Copyright 2024-2025 The Alibaba Wan Team Authors. All rights reserved.
-import gc
-from functools import partial
-
-import torch
-from torch.distributed.fsdp import FullyShardedDataParallel as FSDP
-from torch.distributed.fsdp import MixedPrecision, ShardingStrategy
-from torch.distributed.fsdp.wrap import lambda_auto_wrap_policy
-from torch.distributed.utils import _free_storage
-
-def shard_model(
- model,
- device_id,
- param_dtype=torch.bfloat16,
- reduce_dtype=torch.float32,
- buffer_dtype=torch.float32,
- process_group=None,
- sharding_strategy=ShardingStrategy.FULL_SHARD,
- sync_module_states=True,
-):
- model = FSDP(
- module=model,
- process_group=process_group,
- sharding_strategy=sharding_strategy,
- auto_wrap_policy=partial(
- lambda_auto_wrap_policy, lambda_fn=lambda m: m in model.blocks),
- mixed_precision=MixedPrecision(
- param_dtype=param_dtype,
- reduce_dtype=reduce_dtype,
- buffer_dtype=buffer_dtype),
- device_id=device_id,
- sync_module_states=sync_module_states)
- return model
-
-def free_model(model):
- for m in model.modules():
- if isinstance(m, FSDP):
- _free_storage(m._handle.flat_param.data)
- del model
- gc.collect()
- torch.cuda.empty_cache()
\ No newline at end of file
diff --git a/libs/EchoMimicV3/echomimic_v3_src/dist/wan_xfuser.py b/libs/EchoMimicV3/echomimic_v3_src/dist/wan_xfuser.py
deleted file mode 100644
index 00bfb2e..0000000
--- a/libs/EchoMimicV3/echomimic_v3_src/dist/wan_xfuser.py
+++ /dev/null
@@ -1,105 +0,0 @@
-import torch
-import torch.amp as amp
-
-from ..dist import (get_sequence_parallel_rank,
- get_sequence_parallel_world_size, get_sp_group,
- init_distributed_environment, initialize_model_parallel,
- xFuserLongContextAttention)
-
-
-def pad_freqs(original_tensor, target_len):
- seq_len, s1, s2 = original_tensor.shape
- pad_size = target_len - seq_len
- padding_tensor = torch.ones(
- pad_size,
- s1,
- s2,
- dtype=original_tensor.dtype,
- device=original_tensor.device)
- padded_tensor = torch.cat([original_tensor, padding_tensor], dim=0)
- return padded_tensor
-
-@amp.autocast('cuda', enabled=False)
-def rope_apply(x, grid_sizes, freqs):
- """
- x: [B, L, N, C].
- grid_sizes: [B, 3].
- freqs: [M, C // 2].
- """
- s, n, c = x.size(1), x.size(2), x.size(3) // 2
- # split freqs
- freqs = freqs.split([c - 2 * (c // 3), c // 3, c // 3], dim=1)
-
- # loop over samples
- output = []
- for i, (f, h, w) in enumerate(grid_sizes.tolist()):
- seq_len = f * h * w
-
- # precompute multipliers
- x_i = torch.view_as_complex(x[i, :s].to(torch.float32).reshape(
- s, n, -1, 2))
- freqs_i = torch.cat([
- freqs[0][:f].view(f, 1, 1, -1).expand(f, h, w, -1),
- freqs[1][:h].view(1, h, 1, -1).expand(f, h, w, -1),
- freqs[2][:w].view(1, 1, w, -1).expand(f, h, w, -1)
- ],
- dim=-1).reshape(seq_len, 1, -1)
-
- # apply rotary embedding
- sp_size = get_sequence_parallel_world_size()
- sp_rank = get_sequence_parallel_rank()
- freqs_i = pad_freqs(freqs_i, s * sp_size)
- s_per_rank = s
- freqs_i_rank = freqs_i[(sp_rank * s_per_rank):((sp_rank + 1) *
- s_per_rank), :, :]
- x_i = torch.view_as_real(x_i * freqs_i_rank).flatten(2)
- x_i = torch.cat([x_i, x[i, s:]])
-
- # append to collection
- output.append(x_i)
- return torch.stack(output)
-
-def usp_attn_forward(self,
- x,
- seq_lens,
- grid_sizes,
- freqs,
- dtype=torch.bfloat16):
- b, s, n, d = *x.shape[:2], self.num_heads, self.head_dim
- half_dtypes = (torch.float16, torch.bfloat16)
-
- def half(x):
- return x if x.dtype in half_dtypes else x.to(dtype)
-
- # query, key, value function
- def qkv_fn(x):
- q = self.norm_q(self.q(x)).view(b, s, n, d)
- k = self.norm_k(self.k(x)).view(b, s, n, d)
- v = self.v(x).view(b, s, n, d)
- return q, k, v
-
- q, k, v = qkv_fn(x)
- q = rope_apply(q, grid_sizes, freqs)
- k = rope_apply(k, grid_sizes, freqs)
-
- # TODO: We should use unpaded q,k,v for attention.
- # k_lens = seq_lens // get_sequence_parallel_world_size()
- # if k_lens is not None:
- # q = torch.cat([u[:l] for u, l in zip(q, k_lens)]).unsqueeze(0)
- # k = torch.cat([u[:l] for u, l in zip(k, k_lens)]).unsqueeze(0)
- # v = torch.cat([u[:l] for u, l in zip(v, k_lens)]).unsqueeze(0)
-
- x = xFuserLongContextAttention()(
- None,
- query=half(q),
- key=half(k),
- value=half(v),
- window_size=self.window_size)
-
- # TODO: padding after attention.
- # x = torch.cat([x, x.new_zeros(b, s - x.size(1), n, d)], dim=1)
-
- # output
- x = x.flatten(2)
- x = self.o(x)
- return x
\ No newline at end of file
diff --git a/node.zip b/node.zip
old mode 100755
new mode 100644
diff --git a/official_preset.pt b/official_preset.pt
old mode 100755
new mode 100644
diff --git a/package-lock.json b/package-lock.json
old mode 100755
new mode 100644
diff --git a/personalive/liveportrait/camera.py b/personalive/liveportrait/camera.py
old mode 100755
new mode 100644
diff --git a/personalive/liveportrait/convnextv2.py b/personalive/liveportrait/convnextv2.py
old mode 100755
new mode 100644
diff --git a/personalive/liveportrait/motion_extractor.py b/personalive/liveportrait/motion_extractor.py
old mode 100755
new mode 100644
diff --git a/personalive/liveportrait/util.py b/personalive/liveportrait/util.py
old mode 100755
new mode 100644
diff --git a/personalive/modeling/engine_model.py b/personalive/modeling/engine_model.py
old mode 100755
new mode 100644
diff --git a/personalive/modeling/framed_models.py b/personalive/modeling/framed_models.py
old mode 100755
new mode 100644
diff --git a/personalive/modeling/onnx_export.py b/personalive/modeling/onnx_export.py
old mode 100755
new mode 100644
diff --git a/personalive/models/attention.py b/personalive/models/attention.py
old mode 100755
new mode 100644
diff --git a/personalive/models/motion_encoder/FAN_feature_extractor.py b/personalive/models/motion_encoder/FAN_feature_extractor.py
old mode 100755
new mode 100644
diff --git a/personalive/models/motion_encoder/FAN_temporal_feature_extractor.py b/personalive/models/motion_encoder/FAN_temporal_feature_extractor.py
old mode 100755
new mode 100644
diff --git a/personalive/models/motion_encoder/encoder.py b/personalive/models/motion_encoder/encoder.py
old mode 100755
new mode 100644
diff --git a/personalive/models/motion_module.py b/personalive/models/motion_module.py
old mode 100755
new mode 100644
diff --git a/personalive/models/mutual_self_attention.py b/personalive/models/mutual_self_attention.py
old mode 100755
new mode 100644
diff --git a/personalive/models/pose_guider.py b/personalive/models/pose_guider.py
old mode 100755
new mode 100644
diff --git a/personalive/models/resnet.py b/personalive/models/resnet.py
old mode 100755
new mode 100644
diff --git a/personalive/models/transformer_2d.py b/personalive/models/transformer_2d.py
old mode 100755
new mode 100644
diff --git a/personalive/models/transformer_3d.py b/personalive/models/transformer_3d.py
old mode 100755
new mode 100644
diff --git a/personalive/models/unet_2d_blocks.py b/personalive/models/unet_2d_blocks.py
old mode 100755
new mode 100644
diff --git a/personalive/models/unet_2d_condition.py b/personalive/models/unet_2d_condition.py
old mode 100755
new mode 100644
diff --git a/personalive/models/unet_2d_decoder.py b/personalive/models/unet_2d_decoder.py
old mode 100755
new mode 100644
diff --git a/personalive/models/unet_3d.py b/personalive/models/unet_3d.py
old mode 100755
new mode 100644
diff --git a/personalive/models/unet_3d_blocks.py b/personalive/models/unet_3d_blocks.py
old mode 100755
new mode 100644
diff --git a/personalive/models/unet_3d_explicit_reference.py b/personalive/models/unet_3d_explicit_reference.py
old mode 100755
new mode 100644
diff --git a/personalive/pipelines/context.py b/personalive/pipelines/context.py
old mode 100755
new mode 100644
diff --git a/personalive/pipelines/pipeline_pose2vid.py b/personalive/pipelines/pipeline_pose2vid.py
old mode 100755
new mode 100644
diff --git a/personalive/pipelines/utils.py b/personalive/pipelines/utils.py
old mode 100755
new mode 100644
diff --git a/personalive/scheduler/scheduler_ddim.py b/personalive/scheduler/scheduler_ddim.py
old mode 100755
new mode 100644
diff --git a/personalive/utils/util.py b/personalive/utils/util.py
old mode 100755
new mode 100644
diff --git a/personalive/wrapper.py b/personalive/wrapper.py
old mode 100755
new mode 100644
diff --git a/personalive/wrapper_trt.py b/personalive/wrapper_trt.py
old mode 100755
new mode 100644
diff --git a/pyproject.toml b/pyproject.toml
old mode 100755
new mode 100644
index 8e6cc21..8d33d9d
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -1,29 +1,26 @@
[project]
name = "aiia"
-description = "The Ultimate AI Audio/Video toolkit for ComfyUI. Features an enhanced Ditto (with optimizations that outperform official demos and other SOTA talking head models in lip-sync accuracy and natural motion), EchoMimic V3 & FLOAT, VibeVoice, CosyVoice 3.0 & Qwen3-TTS (Zero-Shot Voice Cloning), Multi-Role Podcast Generation, and a powerful Media Browser."
-version = "1.11.3"
+description = "Advanced AI Audio/Video toolkit for ComfyUI. Features Multi-Role Podcast/Dialogue Generation, High-Fidelity Voice Cloning (CosyVoice/VibeVoice), TTS, Media Management, and efficient Video tools."
+version = "1.8.2"
license = {file = "LICENSE"}
readme = "README.md"
authors = [
- { name = "hawk", email = "hawk@aiia.ai" }
+ { name = "Hawk Lee", email = "l.y@live.cn" }
]
+keywords = ["ComfyUI", "Audio", "TTS", "Voice Cloning", "Podcast", "Dialogue", "CosyVoice", "VibeVoice", "Media Management", "Video"]
dependencies = [
- "torch",
+ "soundfile",
"numpy",
- "Pillow",
- "torchaudio",
- "tqdm",
- "scipy",
- "huggingface_hub",
- "opencv-python",
- "ffmpeg-python",
- "qwen-tts",
+ "matplotlib",
+ "torchaudio"
]
[project.urls]
-Homepage = "https://github.com/havvk/ComfyUI_AIIA"
-Repository = "https://github.com/havvk/ComfyUI_AIIA.git"
+Repository = "https://github.com/havvk/ComfyUI_AIIA"
+# Used by Comfy Registry https://registry.comfy.org
+
+[tool.comfy]
+PublisherId = "hawk"
+DisplayName = "ComfyUI_AIIA"
+Icon = "https://github.com/user-attachments/assets/358b9ca9-59c8-4433-b84c-c150503af04a"
-[build-system]
-requires = ["setuptools>=61.0"]
-build-backend = "setuptools.build_meta"
diff --git a/verify_podcast_nodes.py b/verify_podcast_nodes.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/__init__.py b/vibevoice_core/__init__.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/added_tokens.json b/vibevoice_core/added_tokens.json
old mode 100755
new mode 100644
diff --git a/vibevoice_core/generation_config.json b/vibevoice_core/generation_config.json
old mode 100755
new mode 100644
diff --git a/vibevoice_core/generation_config_0.5B.json b/vibevoice_core/generation_config_0.5B.json
old mode 100755
new mode 100644
diff --git a/vibevoice_core/generation_config_1.5B.json b/vibevoice_core/generation_config_1.5B.json
old mode 100755
new mode 100644
diff --git a/vibevoice_core/generation_config_7B.json b/vibevoice_core/generation_config_7B.json
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/__init__.py b/vibevoice_core/modular/__init__.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/configuration_vibevoice.py b/vibevoice_core/modular/configuration_vibevoice.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/configuration_vibevoice_streaming.py b/vibevoice_core/modular/configuration_vibevoice_streaming.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/modeling_vibevoice.py b/vibevoice_core/modular/modeling_vibevoice.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/modeling_vibevoice_inference.py b/vibevoice_core/modular/modeling_vibevoice_inference.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/modeling_vibevoice_streaming.py b/vibevoice_core/modular/modeling_vibevoice_streaming.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/modeling_vibevoice_streaming_inference.py b/vibevoice_core/modular/modeling_vibevoice_streaming_inference.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/modular_vibevoice_diffusion_head.py b/vibevoice_core/modular/modular_vibevoice_diffusion_head.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/modular_vibevoice_text_tokenizer.py b/vibevoice_core/modular/modular_vibevoice_text_tokenizer.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/modular_vibevoice_tokenizer.py b/vibevoice_core/modular/modular_vibevoice_tokenizer.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/modular/streamer.py b/vibevoice_core/modular/streamer.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/schedule/__init__.py b/vibevoice_core/schedule/__init__.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/schedule/dpm_solver.py b/vibevoice_core/schedule/dpm_solver.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/schedule/timestep_sampler.py b/vibevoice_core/schedule/timestep_sampler.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/special_tokens_map.json b/vibevoice_core/special_tokens_map.json
old mode 100755
new mode 100644
diff --git a/vibevoice_core/vibevoice_processor.py b/vibevoice_core/vibevoice_processor.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/vibevoice_streaming_processor.py b/vibevoice_core/vibevoice_streaming_processor.py
old mode 100755
new mode 100644
diff --git a/vibevoice_core/vibevoice_tokenizer_processor.py b/vibevoice_core/vibevoice_tokenizer_processor.py
old mode 100755
new mode 100644
diff --git a/video_formats_aiia/h264.json b/video_formats_aiia/h264.json
old mode 100755
new mode 100644
diff --git a/video_formats_aiia/h265.json b/video_formats_aiia/h265.json
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/__init__.py b/voxcpm_core/voxcpm/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/cli.py b/voxcpm_core/voxcpm/cli.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/core.py b/voxcpm_core/voxcpm/core.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/model/__init__.py b/voxcpm_core/voxcpm/model/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/model/utils.py b/voxcpm_core/voxcpm/model/utils.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/model/voxcpm.py b/voxcpm_core/voxcpm/model/voxcpm.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/__init__.py b/voxcpm_core/voxcpm/modules/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/audiovae/__init__.py b/voxcpm_core/voxcpm/modules/audiovae/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/audiovae/audio_vae.py b/voxcpm_core/voxcpm/modules/audiovae/audio_vae.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/layers/__init__.py b/voxcpm_core/voxcpm/modules/layers/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/layers/lora.py b/voxcpm_core/voxcpm/modules/layers/lora.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/layers/scalar_quantization_layer.py b/voxcpm_core/voxcpm/modules/layers/scalar_quantization_layer.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/locdit/__init__.py b/voxcpm_core/voxcpm/modules/locdit/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/locdit/local_dit.py b/voxcpm_core/voxcpm/modules/locdit/local_dit.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/locdit/unified_cfm.py b/voxcpm_core/voxcpm/modules/locdit/unified_cfm.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/locenc/__init__.py b/voxcpm_core/voxcpm/modules/locenc/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/locenc/local_encoder.py b/voxcpm_core/voxcpm/modules/locenc/local_encoder.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/minicpm4/__init__.py b/voxcpm_core/voxcpm/modules/minicpm4/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/minicpm4/cache.py b/voxcpm_core/voxcpm/modules/minicpm4/cache.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/minicpm4/config.py b/voxcpm_core/voxcpm/modules/minicpm4/config.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/modules/minicpm4/model.py b/voxcpm_core/voxcpm/modules/minicpm4/model.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/training/__init__.py b/voxcpm_core/voxcpm/training/__init__.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/training/accelerator.py b/voxcpm_core/voxcpm/training/accelerator.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/training/config.py b/voxcpm_core/voxcpm/training/config.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/training/data.py b/voxcpm_core/voxcpm/training/data.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/training/packers.py b/voxcpm_core/voxcpm/training/packers.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/training/state.py b/voxcpm_core/voxcpm/training/state.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/training/tracker.py b/voxcpm_core/voxcpm/training/tracker.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/utils/text_normalize.py b/voxcpm_core/voxcpm/utils/text_normalize.py
old mode 100755
new mode 100644
diff --git a/voxcpm_core/voxcpm/zipenhancer.py b/voxcpm_core/voxcpm/zipenhancer.py
old mode 100755
new mode 100644