Commit Graph
1055 Commits
Author SHA1 Message Date
Hawk Lee 9f046e37d4 fix: 直接从 sortformer_diar_models 导入,绕过 aed_multitask_models 卡死 2026-02-19 14:44:30 +08:00
Hawk Lee c8dcc4d9e2 debug: import hook 追踪 NeMo 导入卡死 2026-02-19 14:34:41 +08:00
Hawk Lee dc31f01b0b debug: 检查 modelscope 和 from_pretrained 状态 2026-02-19 14:03:01 +08:00
Hawk Lee a6ee5569f3 debug: 细化 NeMo 导入诊断 2026-02-19 13:54:11 +08:00
Hawk Lee cf705edaa6 debug: 临时诊断 NeMo restore_from 卡住位置 2026-02-19 13:45:30 +08:00
Hawk Lee 29d64b64d3 fix: 删除 from_pretrained 保存/恢复 hack,修复 classmethod 描述符损坏
保存 bound classmethod 再恢复会破坏描述符:所有子类调用 from_pretrained 时
cls 参数都变成 PreTrainedModel 而不是实际子类。这导致 VibeVoice 创建出空模型
(0个module、空 device_map)并触发 accelerate IndexError。

modelscope 导入已改为 from transformers,不再需要这个 hack。
2026-02-19 13:37:11 +08:00
Hawk Lee 133daf8eee debug: 临时诊断 dispatch_model 空 device_map 2026-02-19 13:19:47 +08:00
Hawk Lee 8119b6cb78 fix: 补丁改为永久应用,不再还原
原来的 _active_transformers_patches 是 context manager,在 finally 中用 delattr
删除补丁。但 IndexTTS 导入链中的模块已经缓存在 sys.modules 里,引用了这些属性。
删除它们会让已缓存的模块持有悬空引用,导致后续加载 VibeVoice 等模型时崩溃。

这些补丁都是添加性的兼容 shim(只添加缺失的属性,不覆盖已有行为),
保留它们不会对其他节点产生副作用。
2026-02-19 12:57:08 +08:00
Hawk Lee f7e96d37ce fix: 用 transformers 替代 modelscope 导入,从根源消除副作用
- infer_v2.py: from modelscope → from transformers import AutoModelForCausalLM
  模型已在本地,不需要 modelscope 的 hub 下载包装
  这样 modelscope 根本不会被 import,消除所有 monkey-patching 副作用
- aiia_vibevoice_nodes.py: 还原 device_map=auto(此文件无需修改)
2026-02-19 12:51:36 +08:00
Hawk Lee f7a549022c fix: VibeVoice device_map=auto 导致 accelerate IndexError
- 与 QwenEmotion 相同的问题,accelerate dispatch_model 获得空 device_map
- 改为手动 .to(device),与 IndexTTS2 加载方式一致
2026-02-19 12:45:03 +08:00
Hawk Lee 8b3b73c2ed fix: QwenEmotion device_map=auto 与 modelscope 不兼容导致 IndexError
- 移除 device_map='auto',改为手动 .to(device)
- 与 IndexTTS2 中其他模型(GPT/BigVGAN/CamPPlus)加载方式一致
- 修复 use_cuda_kernel=False 时的加载失败
2026-02-19 12:39:41 +08:00
Hawk Lee 64a05e4b44 fix: 修复 context manager 双 yield 和 from_pretrained 恢复逻辑
- _active_transformers_patches except 块中的 yield 改为 raise,修复
  'generator didn't stop after throw()' 异常
- from_pretrained 保存/恢复改用 try/finally,确保即使加载失败也能恢复
2026-02-19 12:35:11 +08:00
Hawk Lee 579da6b3e4 fix: 启用所有节点模块导入 2026-02-19 12:08:47 +08:00
Hawk Lee 1cb441eaa9 fix: 修复 IndexTTS 导入导致 NeMo 分段节点卡死的问题
- 根因: indextts 的 infer_v2.py 顶层导入 modelscope,会全局 monkey-patch
  transformers.PreTrainedModel.from_pretrained,导致 NeMo 模型加载卡死
- 修复: 在 aiia_indextts_nodes.py 中保存/恢复 from_pretrained 原始引用
- 清理 libs/ 目录:删除各依赖的文档、示例、测试等非运行时文件
- 新增 libs/index-tts (仅保留 indextts/ 核心包和 LICENSE)
2026-02-19 12:05:07 +08:00
Hawk Lee 1246441500 v1.14.3: Podcast Stitcher - use VAD to cap FA cut_end, prevent tail extending into silence 2026-02-18 00:04:09 +08:00
Hawk Lee 8b556e8435 v1.14.2: Fix NeMo diarization for PyTorch 2.10+, add JSON Extractor/Builder nodes, Qwen3-TTS voice presets & robustness fixes 2026-02-17 23:10:27 +08:00
Hawk Lee fb0c84bfae v1.14.1: Add silence gap between TTS segments & fix Qwen3 TTS stitching
- CosyVoice: replace overlapping crossfade with fade-out + 100ms silence
  gap + fade-in for easier downstream splitting
- Qwen3 TTS: fix bare torch.cat in multi-emotion segment stitching,
  apply same fade-out + silence + fade-in approach
2026-02-17 16:13:08 +08:00
Hawk Lee f95f2a9832 v1.14.0: Fix CosyVoice transformers 4.57+ compatibility & emotion segment crossfade
- Qwen2Encoder: bypass incompatible Qwen2Model.forward() in transformers ≥4.53
  via manual layer iteration + SDPA-compatible mask + POST-norm output
- Add cosine crossfade (50ms) between multi-emotion TTS segments
- Update README compatibility table to reflect full transformers support
2026-02-17 15:49:12 +08:00
Hawk Lee e21a5c9eba fix: restore file permissions to 755 (reverts eff537d mode change)
The eff537d commit changed all files from 755 to 644.
Restores original executable permissions.
2026-02-16 12:15:32 +08:00
Hawk Lee e27987cad6 fix: add llm.base.pt fallback for V3 model switching
The LLM model switch logic only knew about llm.orig.pt as
non-RL fallback. V3 models (Fun-CosyVoice3-0.5B-2512) ship
with llm.base.pt instead of llm.orig.pt.

When use_rl_model=False, the code silently skipped the switch
because llm.orig.pt didn't exist, keeping llm.rl.pt active.
llm.rl.pt is incompatible with seed fallback audio, causing
the max_trials sampling crash.

Now fallback order: llm.orig.pt → llm.base.pt → no-op
2026-02-16 11:49:50 +08:00
Hawk Lee 0435d5c932 fix: restore CosyVoice Chinese UI defaults and revert V3 Base special case
Root cause: UI defaults were changed from Chinese to English,
and V3 Base got unnecessary special-casing. CosyVoice V3 0.5B
actually supports instruct2 and worked fine with Chinese defaults.

Restored:
- Chinese default tts_text/instruct_text/tooltips
- Chinese dialect and emotion option labels
- Original unified instruct formatting (no V3 Base special case)
2026-02-16 11:38:17 +08:00
Hawk Lee 012c4ad7e9 fix: revert V3 Base to inference_instruct2 with minimal instruct
The previous fix (routing V3 Base to inference_zero_shot) broke
CosyVoice V3 because ALL V3 paths require <|endofprompt|> token.

New approach:
- V3 Base: use inference_instruct2 with minimal '<|endofprompt|>'
  (no style instructions that cause sampling failures)
- V3 Instruct/SFT: full instruct formatting as before
- Both paths use the unified inference_instruct2 API
2026-02-16 11:28:35 +08:00
Hawk Lee 6a94b32b1d fix: route CosyVoice V3 Base models to zero-shot path
V3 0.5B Base (Fun-CosyVoice3-0.5B-2512) does NOT support instruct.
Previously all V3 models were forced through inference_instruct2,
causing LLM sampling failures (max_trials EOS error).

Now V3/V2 Base models:
- Skip instruct formatting entirely
- Use inference_zero_shot (with ref audio) instead of instruct2
- Use inference_sft (with speaker ID) for identity path
2026-02-16 11:17:06 +08:00
Hawk Lee 23d9b87ab0 chore: bump version to 1.13.0
New features:
- AIIA Emotion Annotator (LLM-driven emotion tagging)
- AIIA Text Splitter (single-speaker text splitting)
- Emotion tag handling for all TTS engines
- VoiceDesign UI priority fix
2026-02-16 11:06:23 +08:00
Hawk Lee 40bc5627cb feat: add AIIA Text Splitter node (v1.13.0)
New node that splits single-speaker text into dialogue_json format:
- 3 split modes: auto, by_sentence, by_line
- Short sentence merging (min_chars)
- Long sentence splitting at commas/semicolons (max_chars)
- Output compatible with Emotion Annotator and Dialogue TTS

Workflow: Text → Text Splitter → Emotion Annotator → TTS
2026-02-16 11:03:31 +08:00
Hawk Lee 449183c800 fix: VoiceDesign uses UI design description with priority
User's UI instruct (design description) takes priority.
Emotion from upstream tags only used as fallback when
design description is empty.
2026-02-16 10:43:00 +08:00
Hawk Lee 3ef9fd4601 docs: clarify Qwen3 model type emotion support and UI priority
- Split Qwen3-TTS into CustomVoice (supports instruct) and
  Base (strips tags, doesn't support instruct) in compat table
- Add note: UI dropdown emotion takes priority over inline tags
2026-02-16 10:35:47 +08:00
Hawk Lee e01e7c3eff fix: UI emotion dropdown takes priority over inline tags
When user explicitly sets emotion via UI, inline [Emotion] tags
are stripped from text but NOT merged into instruct.
2026-02-16 10:31:43 +08:00
Hawk Lee cf5c7aa78c fix: extract and strip inline emotion tags in standalone Qwen3 TTS
When text contains [Happy]/[Calm] etc (from Splitter output),
extract the emotion into instruct for proper Qwen3 handling,
then strip the tag from text to prevent reading it aloud.
2026-02-16 10:30:51 +08:00
Hawk Lee 87077453f3 fix: include emotion in Dialogue TTS Qwen3 batch grouping key
The grouping key only used (model_id, dialect), so sentences with
different emotions could be batched together. Now includes the
instruct string (which contains emotion) in the grouping key.
2026-02-16 10:27:21 +08:00
Hawk Lee 1f76fae615 docs: clarify VibeVoice auto-strips emotion tags
Updated compatibility table and added note about VibeVoice
Standard/Realtime actively stripping tags during preprocessing.
2026-02-16 10:15:47 +08:00
Hawk Lee 8d374cbf0c docs: promote Emotion Annotator to top-level section 4.2
Previously was 4.1.1 under Script Parser. Now a peer section.
Renumbered all subsequent sections sequentially (4.3-4.11).
2026-02-16 10:14:28 +08:00
Hawk Lee 7f1537b308 fix: strip emotion tags in VibeVoice Realtime node too
Same fix as standard VibeVoice node - strip [Emotion] tags
to prevent reading them aloud.
2026-02-16 10:12:44 +08:00
Hawk Lee 75f11ce028 fix: strip emotion tags before VibeVoice generation
VibeVoice doesn't support [Emotion] tags and would read them
as literal text. Now strips all 24 known emotion tags from
input text during preprocessing.
2026-02-16 10:11:05 +08:00
Hawk Lee 8dae239185 feat: Splitter embeds emotion tags into output text
When Emotion Annotator has tagged dialogue with emotions,
Splitter now outputs '[Happy] text...' format in speaker_A_text
and speaker_B_text, so downstream TTS nodes can consume them.
Also passes emotion field through in split_map.
2026-02-16 10:09:25 +08:00
Hawk Lee 27131510b9 docs: add AIIA Emotion Annotator documentation to README
- Add section 4.1.1 with full parameter table, workflow diagrams,
  24 emotion tags list, and per-engine compatibility table
- Explain pipeline position (Script Parser → Annotator → TTS)
- Document Qwen3 smart batching behavior
2026-02-16 10:02:26 +08:00
Hawk Lee aea2e9b16a fix: split Qwen3 batches on emotion change
Previously all emotions in a batch were merged into one instruct
string (e.g. 'Happy,Calm。'), which is semantically wrong.
Now emotion is included in param_hash, so batches split when
per-sentence emotion changes. Consecutive same-emotion sentences
still batch together for efficiency.
2026-02-16 01:52:18 +08:00
Hawk Lee 69a2d799a1 fix: normalize line index to int and show neutral annotations
- LLM may return line index as string, now cast to int
- Show neutral annotations in log (was hidden before)
- Add raw LLM response debug output
- Track index out-of-bounds warnings
2026-02-16 01:44:31 +08:00
Hawk Lee dc31876404 fix: add User-Agent header to bypass Cloudflare 1010 block
Cloudflare blocks Python's default User-Agent (Python-urllib/3.x)
with error 1010. Using a browser-like UA fixes this.
2026-02-16 01:41:06 +08:00
Hawk Lee 19ea3ec847 fix: add IS_CHANGED to prevent ComfyUI caching
ComfyUI caches node output when inputs are unchanged. Since
emotion annotation depends on external LLM API, we need to
force re-execution each time.
2026-02-16 01:39:30 +08:00
Hawk Lee 68d71bb728 fix: complete EchoMimicV3 dist stubs for single-GPU inference
Add all missing distributed utility stubs:
- get_sequence_parallel_rank (returns 0)
- get_sequence_parallel_world_size (returns 1)
- get_sp_group (returns None)
- set_multi_gpus_devices (no-op)
- xFuserLongContextAttention (stub class)
- wan_xfuser.usp_attn_forward (stub)
2026-02-16 01:37:15 +08:00
Hawk Lee e6e17aec81 fix: add proxy_url input + EchoMimicV3 dist stub
- Replace ~/run.sh reading with proxy_url node input parameter
- Proxy priority: node param > HTTPS_PROXY env var
- Add no-op parallel_magvit_vae stub for single-GPU inference
- Remove environment-specific ~/run.sh dependency
2026-02-16 01:34:21 +08:00
Hawk Lee 7ceb05437d fix: read proxy settings from ~/run.sh when env vars not set
ComfyUI process doesn't source ~/run.sh, so proxy env vars are
missing. Now reads http_proxy/https_proxy from ~/run.sh as fallback.
2026-02-16 01:32:28 +08:00
Hawk Lee 797c477889 docs: mark LLM emotion annotation as completed 2026-02-16 01:25:19 +08:00
Hawk Lee a74a8b20bb feat: add AIIA Emotion Annotator node (LLM-driven)
- New node: 🎭 AIIA Emotion Annotator (LLM)
- Supports Groq (free), Ollama, vLLM via OpenAI-compatible API
- API key from GROQ_API_KEY env var or node parameter
- Custom base URL for local LLM services
- skip_existing / overwrite_all modes
- Robust JSON parsing with markdown code block handling
- Proxy support from environment variables
2026-02-16 01:24:36 +08:00
Hawk Lee 4a7132b82f docs: add manual download instructions for MMS FA model
Two methods: wget from Facebook CDN, or huggingface-cli download.
2026-02-16 00:40:18 +08:00
Hawk Lee d7476f4d42 feat(fa): auto-copy downloaded model to models/mms_fa/ for centralized management
After torchaudio downloads the MMS_FA model to hub cache, automatically
copy it to models/mms_fa/model.pt so future loads use the local copy.
2026-02-16 00:39:21 +08:00
Hawk Lee 1c25316f0d release: v1.12.0 - MMS Forced Alignment integration
- Mark VAD and FA tasks as completed in TODO.md
- Add v1.12.0 changelog entry in README.md
- Bump version to 1.12.0 in pyproject.toml
2026-02-16 00:35:12 +08:00
Hawk Lee 193164f91a fix(fa): add floor guarantee to cut_end clamp - never shorter than FA end
max(fa_end, midpoint) ensures that even if FA reports slightly
overlapping timestamps, we never truncate the current sentence
below its own FA endpoint.
2026-02-16 00:21:55 +08:00
Hawk Lee f0b741a6ab fix(fa): clamp cut_end to prevent overlap with next sentence's FA start
When energy extends cut_end past the next sentence's FA start point,
clamp to midpoint between current FA end and next FA start. Fixes
issue where A[1] tail extension was eating into A[2]'s first char.
2026-02-16 00:19:39 +08:00