diff --git a/README.md b/README.md index 0185347..0be6e73 100755 --- a/README.md +++ b/README.md @@ -1179,11 +1179,16 @@ Script Parser → Emotion Annotator → Dialogue TTS / Qwen Dialogue TTS ##### 典型工作流 -**基础流程**(适合大多数场景): +**对话流程**(适合大多数场景): ``` Script Parser → Emotion Annotator → Dialogue TTS → Video Combine ``` +**单人旁白流程**(配合 Text Splitter): +``` +长文本 → Text Splitter → Emotion Annotator → Qwen3 TTS / CosyVoice +``` + **高级拆分流程**(多引擎混合): ``` Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) @@ -1191,7 +1196,33 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) → Podcast Stitcher ``` -#### 4.3 AIIA Dialogue TTS (对话生成引擎) +#### 4.3 AIIA Text Splitter (文本拆分器) + +**[v1.13.0 New]** 将单人长文本按标点拆分为标准 `dialogue_json`,使其可直接接入 Emotion Annotator → TTS 管线。 + +##### 参数说明 + +| 参数 | 说明 | +|------|------| +| `text` | 待拆分文本(多行/长段落) | +| `speaker_name` | 说话人名称,默认 `Narrator` | +| `split_mode` | `auto`:智能拆分(句末标点 + 短句合并 + 长句再拆);`by_sentence`:仅按句号/问号/感叹号拆分;`by_line`:按换行拆分 | +| `min_chars` | 最小字符数,短于此的句子合并到前一句。默认 `4` | +| `max_chars` | 最大字符数,超长句子在逗号/分号处强制拆分。默认 `100` | + +##### 拆分行为 (`auto` 模式) + +1. **按换行分段** → 保留段落结构 +2. **段内按句末标点拆分** → `。!?!?…` 和省略号 `……` / `...` +3. **短句合并** → 短于 `min_chars` 的句子追加到前一句 +4. **长句拆分** → 超过 `max_chars` 的句子在逗号/分号处再切 + +##### 输出 + +- `dialogue_json`:标准格式,与 Script Parser 输出兼容,可直接接入 Emotion Annotator / Dialogue TTS +- `sentence_count`:拆分后的句子数 + +#### 4.4 AIIA Dialogue TTS (对话生成引擎) 核心调度与生成节点,支持自动角色切换和长音频拼接。 @@ -1214,7 +1245,7 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) - **Emotion Safeguard (New!)**: - **智能检测**: 系统会自动嗅探加载模型的元数据。如果你使用 CosyVoice SFT/Base 或 VibeVoice 等不支持 `Instruct` 功能的模型,系统将自动跳过 `[Emotion]` 标签插入,防止模型读出方括号。 -#### 4.4 AIIA Qwen Dialogue TTS (Qwen 旗舰对话节点) +#### 4.5 AIIA Qwen Dialogue TTS (Qwen 旗舰对话节点) **[v1.11.0 New]** 深度集成 Qwen3-TTS 的多模式特性,支持复杂的混合角色场景。 @@ -1235,7 +1266,7 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) - **Ref Audio**: 当模式为 Clone 时,连接参考音频。 - **特点**: 相对于通用对话节点,此节点能根据每个人的模式自动路由到最合适的 Qwen 引擎,且支持在 UI 直接输入设计描述。 -#### 4.5 AIIA Subtitle Gen (字幕生成器) +#### 4.6 AIIA Subtitle Gen (字幕生成器) **[v1.7.0 New]** 无需 STT,直接从生成过程中提取精准时间轴。 @@ -1250,7 +1281,7 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) - **VibeVoice**: 使用**智能插值算法 (Smart Interpolation)**,根据字符长度自动计算长音频段内的单句时间轴。 - **Qwen3-TTS**: 基于生成的音频振幅精准断句,支持多角色时间轴导出。 -#### 4.6 AIIA Subtitle to Segments (字幕转分段) +#### 4.7 AIIA Subtitle to Segments (字幕转分段) **[v1.10.3 New]** 将现有的 SRT/ASS 字幕文件转换为 `segments_info` 格式,以便进行时间轴重新校准。 @@ -1261,7 +1292,7 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) - `segments_info`: 标准化的 JSON 字符串,可直接输入到 `AIIA Subtitle Gen`。 - **用途**: 结合 `Subtitle Gen` 的 `calibration_info` 输入,可以将**旧的、不准的字幕**自动对齐到**新的、精准的音轨**上。 -#### 4.7 AIIA Subtitle Preview (字幕预览) +#### 4.8 AIIA Subtitle Preview (字幕预览) **[v1.7.1 New]** 实时校验音画同步效果。 @@ -1272,7 +1303,7 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) - **交互式界面**: 提供 Web 播放器,按时间轴滚动显示字幕。 - **ASS 样式渲染**: 尝试还原 ASS 字幕的字体颜色、大小和描边效果。 -#### 4.8 Interactive Teaching (Web Export) (互动式教学导出) +#### 4.9 Interactive Teaching (Web Export) (互动式教学导出) **[v1.8.1 New]** 将播客升级为视听同步的互动网页。支持“读写分离”的缓存优化,修改 Visual 标签无需重跑 TTS。 @@ -1290,7 +1321,7 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) - 支持绝对 URL: `(Visual: https://example.com)` - 支持相对路径: `(Visual: ./slides/01.jpg)` (相对于导出 HTML 的位置) -#### 4.9 AIIA Podcast Splitter (对话拆分器) +#### 4.10 AIIA Podcast Splitter (对话拆分器) **[v1.12.0 New]** 将对话 JSON 按说话人拆分为独立文本列表,用于"**拆分→生成→拼接**"的高级流程。 @@ -1301,7 +1332,7 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) - `text_A`, `text_B`: 分别为说话人 A、B 的纯文本列表(每行一句),可直接送入各自的 TTS 节点独立生成。 - **用途**: 实现对不同说话人使用不同 TTS 引擎/参数生成音频,再通过 Stitcher 精确拼接的高级工作流。 -#### 4.10 AIIA ASR Node (语音识别) +#### 4.11 AIIA ASR Node (语音识别) **[v1.12.0 New]** 基于 FunASR 的语音识别节点,输出带词级时间戳的识别结果,为 Stitcher 提供精确对齐依据。 @@ -1309,7 +1340,7 @@ Script Parser → Emotion Annotator → Podcast Splitter → TTS_A (CosyVoice) - **Output**: `ASR_RESULT`(包含词级时间戳的识别结果)。 - **依赖**: 需要安装 `funasr` 库。模型首次运行时自动下载。 -#### 4.11 AIIA Podcast Stitcher (精确拼接器) +#### 4.12 AIIA Podcast Stitcher (精确拼接器) **[v1.12.0 New]** 将分轨生成的多角色音频按原始对话顺序精确拼接,还原自然对话节奏。 diff --git a/__init__.py b/__init__.py index 5af82c1..e3b6718 100755 --- a/__init__.py +++ b/__init__.py @@ -151,6 +151,9 @@ else: # 30. 处理 aiia_emotion_annotator.py (LLM 情感标注) _load_nodes_from_module(".aiia_emotion_annotator", "aiia_emotion_annotator") + # 31. 处理 aiia_text_splitter.py (单人文本拆分) + _load_nodes_from_module(".aiia_text_splitter", "aiia_text_splitter") + # 告诉 ComfyUI 这个节点包有一个包含网页资源的 'js' 目录 WEB_DIRECTORY = "js" diff --git a/aiia_text_splitter.py b/aiia_text_splitter.py new file mode 100644 index 0000000..4d79453 --- /dev/null +++ b/aiia_text_splitter.py @@ -0,0 +1,232 @@ +""" +AIIA Text Splitter — 单人文本按标点拆分为 dialogue_json +[v1.13.0 New] + +将长文本按句号、问号、感叹号等标点拆分为标准 dialogue_json 格式, +可直接接入 Emotion Annotator → TTS 管线。 + +支持短句合并(避免碎片)和长句拆分(避免 TTS 单句过长)。 +""" + +import json +import re + + +class AIIA_Text_Splitter: + """ + 将纯文本按标点拆分为 dialogue_json 格式。 + + 支持三种拆分模式: + - auto: 中英文自动,按句末标点拆分 + 短句合并 + 长句拆分 + - by_sentence: 仅按句号/问号/感叹号拆分 + - by_line: 按换行拆分 + """ + + NODE_NAME = "AIIA Text Splitter" + + @classmethod + def INPUT_TYPES(cls): + return { + "required": { + "text": ("STRING", { + "multiline": True, + "default": "", + "tooltip": "待拆分的文本。支持多段落、多行。" + }), + "speaker_name": ("STRING", { + "default": "Narrator", + "tooltip": "说话人名称,写入 dialogue_json 的 speaker 字段" + }), + "split_mode": (["auto", "by_sentence", "by_line"], { + "default": "auto", + "tooltip": "拆分模式:\n" + " auto: 按句末标点拆分 + 短句合并 + 长句再拆\n" + " by_sentence: 仅按句号/问号/感叹号拆分\n" + " by_line: 按换行拆分" + }), + }, + "optional": { + "min_chars": ("INT", { + "default": 4, + "min": 1, + "max": 50, + "tooltip": "最小字符数。短于此的句子合并到前一句。" + }), + "max_chars": ("INT", { + "default": 100, + "min": 20, + "max": 500, + "tooltip": "最大字符数。超长句子在逗号/分号处强制拆分。" + }), + } + } + + RETURN_TYPES = ("STRING", "INT") + RETURN_NAMES = ("dialogue_json", "sentence_count") + FUNCTION = "split_text" + CATEGORY = "AIIA/Podcast" + + def split_text(self, text, speaker_name="Narrator", split_mode="auto", + min_chars=4, max_chars=100): + """拆分文本为 dialogue_json 格式。""" + + if not text or not text.strip(): + empty = json.dumps([], ensure_ascii=False) + return (empty, 0) + + text = text.strip() + + if split_mode == "by_line": + raw_sentences = self._split_by_line(text) + elif split_mode == "by_sentence": + raw_sentences = self._split_by_sentence(text) + else: # auto + raw_sentences = self._split_auto(text, min_chars, max_chars) + + # 构建 dialogue_json + dialogue = [] + for sent in raw_sentences: + sent = sent.strip() + if not sent: + continue + dialogue.append({ + "type": "speech", + "speaker": speaker_name, + "text": sent, + "emotion": None + }) + + result = json.dumps(dialogue, ensure_ascii=False, indent=2) + return (result, len(dialogue)) + + def _split_by_line(self, text): + """按换行拆分,空行跳过。""" + return [line.strip() for line in text.split("\n") if line.strip()] + + def _split_by_sentence(self, text): + """仅按句末标点拆分(句号、问号、感叹号、省略号)。""" + # 先按换行分段,再段内按标点拆分 + paragraphs = [p.strip() for p in text.split("\n") if p.strip()] + sentences = [] + for para in paragraphs: + parts = self._split_at_sentence_end(para) + sentences.extend(parts) + return sentences + + def _split_auto(self, text, min_chars, max_chars): + """ + 智能拆分: + 1. 按换行分段 + 2. 段内按句末标点拆分 + 3. 短句合并 + 4. 长句在逗号/分号处再拆 + """ + paragraphs = [p.strip() for p in text.split("\n") if p.strip()] + all_sentences = [] + + for para in paragraphs: + # Step 1: 按句末标点拆分 + raw = self._split_at_sentence_end(para) + + # Step 2: 短句合并 + merged = self._merge_short(raw, min_chars) + + # Step 3: 长句拆分 + final = [] + for sent in merged: + if len(sent) > max_chars: + final.extend(self._split_long(sent, max_chars)) + else: + final.append(sent) + + all_sentences.extend(final) + + return all_sentences + + def _split_at_sentence_end(self, text): + """ + 在句末标点处拆分,保留标点在前一句末尾。 + 支持:。!?!? 以及省略号 …… / ... + """ + # 按句末标点拆分,保留分隔符 + # 匹配:句号/问号/感叹号(中英文),以及省略号 + parts = re.split(r'((?:\.{3}|…{1,2}|[。!?!?]))', text) + + sentences = [] + buffer = "" + for i, part in enumerate(parts): + if i % 2 == 0: + # 正文部分 + buffer += part + else: + # 标点部分,附加到 buffer + buffer += part + if buffer.strip(): + sentences.append(buffer.strip()) + buffer = "" + + # 处理末尾没有标点的残余 + if buffer.strip(): + sentences.append(buffer.strip()) + + return sentences + + def _merge_short(self, sentences, min_chars): + """将短于 min_chars 的句子合并到前一句。""" + if not sentences: + return sentences + + merged = [sentences[0]] + for sent in sentences[1:]: + if len(sent) < min_chars and merged: + # 合并到前一句 + merged[-1] = merged[-1] + sent + else: + merged.append(sent) + + return merged + + def _split_long(self, text, max_chars): + """ + 将超长句子在逗号/分号/顿号处拆分。 + 尽量靠近 max_chars 的位置切,避免太碎。 + """ + # 可选拆分点:逗号、分号、顿号(中英文) + split_points = [] + for i, ch in enumerate(text): + if ch in ',,;;、': + split_points.append(i) + + if not split_points: + # 没有可拆分点,原样返回 + return [text] + + result = [] + start = 0 + for point in split_points: + segment = text[start:point + 1] + if len(segment) >= max_chars // 2: + # 够长了,切一刀 + result.append(segment.strip()) + start = point + 1 + + # 追加剩余部分 + remainder = text[start:].strip() + if remainder: + if result and len(remainder) < max_chars // 4: + # 残余太短,合并到最后一段 + result[-1] = result[-1] + remainder + else: + result.append(remainder) + + return result if result else [text] + + +# === ComfyUI Registration === +NODE_CLASS_MAPPINGS = { + "AIIA_Text_Splitter": AIIA_Text_Splitter, +} + +NODE_DISPLAY_NAME_MAPPINGS = { + "AIIA_Text_Splitter": "💬 AIIA Text Splitter", +}