The LLM model switch logic only knew about llm.orig.pt as
non-RL fallback. V3 models (Fun-CosyVoice3-0.5B-2512) ship
with llm.base.pt instead of llm.orig.pt.
When use_rl_model=False, the code silently skipped the switch
because llm.orig.pt didn't exist, keeping llm.rl.pt active.
llm.rl.pt is incompatible with seed fallback audio, causing
the max_trials sampling crash.
Now fallback order: llm.orig.pt → llm.base.pt → no-op
Root cause: UI defaults were changed from Chinese to English,
and V3 Base got unnecessary special-casing. CosyVoice V3 0.5B
actually supports instruct2 and worked fine with Chinese defaults.
Restored:
- Chinese default tts_text/instruct_text/tooltips
- Chinese dialect and emotion option labels
- Original unified instruct formatting (no V3 Base special case)
The previous fix (routing V3 Base to inference_zero_shot) broke
CosyVoice V3 because ALL V3 paths require <|endofprompt|> token.
New approach:
- V3 Base: use inference_instruct2 with minimal '<|endofprompt|>'
(no style instructions that cause sampling failures)
- V3 Instruct/SFT: full instruct formatting as before
- Both paths use the unified inference_instruct2 API
V3 0.5B Base (Fun-CosyVoice3-0.5B-2512) does NOT support instruct.
Previously all V3 models were forced through inference_instruct2,
causing LLM sampling failures (max_trials EOS error).
Now V3/V2 Base models:
- Skip instruct formatting entirely
- Use inference_zero_shot (with ref audio) instead of instruct2
- Use inference_sft (with speaker ID) for identity path
New features:
- AIIA Emotion Annotator (LLM-driven emotion tagging)
- AIIA Text Splitter (single-speaker text splitting)
- Emotion tag handling for all TTS engines
- VoiceDesign UI priority fix
New node that splits single-speaker text into dialogue_json format:
- 3 split modes: auto, by_sentence, by_line
- Short sentence merging (min_chars)
- Long sentence splitting at commas/semicolons (max_chars)
- Output compatible with Emotion Annotator and Dialogue TTS
Workflow: Text → Text Splitter → Emotion Annotator → TTS
- Split Qwen3-TTS into CustomVoice (supports instruct) and
Base (strips tags, doesn't support instruct) in compat table
- Add note: UI dropdown emotion takes priority over inline tags
When text contains [Happy]/[Calm] etc (from Splitter output),
extract the emotion into instruct for proper Qwen3 handling,
then strip the tag from text to prevent reading it aloud.
The grouping key only used (model_id, dialect), so sentences with
different emotions could be batched together. Now includes the
instruct string (which contains emotion) in the grouping key.
VibeVoice doesn't support [Emotion] tags and would read them
as literal text. Now strips all 24 known emotion tags from
input text during preprocessing.
When Emotion Annotator has tagged dialogue with emotions,
Splitter now outputs '[Happy] text...' format in speaker_A_text
and speaker_B_text, so downstream TTS nodes can consume them.
Also passes emotion field through in split_map.
Previously all emotions in a batch were merged into one instruct
string (e.g. 'Happy,Calm。'), which is semantically wrong.
Now emotion is included in param_hash, so batches split when
per-sentence emotion changes. Consecutive same-emotion sentences
still batch together for efficiency.
- LLM may return line index as string, now cast to int
- Show neutral annotations in log (was hidden before)
- Add raw LLM response debug output
- Track index out-of-bounds warnings
- New node: 🎭 AIIA Emotion Annotator (LLM)
- Supports Groq (free), Ollama, vLLM via OpenAI-compatible API
- API key from GROQ_API_KEY env var or node parameter
- Custom base URL for local LLM services
- skip_existing / overwrite_all modes
- Robust JSON parsing with markdown code block handling
- Proxy support from environment variables
max(fa_end, midpoint) ensures that even if FA reports slightly
overlapping timestamps, we never truncate the current sentence
below its own FA endpoint.
When energy extends cut_end past the next sentence's FA start point,
clamp to midpoint between current FA end and next FA start. Fixes
issue where A[1] tail extension was eating into A[2]'s first char.