- New node: 🎭 AIIA Emotion Annotator (LLM)
- Supports Groq (free), Ollama, vLLM via OpenAI-compatible API
- API key from GROQ_API_KEY env var or node parameter
- Custom base URL for local LLM services
- skip_existing / overwrite_all modes
- Robust JSON parsing with markdown code block handling
- Proxy support from environment variables
max(fa_end, midpoint) ensures that even if FA reports slightly
overlapping timestamps, we never truncate the current sentence
below its own FA endpoint.
When energy extends cut_end past the next sentence's FA start point,
clamp to midpoint between current FA end and next FA start. Fixes
issue where A[1] tail extension was eating into A[2]'s first char.
FA determines precise onset, but cut_end now uses energy valley
detection (direction='after', 150ms radius) near the FA endpoint.
This prevents tail truncation while keeping FA's accurate start.
Fixed: mixed Chinese-English text like 'vscode' caused uppercase 'V'
to be passed to MMS_FA tokenizer which only accepts [a-z, space, ', -].
Now _chinese_to_pinyin() always lowercases output and strips any
characters not in the tokenizer's vocabulary.
- Add 'use_forced_align' boolean toggle (default: off)
- _load_fa_model(): lazy singleton via torchaudio MMS_FA bundle,
auto-symlinks local model from models/mms_fa/model.pt to hub cache
- _chinese_to_pinyin(): pypinyin conversion for MMS_FA compatibility
- _forced_align_sentences(): CTC forced alignment → per-sentence
timestamps with confidence scores
- _compute_iou(): IoU metric for cross-validation scoring
- Priority hierarchy: FA > VAD > Energy
- When FA+VAD both enabled: runs all 3 methods and prints IoU
comparison (FA-VAD, FA-Energy, VAD-Energy) per segment
- Add 'use_vad' boolean toggle to node inputs (default: off)
- _load_vad_model(): lazy singleton via torch.hub (no extra pip install)
- _get_vad_timestamps(): resample to 16kHz, run Silero VAD with tuned
params for TTS audio (threshold=0.3, min_speech=100ms, pad=20ms)
- _refine_with_vad(): maps ASR boundaries to VAD speech intervals
using overlap-based matching with 200ms search margin
- stitch(): conditionally uses VAD or energy detection based on toggle
- Falls back to energy detection if VAD model fails to load
- cut_end expansion reduced from 150ms to 50ms (MAX_EXPAND_END)
- cut_start expansion remains 150ms (MAX_EXPAND_START)
- Padding now only applied to cut_start, NOT cut_end
(fade-out handles the tail smoothly, no need for extra padding)
Two changes:
1. _expand_to_midpoints: cap expansion at 150ms beyond speech boundary
instead of going all the way to the midpoint between sentences.
2. _refine_cut_point for cut_end: change direction from 'after' to
'before' — find where speech actually ends, don't push further
toward the next sentence.
- Add seed input (default 0, -1 for random) to both Standard and Realtime TTS
- Set torch.manual_seed + cuda.manual_seed_all before model.generate()
- Ensures same seed + same text = identical audio output
Track each speaker's previous cut_end. After applying padding,
clamp cut_start to never be earlier than the speaker's previous
cut_end. This eliminates the 200ms overlap that caused A's tail
audio to replay when switching A→B→A.
- Replace 5ms linear fade with 30ms cosine fade (user-adjustable via fade_ms)
- Use low-level noise floor in speaker gaps instead of dead silence
- Increase default padding 50ms→100ms, reduce gap 300ms→250ms
- Cosine curve provides smoother energy transition than linear
- _refine_cut_point now accepts direction parameter (before/after/both)
- cut_start only searches backward (away from speech onset)
- cut_end only searches forward (away from speech end)
- Increased default padding from 50ms to 100ms for safety margin
Restored from ef2c434:
- deleteItem() method and Delete key binding
- close() method with proper cleanup
- outsideClickListener (click-outside-to-close)
- Escape key to close browser dialog
- IntersectionObserver root: null for proper icon lazy loading
- Force re-render on dialog show
- Add voice-cloning-tts.json, voice-conversion.json, ditto-talking-head.json, podcast-dialogue.json
- Fix(JS): restore onConfigure callback in aiia_video_nodes.js that was accidentally removed in eff537d, causing pix_fmt to reset to h264 defaults on page refresh
- All ToDisk nodes now write .aiia_temp marker on directory creation
- VideoCombine cleanup_frames only deletes directories with this marker
- Affected nodes: FloatProcess_ToDisk, DittoSampler, PersonaLive_ToDisk, BodySway
- Changed fps from INT to FLOAT in EchoMimic and Ditto nodes
- Added step=0.001 to support non-integer frame rates (23.976, 29.97, etc.)
- Float nodes already used FLOAT, no change needed
- Stitcher boundary refinement with energy-based cut point detection
- Processor _parse_script: accept both [N]: and legacy Speaker N: input
- Processor _process_single: tokenize as [N]: format (1-based)
- Processor _create_voice_prompt: use [N]: prefix for voice prompts
- Node _normalize_roles: convert custom roles and Speaker N: to [N]: format
- Node auto-wrap: plain text becomes [1]: text
- Dialogue node: output [N]: format for batch TTS