Commit Graph
1014 Commits
Author SHA1 Message Date
Hawk Lee e6e17aec81 fix: add proxy_url input + EchoMimicV3 dist stub
- Replace ~/run.sh reading with proxy_url node input parameter
- Proxy priority: node param > HTTPS_PROXY env var
- Add no-op parallel_magvit_vae stub for single-GPU inference
- Remove environment-specific ~/run.sh dependency
2026-02-16 01:34:21 +08:00
Hawk Lee 7ceb05437d fix: read proxy settings from ~/run.sh when env vars not set
ComfyUI process doesn't source ~/run.sh, so proxy env vars are
missing. Now reads http_proxy/https_proxy from ~/run.sh as fallback.
2026-02-16 01:32:28 +08:00
Hawk Lee 797c477889 docs: mark LLM emotion annotation as completed 2026-02-16 01:25:19 +08:00
Hawk Lee a74a8b20bb feat: add AIIA Emotion Annotator node (LLM-driven)
- New node: 🎭 AIIA Emotion Annotator (LLM)
- Supports Groq (free), Ollama, vLLM via OpenAI-compatible API
- API key from GROQ_API_KEY env var or node parameter
- Custom base URL for local LLM services
- skip_existing / overwrite_all modes
- Robust JSON parsing with markdown code block handling
- Proxy support from environment variables
2026-02-16 01:24:36 +08:00
Hawk Lee 4a7132b82f docs: add manual download instructions for MMS FA model
Two methods: wget from Facebook CDN, or huggingface-cli download.
2026-02-16 00:40:18 +08:00
Hawk Lee d7476f4d42 feat(fa): auto-copy downloaded model to models/mms_fa/ for centralized management
After torchaudio downloads the MMS_FA model to hub cache, automatically
copy it to models/mms_fa/model.pt so future loads use the local copy.
2026-02-16 00:39:21 +08:00
Hawk Lee 1c25316f0d release: v1.12.0 - MMS Forced Alignment integration
- Mark VAD and FA tasks as completed in TODO.md
- Add v1.12.0 changelog entry in README.md
- Bump version to 1.12.0 in pyproject.toml
2026-02-16 00:35:12 +08:00
Hawk Lee 193164f91a fix(fa): add floor guarantee to cut_end clamp - never shorter than FA end
max(fa_end, midpoint) ensures that even if FA reports slightly
overlapping timestamps, we never truncate the current sentence
below its own FA endpoint.
2026-02-16 00:21:55 +08:00
Hawk Lee f0b741a6ab fix(fa): clamp cut_end to prevent overlap with next sentence's FA start
When energy extends cut_end past the next sentence's FA start point,
clamp to midpoint between current FA end and next FA start. Fixes
issue where A[1] tail extension was eating into A[2]'s first char.
2026-02-16 00:19:39 +08:00
Hawk Lee c354adbca8 feat(log): show actual hybrid cut points (USED=) in cross-validation log 2026-02-16 00:13:15 +08:00
Hawk Lee 0fe63bfa65 fix(fa): hybrid FA start + energy valley end to preserve tail resonance
FA determines precise onset, but cut_end now uses energy valley
detection (direction='after', 150ms radius) near the FA endpoint.
This prevents tail truncation while keeping FA's accurate start.
2026-02-16 00:05:25 +08:00
Hawk Lee ab932a030f fix(fa): lowercase all pinyin output and strip non-tokenizable chars
Fixed: mixed Chinese-English text like 'vscode' caused uppercase 'V'
to be passed to MMS_FA tokenizer which only accepts [a-z, space, ', -].
Now _chinese_to_pinyin() always lowercases output and strips any
characters not in the tokenizer's vocabulary.
2026-02-15 23:53:41 +08:00
Hawk Lee fece4c171a feat(stitcher): integrate MMS Forced Alignment with cross-validation
- Add 'use_forced_align' boolean toggle (default: off)
- _load_fa_model(): lazy singleton via torchaudio MMS_FA bundle,
  auto-symlinks local model from models/mms_fa/model.pt to hub cache
- _chinese_to_pinyin(): pypinyin conversion for MMS_FA compatibility
- _forced_align_sentences(): CTC forced alignment → per-sentence
  timestamps with confidence scores
- _compute_iou(): IoU metric for cross-validation scoring
- Priority hierarchy: FA > VAD > Energy
- When FA+VAD both enabled: runs all 3 methods and prints IoU
  comparison (FA-VAD, FA-Energy, VAD-Energy) per segment
2026-02-15 23:43:08 +08:00
Hawk Lee 708c7d8117 docs: add Splitter/ASR/Stitcher nodes to README
- 4.8 Podcast Splitter: dialogue JSON → per-speaker text lists
- 4.9 ASR Node: FunASR word-level timestamps for alignment
- 4.10 Podcast Stitcher: precise multi-speaker stitching with
  optional Silero VAD boundary detection toggle
2026-02-15 23:24:50 +08:00
Hawk Lee 56fcaef331 feat(stitcher): integrate Silero VAD for precise speech boundary detection
- Add 'use_vad' boolean toggle to node inputs (default: off)
- _load_vad_model(): lazy singleton via torch.hub (no extra pip install)
- _get_vad_timestamps(): resample to 16kHz, run Silero VAD with tuned
  params for TTS audio (threshold=0.3, min_speech=100ms, pad=20ms)
- _refine_with_vad(): maps ASR boundaries to VAD speech intervals
  using overlap-based matching with 200ms search margin
- stitch(): conditionally uses VAD or energy detection based on toggle
- Falls back to energy detection if VAD model fails to load
2026-02-15 23:16:34 +08:00
Hawk Lee 8a2dc5c53d feat(stitcher): smoothed energy envelope + balanced cut boundaries
_refine_cut_point:
- Window: 10ms → 20ms (spans brief consonant closures)
- Step: 10ms (overlapping for precision)
- 5-point sliding average smoothing (avoids transient low-energy
  traps from fricatives like s/sh/f)
- Min search region raised to 50ms

_expand_to_midpoints:
- MAX_EXPAND_END: 50ms → 100ms (compensates ASR early-exit)

stitch():
- cut_end direction: 'before' → 'both' (±100ms balanced search)
- cut_end padding: 0 → padding*0.3 (preserves tail resonance)

Also updated TODO.md with Silero VAD and Forced Alignment ideas.
2026-02-15 22:11:03 +08:00
Hawk Lee 8b351485ca chore: add TODO.md with LLM emotion tagging idea 2026-02-15 22:05:05 +08:00
Hawk Lee 210b8d5396 fix(podcast): ensure TTS produces natural sentence-ending pauses
- Auto-append 。to lines missing sentence-final punctuation
  (prevents TTS from running sentences together)
- Use \n\n separator instead of \n in VibeVoice batch text
  (forces stronger pause cue between dialogue lines)
2026-02-15 21:37:50 +08:00
Hawk Lee f84c0b2552 fix(stitcher): asymmetric cut boundaries to prevent tail bleed
- cut_end expansion reduced from 150ms to 50ms (MAX_EXPAND_END)
- cut_start expansion remains 150ms (MAX_EXPAND_START)
- Padding now only applied to cut_start, NOT cut_end
  (fade-out handles the tail smoothly, no need for extra padding)
2026-02-15 21:19:30 +08:00
Hawk Lee 308fc8cb1b fix(stitcher): prevent cut_end from bleeding into next sentence
Two changes:
1. _expand_to_midpoints: cap expansion at 150ms beyond speech boundary
   instead of going all the way to the midpoint between sentences.
2. _refine_cut_point for cut_end: change direction from 'after' to
   'before' — find where speech actually ends, don't push further
   toward the next sentence.
2026-02-15 20:27:23 +08:00
Hawk Lee 63d05a8c5e feat(vibevoice): add seed parameter for reproducible speech generation
- Add seed input (default 0, -1 for random) to both Standard and Realtime TTS
- Set torch.manual_seed + cuda.manual_seed_all before model.generate()
- Ensures same seed + same text = identical audio output
2026-02-15 20:20:51 +08:00
Hawk Lee 96fce4e28d fix(stitcher): prevent repeated audio from padding-induced segment overlap
Track each speaker's previous cut_end. After applying padding,
clamp cut_start to never be earlier than the speaker's previous
cut_end. This eliminates the 200ms overlap that caused A's tail
audio to replay when switching A→B→A.
2026-02-15 20:05:45 +08:00
Hawk Lee e33c0b37ea fix(example): correct AudioPostProcess node name to AIIA_Audio_PostProcess 2026-02-15 19:55:48 +08:00
Hawk Lee 29e29a2393 fix(example): correct SubtitleGen node name to AIIA_Subtitle_Gen 2026-02-15 19:54:17 +08:00
Hawk Lee ed329bc23d fix(stitcher): improve transition naturalness with cosine fade and noise floor
- Replace 5ms linear fade with 30ms cosine fade (user-adjustable via fade_ms)
- Use low-level noise floor in speaker gaps instead of dead silence
- Increase default padding 50ms→100ms, reduce gap 300ms→250ms
- Cosine curve provides smoother energy transition than linear
2026-02-15 19:51:52 +08:00
Hawk Lee 4ca23f2d9d fix(stitcher): directional cut points to prevent speech onset clipping
- _refine_cut_point now accepts direction parameter (before/after/both)
- cut_start only searches backward (away from speech onset)
- cut_end only searches forward (away from speech end)
- Increased default padding from 50ms to 100ms for safety margin
2026-02-15 19:20:01 +08:00
Hawk Lee e91c905358 fix(browser): restore features lost in eff537d overwrite
Restored from ef2c434:
- deleteItem() method and Delete key binding
- close() method with proper cleanup
- outsideClickListener (click-outside-to-close)
- Escape key to close browser dialog
- IntersectionObserver root: null for proper icon lazy loading
- Force re-render on dialog show
2026-02-15 18:18:23 +08:00
Hawk Lee 536410b359 feat: add 4 workflow examples + fix pix_fmt regression
- Add voice-cloning-tts.json, voice-conversion.json, ditto-talking-head.json, podcast-dialogue.json
- Fix(JS): restore onConfigure callback in aiia_video_nodes.js that was accidentally removed in eff537d, causing pix_fmt to reset to h264 defaults on page refresh
2026-02-15 18:13:46 +08:00
Hawk Lee d6d6e76a3d docs: add cleanup_frames documentation and v1.11.1 changelog to README 2026-02-15 17:53:19 +08:00
Hawk Lee 98cff2b74e docs: sync version to 1.11.1, add cleanup_frames documentation and changelog entry 2026-02-15 17:51:58 +08:00
Hawk Lee c1b5f83cb9 feat(safety): add .aiia_temp marker file to prevent accidental deletion of user frame directories
- All ToDisk nodes now write .aiia_temp marker on directory creation
- VideoCombine cleanup_frames only deletes directories with this marker
- Affected nodes: FloatProcess_ToDisk, DittoSampler, PersonaLive_ToDisk, BodySway
2026-02-15 17:46:13 +08:00
Hawk Lee f02eab2669 feat(video): add cleanup_frames option to delete input frames directory after successful merge 2026-02-15 17:43:11 +08:00
Hawk Lee 91adc48bf2 fix: unify fps parameter to FLOAT across all nodes; bump to v1.8.3
- Changed fps from INT to FLOAT in EchoMimic and Ditto nodes
- Added step=0.001 to support non-integer frame rates (23.976, 29.97, etc.)
- Float nodes already used FLOAT, no change needed
- Stitcher boundary refinement with energy-based cut point detection
2026-02-15 17:38:15 +08:00
Hawk Lee b8366be6ea fix(assets): replace seed_male_hq.wav with NotebookLM male voice (8s, 22050Hz) 2026-02-15 15:45:11 +08:00
Hawk Lee 28877f4e30 fix(stitcher): add 5ms fade-in/fade-out to eliminate pops at splice points 2026-02-15 14:30:30 +08:00
Hawk Lee f8ba18ef0c fix: auto-wrap each line with [1]: for multi-line plain text input 2026-02-15 14:14:35 +08:00
Hawk Lee feb7938ade revert(splitter): output pure text, VibeVoice auto-wrap handles [1]: prefix 2026-02-15 14:12:02 +08:00
Hawk Lee 3df8734b14 fix(splitter): add [1]: speaker prefix to output text for VibeVoice format 2026-02-15 13:45:30 +08:00
Hawk Lee d25f0e3ac9 fix(dialogue): use [N]: format and add voice_preset param for VibeVoice 2026-02-15 13:26:33 +08:00
Hawk Lee c0bf61185e fix: use official VibeVoice [N]: speaker format instead of Speaker N:
- Processor _parse_script: accept both [N]: and legacy Speaker N: input
- Processor _process_single: tokenize as [N]: format (1-based)
- Processor _create_voice_prompt: use [N]: prefix for voice prompts
- Node _normalize_roles: convert custom roles and Speaker N: to [N]: format
- Node auto-wrap: plain text becomes [1]: text
- Dialogue node: output [N]: format for batch TTS
2026-02-15 13:18:18 +08:00
Hawk Lee ee51d51e38 fix(vibevoice): auto-wrap plain text as 'Speaker 1:' for processor compatibility 2026-02-15 13:10:13 +08:00
Hawk Lee 88710ae215 feat(vibevoice): add voice_preset dropdown for built-in reference voices
- Add voice_preset selector (Female_HQ, Male_HQ, Female, Male) to required inputs
- Priority: user-provided reference_audio > voice_preset dropdown
- Single speaker uses selected preset; multi-speaker starts with selected, cycles through others
2026-02-15 13:06:09 +08:00
Hawk Lee c6c747b424 fix: restore overwritten remote files and add new ASR/Splitter/Stitcher nodes
- Restore 12 files from 37cd909 that were accidentally overwritten
  by older local copies (includes BodySway, expanded subtitle/vibevoice nodes, etc.)
- Re-add ASR, Podcast Splitter, Podcast Stitcher registrations in __init__.py
- Keep new files: aiia_asr_nodes.py, aiia_podcast_splitter.py, aiia_podcast_stitcher.py
2026-02-15 12:20:55 +08:00
Hawk Lee eff537d9dd feat: add ASR, Podcast Splitter, and Podcast Stitcher nodes for anti-leakage pipeline
- Add AIIA_ASR node with FunASR (paraformer-zh + SenseVoiceSmall)
- Add AIIA_Podcast_Splitter for per-speaker text splitting
- Add AIIA_Podcast_Stitcher with 3-tier fuzzy alignment (exact, Levenshtein, fallback)
- Register 5 previously missing modules in __init__.py
- Update README with installation guide and node documentation
2026-02-15 12:14:45 +08:00
Hawk 37cd9096f2 Delete __pycache__ directory 2026-02-15 01:01:48 +08:00
Hawk Lee 5343666e21 bump: version 1.11.2 -> 1.11.3 2026-02-04 20:05:23 +08:00
Hawk Lee aa601cdeb0 fix(qwen): add fade-in/out to dialogue segments to prevent audio clicks 2026-02-04 19:48:12 +08:00
Hawk Lee 0859752cc9 feat(qwen): implement RMS normalization and reference audio volume scaling for Qwen3-TTS nodes 2026-02-04 19:30:12 +08:00
Hawk Lee 09d68fffa6 feat: Add more descriptive micro-expressions like 'With a hint of a smile' 2026-02-04 19:07:11 +08:00
Hawk Lee 9ae80b0be6 fix: Restore missing 'Gentle (温柔)' emotion to lists 2026-02-04 19:04:19 +08:00