diff --git a/CHANGELOG.md b/CHANGELOG.md index 994c719..1278bd0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,19 @@ # IAMCCS Nodes - Changelog +## πŸ†• 2026-02-24 β€” πŸ†• Version 1.3.5 WanImageMotionPro + Motion Safety Preset + +Changes: +- Added new video node: `WanImageMotionPro` (Motion + FLF End Lock) + - Optional `end_samples` to lock the ending latent slots (FLF-style end control) +- Added `safety_preset` to motion nodes (`IAMCCS_WanImageMotion` and `WanImageMotionPro`) + - `safe` (default): enables stabilizations only when `motion > 1.15` + - `safer`: stronger stabilization for higher motion values + - `legacy`: keeps the older behavior + +Docs: +- Added `docs/wanimagemotion_instructions.md` (Simple + Pro guide + example recipes) +- Updated `docs/WanImageMotion.md` + ## πŸ†• Version 1.3.4 β€” Video Performance + Low-RAM Tools Date: 2026-02-01 diff --git a/IAMCCS_HwSupporter.md b/IAMCCS_HwSupporter.md deleted file mode 100644 index 7bbd39f..0000000 --- a/IAMCCS_HwSupporter.md +++ /dev/null @@ -1,99 +0,0 @@ -# IAMCCS_HwSupporter - -Node pack per ComfyUI che applica in modo β€œauto / preset / manual” alcune impostazioni anti-OOM e speed knobs, con un report JSON in output. - -## Nodi - -### 1) HW Supporter (auto VRAM/attention/torch knobs) -- File: `iamccs_hw_supporter.py` (`IAMCCS_HwSupporter`) -- Input principale: `model` (MODEL) -- Output: `model`, `clip` (passthrough), `vae` (passthrough), `report_json` - -Posizionamento consigliato: -- Mettilo subito dopo il nodo che crea/carica il `MODEL` (e prima di LoRA/sampling). -- Se vuoi anche `vae_tiling_suggestion` nel report, collega anche `vae` in input (opzionale). - -Cosa fa: -- VRAM reserve: imposta `comfy.model_management.EXTRA_RESERVED_VRAM` (simile al nodo reservedvram). -- SageAttention: se installato, patcha l’attenzione del modello via `model.model_options["transformer_options"]["optimized_attention_override"]`. -- PyTorch knobs: `torch.backends.cuda.matmul.allow_fp16_accumulation`, TF32. -- (Opzionale) `torch.compile`: prova a compilare `model.model.diffusion_model` (attenzione: puΓ² aumentare picco VRAM al primo run). -- Nel `report_json` include anche `vae_tiling_suggestion` (tile_size/overlap consigliati) basati su VRAM rilevata. -- Se `console_log=true` stampa una riga riassuntiva nel terminale (e i warning). - -### 2) VRAM Cleanup (unload + empty cache) -- File: `iamccs_hw_supporter.py` (`IAMCCS_VRAMCleanup`) -- Utility node per forzare `unload_all_models()` + `soft_empty_cache()` (piΓΉ `gc.collect()` e `torch.cuda.empty_cache()`). - -### 3) VAE Decode Tiled (safe, optional cleanup) -- File: `iamccs_hw_supporter.py` (`IAMCCS_VAEDecodeTiledSafe`) -- Wrapper di `vae.decode_tiled(...)` con tile/overlap e supporto chunk temporale (video VAE). -- Opzione `cleanup_before_decode` per ridurre i picchi VRAM quando il decode arriva dopo il sampling. -- Nuova opzione `tiling_mode`: - - `auto`: sceglie automaticamente `tile_size` e `overlap` in base alla VRAM rilevata (conservativo, anti-OOM) - - `manual`: usa i valori inseriti a mano - -## Preset consigliati (12GB VRAM / 32GB RAM) -Impostazione pratica (conservativa): -- `profile`: `12GB_VRAM_32GB_RAM` -- `reserved_vram_gb`: 1.25 (oppure 1.5 se spesso in OOM) -- `sage_attention`: `auto` (se disponibile) -- `torch_compile_mode`: `off` (in genere piΓΉ stabile su low-vram/offload) -- `fp16_accumulation`: `auto` -- `tf32`: `auto` - -## Note importanti -- `PYTORCH_CUDA_ALLOC_CONF`: in genere va impostato **prima** di avviare ComfyUI per influenzare l’allocator. Il nodo riporta un warning/nota, ma non β€œgarantisce” di cambiare l’allocator a runtime. -- `torch.compile`: in molti setup low-vram/offload puΓ² dare instabilitΓ  o aumentare il picco VRAM (soprattutto al primo run). Usalo solo se hai margine. - -## Suggerimento pratico (pipeline 12GB) -- Sampling β†’ (opzionale) `VRAM Cleanup` β†’ `VAE Decode Tiled (safe)` con `tiling_mode=auto` e `cleanup_before_decode=true` se sei al limite. - -## Debug -Se qualcosa non funziona: -- guarda `report_json` (warnings + applied). -- prova a disabilitare SageAttention o `torch.compile`. -- inserisci `VRAM Cleanup` tra fasi pesanti (es. prima del VAE decode). - -## Crash Triton su Windows (libtriton.pyd / 0x80000003) -Se vedi un hard-crash tipo `libtriton.pyd` + `Exception Code: 0x80000003`, non Γ¨ un OOM: di solito Γ¨ un crash interno Triton/MLIR. - -Mitigazioni consigliate: -- In `IAMCCS_HwSupporter`: `torch_compile_mode = off`. -- In `IAMCCS_HwSupporter`: evita modalitΓ  SageAttention basate su Triton. - - usa `sageattn_qk_int8_pv_fp16_cuda` (consigliato) oppure `disabled`. -- Riavvia ComfyUI dopo i cambi (i crash Triton non sono β€œrecoverable”). - ---- - -# HW Probe & Apply (English) - -IAMCCS provides a **Hardware Probe** endpoint and UI buttons to automatically recommend and apply settings. - -## What you get - -- Backend endpoint: `GET /api/iamccs/hw_probe` -- Optional query params (best-effort context): `width`, `height`, `frames`, `fps` -- Frontend buttons (added to several IAMCCS nodes): - - **Probe HW & Apply**: updates widgets immediately (visible in real-time) - - **Copy HW report**: copies the full JSON report - -## Nodes supported by the button - -- `IAMCCS_HwSupporter` -- `IAMCCS_HwSupporterAny` -- `IAMCCS_SamplerCustomAdvancedWindowed` -- `IAMCCS_VAEDecodeTiledSafe` - -## Tips - -- The hw probe uses heuristics; best values still depend on your resolution and clip length. -- For long videos, the most important VRAM lever is **temporal chunking** (`temporal_size`). - -### torch.compile on Windows - -- Default is `torch_compile_mode=off` (safest). -- If you set `torch_compile_mode=auto`, the node will attempt compilation (internally uses a conservative mode, typically `reduce-overhead`). -- On Windows, torch.compile may still hard-crash depending on Torch/Inductor/driver; if you get hard crashes, switch back to `off`. - -See `LOW_VRAM_VIDEO_TIPS.md` for practical guidance. diff --git a/README.md b/README.md index b90847d..39d416f 100644 --- a/README.md +++ b/README.md @@ -9,7 +9,21 @@ ### Category: ComfyUI Custom Nodes ### Main Feature: Fix for LoRA loading in native WANAnimate workflows + general nodes 4 ComfyUI -Version: 1.3.4 +Version: 1.3.5 + +## πŸ†• Motion Nodes Update (2026-02-24) + +This update extends the WAN SVI Pro motion toolset: + +- New node: `WanImageMotionPro (Motion + FLF End Lock)` + - Adds optional `end_samples` end-lock (FLF-style) on top of motion continuity. +- New artifact-mitigation widget on both motion nodes: `safety_preset` + - `safe` (default): activates stabilizations only when `motion > 1.15` + - `safer`: stronger stabilization for higher motion values + - `legacy`: keeps the older behavior + +![[Node piece](assets/wanimagemotionpro.png)](https://github.com/IAMCCS/IAMCCS-nodes/blob/main/assets/wanimagemotionpro.png) + # UPDATE VERSION 1-3-4 @@ -35,16 +49,13 @@ Highlights (EN): - Backward-compatible input ordering preserved for older workflows. - Frontend quality-of-life: - - Bus Group β€œHide options” now persists across sessions. + - Bus Group with MACRO settings. - HW probe apply is user-controlled (overwrite vs fill-missing) and preset sync can be disabled to keep manual tuning. - MultiSwitch (frontend + workflow UX): `MultiSwitch (dynamic inputs)` (`IAMCCS_MultiSwitch`) - Active-link indicator: visually shows which input is currently connected/used. - Input rename: you can rename inputs to keep complex graphs readable (especially when routing MANY signals). -Docs: -- Low VRAM Video Tips: `LOW_VRAM_VIDEO_TIPS.md` - --- # UPDATE VERSION 1-3-3 @@ -62,10 +73,6 @@ Highlights (EN): ![[Node piece](assets/extension.png)](https://github.com/IAMCCS/IAMCCS-nodes/blob/main/assets/extension.png) -Docs: - -- LTX-2 Extension Module (EN/IT): `LTX2_EXTENSION_MODULE_README.md` -- LTX-2 Nodes Guide: `LTX2_EXTENSION_NODES_GUIDE_EN.md` GGUF / OOM tips: - If you use `IAMCCS_GGUF_accelerator` and you are close to the VRAM limit, consider PyTorch allocator tuning to reduce fragmentation (must be set **before** launching ComfyUI). @@ -140,8 +147,9 @@ Highlights: - Motion modes: apply boost to `prev_samples` only or all non-first latents. - VRAM profiles: normal / chunked / per-frame loop / CPU offload for memory-constrained systems. - `include_padding_in_motion` toggle: enables motion boost on padded frames when anchor has single frame (T=1). +- `safety_preset` (safe defaults for higher motion): helps reduce color artifacts and seam degradation when pushing `motion`. - Comprehensive logging with warnings when motion_range is empty. -- Full documentation: `WanImageMotion.md` +- Full documentation: `docs/WanImageMotion.md` and `docs/wanimagemotion_instructions.md` - Removed the previously included external-model LoRA loader node and related documentation. ### New Node: IAMCCS WanImageMotion @@ -158,6 +166,7 @@ Inputs: - `include_padding_in_motion`: enable to apply motion on padded frames - `vram_profile`: memory optimization strategy - `latent_precision`: dtype control (auto/fp16/fp32) +- `safety_preset`: `safe` / `safer` / `legacy` (artifact mitigation when `motion > 1.15`) - `add_reference_latents`: optional conditioning stabilization - Optional `prev_samples`: previous latents for motion continuity diff --git a/__init__.py b/__init__.py index b431cae..3d10113 100644 --- a/__init__.py +++ b/__init__.py @@ -43,10 +43,12 @@ from .iamccs_ltx2_extension_module import ( IAMCCS_LTX2_ReferenceImageSwitch, IAMCCS_LTX2_ReferenceStartFramesInjector, IAMCCS_LTX2_FrameCountValidator, + IAMCCS_LTX2_FirstLastFramesController, ) from .iamccs_wan_svipro_motion import ( IAMCCS_WanImageMotion, + WanImageMotionPro, ) from .iamccs_autolink import ( @@ -84,6 +86,11 @@ from .iamccs_hw_probe_node import ( IAMCCS_HWProbeRecommendations, ) +from .iamccs_qwen_vl_flf import ( + IAMCCS_QWEN_VL_FLF, + IAMCCS_QWEN_VL_FLF_Advanced, +) + # Nodi principali NODE_CLASS_MAPPINGS = { "IAMCCS_WanLoRAStack": IAMCCS_WanLoRAStack, @@ -113,7 +120,10 @@ NODE_CLASS_MAPPINGS = { "IAMCCS_LTX2_ReferenceImageSwitch": IAMCCS_LTX2_ReferenceImageSwitch, "IAMCCS_LTX2_ReferenceStartFramesInjector": IAMCCS_LTX2_ReferenceStartFramesInjector, "IAMCCS_LTX2_FrameCountValidator": IAMCCS_LTX2_FrameCountValidator, + "IAMCCS_LTX2_FirstLastFramesController": IAMCCS_LTX2_FirstLastFramesController, "IAMCCS_WanImageMotion": IAMCCS_WanImageMotion, + "WanImageMotionPro": WanImageMotionPro, + "IAMCCS_WanImageMotionPro": WanImageMotionPro, "IAMCCS_SetAutoLink": IAMCCS_SetAutoLink, "IAMCCS_GetAutoLink": IAMCCS_GetAutoLink, @@ -135,6 +145,12 @@ NODE_CLASS_MAPPINGS = { "IAMCCS_VAEDecodeToDisk": IAMCCS_VAEDecodeToDisk, "IAMCCS_HWProbeRecommendations": IAMCCS_HWProbeRecommendations, + # QwenVL First/Last Frame (registered only if QwenVL is installed) + **({ + "IAMCCS_QWEN_VL_FLF": IAMCCS_QWEN_VL_FLF, + "IAMCCS_QWEN_VL_FLF_Advanced": IAMCCS_QWEN_VL_FLF_Advanced, + } if IAMCCS_QWEN_VL_FLF is not None else {}), + } NODE_DISPLAY_NAME_MAPPINGS = { @@ -163,7 +179,10 @@ NODE_DISPLAY_NAME_MAPPINGS = { "IAMCCS_LTX2_ReferenceImageSwitch": "LTX-2 Reference Image Switch 🧷", "IAMCCS_LTX2_ReferenceStartFramesInjector": "LTX-2 Inject Reference Into Start Frames 🧬", "IAMCCS_LTX2_FrameCountValidator": "LTX-2 Frame Count Validator βœ… (8n+1)", + "IAMCCS_LTX2_FirstLastFramesController": "LTX-2 First/Last Frames Controller 🧲", "IAMCCS_WanImageMotion": "WanImageMotion", + "WanImageMotionPro": "WanImageMotionPro (Motion + FLF End Lock)", + "IAMCCS_WanImageMotionPro": "WanImageMotionPro (Motion + FLF End Lock)", "IAMCCS_SetAutoLink": "Set AutoLink", "IAMCCS_GetAutoLink": "Get AutoLink", @@ -185,6 +204,12 @@ NODE_DISPLAY_NAME_MAPPINGS = { "IAMCCS_VAEDecodeToDisk": "VAE Decode β†’ Disk (frames, low RAM)", "IAMCCS_HWProbeRecommendations": "HW Probe Recommendations (JSON)", + # QwenVL FLF + **({ + "IAMCCS_QWEN_VL_FLF": "QwenVL FLF β€” First/Last Frame Prompt 🎬", + "IAMCCS_QWEN_VL_FLF_Advanced": "QwenVL FLF β€” First/Last Frame Prompt (Advanced) 🎬", + } if IAMCCS_QWEN_VL_FLF is not None else {}), + } WEB_DIRECTORY = "./web" diff --git a/assets/wanimagemotionpro.png b/assets/wanimagemotionpro.png new file mode 100644 index 0000000..552a358 Binary files /dev/null and b/assets/wanimagemotionpro.png differ diff --git a/docs/AUTOLINK_PAPER.md b/docs/AUTOLINK_PAPER.md deleted file mode 100644 index d71f4b9..0000000 --- a/docs/AUTOLINK_PAPER.md +++ /dev/null @@ -1,185 +0,0 @@ -# IAMCCS AutoLink β€” Paper & Usage Instructions (EN/IT) - -## English - -### 1) What is AutoLink? -AutoLink is a **Set/Get** workflow tool designed to keep ComfyUI graphs clean and maintainable. - -Instead of long cables across the canvas, AutoLink lets you: -- Convert direct connections into **Set** (source) + **Get** (destination) pairs -- Restore the original direct connections when needed -- Apply repeatable filters (groups/blacklist), layout rules, and colors - -Everything is controlled by a dedicated β€œtool” node that operates on the canvas. - -### 2) Components -AutoLink is made of four logical elements: - -1. **AutoLink Converter** - - Buttons to convert/restore links. -2. **AutoLink Arguments** - - Central configuration: group filters, alignment/layout, packing/anti-overlap, colors, blacklist. -3. **AutoLink Set** - - Created near the source node: captures an output and exposes it under a key. -4. **AutoLink Get** - - Created near the destination node: retrieves the key and feeds the target input. - -### 3) Quickstart -1. Add to the canvas: - - **AutoLink Arguments** - - **AutoLink Converter** -2. Connect **AutoLink Arguments** output to the Converter `arg` input. -3. Adjust options (or keep defaults). -4. Click **Convert All Links**. - -To revert: -- Click **Restore Direct Links**. - -### 4) v1.3.3 reliability updates (important) -AutoLink Set/Get nodes are **UI tools** and are treated as **virtual** nodes. To prevent β€œmissing required input” prompt errors, the extension automatically: -- Materializes direct links **only during prompt serialization/queue**, then restores the AutoLink wiring -- Supports nested graphs/subgraphs -- Truncates long AutoLink titles with an ellipsis (`…`) so they stay inside the node header - ---- - -## Italiano - -### 1) Cos’è AutoLink -AutoLink Γ¨ un sistema **Set/Get** pensato per rendere i workflow ComfyUI piΓΉ ordinati, leggibili e facili da mantenere. - -Invece di avere cavi lunghi che attraversano la canvas, AutoLink permette di: -- Convertire automaticamente collegamenti diretti in coppie **Set** (sorgente) + **Get** (destinazione) -- Ripristinare i collegamenti originali quando serve -- Gestire filtri, gruppi, layout e colori in modo ripetibile - -Il tutto Γ¨ controllato da un nodo β€œtool” che opera sulla canvas. - -### 2) I nodi coinvolti -AutoLink Γ¨ composto da quattro elementi logici: - -1. **AutoLink Converter** - - Contiene i pulsanti per convertire/ripristinare i collegamenti. -2. **AutoLink Arguments** - - Contiene tutte le opzioni: filtri per gruppi, layout, packing/anti-overlap, colori, blacklist. -3. **AutoLink Set** - - Viene creato vicino al nodo sorgente: cattura un output e lo espone con una chiave. -4. **AutoLink Get** - - Viene creato vicino al nodo destinazione: recupera la chiave del Set e alimenta l’input. - -### 3) Quickstart (workflow consigliato) -1. Aggiungi in canvas: - - **AutoLink Arguments** - - **AutoLink Converter** -2. Collega l’output di **AutoLink Arguments** all’input `arg` di **AutoLink Converter**. -3. Imposta le opzioni nel nodo **AutoLink Arguments** (anche lasciando i default). -4. Premi **Convert All Links** nel nodo **AutoLink Converter**. - -Per tornare indietro: -- Premi **Restore Direct Links** nel Converter. - -### 4) Aggiornamenti affidabilitΓ  v1.3.3 (importante) -I nodi Set/Get di AutoLink sono strumenti **lato UI** e vengono trattati come nodi **virtuali**. Per evitare errori di prompt del tipo β€œrequired input missing”, l’estensione: -- Materializza i link diretti **solo durante la queue/serializzazione del prompt**, poi ripristina il wiring AutoLink -- Supporta grafi annidati/subgraph -- Tronca i titoli AutoLink troppo lunghi con ellissi (`…`) per non farli uscire dal nodo - ---- - -## 5) Opzioni principali (Arguments) - -### 4.1 GroupExclude -- Se abilitato, **non converte** i collegamenti tra due nodi che stanno **dentro lo stesso group**. -- I collegamenti che **entrano** o **escono** dal group possono comunque essere convertiti (dipende anche da GroupInOutExclude). - -Quando usarlo: -- Se un group rappresenta un β€œblocco logico” che vuoi tenere cablato internamente. - -### 4.2 GroupInOutExclude -Gestisce i link che attraversano un confine di group: -- `None`: nessuna esclusione. -- `ExcludeEnter`: non converte i link che **entrano** in un group. -- `ExcludeExit`: non converte i link che **escono** da un group. -- `ExcludeBoth`: combina entrambe. - -### 4.3 Align mode -Determina come vengono posizionati Set/Get dopo la conversione e quando fai relayout. - -Opzioni principali: -- `TopToDown`, `BottomToTop`, `CenterUpDown`, `CenterDownUp` -- `AlignX_Right`, `AlignX_Left` -- `Columns_Down`, `Columns_Up` -- `Rake_Down`, `Rake_Up` -- **`Proportional`** (consigliato per layout β€œcome i cavi”) - -#### Align = Proportional (come nell’immagine) -Con `Proportional`, Set e Get vengono agganciati alla **stessa altezza (Y)** del relativo connettore (slot) del nodo: -- Set: si allinea alla Y dello **slot di output** sorgente -- Get: si allinea alla Y dello **slot di input** destinazione - -In caso di collisioni, mantiene la Y e cerca spazio spostandosi orizzontalmente. - -### 4.4 Packing mode -Controlla l’anti-overlap durante posizionamento e relayout: -- `AvoidAll`: evita sovrapposizioni con tutti i nodi. -- `AvoidNonAutoLink`: evita solo i nodi non-AutoLink (Set/Get possono compattarsi fra loro). - -### 4.5 SeparateCol + colori -- `SeparateCol`: se attivo, permette di usare colori diversi per Set e Get. -- `AutoLinkColor`: colore base (Set). -- `AutoLinkColorGet`: colore dei Get (solo se SeparateCol Γ¨ attivo). - -### 4.6 ColorTitles -Cambia il colore del testo del titolo dei nodi AutoLink: -- `White` -- `Black` -- `Auto` - -### 4.7 Blacklist (ID e Types) -AutoLink permette di escludere nodi dalla conversione: - -- `all_nodes_sel`: - - OFF: la blacklist lavora per **tipo** (`[TYPE] ...`) - - ON: la blacklist lavora per **ID singolo nodo** - -- `add_to_blacklist`: - - Scegli un nodo (ID) o un tipo. - -- `blacklist_mode` (solo per nodi singoli): - - `both`: esclude link dove il nodo Γ¨ sorgente o destinazione - - `only_output`: esclude solo quando il nodo Γ¨ sorgente (output) - - `only_input`: esclude solo quando il nodo Γ¨ destinazione (input) - -- `EXECUTE`: - - Applica davvero l’inserimento (o l’update della modalitΓ ) e poi pulisce i widget. - -- `blacklist_view`: - - Elenco leggibile: `id - nome nodo - (modalitΓ )` e `[TYPE] ...`. - - Selezionare una voce **non rimuove nulla**. - -- `remove_blacklist`: - - Rimuove la voce attualmente selezionata in `blacklist_view`. - ---- - -## 6) Best practices -- Prima di convertire β€œtutto”, imposta la blacklist per escludere nodi che vuoi lasciare cablati. -- Usa `GroupExclude` per mantenere β€œblocchi” interni puliti. -- Usa `Proportional` quando vuoi un layout che segua visivamente l’ordine degli slot (come routing naturale dei cavi). -- Se la canvas Γ¨ molto piena, prova `PackingMode = AvoidAll`. - ---- - -## 7) Troubleshooting -- **Convert All Links non sembra fare nulla**: - - Verifica che `AutoLink Arguments` sia collegato all’input `arg` del Converter. - - Controlla blacklist e filtri group. -- **Nodi sovrapposti**: - - Prova `PackingMode = AvoidAll`. - - Cambia align mode o usa relayout cambiando `align_mode`. - ---- - -## 8) Documentazione correlata -- AUTOLINK_README.md -- AUTOLINK_TECHNICAL_PAPER.md diff --git a/docs/LOW_VRAM_VIDEO_TIPS.md b/docs/LOW_VRAM_VIDEO_TIPS.md deleted file mode 100644 index dab699a..0000000 --- a/docs/LOW_VRAM_VIDEO_TIPS.md +++ /dev/null @@ -1,114 +0,0 @@ -# IAMCCS Nodes – Low VRAM Video Tips - -This doc describes the low-VRAM features added to IAMCCS nodes for LTX video workflows. - -## 1) Hardware Probe + One-Click Apply - -IAMCCS exposes a small backend endpoint: - -- `GET /api/iamccs/hw_probe` -- Optional query params: `width`, `height`, `frames`, `fps` - -The IAMCCS UI extension adds buttons to several nodes: - -- **Probe HW & Apply** – reads your current GPU/RAM and (best-effort) reads the workflow context (width/height/frames/fps). It then applies recommended widget values immediately. -- **Copy HW report** – copies the full JSON report to clipboard. - -Notes: -- Recommendations are heuristics. Final best values depend on the model, resolution, and clip length. - -Frontend control (not rigid): -- **HW probe apply mode** - - `overwrite`: always overwrite widgets with recommended values - - `fill_missing`: only fills empty fields (does not clobber manual tuning) -- **Preset sync (profile β†’ widgets)** (on `IAMCCS_HwSupporter` / `IAMCCS_HwSupporterAny`) - - When ON: changing `profile` updates the other widgets to match the preset. - - When OFF: you keep full manual control; profile changes won’t overwrite your values. - -## 2) VAE Decode Tiled Safe (Video) - -Node: -- `VAE Decode Tiled (safe, optional cleanup)` (`IAMCCS_VAEDecodeTiledSafe`) - -Tips: -- For long videos, the most important VRAM control is **temporal chunking** (`temporal_size`). -- If you see CUDA OOM during decode, reduce: - - `tile_size` - - `temporal_size` - - keep `overlap` and `temporal_overlap` small but non-zero - -The HW probe can also recommend values for VAE decode based on: -- GPU VRAM -- width/height -- frames/fps (if detected) - -## 3) Debug / Verification - -Where to look: -- **ComfyUI server console**: - - `/api/iamccs/hw_probe` logs a short line whenever the button is used. -- **Browser devtools console**: - - the UI prints the full hw probe JSON under `[IAMCCS HW Probe]`. - -If the button updates widgets but values get overwritten: -- ensure you clicked the button last (after changing profile/preset), -- or disable any profile auto-sync if you prefer manual tuning. - -## 4) Recommended Workflow Pattern (Low VRAM) - -Typical ordering: -- GGUF model loader -- `IAMCCS_GGUF_accelerator` -- `IAMCCS_HwSupporter` (or `IAMCCS_HwSupporterAny`) -- sampler -- VAE decode tiled safe - -## 5) VAE Decode β†’ Disk (True Low-RAM Mode) - -New node: -- `VAE Decode β†’ Disk (frames, low RAM)` (`IAMCCS_VAEDecodeToDisk`) - -What it does: -- Decodes **one frame at a time** and writes frames to disk, instead of keeping the full `IMAGE` batch in RAM. -- This is the most reliable way to avoid CPU OOM on long clips when you still want full-resolution outputs. - -When to use it: -- Very long videos (hundreds of frames) -- Low system RAM (or heavy multitasking) -- When `VAEDecodeTiled` still spikes CPU allocator memory - -Tip: -- Keep `cleanup_between_frames=true` if you’re tight on VRAM. -- Use PNG for best quality; use JPG if disk size is a problem. - -## 6) GGUF Accelerator – Safer β€œmove_patches_now” - -`IAMCCS_GGUF_accelerator` now supports: -- `move_policy`: `all_or_nothing` / `partial_small_first` / `partial_large_first` -- `leave_free_vram_mb`: how much VRAM to keep free during eager patch moves - -Practical guidance: -- **8GB VRAM**: `move_policy=partial_small_first`, `leave_free_vram_mb=1500` (best chance to avoid OOM) -- **12–16GB VRAM**: `all_or_nothing`, `leave_free_vram_mb=1200` -- **24GB+ VRAM**: `all_or_nothing`, `leave_free_vram_mb=1024` (fastest) - -## 7) Presets (Low / Normal / High) - -These are sane starting points for LTX-style video workflows (no windowing): - -### Low (8GB VRAM or low RAM) -- Sampler: `IAMCCS_SamplerAdvancedVersion1` with `disable_progress=true`, `cleanup=true` -- GGUF: `mode=auto_oom_safe`, `patch_on_device=true`, `move_patches_now=true`, `move_policy=partial_small_first`, `leave_free_vram_mb=1500` -- VAE: prefer `IAMCCS_VAEDecodeTiledSafe` with smaller `tile_size` and `temporal_size=64` -- If CPU RAM is the limiter: use `IAMCCS_VAEDecodeToDisk` - -### Normal (12–16GB VRAM, 32GB RAM) -- Sampler: `disable_progress=true`, `cleanup=false` -- GGUF: `move_policy=all_or_nothing`, `leave_free_vram_mb=1200` -- VAE: `IAMCCS_VAEDecodeTiledSafe` with `tiling_mode=auto` (or manual: `tile_sizeβ‰ˆ384–512`, `temporal_size=64–96`) - -### High (24GB+ VRAM, 64GB+ RAM) -- Sampler: `disable_progress=true`, `cleanup=false` -- GGUF: `all_or_nothing`, `leave_free_vram_mb=1024` -- VAE: you can often increase `tile_size` and `temporal_size=128` for faster decode - diff --git a/docs/LTX2_EXTENSION_MODULE_COMPLETE_GUIDE.md b/docs/LTX2_EXTENSION_MODULE_COMPLETE_GUIDE.md deleted file mode 100644 index 6ec34cc..0000000 --- a/docs/LTX2_EXTENSION_MODULE_COMPLETE_GUIDE.md +++ /dev/null @@ -1,851 +0,0 @@ -# LTX-2 Extension Module - Complete Technical Guide - -## Table of Contents -1. [Overview](#overview) -2. [Architecture & Workflow](#architecture--workflow) -3. [Parameters Reference](#parameters-reference) -4. [Usage Scenarios](#usage-scenarios) -5. [Advanced Features](#advanced-features) -6. [Troubleshooting](#troubleshooting) -7. [Best Practices](#best-practices) - ---- - -## Overview - -The **IAMCCS LTX-2 Extension Module** is an all-in-one node designed for iterative video extension workflows with the LTX-2 model. It combines multiple operations into a single, efficient node: - -- **Image batch merging** with configurable overlap -- **Multiple blending modes** for smooth transitions -- **Automatic frame calculations** with built-in math operations -- **LTX-2 8n+1 conformance** for start_images -- **Advanced quality features** (color matching, seam search) - -### Key Benefits -- βœ… Eliminates need for multiple separate nodes (GetImageRange, ImageBatchExtend, SimpleMath, etc.) -- βœ… Automatic 8n+1 validation prevents encoding errors -- βœ… Seamless video segment concatenation with no visible cuts -- βœ… Flexible overlap strategies for different content types -- βœ… Built-in quality enhancement features - ---- - -## Architecture & Workflow - -### Basic Extension Flow - -```mermaid -graph TB - A[Generation 1
121 frames] --> B[Extension Module] - C[Generation 2
121 frames] --> B - B --> D[extended_images
217 frames] - B --> E[start_images
17 frames 8n+1] - E --> F[Next Generation Input] - - style B fill:#2a363b,stroke:#3f5159,color:#fff - style E fill:#233,stroke:#355,color:#fff -``` - -### Complete Multi-Segment Workflow - -``` -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ ITERATIVE EXTENSION LOOP β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - -Iteration 1: Initial Generation -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ Initial Image β”‚ 1 frame -β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - v -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ LTX Sampler β”‚ Generate 121 frames -β”‚ (8Γ—15 + 1) β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - v -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ VAE Decode β”‚ Latent β†’ Images -β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - v - source_images (121 frames) - β”‚ - └──────────────────────────────────┐ - β”‚ -Iteration 2: First Extension β”‚ -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ -β”‚ Extension β”‚β—„β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ -β”‚ Module │◄── new_images (121 frames from Gen 2) -β”‚ overlap=25 β”‚ -β”‚ mode=linear β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - β”œβ”€β”€β–Ί extended_images (217 frames) - β”‚ 121 - 25 + 121 = 217 - β”‚ - └──► start_images (17 frames) - 25 β†’ 24 (math: a-1) β†’ 17 (8n+1 conform) - β”‚ - v - β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” - β”‚ LTX Sampler β”‚ Gen 3 (121 frames) - β”‚ uses 17 frames β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - v - new_images - β”‚ - └──► Loop continues... - -Final Output: -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ Video Segments β”‚ -β”‚ 217 + 217 + ... β”‚ -β”‚ Seamless Concat β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ -``` - -### Internal Processing Flow - -``` -INPUT IMAGES - β”‚ - β”œβ”€β”€β”€ source_images (previous generation) - β”‚ β”‚ - β”‚ └─── Last 25 frames ──┐ - β”‚ β”‚ - └─── new_images (current generation) - β”‚ β”‚ - └─── First 25 frames ──── - β”‚ - β”Œβ”€β”€β”€β”€β”€β”€β”€β”€v────────┐ - β”‚ OVERLAP ZONE β”‚ - β”‚ 25 frames β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - β”Œβ”€β”€β”€β”€β”€β”€β”€β”€v────────┐ - β”‚ BLENDING β”‚ - β”‚ linear_blend β”‚ - β”‚ Alpha: 0β†’1 β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” - β”‚ β”‚ - β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€v─────────┐ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€v──────────┐ - β”‚ extended_images β”‚ β”‚ start_images β”‚ - β”‚ Full merged batch β”‚ β”‚ For next iteration β”‚ - β”‚ (source-25+new) β”‚ β”‚ With 8n+1 conform β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ -``` - ---- - -## Parameters Reference - -### Core Parameters - -#### `overlap_frames` (INT) -- **Default**: 10 -- **Range**: 1-256 -- **Recommended**: 25-40 for smooth transitions -- **Purpose**: Number of frames to overlap and blend between segments - -**Impact**: -- **Low (8-15)**: Fast processing, visible seams possible -- **Medium (20-30)**: βœ… **Recommended** - Good balance -- **High (40-60)**: Very smooth, but higher computational cost - -**Formula**: `extended_length = source_count - overlap + new_count` - -Example with overlap=25: -``` -source: [1...121] -new: [1...121] -overlap: 25 frames -extended: 121 - 25 + 121 = 217 frames -``` - ---- - -#### `overlap_side` (DROPDOWN) -- **Options**: `source` | `new_images` -- **Default**: `source` -- **Purpose**: Which batch to take overlap frames from - -``` -overlap_side = "source": - Take last 25 from source - Take first 25 from new - Blend sourceβ†’new (recommended) - -overlap_side = "new_images": - Take first 25 from new - Take last 25 from source - Blend newβ†’source (reverse) -``` - -**Use Cases**: -- `source`: βœ… **Standard** - Smooth forward progression -- `new_images`: Experimental - reverse blending effect - ---- - -#### `overlap_mode` (DROPDOWN) -- **Options**: `cut` | `linear_blend` | `ease_in_out` | `filmic_crossfade` | `perceptual_crossfade` -- **Default**: `linear_blend` - -### Blending Modes Comparison - -| Mode | Speed | Quality | Use Case | Formula | -|------|-------|---------|----------|---------| -| **cut** | ⚑⚑⚑ | ⭐ | Testing, no blend needed | Direct concatenation | -| **linear_blend** | ⚑⚑ | ⭐⭐⭐⭐ | βœ… **General use** | `(1-t)Γ—src + tΓ—dst` | -| **ease_in_out** | ⚑⚑ | ⭐⭐⭐⭐⭐ | Smooth artistic transitions | `3tΒ² - 2tΒ³` | -| **filmic_crossfade** | ⚑ | ⭐⭐⭐⭐⭐ | Color-accurate blending | Gamma 2.2 correction | -| **perceptual_crossfade** | ⚑ | ⭐⭐⭐⭐⭐ | Best quality (needs Kornia) | LAB color space blend | - -**Visual Comparison**: -``` -Alpha progression over 25 frames: - -linear_blend: -0.0 β–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 1.0 - β”‚ β”‚ - Linear interpolation - -ease_in_out: -0.0 β–ˆβ–ˆβ–“β–“β–’β–’β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–’β–’β–“β–“β–ˆβ–ˆβ–ˆβ–ˆ 1.0 - β”‚ Slowβ†’Fastβ†’Slow β”‚ - Smooth S-curve - -filmic_crossfade: -0.0 β–ˆβ–ˆβ–ˆβ–“β–“β–’β–’β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–’β–“β–“β–ˆβ–ˆβ–ˆ 1.0 - β”‚ Gamma-corrected β”‚ - Perceptually uniform -``` - -**Recommendations**: -- **General video**: `linear_blend` (fast, reliable) -- **High quality**: `ease_in_out` (smooth, cinematic) -- **Color-critical**: `filmic_crossfade` or `perceptual_crossfade` -- **Testing/Debug**: `cut` (no blending overhead) - ---- - -#### `enable_math` (BOOLEAN) -- **Default**: `true` -- **Purpose**: Enable mathematical operations on overlap value for start_images calculation - -When enabled, applies `math_operation` to calculate the number of frames for `start_images`. - ---- - -#### `math_operation` (DROPDOWN) -- **Options**: `none` | `a-b` | `a-1` | `a+b` | `a*b` | `a/b` | `min(a,b)` | `max(a,b)` -- **Default**: `a-b` -- **Variables**: - - `a` = overlap_frames - - `b` = math_value_b (optional input) - -**Common Use Cases**: - -| Operation | Example | Result | Use Case | -|-----------|---------|--------|----------| -| `none` | overlap=25 | 25 | Direct use of overlap | -| `a-1` | 25-1 | 24 | βœ… **Standard** - LTX-2 workflow | -| `a-b` | 25-15 | 10 | Custom frame count | -| `a/b` | 25/2.5 | 10 | Proportional reduction | - -**Recommended Configuration**: -```json -{ - "overlap_frames": 25, - "enable_math": true, - "math_operation": "a-1" -} -``` -Result: 25 - 1 = 24 frames β†’ 17 frames (after 8n+1 conform) - ---- - -#### `start_frames_rule` (DROPDOWN) -- **Options**: `none` | `ltx2_round_down` | `ltx2_nearest` -- **Default**: `none` -- **Purpose**: Enforce LTX-2 8n+1 rule for VideoVAE encoding - -### LTX-2 Frame Count Rule - -LTX-2 VideoVAE requires frame counts following the formula: **`frames = 8n + 1`** - -Valid frame counts: `1, 9, 17, 25, 33, 41, 49, 57, 65, 73, 81, 89, 97, 105, 113, 121...` - -**Examples**: - -| Input | ltx2_round_down | ltx2_nearest | none | -|-------|----------------|--------------|------| -| 24 | 17 (8Γ—2+1) | 17 (closer) | 24 ❌ | -| 26 | 25 (8Γ—3+1) | 25 (closer) | 26 ❌ | -| 30 | 25 (8Γ—3+1) | 33 (closer) | 30 ❌ | -| 17 | 17 βœ… | 17 βœ… | 17 βœ… | - -**When to Use**: -- βœ… **Always use** `ltx2_round_down` or `ltx2_nearest` when start_images feeds into a sampler -- ❌ **Never use** when output is only for preview/saving (not encoding) - -**Critical**: Without this, you'll get errors like: -``` -Error: Expected frame count 8n+1, got 24 -``` - ---- - -### Advanced Quality Parameters - -#### `color_match_mode` (DROPDOWN) -- **Options**: `none` | `luma_only` | `per_channel` -- **Default**: `none` -- **Purpose**: Match color/exposure of new_images to source_images tail - -**Use Cases**: -- **Lighting changes**: Different segments with varying brightness -- **Color shifts**: Camera auto-balance between shots -- **Consistency**: Maintain uniform look across segments - -``` -none: - source: β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–“β–“β–“β–“β–“ (bright end) - new: β–’β–’β–’β–’β–’β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ (dark start) - β†’ Visible seam - -luma_only: - Match overall brightness only - β†’ Quick, preserves color tone - -per_channel: - Match R, G, B independently - β†’ Best quality, may shift colors -``` - ---- - -#### `color_match_strength` (FLOAT) -- **Range**: 0.0-1.0 -- **Default**: 1.0 -- **Purpose**: Blend factor for color matching - -``` -strength = 0.0: No correction -strength = 0.5: Partial correction -strength = 1.0: Full correction -``` - ---- - -#### `seam_search_mode` (DROPDOWN) -- **Options**: `none` | `best_of_k` -- **Default**: `none` -- **Purpose**: Search for optimal seam position within overlap zone - -**How It Works**: -``` -Standard overlap (offset=0): -source: β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–“β–“β–“β–“β–“ -new: β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ - ↑ Potential seam - -Best-of-k search (k=8): -Tries offsets 0-8: -offset=0: β–“β–“β–“β–“β–“ vs β–‘β–‘β–‘β–‘β–‘ β†’ score: 0.85 -offset=1: β–“β–“β–“β–“β–“ vs β–‘β–‘β–‘β–‘β–‘ β†’ score: 0.72 -offset=2: β–“β–“β–“β–“β–“ vs β–‘β–‘β–‘β–‘β–‘ β†’ score: 0.65 βœ… Best! -... -Chooses offset=2 (lowest discontinuity) -``` - -**Scoring Metrics**: -- Color/luma continuity (weighted by `metric_weight_color`) -- Edge continuity (weighted by `metric_weight_edges`) - -**Trade-offs**: -- βœ… Reduces visible seams -- βœ… Handles motion/camera cuts better -- ❌ Slower (tests k candidates) -- ❌ May "skip" frames from new_images - ---- - -## Usage Scenarios - -### Scenario 1: Standard Video Extension (Recommended) - -**Goal**: Extend a video smoothly without visible seams - -**Configuration**: -```json -{ - "overlap_frames": 25, - "overlap_side": "source", - "overlap_mode": "linear_blend", - "enable_math": true, - "math_operation": "a-1", - "start_frames_rule": "ltx2_round_down", - "color_match_mode": "none", - "seam_search_mode": "none" -} -``` - -**Workflow**: -1. Generate segment 1 (121 frames) -2. Extract last 17 frames (8Γ—2+1) -3. Generate segment 2 with those 17 frames as reference -4. Extension Module merges with 25-frame overlap -5. Repeat - -**Output**: Seamless 217-frame video (then 313, 409, etc.) - ---- - -### Scenario 2: High-Quality Cinematic Extension - -**Goal**: Maximum quality with perceptual blending - -**Configuration**: -```json -{ - "overlap_frames": 40, - "overlap_side": "source", - "overlap_mode": "perceptual_crossfade", - "enable_math": true, - "math_operation": "a-1", - "start_frames_rule": "ltx2_nearest", - "color_match_mode": "per_channel", - "color_match_strength": 0.8, - "seam_search_mode": "best_of_k", - "k_search": 16 -} -``` - -**Best For**: -- Film production -- High-resolution output -- Color-critical content -- Complex lighting scenarios - ---- - -### Scenario 3: Fast Preview / Testing - -**Goal**: Quick iteration, minimal processing - -**Configuration**: -```json -{ - "overlap_frames": 10, - "overlap_side": "source", - "overlap_mode": "cut", - "enable_math": true, - "math_operation": "a-1", - "start_frames_rule": "ltx2_round_down", - "color_match_mode": "none", - "seam_search_mode": "none" -} -``` - -**Best For**: -- Testing prompts -- Workflow debugging -- Quick previews - ---- - -### Scenario 4: Lighting-Corrected Extension - -**Goal**: Handle varying lighting between segments - -**Configuration**: -```json -{ - "overlap_frames": 30, - "overlap_side": "source", - "overlap_mode": "ease_in_out", - "enable_math": true, - "math_operation": "a-1", - "start_frames_rule": "ltx2_round_down", - "color_match_mode": "luma_only", - "color_match_strength": 1.0, - "color_reference_window": 12 -} -``` - -**Best For**: -- Outdoor scenes (sun changes) -- Mixed lighting conditions -- Auto-exposure variations - ---- - -## Advanced Features - -### Two-Stage Overlap Strategy - -Replicating the "early version" workflow behavior with separate overlap values: - -```python -# Early version used: -# - overlap=10 for frame extraction -# - overlap=25 for blending - -# Extension Module equivalent: -{ - "overlap_frames": 25, # For blending - "math_operation": "a/b", # Calculate extraction - "math_value_b": 2.5, # 25/2.5 = 10 - "start_frames_rule": "ltx2_round_down" -} - -# Result: -# - Blending uses 25 frames (smooth) -# - start_images calculated from 10 β†’ 9 β†’ 9 frames (8Γ—1+1) -``` - ---- - -### Custom Frame Count Calculation - -**Example**: Generate 33 frames for next iteration (8Γ—4+1) - -```json -{ - "overlap_frames": 25, - "math_operation": "a+b", - "math_value_b": 9, // 25 + 9 = 34 - "start_frames_rule": "ltx2_round_down" // 34 β†’ 33 -} -``` - ---- - -### Adaptive Overlap with AutoLink - -When using AutoLink for iterative loops: - -```json -{ - "overlap_frames": 25, - "autolink_overlap_in": 0, // Override if > 0 from AutoLink - // ... other params ... -} - -// Extension Module outputs: -// autolink_overlap_out β†’ feeds next iteration's autolink_overlap_in -``` - ---- - -## Troubleshooting - -### Problem: Visible seams between segments - -**Symptoms**: Hard cuts, color shifts, motion jumps - -**Solutions**: -1. βœ… Increase `overlap_frames` to 25-40 -2. βœ… Change to `ease_in_out` or `filmic_crossfade` -3. βœ… Enable `color_match_mode = "luma_only"` -4. βœ… Try `seam_search_mode = "best_of_k"` with `k_search = 8` - ---- - -### Problem: Error "Expected 8n+1 frames" - -**Symptoms**: Workflow fails at sampler/encoder - -**Solutions**: -1. βœ… Set `start_frames_rule = "ltx2_round_down"` -2. βœ… Verify `enable_math = true` -3. βœ… Check math formula produces reasonable values -4. ❌ Don't use `start_frames_rule` if output is for preview only - ---- - -### Problem: Videos too long / memory issues - -**Symptoms**: Out of memory, slow processing - -**Solutions**: -1. βœ… Reduce `overlap_frames` to 15-20 -2. βœ… Use `overlap_mode = "linear_blend"` (faster) -3. βœ… Disable `seam_search_mode` -4. βœ… Process in smaller batches - ---- - -### Problem: Color mismatch at seams - -**Symptoms**: Brightness/hue shifts visible - -**Solutions**: -1. βœ… Enable `color_match_mode = "per_channel"` -2. βœ… Set `color_match_strength = 0.8-1.0` -3. βœ… Increase `color_reference_window` to 16-24 -4. βœ… Use `filmic_crossfade` for gamma-correct blending - ---- - -## Best Practices - -### 1. Start with Recommended Defaults - -```json -{ - "overlap_frames": 25, - "overlap_side": "source", - "overlap_mode": "linear_blend", - "enable_math": true, - "math_operation": "a-1", - "start_frames_rule": "ltx2_round_down", - "color_match_mode": "none", - "seam_search_mode": "none" -} -``` - -Then optimize based on your specific needs. - ---- - -### 2. Overlap Guidelines by Content Type - -| Content Type | Overlap | Blend Mode | Reason | -|--------------|---------|------------|--------| -| **Static scenes** | 15-20 | linear_blend | Less motion, simpler blend | -| **Camera movement** | 25-40 | ease_in_out | Smooth motion transition | -| **Fast action** | 30-50 | filmic_crossfade | Avoid motion artifacts | -| **Talking heads** | 20-30 | linear_blend | Consistent framing | -| **Nature/landscape** | 25-35 | perceptual_crossfade | Color accuracy | - ---- - -### 3. Processing Order - -Always follow this order in your workflow: - -``` -1. Initial Image - ↓ -2. LTX Sampler (8n+1 frames) - ↓ -3. VAE Decode - ↓ -4. Extension Module - β”œβ”€β†’ extended_images (for final output) - └─→ start_images (for next iteration) - ↓ -5. Loop back to step 2 -``` - -**Critical**: Never feed `extended_images` back into the sampler directly - always use `start_images` (conformant to 8n+1). - ---- - -### 4. Testing Workflow - -Before full production: - -1. Test with `overlap=10`, `mode=cut` (fast preview) -2. Verify no errors with `start_frames_rule = "ltx2_round_down"` -3. Increase overlap to 25, switch to `linear_blend` -4. Fine-tune with quality features if needed - ---- - -### 5. Output Validation - -Check the `report` output for each iteration: - -``` -Source: 121 frames | -Overlap (effective): 25 frames | -Start range: start_index=96, num_frames=17 | -Math: a-1 | -Start frames rule: ltx2_round_down | -Extended: 217 frames | -Extension delta: +96 frames | -Blend mode: linear_blend -``` - -Verify: -- βœ… `num_frames` is 8n+1 (9, 17, 25, 33, etc.) -- βœ… `Extension delta` is positive -- βœ… No warnings in console - ---- - -## Performance Optimization - -### Memory Usage - -| Configuration | Memory Impact | Speed | -|---------------|---------------|-------| -| overlap=10, cut | Low | ⚑⚑⚑ | -| overlap=25, linear | Medium | ⚑⚑ | -| overlap=40, ease_in_out | Medium-High | ⚑⚑ | -| overlap=40, perceptual + seam search | High | ⚑ | - ---- - -### Batch Processing Tips - -For very long videos (10+ segments): - -1. **Save intermediate results**: - ``` - Segment 1 β†’ Save - Segment 2 β†’ Save - ... - Final concatenation separately - ``` - -2. **Use progressive overlap**: - ``` - Segments 1-3: overlap=25 (quality) - Segments 4+: overlap=15 (speed) - ``` - -3. **Monitor VRAM**: - - Each 121-frame batch β‰ˆ 4-8GB VRAM - - Reduce resolution if needed - ---- - -## Workflow Diagrams - -### Complete Extension Pipeline - -``` -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ INITIALIZATION β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ Load Model │────>β”‚ Load VAE │────>β”‚ Load CLIP β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ β”‚ β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - v -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ GENERATION LOOP START β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - -Iteration N: -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ start_imagesβ”‚ (17 frames, 8Γ—2+1) -β”‚ from prev β”‚ -β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - v -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ SUBGRAPH: Samplers β”‚ -β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ -β”‚ β”‚ VAE Encode │────>β”‚ LTX Sampler │────>β”‚ VAE Decode β”‚ β”‚ -β”‚ β”‚ (to latent) β”‚ β”‚ (121 frames)β”‚ β”‚ (to images) β”‚ β”‚ -β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - v - new_images (121 frames) - β”‚ - └──────────────────────────┐ - β”‚ -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€v──────────────────────────┐ -β”‚ Extension Module β”‚ -β”‚ β”‚ -β”‚ source_images (121) + new_images (121) β”‚ -β”‚ β”‚ β”‚ -β”‚ v β”‚ -β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ -β”‚ β”‚ Overlap Extraction β”‚ β”‚ -β”‚ β”‚ Last 25 from source β”‚ β”‚ -β”‚ β”‚ First 25 from new β”‚ β”‚ -β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ -β”‚ β”‚ β”‚ -β”‚ v β”‚ -β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ -β”‚ β”‚ Blending β”‚ β”‚ -β”‚ β”‚ Mode: linear_blend β”‚ β”‚ -β”‚ β”‚ Alpha: 0β†’1 over 25 β”‚ β”‚ -β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ -β”‚ β”‚ β”‚ -β”‚ v β”‚ -β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ -β”‚ β”‚ Concatenation β”‚ β”‚ -β”‚ β”‚ [prefix][blend][suffix]β”‚ β”‚ -β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ -β”‚ β”‚ β”‚ -β”‚ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ -β”‚ β”‚ β”‚ β”‚ -β”‚ v v β”‚ -β”‚ extended_images (217) start_images (17, 8n+1) β”‚ -β”‚ β”‚ β”‚ β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ β”‚ - v └─> Next Iteration - β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” - β”‚ CreateVideo β”‚ - β”‚ Concatenate β”‚ - β”‚ with Audio β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - v - β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” - β”‚ SaveVideo β”‚ - β”‚ Final Output β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ -``` - ---- - -### Overlap Blending Visualization - -``` -Source Batch (121 frames): -[β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“β–“] - └─ Last 25 frames β”€β”˜ - -New Batch (121 frames): - [β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ] - └─ First 25 frames β”€β”˜ - -Blending Zone (25 frames with linear alpha): -Frame: 1 2 3 4 5 ... 23 24 25 -Alpha: 0.00 0.04 0.08 0.12 0.16 ... 0.92 0.96 1.00 - β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–“ β–ˆβ–ˆβ–ˆβ–’ β–ˆβ–ˆβ–’β–‘ β–ˆβ–ˆβ–‘β–‘ ... β–‘β–’β–ˆβ–ˆ β–‘β–“β–ˆβ–ˆβ–ˆ β–‘β–ˆβ–ˆβ–ˆ - -Blended: (1-Ξ±)Γ—source + Ξ±Γ—new - -Extended Result (217 frames): -[β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–“β–“β–’β–’β–‘β–‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ] - └─ Smooth transition β”€β”˜ -``` - ---- - -## Conclusion - -The Extension Module provides a powerful, flexible solution for iterative video generation with LTX-2. Key takeaways: - -1. **Always use 8n+1 conformance** (`ltx2_round_down`) when feeding samplers -2. **Start with overlap=25** and `linear_blend` for best results -3. **Enable quality features** (color match, seam search) only when needed -4. **Monitor the report output** to verify correct operation -5. **Test with simple configs first**, then optimize - -For support and updates, see the [IAMCCS-nodes repository](https://github.com/IAMCCS/IAMCCS-nodes). - ---- - -*Document Version: 1.0* -*Last Updated: January 2026* -*Extension Module Version: 87665e5* diff --git a/docs/LTX2_EXTENSION_NODES_GUIDE_EN.md b/docs/LTX2_EXTENSION_NODES_GUIDE_EN.md deleted file mode 100644 index 795ee80..0000000 --- a/docs/LTX2_EXTENSION_NODES_GUIDE_EN.md +++ /dev/null @@ -1,394 +0,0 @@ -# IAMCCS LTX-2 Extension Nodes β€” Final Guide (EN) - -This document explains how to use the IAMCCS LTX-2 nodes for **long-length / multi-segment video generation and extension** in ComfyUI, including the purpose of each widget and recommended usage patterns. - -## What problems these nodes solve - -1. **Seam artifacts between segments** (visible cut, flicker, exposure shift) -2. **Bad seam position** (the extension starts at an awkward frame) -3. **LTX VideoVAE frame-count constraint**: some encode paths require the number of guide frames to be of the form: - -$$N = 1 + 8k$$ - -4. **Workflow simplification**: reduce reliance on multiple helper nodes for overlap math, ranges, etc. - ---- - -## Quick decision guide (what to touch first) - -### If you get the LTX VideoVAE error: β€œEncode input must have 1 + 8 * x frames” - -This is the **`8n+1`** rule: the number of frames going into certain LTX/LTXV encode paths must be: - -$$N = 1 + 8k$$ - -In iterative extension workflows, this usually affects **the guide/start frames** you feed into the next segment. - -Use these fixes in this order: - -1) **Set `safe_mode = native_workflow_safe`** in `IAMCCS_LTX2_ExtensionModule` - - This extracts the start frames exactly like the original stable workflow: - - `start_images = extended_images[-overlap_frames:-1]` - -2) If you still need a strict `8n+1` count, set **`start_frames_rule`**: - - `ltx2_round_down`: most predictable and β€œnever increases” the frame count. - - `ltx2_nearest`: useful if you want the closest valid count (may go up or down). - -3) If the frame rule is needed elsewhere (not on the extension module), use `IAMCCS_LTX2_FrameCountValidator` on the integer driving that node. - -### If the seam is visible (hard cut / flicker) - -- Start with: - - `overlap_mode = ease_in_out` (or `linear_blend` if you want the simplest behavior) - - keep overlap modest (common range: ~8–24 frames; larger overlap can help but costs compute/time) - -### If exposure/white balance shifts at the seam - -- Enable color matching: - - `color_match_mode = luma_only` (usually the safest) - - `color_match_strength = 0.3..0.7` - - `color_reference_window = 6..12` - -### If the seam β€œrestarts weirdly” (bad timing / rewind) - -- Use seam search (try only after overlap/blend): - - `seam_search_mode = best_of_k` - - `k_search = 8..24` - -### If you are using AutoLink overlap loops - -- Prefer wiring `autolink_overlap_in` / `autolink_overlap_out` so each iteration can override overlap cleanly. - -### About `IAMCCS_LTX2_ExtensionModule_simple` - -- `IAMCCS_LTX2_ExtensionModule_simple` is the **minimal** variant of the Extension Module. -- It exposes only the core overlap/blend/math widgets (no color match, seam search, metrics). -- It does **not** expose `safe_mode` or `start_frames_rule` as widgets. -- It **always enforces** the LTX-2 start-frame rule $N = 1 + 8k$ automatically (round-down), to avoid VideoVAE encode frame-count errors. - -## Nodes overview - -- `IAMCCS_LTX2_ExtensionModule` - - Merges the previous segment (`source_images`) with the new segment (`new_images`) using overlap/blend. - - Outputs `extended_images` (merged batch) and `start_images` (frames used to guide the next segment). - - Optional seam improvements: exposure/color matching and best-of-k seam selection. - - Optional β€œnative safe” extraction that matches the original stable workflow behavior. - - Optional AutoLink overlap loop I/O: - - `autolink_overlap_in` (override overlap when > 0) - - `autolink_overlap_out` (feed the next iteration) - -- `IAMCCS_LTX2_GetImageFromBatch` - - Extracts frames from the start/end of an image batch, or by an explicit range. - - Adds optional auto-count and diagnostics outputs. - - Optional β€œnative safe” mode matching `images[-count:-1]` in from-end mode. - -- `IAMCCS_LTX2_ReferenceImageSwitch` - - Safe way to inject a **reference image** to improve identity/style consistency **without breaking overlap continuity**. - - Default is `none`, so existing workflows are unchanged. - -- `IAMCCS_LTX2_ReferenceStartFramesInjector` - - (New) Injects/blends the reference directly into the **guide/conditioning frames** (`start_images` / segment `images`). - - Useful when feeding the reference into `image_1` (empty latent image) has **weak or no identity effect**. - - Can be applied to **only one segment** (e.g. segment 3 only). - -- `IAMCCS_LTX2_FrameCountValidator` - - Helper to validate/correct an integer frame count to the `1 + 8*k` rule. - ---- - -## 1) IAMCCS_LTX2_ExtensionModule - -### Inputs - -**Required** - -- `source_images` (IMAGE) - - The current accumulated batch (previous segment output). -- `overlap_frames` (INT) - - How many frames overlap between segments. -- `overlap_side` (dropdown) - - `source`: overlap uses the tail of `source_images` against the head of `new_images`. - - `new_images`: swaps which side is treated as source/destination for blending. -- `overlap_mode` (dropdown) - - `cut`: hard cut (fastest, most visible seam). - - `linear_blend`: linear crossfade. - - `ease_in_out`: smoother crossfade. - - `filmic_crossfade`: gamma-aware blend (often smoother in highlights). - - `perceptual_crossfade`: LAB blend via Kornia (falls back if Kornia not installed). -- `enable_math` (BOOLEAN) - - Enables the built-in β€œhow many start frames to output” calculation. -- `math_operation` (dropdown) - - Applies to `overlap_frames` (as `a`) and `math_value_b` (as `b`) when computing how many frames to output as `start_images`. - - Typical: `a-b` or `a-1`. - -**Safety / LTX rule** - -- `safe_mode` (dropdown) - - `none`: uses the node’s normal start-images logic. - - `native_workflow_safe`: extracts start images exactly like the proven stable graph: - - `start_images = extended_images[-overlap_frames:-1]` - - Use this if you are hitting the LTX VideoVAE error β€œEncode input must have 1 + 8 * x frames”. - -- `start_frames_rule` (dropdown) - - `none`: do not modify the calculated number of start frames. - - `ltx2_round_down`: force the count down to the nearest valid `1 + 8*k`. - - `ltx2_nearest`: choose the nearest valid `1 + 8*k` within bounds. - - Use this when a downstream node (VideoVAE encode/guide) requires `1 + 8*k` frame counts. - -**Quality upgrades (defaults are safe/off)** - -- `color_match_mode` (dropdown) - - `none`: no change (original behavior). - - `luma_only`: match exposure/contrast on luma. - - `per_channel`: match mean/std per RGB channel. -- `color_match_strength` (FLOAT 0..1) - - Blend between original and matched. -- `color_reference_window` (INT) - - Number of frames used from tail/head for statistics. - -- `seam_search_mode` (dropdown) - - `none`: no seam search. - - `best_of_k`: search for a better seam by testing candidate offsets. -- `k_search` (INT) - - How many candidate offsets to test (0 disables). -- `metric_weight_color` (FLOAT) - - Weight of luma continuity in the seam score. -- `metric_weight_edges` (FLOAT) - - Weight of edge continuity in the seam score. - -**Optional** - -- `new_images` (IMAGE) - - The newly generated segment. - - If omitted, the node can be used as a β€œprep” node (it will still output `start_images` from the current batch). -- `math_value_b` (INT) - - Used by `math_operation`. - -### Outputs - -- `source_images` (IMAGE) β€” passthrough -- `start_images` (IMAGE) β€” frames to feed as guide for the next segment -- `extended_images` (IMAGE) β€” merged batch -- `overlap_frames` (INT) -- `calculated_frames` (INT) β€” actual number of frames output in `start_images` -- `extension_frames` (INT) β€” how many frames were added -- `report` (STRING) - -### Recommended settings - -- Most stable: `safe_mode = native_workflow_safe`, `overlap_mode = ease_in_out` (or `linear_blend`) -- If you see exposure shift: `color_match_mode = luma_only`, `strength = 0.3..0.7` -- If you see weird seam timing: `seam_search_mode = best_of_k`, `k_search = 8..24` - ---- - -## 2) IAMCCS_LTX2_GetImageFromBatch - -### Purpose -A small helper to extract frames for the next segment or for debugging. - -### Inputs - -- `images` (IMAGE) -- `mode` (dropdown) - - `from_start`: take the first `count` frames - - `from_end`: take the last `count` frames - - `range`: take `[start_index:end_index)` -- `count` (INT) - -**Upgrades** - -- `auto_count_mode` (dropdown) - - `none`: use `count` widget. - - `prefer_input`: use `count_in` if connected. - - `use_widget`: explicitly use the widget value. -- `diagnostics` (dropdown) - - `none`: normal behavior. - - `basic`: exposes `start_index` and `end_index` outputs. - -**Safety / LTX rule** - -- `count_rule` (dropdown) - - `none` / `ltx2_round_down` / `ltx2_nearest` for `1 + 8*k`. -- `safe_mode` (dropdown) - - `none`: normal extraction. - - `native_workflow_safe`: for `from_end` uses `images[-count:-1]`. - -**Optional** - -- `count_in` (INT) -- `start_index` / `end_index` (INT) for `range` mode. - -### Outputs - -- `images` (IMAGE) -- `count` (INT) -- `report` (STRING) -- `start_index`, `end_index` (INT) - ---- - -## 3) IAMCCS_LTX2_ReferenceImageSwitch - -### Why this node exists -In long-length generation, you typically want: -- **Continuity** driven by overlap/start frames -- **Identity/style consistency** reinforced by a stable reference image - -This node lets you add a reference image **without replacing** the overlap continuity input. - -### Inputs - -- `default_image` (IMAGE) - - What the workflow already used before (pass-through by default). -- `mode` (dropdown) - - `none`: output `default_image` (fully backward-compatible). - - `use_reference`: output `reference_image`. - - `blend`: output mix of `default_image` and `reference_image`. -- `blend_strength` (FLOAT) - - Only for `blend` mode. -- `reference_image` (optional IMAGE) - - If not connected, the node behaves like `none`. - -### Output - -- `image` (IMAGE) -- `report` (STRING) - -### Practical usage - -- Insert it on the **auxiliary** image input of your segment sampler (often called `image_1`). -- Keep overlap/start frames connected exactly as before. -- If you enable `use_reference`/`blend`, the reference is **automatically resized** to match `default_image` (more stable for downstream nodes). - -Note: in many LTX/LTXV workflows, feeding the reference into `image_1` (empty latent image) may not be enough to β€œlock” identity when a face is revealed later in the segment. In that case, use the node below. - ---- - -## 3b) IAMCCS_LTX2_ReferenceStartFramesInjector - -### Why it exists -If identity drifts even with a reference, it often means the reference is connected to an input that the model barely uses. This node modifies the actual guide/conditioning frames. - -### Inputs - -- `start_images` (IMAGE) - - The guide frames that feed the segment (typically `start_images` from the extension module, or the sampler’s `images` input). -- `mode` - - `none`: passthrough. - - `inject`: replaces the selected frames with the reference. - - `blend`: mixes reference and original frames. -- `blend_strength` (0..1) - - Only used for `blend` (0 = no effect, 1 = full reference). In `inject` it behaves like 1. -- `frames_to_inject` (INT) - - How many guide frames to modify. -- `ramp` (BOOLEAN) - - If `true`, applies a gradual ramp across the injected frames. -- `position` - - `tail`: last K frames (usually best, closest to the seam). - - `head`: first K frames. -- `reference_image` (optional IMAGE) - - Usually the output of `IAMCCS_LTX2_ReferenceImageSwitch`. - -### Outputs - -- `start_images` (IMAGE) -- `report` (STRING) - -### Recommended starter settings - -- If identity is not sticking but you want to preserve continuity: - - `mode = blend` - - `frames_to_inject = 3..6` - - `blend_strength = 0.5..0.85` - - `ramp = true` - - `position = tail` - -If you see seam discontinuity, lower `blend_strength` and/or reduce `frames_to_inject`. - ---- - -## How to decide when/where to use a reference - -Quick checklist: - -1. **Is the face/identity visible in the first frames of the segment?** - - Yes β†’ a reference can work well. - - No (reveal happens mid/late segment) β†’ the reference may have little leverage: consider cutting segments so the reveal starts at the segment boundary, or use `ReferenceStartFramesInjector` (and/or dedicated tools like FaceID/IPAdapter if compatible). - -2. **What are you stabilizing?** - - Style / global look β†’ `ReferenceImageSwitch` (or `color_match_mode` in ExtensionModule) is often enough. - - Identity (specific face) β†’ `ReferenceStartFramesInjector` is more likely required. - -3. **Where to wire it?** - - `image_1` / empty latent image: can be a hint, not guaranteed. - - `images` / start frames (conditioning): highest impact. - -4. **How to limit it to one segment (e.g. segment 3 only)** - - Place `ReferenceStartFramesInjector` only in the path feeding that segment’s `images` / `start_images`. - - Leave other segments untouched (no injector). - -## 4) IAMCCS_LTX2_FrameCountValidator - -### Inputs - -- `frame_count` (INT) -- `auto_correct` (BOOLEAN) -- `correction_mode` (`nearest` / `round_up` / `round_down`) - -### Outputs - -- `validated_count` (INT) -- `is_valid` (BOOLEAN) -- `nearest_valid` (INT) -- `report` (STRING) - ---- - -## Common workflows / use cases - -### A) Long-length extension (multi segment) -1. Generate segment 1. -2. Use `IAMCCS_LTX2_ExtensionModule` to compute `start_images` and merge segments. -3. Feed `start_images` into the next segment guide/conditioning. -4. Repeat. - -Recommended: enable `safe_mode = native_workflow_safe` if you see LTX frame-count errors. - -### B) Reduce seams -- Prefer `ease_in_out` or `filmic_crossfade`. -- Use `color_match_mode` if you see exposure shifts. -- Use `best_of_k` seam search if the seam starts at a bad moment. - -### C) Improve identity consistency -- Add `IAMCCS_LTX2_ReferenceImageSwitch` to `image_1`. -- Connect a single reference image and set mode to `blend` (start at 0.2..0.4). - ---- - -## Troubleshooting - -- **β€œIAMCCS_LTX2_ReferenceImageSwitch not found”** - - Ensure you updated the IAMCCS nodes and restart ComfyUI. - - The node must be exported in the package registry (`__init__.py`). - -- **β€œEncode input must have 1 + 8 * x frames”** - - Use `safe_mode = native_workflow_safe` or set `start_frames_rule/count_rule` to enforce `1 + 8*k`. - -- **Border motion artifacts (edge warping / flicker)** - - Note: `metric_weight_edges` and `best_of_k` improve seam selection *inside the overlap* between segments; they do not automatically β€œfix” frame borders. - - Common improvements: - - Avoid changing resize/crop between segments; keep one resolution end-to-end. - - Prefer β€œclean” resolutions (multiples of 64 where possible) to reduce VAE boundary artifacts. - - Quick workaround: apply a small crop (e.g., 8–16 px per side) then resize back. - - Helpful nodes (IAMCCS): - - `IAMCCS_LTX2_ImageBatchPadReflect`: adds a reflect border (increases resolution). - - `IAMCCS_LTX2_ImageBatchCropByPad`: removes that border (back to target resolution). - - Recommended usage (when you want the model to have more border context): - - Pick `pad_x/pad_y` (e.g., 16). - - Generate at a higher resolution: `W_pad = W + 2*pad_x`, `H_pad = H + 2*pad_y` (including `EmptyImage`). - - If you have β€œinitial”/reference images at the old resolution, run them through `PadReflect` to reach `W_pad x H_pad`. - - At the end (before `CreateVideo`), run `CropByPad` with the same `pad_x/pad_y` to return to `W x H`. - -- **Reference image causes a resolution error** - - Resize/crop the reference to match your workflow resolution before feeding it. diff --git a/docs/WanImageMotion.md b/docs/WanImageMotion.md deleted file mode 100644 index 0426597..0000000 --- a/docs/WanImageMotion.md +++ /dev/null @@ -1,136 +0,0 @@ -# IAMCCS WanImageMotion - -`IAMCCS_WanImageMotion` is a **drop-in replacement** for the SVIPro latent-conditioning node used in WAN image-to-video workflows. Its purpose is to build the *conditioning* fields required by the WAN I2V pipeline while optionally boosting perceived motion via a controllable `motion` parameter. - -This node **does not perform sampling**. It only: -- prepares an β€œempty” latent sequence to be denoised by the sampler, and -- injects `concat_latent_image` and `concat_mask` into both positive/negative conditioning. - ---- - -## Inputs - -Required: -- `positive` / `negative` (`CONDITIONING`): conditioning streams to be augmented. -- `length` (`INT`): number of frames in the video. Internally converted to latent-frame count: - $T = \left\lfloor\frac{length-1}{4}\right\rfloor + 1$. -- `anchor_samples` (`LATENT`): the β€œanchor” latent(s), typically representing the initial visual content. -- `motion_latent_count` (`INT`): how many latent frames to take from `prev_samples` (if present) to seed motion. -- `motion` (`FLOAT`): motion amplification factor. `1.0` means β€œno change”. Values > `1.0` increase motion. -- `motion_mode` (dropdown): chooses *where* the motion boost is applied. -- `latent_precision` (dropdown): controls the dtype used for the **empty latent** allocation (quality vs VRAM). - - `auto`: matches anchor samples dtype - - `fp16`: half precision (lower VRAM, slight quality loss) - - `fp32`: full precision (higher VRAM, maximum quality) -- `vram_profile` (dropdown): chooses *how* the motion boost is computed to reduce peak VRAM. - - `normal`: process all frames at once (fastest, highest VRAM) - - `chunked_blocks_2` / `chunked_blocks_4`: process in chunks (balanced) - - `loop_per_frame (lowest_vram)`: process one frame at a time - - `cpu_offload (slowest)`: offload computation to CPU (extreme low VRAM) -- `include_padding_in_motion` (`BOOLEAN`): if enabled, the motion boost may also affect padded latent frames. - - **Critical for single-frame anchors**: when `anchor_samples` has only `T=1` and there are no `prev_samples`, this must be `True` to apply any motion boost. - - The node will log a warning if motion_range is empty and suggest enabling this option. - -Optional: -- `prev_samples` (`LATENT`): previous latent sequence; when provided, the last `motion_latent_count` latent frames are appended after the anchor to seed motion. - ---- - -## Outputs - -- `positive` / `negative` (`CONDITIONING`): same as input, but with added conditioning keys: - - `concat_latent_image` - - `concat_mask` -- `latent` (`LATENT`): an **empty latent sequence** shaped like the target video latents. This is what the sampler will denoise. - ---- - -## Core Logic - -### 1) Create the empty latent sequence -The node allocates an empty latent tensor with shape: -- `[B, 16, T, H, W]` where `T` is derived from `length`. - -This tensor is intentionally initialized to zeros. - -`latent_precision` affects only this allocation: -- `auto`: matches the dtype of `anchor_samples` (recommended). -- `fp16`: forces FP16 (lower VRAM, can be slightly less stable). -- `fp32`: forces FP32 (higher VRAM, can be slightly more stable). - -### 2) Build `concat_latent_image` -The node builds a latent conditioning sequence (`image_cond_latent`) by concatenating: -1. `anchor_samples["samples"]` (anchor latents) -2. the last `motion_latent_count` frames from `prev_samples["samples"]` (only if provided) -3. zero padding to reach exactly `T` latent frames - -Padding is processed with `Wan21().process_out(...)` to match expected latent formatting. - -### 3) Build `concat_mask` -A mask is created with shape `[1, 1, T, H, W]`. -- The first latent frame is unmasked: `mask[:, :, :1] = 0.0` -- All subsequent latent frames are masked: `1.0` - -### 4) Inject into conditioning -The node injects: -- `concat_latent_image = image_cond_latent` -- `concat_mask = mask` - -into **both** `positive` and `negative` conditioning. - ---- - -## Motion Boost (`motion`) - -When `motion > 1.0`, the node amplifies motion by modifying selected latent frames while preserving the per-frame mean offset to reduce brightness/shift artifacts. - -Let: -- `base` be the first latent frame `image_cond_latent[:, :, 0:1]` -- `x` be the target latent frames to be modified - -The transformation is: -1. `diff = x - base` -2. `mean = mean(diff over C,H,W)` (per-batch/per-time) -3. `diff_centered = diff - mean` -4. `scaled = base + diff_centered * motion + mean` -5. clamp to a safe range: `[-6, 6]` - -By default, the node **does not modify padding frames**. - -If `include_padding_in_motion = true`, the node may treat padded frames as motion targets. This can help when `anchor_samples` provides only a single latent frame (e.g. `T=1`) and there are no motion latents from `prev_samples`. - ---- - -## Motion Mode (two modes) - -### `motion_only (prev_samples)` -- Applies the motion boost **only** to the latent frames coming from `prev_samples`. -- Conservative: changes less of the anchor content. -- Recommended when you want motion injection without destabilizing the initial anchor. - -### `all_nonfirst (anchor+motion)` -- Applies the motion boost to **all real latent frames except the first** (anchor + motion latents). -- More aggressive: stronger motion effect, but can change the look more. - ---- - -## VRAM Profile - -These profiles only change *how the motion boost is computed* (peak memory vs speed). They do not change the rest of the pipeline. - -- `normal`: processes the selected time range in one tensor block (fastest, highest peak VRAM). -- `chunked_blocks_2`: processes 2 latent frames at a time (lower peak VRAM). -- `chunked_blocks_4`: processes 4 latent frames at a time (middle ground). -- `loop_per_frame (lowest_vram)`: processes 1 latent frame at a time (lowest peak VRAM, slower). -- `cpu_offload (slowest)`: moves the targeted slice to CPU for the computation, then copies back (lowest GPU peak, highest runtime cost). - ---- - -## Notes / Troubleshooting - -- If you are hitting CUDA OOM at high resolutions, try: - 1) `vram_profile = chunked_blocks_2` - 2) then `loop_per_frame (lowest_vram)` - 3) then (only if necessary) `cpu_offload (slowest)` - -- If you want to isolate whether OOM is caused by motion scaling vs sampling, set `motion = 1.0` temporarily. diff --git a/iamccs_ltx2_extension_module.py b/iamccs_ltx2_extension_module.py index e3fabf1..cc7a218 100644 --- a/iamccs_ltx2_extension_module.py +++ b/iamccs_ltx2_extension_module.py @@ -1141,6 +1141,181 @@ class IAMCCS_LTX2_ExtensionModule_simple(IAMCCS_LTX2_ExtensionModule): ) +class IAMCCS_LTX2_FirstLastFramesController: + """ + First-Last Frame (FLF) controller for LTX-2 image conditioning. + + Injects a reference first_frame and/or last_frame directly into the + `images` conditioning tensor used by the sampler. Works on the + 'MISTO' pattern: the tensor already contains both external images and + generated frames β€” this node simply overwrites / blends the head and/or + tail K frames with the supplied references. + + Modes + ----- + hard_lock : replace the K frames completely with the reference + linear_blend: weighted blend (reference * strength + original * (1-strength)) + ramp : progressive blend, strength ramps from 0 β†’ strength over K frames + (for head: 0β†’strength left-to-right; for tail: strengthβ†’0 left-to-right) + + Positions + --------- + head : operate on first K frames only + tail : operate on last K frames only + both : operate on both ends simultaneously + """ + + @classmethod + def INPUT_TYPES(cls): + return { + "required": { + "images": ("IMAGE", { + "tooltip": "Conditioning image batch (the 'images' input to the sampler)" + }), + "k_frames": ("INT", { + "default": 4, + "min": 1, + "max": 64, + "step": 1, + "tooltip": "Number of frames to affect at each injection site" + }), + "mode": (["hard_lock", "linear_blend", "ramp"], { + "default": "hard_lock", + "tooltip": ( + "hard_lock: full replace | " + "linear_blend: uniform blend at given strength | " + "ramp: progressive blend from 0 to strength" + ), + }), + "position": (["head", "tail", "both"], { + "default": "both", + "tooltip": "Where to inject references (head=first K, tail=last K, both=head+tail)", + }), + "blend_strength": ("FLOAT", { + "default": 1.0, + "min": 0.0, + "max": 1.0, + "step": 0.05, + "tooltip": "Max blend weight (ignored for hard_lock which always uses 1.0)" + }), + }, + "optional": { + "first_frame": ("IMAGE", { + "tooltip": "Reference image to inject at the HEAD of the batch (ignored if position=tail)" + }), + "last_frame": ("IMAGE", { + "tooltip": "Reference image to inject at the TAIL of the batch (ignored if position=head)" + }), + }, + } + + RETURN_TYPES = ("IMAGE", "STRING") + RETURN_NAMES = ("images", "report") + FUNCTION = "apply" + CATEGORY = "IAMCCS/LTX-2" + + # ------------------------------------------------------------------ + # helpers + # ------------------------------------------------------------------ + @staticmethod + def _resize_to(image: torch.Tensor, target_h: int, target_w: int) -> torch.Tensor: + """Resize image tensor [N,H,W,C] to (target_h, target_w).""" + if int(image.shape[1]) == target_h and int(image.shape[2]) == target_w: + return image + x = image.permute(0, 3, 1, 2) + x = F.interpolate(x.float(), size=(target_h, target_w), mode="bilinear", align_corners=False) + return x.permute(0, 2, 3, 1).clamp(0.0, 1.0).to(image.dtype) + + @staticmethod + def _broadcast_ref(ref: torch.Tensor, k: int) -> torch.Tensor: + """Ensure ref has exactly k frames (repeat single-frame or crop).""" + n = int(ref.shape[0]) + if n == k: + return ref + if n == 1: + return ref.repeat(k, 1, 1, 1) + return ref[:k] + + @staticmethod + def _blend_weights(k: int, mode: str, max_s: float, ramp_direction: str) -> list: + """ + Returns list of k blend weights. + ramp_direction: 'up' = 0β†’max_s, 'down' = max_sβ†’0 + """ + if mode == "hard_lock": + return [1.0] * k + if mode == "linear_blend": + return [max_s] * k + # ramp + if k == 1: + return [max_s] + if ramp_direction == "up": + return [max_s * float(i + 1) / float(k) for i in range(k)] + else: # down + return [max_s * float(k - i) / float(k) for i in range(k)] + + def _inject( + self, + out: torch.Tensor, + ref: torch.Tensor, + idxs: list, + weights: list, + ) -> torch.Tensor: + """Blend ref frames into out at given indices with given per-frame weights.""" + h, w = int(out.shape[1]), int(out.shape[2]) + ref_r = self._resize_to(ref, h, w) + ref_r = self._broadcast_ref(ref_r, len(idxs)) + for j, i in enumerate(idxs): + s = float(weights[j]) + out[i] = ((1.0 - s) * out[i].float() + s * ref_r[j].float()).clamp(0.0, 1.0).to(out.dtype) + return out + + # ------------------------------------------------------------------ + # main + # ------------------------------------------------------------------ + def apply( + self, + images: torch.Tensor, + k_frames: int, + mode: str, + position: str, + blend_strength: float, + first_frame: Optional[torch.Tensor] = None, + last_frame: Optional[torch.Tensor] = None, + ): + total = int(images.shape[0]) + k = max(1, min(int(k_frames), total // 2 if total > 1 else 1)) + max_s = 1.0 if mode == "hard_lock" else float(max(0.0, min(1.0, blend_strength))) + + out = images.clone() + ops = [] + + do_head = position in ("head", "both") + do_tail = position in ("tail", "both") + + if do_head and first_frame is not None: + idxs = list(range(0, k)) + # ramp up: 0 β†’ max_s (anchor gets full weight at the end) + weights = self._blend_weights(k, mode, max_s, "up") + out = self._inject(out, first_frame, idxs, weights) + ops.append(f"head(k={k},mode={mode},s={max_s:.2f})") + + if do_tail and last_frame is not None: + idxs = list(range(total - k, total)) + # ramp down: max_s β†’ 0 (anchor gets full weight at the start) + weights = self._blend_weights(k, mode, max_s, "down") + out = self._inject(out, last_frame, idxs, weights) + ops.append(f"tail(k={k},mode={mode},s={max_s:.2f})") + + if not ops: + report = f"FLF Controller: no-op (position={position}, first_frame={'yes' if first_frame is not None else 'no'}, last_frame={'yes' if last_frame is not None else 'no'})" + else: + report = "FLF Controller: " + " + ".join(ops) + f" | total_frames={total}" + + _log.debug(report) + return (out, report) + + # Node registration NODE_CLASS_MAPPINGS = { "IAMCCS_LTX2_ExtensionModule": IAMCCS_LTX2_ExtensionModule, @@ -1149,6 +1324,7 @@ NODE_CLASS_MAPPINGS = { "IAMCCS_LTX2_ReferenceImageSwitch": IAMCCS_LTX2_ReferenceImageSwitch, "IAMCCS_LTX2_ReferenceStartFramesInjector": IAMCCS_LTX2_ReferenceStartFramesInjector, "IAMCCS_LTX2_FrameCountValidator": IAMCCS_LTX2_FrameCountValidator, + "IAMCCS_LTX2_FirstLastFramesController": IAMCCS_LTX2_FirstLastFramesController, } NODE_DISPLAY_NAME_MAPPINGS = { @@ -1158,4 +1334,5 @@ NODE_DISPLAY_NAME_MAPPINGS = { "IAMCCS_LTX2_ReferenceImageSwitch": "LTX-2 Reference Image Switch 🧷", "IAMCCS_LTX2_ReferenceStartFramesInjector": "LTX-2 Inject Reference Into Start Frames 🧬", "IAMCCS_LTX2_FrameCountValidator": "LTX-2 Frame Count Validator βœ… (8n+1)", + "IAMCCS_LTX2_FirstLastFramesController": "LTX-2 First-Last Frames Controller 🎯", } diff --git a/iamccs_qwen_vl_flf.py b/iamccs_qwen_vl_flf.py new file mode 100644 index 0000000..8d80fc0 --- /dev/null +++ b/iamccs_qwen_vl_flf.py @@ -0,0 +1,524 @@ +# ========================================================== +# iamccs_qwen_vl_flf.py β€” IAMCCS QwenVL First/Last Frame +# ========================================================== +# Dual-image QwenVL node: accepts a FIRST FRAME and a LAST FRAME, +# then queries QwenVL to describe the motion/action occurring +# between the two frames β€” the ideal prompt for FLF video generators +# (WAN SVI Pro, LTX-2 FLF, etc.). +# +# This is a 1:1 extension of AILab_QwenVL (ComfyUI-QwenVL) +# with the image input replaced by two independent IMAGE inputs. +# +# Author : IAMCCS (carminecristalloscalzi.com / patreon.com/IAMCCS) +# License: GPL-3.0 +# ========================================================== + +import importlib +import os +import sys +from pathlib import Path + +import numpy as np +import torch + + +# --------------------------------------------------------------------------- +# Dynamic import of QwenVLBase from the ComfyUI-QwenVL custom node +# --------------------------------------------------------------------------- + +def _import_qwen_base(): + """Locate and import QwenVLBase from ComfyUI-QwenVL, however it was loaded.""" + + # 1) Already loaded by ComfyUI's module system? + for module_name, module in sys.modules.items(): + if "AILab_QwenVL" in module_name: + if hasattr(module, "QwenVLBase"): + return module.QwenVLBase + + # 2) Look at sibling custom_node directories + this_dir = Path(__file__).resolve().parent # …/IAMCCS-nodes + custom_nodes_dir = this_dir.parent # …/custom_nodes + + candidates = [ + custom_nodes_dir / "ComfyUI-QwenVL" / "AILab_QwenVL.py", + custom_nodes_dir / "comfyui-qwenvl" / "AILab_QwenVL.py", + ] + for candidate in candidates: + if candidate.exists(): + spec = importlib.util.spec_from_file_location("AILab_QwenVL_ext", str(candidate)) + mod = importlib.util.module_from_spec(spec) + sys.modules["AILab_QwenVL_ext"] = mod + spec.loader.exec_module(mod) + return mod.QwenVLBase + + raise ImportError( + "[IAMCCS_QWEN_VL_FLF] Cannot find QwenVLBase. " + "Make sure ComfyUI-QwenVL is installed under custom_nodes/ComfyUI-QwenVL." + ) + + +# Lazy-load so the import error is surfaced only when the node is used +_QwenVLBase = None + +def _get_base(): + global _QwenVLBase + if _QwenVLBase is None: + _QwenVLBase = _import_qwen_base() + return _QwenVLBase + + +# --------------------------------------------------------------------------- +# FLF-specific prompt presets +# --------------------------------------------------------------------------- + +FLF_PRESET_PROMPTS = [ + "🎬 Video Action Description (FLF)", + "πŸŽ₯ Cinematic Motion Prompt (FLF)", + "πŸƒ Subject Movement & Camera (FLF)", + "πŸŒ€ Scene Transition Description (FLF)", + "πŸ“· Static Shot Action Prompt (FLF)", + "🌊 WAN 2.2 SVI Pro 2 β€” FLF Prompt", + "⚑ LTX-2 FLF Prompt", +] + +FLF_SYSTEM_PROMPTS = { + "🎬 Video Action Description (FLF)": ( + "You are given two images: the FIRST FRAME and the LAST FRAME of a video clip. " + "Your task is to write a single, concise video-generation prompt (2-4 sentences) that describes " + "the motion, action, and visual transformation occurring between these two frames. " + "Include: subject actions, camera movement (pan, tilt, zoom, static, etc.), environmental changes, " + "lighting shifts, and any notable visual effects. " + "Write in present tense, imperative style, as if directing an AI video generator. " + "Do NOT describe what is in the images statically β€” focus entirely on the MOTION and TRANSITION." + ), + "πŸŽ₯ Cinematic Motion Prompt (FLF)": ( + "You are given the FIRST FRAME and the LAST FRAME of a cinematic video shot. " + "Describe the complete camera movement and subject action as a professional cinematography prompt. " + "Include: shot type (close-up, wide, medium), camera movement (dolly, pan, handheld shake, etc.), " + "subject movement direction and speed, focus changes, and mood/lighting evolution. " + "Output a single fluid paragraph suitable for an AI video generator." + ), + "πŸƒ Subject Movement & Camera (FLF)": ( + "Compare the first frame and the last frame provided. " + "Write a detailed motion description focused on: " + "1) How the main subject(s) move between the two frames (direction, speed, posture changes), " + "2) Camera behavior (static, following, pulling back, zooming in/out), " + "3) Background/environment changes. " + "Summarize in 2-3 sentences optimized for AI video generation input." + ), + "πŸŒ€ Scene Transition Description (FLF)": ( + "You are shown the opening frame and the closing frame of a video sequence. " + "Analyse the differences and infer what visual narrative connects them. " + "Write a prompt that describes the scene transition: object positions, lighting evolution, " + "atmospheric changes, and any implied motion. Be specific and concise (2-3 sentences). " + "The output should work as a direct input for an AI video generator." + ), + "πŸ“· Static Shot Action Prompt (FLF)": ( + "Given the first and last frame of a static-camera video clip, " + "describe only the subject's actions and movements within the fixed frame. " + "Mention entry/exit directions, gestures, expressions, interaction with objects, " + "and any notable background activity. " + "Output a crisp 1-3 sentence prompt for an AI video generator." + ), + + "🌊 WAN 2.2 SVI Pro 2 β€” FLF Prompt": ( + "You are an AI video prompt expert for the WAN 2.2 SVI Pro 2 model in ComfyUI. " + "I will give you two images: the FIRST FRAME and the LAST FRAME of a video clip. " + "Your job is to write one single, detailed prompt in clear English " + "that describes the motion and transformation occurring between these two frames, " + "suitable for use directly with WAN 2.2 SVI Pro 2. " + "Rules: " + "Write normal sentences, not JSON, not a list. " + "Include: subject action and movement, camera motion (pan, tilt, zoom, dolly, static), " + "environment and background evolution, lighting and atmosphere changes, " + "and overall motion style (slow, fast, smooth, handheld). " + "Focus entirely on the MOTION and TRANSITION between the two frames β€” " + "do NOT describe the frames as static images. " + "Keep it under 4 sentences. " + "Do not mention these rules in your answer." + ), + + "⚑ LTX-2 FLF Prompt": ( + "You are an AI video prompt expert for the LTX-2 First/Last Frame (FLF) model in ComfyUI. " + "I will give you two images: the FIRST FRAME and the LAST FRAME of a video clip. " + "Your job is to write one single, detailed prompt in clear English " + "that describes the visual and motion continuity connecting these two frames, " + "optimised for LTX-2 FLF video generation. " + "Rules: " + "Write normal sentences, not JSON, not a list. " + "Include: subject description and action, precise camera movement, " + "spatial transitions (near-to-far, left-to-right, etc.), " + "lighting and color mood evolution between frames, " + "and motion speed/smoothness (e.g. slow drift, rapid motion, gradual zoom). " + "LTX-2 responds best to prompts that are visually rich and temporally explicit β€” " + "describe what changes and how it changes, not just what is visible. " + "Keep it under 4 sentences. " + "Do not mention these rules in your answer." + ), +} + +FLF_TOOLTIPS = { + "first_frame": "The opening frame of the video clip (frame 0).", + "last_frame": "The closing frame of the video clip (last frame).", + "preset_prompt": "Built-in FLF instruction set for QwenVL. Each preset focuses on a different aspect of motion description.", + "custom_prompt": "Optional override β€” replaces the preset completely when filled in.", + "model_name": "Pick the Qwen-VL checkpoint. First run downloads weights into models/LLM/Qwen-VL.", + "quantization": "Precision vs VRAM. FP16 = best quality; 8-bit = 8-16 GB GPUs; 4-bit = 6 GB or lower.", + "attention_mode": "auto tries SageAttention / Flash-Attn v2 and falls back to SDPA.", + "max_tokens": "Maximum tokens to generate. 256-512 is usually sufficient for motion prompts.", + "keep_model_loaded": "Keep model in VRAM after generation to skip reloading on next run.", + "seed": "Random seed β€” reuse to reproduce the same description.", + "use_torch_compile": "Enable torch.compile (reduce-overhead) on supported CUDA/Torch 2.1+ builds.", + "device": "Inference device: auto, cpu, mps, or cuda:N.", + "temperature": "Sampling randomness (when num_beams=1). 0.2-0.4 focused, 0.7+ creative.", + "top_p": "Nucleus sampling cutoff (when num_beams=1).", + "num_beams": "Beam-search width. >1 disables temperature/top_p for more stable output.", + "repetition_penalty": "Values >1 penalise repeated phrases (1.1-1.3 recommended).", +} + + +# --------------------------------------------------------------------------- +# FLF mixin β€” overrides generate() to accept two frames +# --------------------------------------------------------------------------- + +class _FLFMixin: + """Mixin that provides dual-image (first/last frame) generation.""" + + @staticmethod + def tensor_to_pil(tensor): + if tensor is None: + return None + if tensor.dim() == 4: + tensor = tensor[0] + array = (tensor.cpu().numpy() * 255).clip(0, 255).astype(np.uint8) + from PIL import Image + return Image.fromarray(array) + + @torch.no_grad() + def generate_flf( + self, + prompt_text, + first_frame, + last_frame, + max_tokens, + temperature, + top_p, + num_beams, + repetition_penalty, + ): + """Build a two-image conversation: [first_frame, last_frame, text prompt].""" + content = [] + + img1 = self.tensor_to_pil(first_frame) + img2 = self.tensor_to_pil(last_frame) + + if img1 is not None: + content.append({"type": "image", "image": img1}) + if img2 is not None: + content.append({"type": "image", "image": img2}) + + content.append({"type": "text", "text": prompt_text}) + + conversation = [{"role": "user", "content": content}] + + chat = self.processor.apply_chat_template( + conversation, tokenize=False, add_generation_prompt=True + ) + images = [item["image"] for item in content if item["type"] == "image"] + processed = self.processor( + text=chat, + images=images or None, + videos=None, + return_tensors="pt", + ) + + model_device = next(self.model.parameters()).device + model_inputs = { + k: v.to(model_device) if torch.is_tensor(v) else v + for k, v in processed.items() + } + + stop_tokens = [self.tokenizer.eos_token_id] + if hasattr(self.tokenizer, "eot_id") and self.tokenizer.eot_id is not None: + stop_tokens.append(self.tokenizer.eot_id) + + kwargs = { + "max_new_tokens": max_tokens, + "repetition_penalty": repetition_penalty, + "num_beams": num_beams, + "eos_token_id": stop_tokens, + "pad_token_id": self.tokenizer.pad_token_id, + } + if num_beams == 1: + kwargs.update({"do_sample": True, "temperature": temperature, "top_p": top_p}) + else: + kwargs["do_sample"] = False + + outputs = self.model.generate(**model_inputs, **kwargs) + if torch.cuda.is_available(): + torch.cuda.synchronize() + + input_len = model_inputs["input_ids"].shape[-1] + text = self.tokenizer.decode(outputs[0, input_len:], skip_special_tokens=True) + return text.strip() + + def run_flf( + self, + model_name, + quantization, + preset_prompt, + custom_prompt, + first_frame, + last_frame, + max_tokens, + temperature, + top_p, + num_beams, + repetition_penalty, + seed, + keep_model_loaded, + attention_mode, + use_torch_compile, + device, + ): + from comfy.utils import ProgressBar + pbar = ProgressBar(3) + + torch.manual_seed(seed) + prompt = FLF_SYSTEM_PROMPTS.get(preset_prompt, preset_prompt) + if custom_prompt and custom_prompt.strip(): + prompt = custom_prompt.strip() + + pbar.update_absolute(1, 3, None) + + self.load_model( + model_name, + quantization, + attention_mode, + use_torch_compile, + device, + keep_model_loaded, + ) + + pbar.update_absolute(2, 3, None) + + try: + text = self.generate_flf( + prompt, + first_frame, + last_frame, + max_tokens, + temperature, + top_p, + num_beams, + repetition_penalty, + ) + pbar.update_absolute(3, 3, None) + return (text,) + finally: + if not keep_model_loaded: + self.clear() + + +# --------------------------------------------------------------------------- +# Node class factory (deferred because QwenVLBase is lazy-loaded) +# --------------------------------------------------------------------------- + +def _build_node_classes(): + """Return (IAMCCS_QWEN_VL_FLF, IAMCCS_QWEN_VL_FLF_Advanced) after QwenVLBase loads.""" + Base = _get_base() + + # Import Quantization enum from the same module as Base + import sys + qwen_mod = sys.modules.get("AILab_QwenVL") or sys.modules.get("AILab_QwenVL_ext") + if qwen_mod is None: + # The module might be registered under a different key + for k, v in sys.modules.items(): + if "AILab_QwenVL" in k and hasattr(v, "Quantization"): + qwen_mod = v + break + if qwen_mod is None: + raise ImportError("[IAMCCS_QWEN_VL_FLF] Could not locate Quantization enum in QwenVL module.") + + Quantization = qwen_mod.Quantization + ATTENTION_MODES = qwen_mod.ATTENTION_MODES + HF_VL_MODELS = qwen_mod.HF_VL_MODELS + + # ------------------------------------------------------------------ + # Simple version + # ------------------------------------------------------------------ + class IAMCCS_QWEN_VL_FLF(_FLFMixin, Base): + """QwenVL node with FIRST FRAME + LAST FRAME inputs for FLF video generation.""" + + @classmethod + def INPUT_TYPES(cls): + # Refresh model list at call time (models may be downloaded after startup) + models = list(HF_VL_MODELS.keys()) + default_model = models[0] if models else "Qwen2.5-VL-3B-Instruct" + default_prompt = ( + "🎬 Video Action Description (FLF)" + if "🎬 Video Action Description (FLF)" in FLF_PRESET_PROMPTS + else FLF_PRESET_PROMPTS[0] + ) + return { + "required": { + "model_name": (models, {"default": default_model, "tooltip": FLF_TOOLTIPS["model_name"]}), + "quantization": (Quantization.get_values(), {"default": Quantization.FP16.value, "tooltip": FLF_TOOLTIPS["quantization"]}), + "attention_mode": (ATTENTION_MODES, {"default": "auto", "tooltip": FLF_TOOLTIPS["attention_mode"]}), + "preset_prompt": (FLF_PRESET_PROMPTS, {"default": default_prompt, "tooltip": FLF_TOOLTIPS["preset_prompt"]}), + "custom_prompt": ("STRING", {"default": "", "multiline": True, "tooltip": FLF_TOOLTIPS["custom_prompt"]}), + "max_tokens": ("INT", {"default": 384, "min": 64, "max": 2048, "tooltip": FLF_TOOLTIPS["max_tokens"]}), + "keep_model_loaded": ("BOOLEAN", {"default": True, "tooltip": FLF_TOOLTIPS["keep_model_loaded"]}), + "seed": ("INT", {"default": 1, "min": 1, "max": 2**32 - 1, "tooltip": FLF_TOOLTIPS["seed"]}), + }, + "optional": { + "first_frame": ("IMAGE", {"tooltip": FLF_TOOLTIPS["first_frame"]}), + "last_frame": ("IMAGE", {"tooltip": FLF_TOOLTIPS["last_frame"]}), + }, + } + + RETURN_TYPES = ("STRING",) + RETURN_NAMES = ("FLF_PROMPT",) + FUNCTION = "process" + CATEGORY = "IAMCCS/QwenVL" + DESCRIPTION = ( + "Uses QwenVL to analyse the FIRST and LAST frame of a video clip " + "and generate a motion/action description prompt for FLF video generators " + "(WAN SVI Pro, LTX-2 FLF, Wan2.1 i2v, etc.)." + ) + + def process( + self, + model_name, + quantization, + attention_mode, + preset_prompt, + custom_prompt, + max_tokens, + keep_model_loaded, + seed, + first_frame=None, + last_frame=None, + ): + return self.run_flf( + model_name, quantization, preset_prompt, custom_prompt, + first_frame, last_frame, + max_tokens, + temperature=0.6, top_p=0.9, num_beams=1, + repetition_penalty=1.2, + seed=seed, + keep_model_loaded=keep_model_loaded, + attention_mode=attention_mode, + use_torch_compile=False, + device="auto", + ) + + # ------------------------------------------------------------------ + # Advanced version + # ------------------------------------------------------------------ + class IAMCCS_QWEN_VL_FLF_Advanced(_FLFMixin, Base): + """Advanced version of IAMCCS_QWEN_VL_FLF with full parameter control.""" + + @classmethod + def INPUT_TYPES(cls): + models = list(HF_VL_MODELS.keys()) + default_model = models[0] if models else "Qwen2.5-VL-3B-Instruct" + default_prompt = ( + "🎬 Video Action Description (FLF)" + if "🎬 Video Action Description (FLF)" in FLF_PRESET_PROMPTS + else FLF_PRESET_PROMPTS[0] + ) + + num_gpus = torch.cuda.device_count() + gpu_list = [f"cuda:{i}" for i in range(num_gpus)] + device_options = ["auto", "cpu", "mps"] + gpu_list + + return { + "required": { + "model_name": (models, {"default": default_model, "tooltip": FLF_TOOLTIPS["model_name"]}), + "quantization": (Quantization.get_values(), {"default": Quantization.FP16.value, "tooltip": FLF_TOOLTIPS["quantization"]}), + "attention_mode": (ATTENTION_MODES, {"default": "auto", "tooltip": FLF_TOOLTIPS["attention_mode"]}), + "use_torch_compile":("BOOLEAN", {"default": False, "tooltip": FLF_TOOLTIPS["use_torch_compile"]}), + "device": (device_options, {"default": "auto", "tooltip": FLF_TOOLTIPS["device"]}), + "preset_prompt": (FLF_PRESET_PROMPTS, {"default": default_prompt, "tooltip": FLF_TOOLTIPS["preset_prompt"]}), + "custom_prompt": ("STRING", {"default": "", "multiline": True, "tooltip": FLF_TOOLTIPS["custom_prompt"]}), + "max_tokens": ("INT", {"default": 512, "min": 64, "max": 4096, "tooltip": FLF_TOOLTIPS["max_tokens"]}), + "temperature": ("FLOAT", {"default": 0.6, "min": 0.1, "max": 1.0, "step": 0.05, "tooltip": FLF_TOOLTIPS["temperature"]}), + "top_p": ("FLOAT", {"default": 0.9, "min": 0.0, "max": 1.0, "step": 0.05, "tooltip": FLF_TOOLTIPS["top_p"]}), + "num_beams": ("INT", {"default": 1, "min": 1, "max": 8, "tooltip": FLF_TOOLTIPS["num_beams"]}), + "repetition_penalty": ("FLOAT", {"default": 1.2, "min": 0.5, "max": 2.0, "step": 0.05, "tooltip": FLF_TOOLTIPS["repetition_penalty"]}), + "keep_model_loaded":("BOOLEAN", {"default": True, "tooltip": FLF_TOOLTIPS["keep_model_loaded"]}), + "seed": ("INT", {"default": 1, "min": 1, "max": 2**32 - 1, "tooltip": FLF_TOOLTIPS["seed"]}), + }, + "optional": { + "first_frame": ("IMAGE", {"tooltip": FLF_TOOLTIPS["first_frame"]}), + "last_frame": ("IMAGE", {"tooltip": FLF_TOOLTIPS["last_frame"]}), + }, + } + + RETURN_TYPES = ("STRING",) + RETURN_NAMES = ("FLF_PROMPT",) + FUNCTION = "process" + CATEGORY = "IAMCCS/QwenVL" + DESCRIPTION = ( + "Advanced version of IAMCCS QwenVL FLF node with full control over " + "generation parameters. Accepts FIRST FRAME + LAST FRAME and outputs " + "an action/motion description prompt for AI video generators." + ) + + def process( + self, + model_name, + quantization, + attention_mode, + use_torch_compile, + device, + preset_prompt, + custom_prompt, + max_tokens, + temperature, + top_p, + num_beams, + repetition_penalty, + keep_model_loaded, + seed, + first_frame=None, + last_frame=None, + ): + return self.run_flf( + model_name, quantization, preset_prompt, custom_prompt, + first_frame, last_frame, + max_tokens, temperature, top_p, num_beams, repetition_penalty, + seed, keep_model_loaded, attention_mode, use_torch_compile, device, + ) + + return IAMCCS_QWEN_VL_FLF, IAMCCS_QWEN_VL_FLF_Advanced + + +# --------------------------------------------------------------------------- +# Module-level instantiation (deferred, with graceful fallback) +# --------------------------------------------------------------------------- + +try: + IAMCCS_QWEN_VL_FLF, IAMCCS_QWEN_VL_FLF_Advanced = _build_node_classes() + + NODE_CLASS_MAPPINGS = { + "IAMCCS_QWEN_VL_FLF": IAMCCS_QWEN_VL_FLF, + "IAMCCS_QWEN_VL_FLF_Advanced": IAMCCS_QWEN_VL_FLF_Advanced, + } + + NODE_DISPLAY_NAME_MAPPINGS = { + "IAMCCS_QWEN_VL_FLF": "QwenVL FLF β€” First/Last Frame Prompt 🎬", + "IAMCCS_QWEN_VL_FLF_Advanced": "QwenVL FLF β€” First/Last Frame Prompt (Advanced) 🎬", + } + + print("[IAMCCS] IAMCCS_QWEN_VL_FLF nodes loaded OK") + +except Exception as _err: + print(f"[IAMCCS] WARNING: IAMCCS_QWEN_VL_FLF could not load β€” {_err}") + print("[IAMCCS] Make sure ComfyUI-QwenVL is installed in custom_nodes/ComfyUI-QwenVL") + + NODE_CLASS_MAPPINGS = {} + NODE_DISPLAY_NAME_MAPPINGS = {} + IAMCCS_QWEN_VL_FLF = None + IAMCCS_QWEN_VL_FLF_Advanced = None diff --git a/iamccs_wan_svipro_motion.py b/iamccs_wan_svipro_motion.py index 532de55..61295ab 100644 --- a/iamccs_wan_svipro_motion.py +++ b/iamccs_wan_svipro_motion.py @@ -11,6 +11,79 @@ import comfy.latent_formats import node_helpers +def _smoothstep(x: torch.Tensor) -> torch.Tensor: + # x in [0,1] + return x * x * (3.0 - 2.0 * x) + + +def _apply_soft_limiter(x: torch.Tensor, *, mode: str, limit: float) -> torch.Tensor: + if mode == "hard": + return x.clamp_(-limit, limit) + if mode == "tanh": + # Smooth limiter: prevents hard saturation artifacts. + # For small values, tanh(x/limit) β‰ˆ x/limit. + return x.div_(limit).tanh_().mul_(limit) + # Fallback + return x.clamp_(-limit, limit) + + +def _center_diff(diff: torch.Tensor, *, mean_mode: str) -> tuple[torch.Tensor, torch.Tensor]: + """Return (diff_centered, diff_mean). + + mean_mode: + - frame_scalar: legacy behavior (mean over channels+spatial). + - per_channel: mean per channel (mean over spatial only). Helps reduce hue shifts. + """ + + if mean_mode == "per_channel": + diff_mean = diff.mean(dim=(3, 4), keepdim=True) + else: + # legacy + diff_mean = diff.mean(dim=(1, 3, 4), keepdim=True) + diff_centered = diff - diff_mean + return diff_centered, diff_mean + + +def _preset_params(safety_preset: str, motion_amplitude: float) -> dict: + # Keep changes non-invasive unless motion is pushed above the common safe zone. + only_if_gt = 1.15 + + legacy = { + "enabled": True, + "only_if_gt": float("inf"), + "mean_mode": "frame_scalar", + "limiter_mode": "hard", + "limiter_limit": 6.0, + "ramp_frames": 0, + } + + safe = { + "enabled": True, + "only_if_gt": only_if_gt, + "mean_mode": "per_channel", + "limiter_mode": "tanh", + "limiter_limit": 6.0, + "ramp_frames": 2, + } + + safer = { + "enabled": True, + "only_if_gt": only_if_gt, + "mean_mode": "per_channel", + "limiter_mode": "tanh", + # slightly tighter limiter to avoid outliers at high motion + "limiter_limit": 5.5, + "ramp_frames": 4, + } + + if safety_preset == "legacy": + return legacy + if safety_preset == "safer": + return safer + # default + return safe + + class IAMCCS_WanImageMotion: @classmethod def INPUT_TYPES(cls): @@ -20,7 +93,8 @@ class IAMCCS_WanImageMotion: "negative": ("CONDITIONING",), "length": ("INT", {"default": 81, "min": 1, "max": 16384, "step": 4}), "anchor_samples": ("LATENT",), - "motion_latent_count": ("INT", {"default": 1, "min": 0, "max": 128, "step": 1}), + # Match FLF/SVI Pro semantics: typical 0-16. + "motion_latent_count": ("INT", {"default": 1, "min": 0, "max": 16, "step": 1}), "motion": ("FLOAT", {"default": 1.15, "min": 1.0, "max": 2.0, "step": 0.05}), "motion_mode": ( [ @@ -37,7 +111,8 @@ class IAMCCS_WanImageMotion: "fp32", "normal", ], - {"default": "auto"}, + # FLF reference node allocates empty latent as fp32 by default. + {"default": "fp32"}, ), "vram_profile": ( [ @@ -50,6 +125,22 @@ class IAMCCS_WanImageMotion: {"default": "normal"}, ), "include_padding_in_motion": ("BOOLEAN", {"default": False}), + # Keep this at the end to preserve existing widgets_values indexing in saved workflows. + "safety_preset": ( + [ + "safe", + "safer", + "legacy", + ], + { + "default": "safe", + "tooltip": ( + "Safe preset reduces color/seam artifacts when motion > 1.15. " + "It applies per-channel stabilization, smooth limiter, and a short ramp. " + "Set legacy to use the original hard-clamp behavior." + ), + }, + ), }, "optional": { "prev_samples": ("LATENT",), @@ -67,6 +158,9 @@ class IAMCCS_WanImageMotion: if latent_precision == "normal": # Backward-compat alias for older workflows. return anchor_dtype + if latent_precision == "auto": + # Prefer fp32 for 1:1 compatibility with the FLF reference node. + return torch.float32 if latent_precision == "fp32": return torch.float32 if latent_precision == "fp16": @@ -75,7 +169,7 @@ class IAMCCS_WanImageMotion: def _apply_motion_amplitude(self, image_cond_latent: torch.Tensor, *, real_latents: int, anchor_latents: int, motion_latents: int, motion_amplitude: float, motion_mode: str, - vram_profile: str) -> torch.Tensor: + vram_profile: str, safety_preset: str = "safe") -> torch.Tensor: if motion_amplitude is None or motion_amplitude <= 1.0: return image_cond_latent @@ -84,18 +178,41 @@ class IAMCCS_WanImageMotion: base_latent = image_cond_latent[:, :, 0:1] # first latent frame - def _scale_slice_gpu(view_slice: torch.Tensor) -> torch.Tensor: + preset = _preset_params(safety_preset, motion_amplitude) + # Non-invasive: if motion is within the usual safe zone, keep legacy behavior. + # This preserves 1:1 results for typical workflows. + if motion_amplitude <= preset["only_if_gt"]: + preset = _preset_params("legacy", motion_amplitude) + + def _scale_slice_gpu(view_slice: torch.Tensor, *, gain: torch.Tensor | float) -> torch.Tensor: # view_slice: [B,C,T,H,W] # VRAM-optimized variant: keep only one full-sized temporary tensor. with torch.no_grad(): diff = view_slice - base_latent - diff_mean = diff.mean(dim=(1, 3, 4), keepdim=True) - diff.sub_(diff_mean) - diff.mul_(motion_amplitude) - diff.add_(diff_mean) - diff.add_(base_latent) - diff.clamp_(-6, 6) - return diff + diff_centered, diff_mean = _center_diff(diff, mean_mode=preset["mean_mode"]) + diff_centered.mul_(gain) + out_local = diff_centered.add_(diff_mean).add_(base_latent) + _apply_soft_limiter(out_local, mode=preset["limiter_mode"], limit=float(preset["limiter_limit"])) + return out_local + + def _gain_weights(start: int, end: int) -> torch.Tensor | float: + # Returns broadcastable gain weights for the slice. + # gain = 1 + (motion-1)*w(t) + ramp_frames = int(preset["ramp_frames"]) + if ramp_frames <= 0: + return float(motion_amplitude) + tcount = max(0, end - start) + if tcount <= 0: + return float(motion_amplitude) + # ramp up only at the beginning of the boosted range + ramp = min(ramp_frames, tcount) + w = torch.ones((tcount,), device=image_cond_latent.device, dtype=image_cond_latent.dtype) + if ramp > 0: + # 0..1 over ramp + x = torch.linspace(0.0, 1.0, steps=ramp, device=w.device, dtype=w.dtype) + w[:ramp] = _smoothstep(x) + gain = 1.0 + (motion_amplitude - 1.0) * w + return gain.view(1, 1, tcount, 1, 1) def _apply_to_range(start: int, end: int) -> torch.Tensor: # Applies scaling to out[:, :, start:end] according to VRAM profile. @@ -105,7 +222,8 @@ class IAMCCS_WanImageMotion: return out if vram_profile == "normal": - out[:, :, start:end] = _scale_slice_gpu(out[:, :, start:end]) + gain = _gain_weights(start, end) + out[:, :, start:end] = _scale_slice_gpu(out[:, :, start:end], gain=gain) return out if vram_profile in ("chunked_blocks_2", "chunked_blocks_4"): @@ -113,13 +231,15 @@ class IAMCCS_WanImageMotion: t = start while t < end: t2 = min(end, t + block) - out[:, :, t:t2] = _scale_slice_gpu(out[:, :, t:t2]) + gain = _gain_weights(t, t2) + out[:, :, t:t2] = _scale_slice_gpu(out[:, :, t:t2], gain=gain) t = t2 return out if vram_profile == "loop_per_frame (lowest_vram)": for t in range(start, end): - out[:, :, t:t+1] = _scale_slice_gpu(out[:, :, t:t+1]) + gain = _gain_weights(t, t + 1) + out[:, :, t:t+1] = _scale_slice_gpu(out[:, :, t:t+1], gain=gain) return out if vram_profile == "cpu_offload (slowest)": @@ -130,17 +250,27 @@ class IAMCCS_WanImageMotion: base_cpu = base_latent.detach().to("cpu") slice_cpu = out[:, :, start:end].detach().to("cpu") diff = slice_cpu - base_cpu - diff_mean = diff.mean(dim=(1, 3, 4), keepdim=True) - diff.sub_(diff_mean) - diff.mul_(motion_amplitude) - diff.add_(diff_mean) - diff.add_(base_cpu) - diff.clamp_(-6, 6) - out[:, :, start:end] = diff.to(device) + diff_centered, diff_mean = _center_diff(diff, mean_mode=preset["mean_mode"]) + # gain weights are computed on the target device; rebuild on CPU + ramp_frames = int(preset["ramp_frames"]) + tcount = max(0, end - start) + if ramp_frames > 0 and tcount > 0: + ramp = min(ramp_frames, tcount) + w = torch.ones((tcount,), device=diff_centered.device, dtype=diff_centered.dtype) + x = torch.linspace(0.0, 1.0, steps=ramp, device=w.device, dtype=w.dtype) + w[:ramp] = _smoothstep(x) + gain = (1.0 + (motion_amplitude - 1.0) * w).view(1, 1, tcount, 1, 1) + else: + gain = float(motion_amplitude) + diff_centered.mul_(gain) + out_cpu = diff_centered.add_(diff_mean).add_(base_cpu) + _apply_soft_limiter(out_cpu, mode=preset["limiter_mode"], limit=float(preset["limiter_limit"])) + out[:, :, start:end] = out_cpu.to(device) return out # Fallback - out[:, :, start:end] = _scale_slice_gpu(out[:, :, start:end]) + gain = _gain_weights(start, end) + out[:, :, start:end] = _scale_slice_gpu(out[:, :, start:end], gain=gain) return out # Avoid touching padding: operate only within [0:real_latents) @@ -168,9 +298,10 @@ class IAMCCS_WanImageMotion: def apply(self, positive, negative, length, anchor_samples, motion_latent_count, motion, motion_mode, add_reference_latents, latent_precision, vram_profile, include_padding_in_motion, - prev_samples=None): + safety_preset="safe", prev_samples=None): with torch.no_grad(): - anchor_latent = anchor_samples["samples"] + # Clone to prevent in-place motion amplitude writes from corrupting the caller's tensor. + anchor_latent = anchor_samples["samples"].clone() B, C, T_anchor, H, W = anchor_latent.shape @@ -201,9 +332,18 @@ class IAMCCS_WanImageMotion: image_cond_latent = torch.cat([anchor_latent, motion_latent], dim=2) padding_size = max(0, padding_size) - padding = torch.zeros(1, C, padding_size, H, W, dtype=dtype, device=device) - padding = comfy.latent_formats.Wan21().process_out(padding) - image_cond_latent = torch.cat([image_cond_latent, padding], dim=2) + if padding_size > 0: + padding = torch.zeros(B, C, padding_size, H, W, dtype=dtype, device=device) + padding = comfy.latent_formats.Wan21().process_out(padding) + image_cond_latent = torch.cat([image_cond_latent, padding], dim=2) + + # FLF/SVI reference behavior: ensure exact temporal length. + if image_cond_latent.shape[2] > total_latents: + image_cond_latent = image_cond_latent[:, :, :total_latents] + elif image_cond_latent.shape[2] < total_latents: + # Safety: if something went off, truncate/pad has already handled it, + # but keep a hard guard. + image_cond_latent = image_cond_latent[:, :, :total_latents] # Apply motion amplitude before injecting into conditioning effective_latents = total_latents if include_padding_in_motion else min(total_latents, T_anchor + T_motion) @@ -285,6 +425,7 @@ class IAMCCS_WanImageMotion: motion_amplitude=motion, motion_mode=motion_mode_effective, vram_profile=vram_profile, + safety_preset=safety_preset, ) mask = torch.ones((1, 1, empty_latent.shape[2], H, W), device=device, dtype=dtype) @@ -312,10 +453,440 @@ class IAMCCS_WanImageMotion: return (positive, negative, out_latent) +class WanImageMotionPro: + """WanImageMotionPro + + Combines IAMCCS_WanImageMotion motion amplitude control with FLF-style + (First/Last Frame) hard lock via optional end_samples. + + Behavior: + - Start: anchor_samples + optional motion tail from prev_samples. + - Motion: apply motion amplitude scaling (VRAM-aware) as in IAMCCS_WanImageMotion. + - End: overwrite last temporal latent slots with end_samples, then lock them + via concat_mask (FLF-style control). + """ + + @classmethod + def INPUT_TYPES(cls): + return { + "required": { + # Keep the same socket-style inputs as WanImageToVideoSVIProFLF + # so existing FLF workflows can be migrated with minimal friction. + "positive": ("CONDITIONING",), + "negative": ("CONDITIONING",), + "length": ("INT", {"default": 81, "min": 1, "max": 16384, "step": 4}), + "anchor_samples": ("LATENT",), + # Match FLF reference node range. + "motion_latent_count": ("INT", {"default": 1, "min": 0, "max": 16, "step": 1}), + "motion": ("FLOAT", {"default": 1.15, "min": 1.0, "max": 2.0, "step": 0.05}), + "motion_mode": ( + [ + "motion_only (prev_samples)", + "all_nonfirst (anchor+motion)", + ], + {"default": "motion_only (prev_samples)"}, + ), + "add_reference_latents": ("BOOLEAN", {"default": False}), + "latent_precision": ( + [ + "auto", + "fp16", + "fp32", + "normal", + ], + {"default": "fp32"}, + ), + "vram_profile": ( + [ + "normal", + "chunked_blocks_2", + "chunked_blocks_4", + "loop_per_frame (lowest_vram)", + "cpu_offload (slowest)", + ], + {"default": "normal"}, + ), + "include_padding_in_motion": ("BOOLEAN", {"default": False}), + # Keep this at the end to preserve existing widgets_values indexing in saved workflows. + "safety_preset": ( + [ + "safe", + "safer", + "legacy", + ], + { + "default": "safe", + "tooltip": ( + "Safe preset reduces color/seam artifacts when motion > 1.15. " + "Set legacy to use original hard-clamp behavior." + ), + }, + ), + }, + "optional": { + # prev_samples is optional – mirrors original FLF node and IAMCCS_WanImageMotion. + # apply() already handles None gracefully. + "prev_samples": ("LATENT",), + "end_samples": ("LATENT",), + }, + } + + RETURN_TYPES = ("CONDITIONING", "CONDITIONING", "LATENT") + RETURN_NAMES = ("positive", "negative", "latent") + FUNCTION = "apply" + CATEGORY = "IAMCCS/video" + + _log = logging.getLogger("IAMCCS.WanImageMotionPro") + + def _pick_empty_latent_dtype(self, anchor_dtype: torch.dtype, latent_precision: str) -> torch.dtype: + # Keep 1:1 behavior with IAMCCS_WanImageMotion. + if latent_precision == "normal": + return anchor_dtype + if latent_precision == "auto": + return torch.float32 + if latent_precision == "fp32": + return torch.float32 + if latent_precision == "fp16": + return torch.float16 + return anchor_dtype + + def _apply_motion_amplitude( + self, + image_cond_latent: torch.Tensor, + *, + real_latents: int, + anchor_latents: int, + motion_latents: int, + motion_amplitude: float, + motion_mode: str, + vram_profile: str, + safety_preset: str = "safe", + ) -> torch.Tensor: + # Reuse the exact logic from IAMCCS_WanImageMotion (copy to keep node self-contained). + if motion_amplitude is None or motion_amplitude <= 1.0: + return image_cond_latent + + if image_cond_latent.shape[2] <= 1: + return image_cond_latent + + base_latent = image_cond_latent[:, :, 0:1] + + preset = _preset_params(safety_preset, motion_amplitude) + if motion_amplitude <= preset["only_if_gt"]: + preset = _preset_params("legacy", motion_amplitude) + + def _scale_slice_gpu(view_slice: torch.Tensor, *, gain: torch.Tensor | float) -> torch.Tensor: + with torch.no_grad(): + diff = view_slice - base_latent + diff_centered, diff_mean = _center_diff(diff, mean_mode=preset["mean_mode"]) + diff_centered.mul_(gain) + out_local = diff_centered.add_(diff_mean).add_(base_latent) + _apply_soft_limiter(out_local, mode=preset["limiter_mode"], limit=float(preset["limiter_limit"])) + return out_local + + def _gain_weights(start: int, end: int) -> torch.Tensor | float: + ramp_frames = int(preset["ramp_frames"]) + if ramp_frames <= 0: + return float(motion_amplitude) + tcount = max(0, end - start) + if tcount <= 0: + return float(motion_amplitude) + ramp = min(ramp_frames, tcount) + w = torch.ones((tcount,), device=image_cond_latent.device, dtype=image_cond_latent.dtype) + if ramp > 0: + x = torch.linspace(0.0, 1.0, steps=ramp, device=w.device, dtype=w.dtype) + w[:ramp] = _smoothstep(x) + gain = 1.0 + (motion_amplitude - 1.0) * w + return gain.view(1, 1, tcount, 1, 1) + + def _apply_to_range(start: int, end: int) -> torch.Tensor: + if end <= start: + return out + + if vram_profile == "normal": + gain = _gain_weights(start, end) + out[:, :, start:end] = _scale_slice_gpu(out[:, :, start:end], gain=gain) + return out + + if vram_profile in ("chunked_blocks_2", "chunked_blocks_4"): + block = 2 if vram_profile == "chunked_blocks_2" else 4 + t = start + while t < end: + t2 = min(end, t + block) + gain = _gain_weights(t, t2) + out[:, :, t:t2] = _scale_slice_gpu(out[:, :, t:t2], gain=gain) + t = t2 + return out + + if vram_profile == "loop_per_frame (lowest_vram)": + for t in range(start, end): + gain = _gain_weights(t, t + 1) + out[:, :, t:t + 1] = _scale_slice_gpu(out[:, :, t:t + 1], gain=gain) + return out + + if vram_profile == "cpu_offload (slowest)": + with torch.no_grad(): + device = out.device + base_cpu = base_latent.detach().to("cpu") + slice_cpu = out[:, :, start:end].detach().to("cpu") + diff = slice_cpu - base_cpu + diff_centered, diff_mean = _center_diff(diff, mean_mode=preset["mean_mode"]) + ramp_frames = int(preset["ramp_frames"]) + tcount = max(0, end - start) + if ramp_frames > 0 and tcount > 0: + ramp = min(ramp_frames, tcount) + w = torch.ones((tcount,), device=diff_centered.device, dtype=diff_centered.dtype) + x = torch.linspace(0.0, 1.0, steps=ramp, device=w.device, dtype=w.dtype) + w[:ramp] = _smoothstep(x) + gain = (1.0 + (motion_amplitude - 1.0) * w).view(1, 1, tcount, 1, 1) + else: + gain = float(motion_amplitude) + diff_centered.mul_(gain) + out_cpu = diff_centered.add_(diff_mean).add_(base_cpu) + _apply_soft_limiter(out_cpu, mode=preset["limiter_mode"], limit=float(preset["limiter_limit"])) + out[:, :, start:end] = out_cpu.to(device) + return out + + gain = _gain_weights(start, end) + out[:, :, start:end] = _scale_slice_gpu(out[:, :, start:end], gain=gain) + return out + + real_latents = max(0, min(real_latents, image_cond_latent.shape[2])) + if real_latents <= 1: + return image_cond_latent + + out = image_cond_latent + + if motion_mode == "motion_only (prev_samples)": + if motion_latents <= 0: + return out + + start = anchor_latents + end = min(anchor_latents + motion_latents, real_latents) + if end <= start: + return out + + return _apply_to_range(start, end) + + start = 1 + end = real_latents + return _apply_to_range(start, end) + + def apply( + self, + positive, + negative, + length, + anchor_samples, + motion_latent_count, + motion, + motion_mode, + add_reference_latents, + latent_precision, + vram_profile, + include_padding_in_motion, + safety_preset="safe", + prev_samples=None, + end_samples=None, + ): + with torch.no_grad(): + # Clone to prevent in-place motion amplitude writes from corrupting the caller's tensor. + anchor_latent = anchor_samples["samples"].clone() + B, C, T_anchor, H, W = anchor_latent.shape + + total_latents = (length - 1) // 4 + 1 + + device = anchor_latent.device + dtype = anchor_latent.dtype + + empty_latent_dtype = self._pick_empty_latent_dtype(dtype, latent_precision) + empty_latent = torch.zeros( + [B, 16, total_latents, H, W], + device=comfy.model_management.intermediate_device(), + dtype=empty_latent_dtype, + ) + + motion_latent = None + T_motion = 0 + # In the original FLF node, prev_samples is a required socket. + # If a workflow leaves it disconnected, ComfyUI may pass None. + has_prev = prev_samples is not None and motion_latent_count != 0 + + if prev_samples is None or motion_latent_count == 0: + padding_size = total_latents - T_anchor + image_cond_latent = anchor_latent + else: + motion_latent = prev_samples["samples"][:, :, -motion_latent_count:] + T_motion = motion_latent.shape[2] + padding_size = total_latents - T_anchor - T_motion + image_cond_latent = torch.cat([anchor_latent, motion_latent], dim=2) + + padding_size = max(0, padding_size) + if padding_size > 0: + padding = torch.zeros(B, C, padding_size, H, W, dtype=dtype, device=device) + padding = comfy.latent_formats.Wan21().process_out(padding) + image_cond_latent = torch.cat([image_cond_latent, padding], dim=2) + + # FLF/SVI reference behavior: enforce exact temporal length. + if image_cond_latent.shape[2] > total_latents: + image_cond_latent = image_cond_latent[:, :, :total_latents] + elif image_cond_latent.shape[2] < total_latents: + image_cond_latent = image_cond_latent[:, :, :total_latents] + + # Pre-compute end_t_fix so we can exclude the end-locked zone from motion amplitude. + # Motion should NOT touch slots that will be hard-locked to end_samples: scaling those + # intermediate latents would generate noise that hurts the model's firstβ†’last interpolation. + end_t_fix_early = 0 + if end_samples is not None: + _e = end_samples["samples"] + if ( + _e.shape[1] == C + and _e.shape[3] == H + and _e.shape[4] == W + ): + end_t_fix_early = min(_e.shape[2], total_latents) + + # Motion boost applied before FLF overwrite. + # Cap effective_latents so motion never reaches into the end-locked zone. + effective_latents_base = total_latents if include_padding_in_motion else min(total_latents, T_anchor + T_motion) + effective_latents = max(1, min(effective_latents_base, total_latents - end_t_fix_early)) + + motion_mode_effective = motion_mode + if motion_mode == "motion_only (prev_samples)" and T_motion == 0 and include_padding_in_motion: + motion_mode_effective = "all_nonfirst (anchor+motion)" + + try: + free_vram, total_vram = comfy.model_management.get_free_memory(device) + except Exception: + free_vram, total_vram = None, None + + self._log.info( + "[WanImageMotionPro] length=%s -> total_latents=%s | motion=%s | mode=%s | vram_profile=%s | latent_precision=%s | add_reference_latents=%s | include_padding_in_motion=%s", + length, + total_latents, + motion, + motion_mode, + vram_profile, + latent_precision, + add_reference_latents, + include_padding_in_motion, + ) + self._log.info( + "[WanImageMotionPro] anchor: B=%s C=%s T=%s H=%s W=%s dtype=%s device=%s | prev=%s motion_latent_count=%s T_motion=%s | padding_size=%s | end_samples=%s", + B, + C, + T_anchor, + H, + W, + str(dtype).replace("torch.", ""), + str(device), + has_prev, + motion_latent_count, + T_motion, + padding_size, + end_samples is not None, + ) + if free_vram is not None: + self._log.info("[WanImageMotionPro] free_vram=%s total_vram=%s", free_vram, total_vram) + + if motion_mode_effective == "motion_only (prev_samples)": + motion_start = T_anchor + motion_end = min(T_anchor + T_motion, effective_latents) + else: + motion_start = 1 + motion_end = effective_latents + + motion_frames_count = max(0, motion_end - motion_start) + self._log.info( + "[WanImageMotionPro] motion_range=[%s:%s] (effective_latents=%s) padding_included=%s", + motion_start, + motion_end, + effective_latents, + include_padding_in_motion, + ) + if motion_frames_count == 0: + self._log.warning( + "[WanImageMotionPro] WARNING: motion_range is EMPTY (no frames will be modified). " + "Enable include_padding_in_motion or provide prev_samples with motion_latent_count > 0." + ) + else: + self._log.info( + "[WanImageMotionPro] Motion boost applies to %s frame(s) amplitude=%.2f", + motion_frames_count, + motion, + ) + + image_cond_latent = self._apply_motion_amplitude( + image_cond_latent, + real_latents=effective_latents, + anchor_latents=T_anchor, + motion_latents=T_motion, + motion_amplitude=motion, + motion_mode=motion_mode_effective, + vram_profile=vram_profile, + safety_preset=safety_preset, + ) + + # FLF end lock: overwrite last slots with end_samples (if provided). + end_t_fix = 0 + if end_samples is not None: + # Clone to prevent mutations from affecting the caller's tensor. + end_latent = end_samples["samples"].clone() + + if end_latent.shape[0] == 1 and B > 1: + end_latent = end_latent.repeat(B, 1, 1, 1, 1) + + if ( + end_latent.shape[1] == C + and end_latent.shape[3] == H + and end_latent.shape[4] == W + ): + T_end = end_latent.shape[2] + end_t_fix = min(T_end, total_latents) + if end_t_fix > 0: + image_cond_latent[:, :, -end_t_fix:] = end_latent[:, :, -end_t_fix:] + else: + end_t_fix = 0 + self._log.warning( + "[WanImageMotionPro] end_samples shape mismatch, skipping end lock. end=%s anchor=%s", + tuple(end_latent.shape), + tuple(anchor_latent.shape), + ) + + # Mask: lock first slot + lock last end_t_fix slots. + mask = torch.ones((1, 1, total_latents, H, W), device=device, dtype=dtype) + mask[:, :, :1] = 0.0 + if end_t_fix > 0: + mask[:, :, -end_t_fix:] = 0.0 + + positive = node_helpers.conditioning_set_values( + positive, {"concat_latent_image": image_cond_latent, "concat_mask": mask} + ) + negative = node_helpers.conditioning_set_values( + negative, {"concat_latent_image": image_cond_latent, "concat_mask": mask} + ) + + if add_reference_latents: + ref_latent = anchor_latent[:, :, 0:1] + positive = node_helpers.conditioning_set_values( + positive, {"reference_latents": [ref_latent]}, append=True + ) + negative = node_helpers.conditioning_set_values( + negative, {"reference_latents": [torch.zeros_like(ref_latent)]}, append=True + ) + + out_latent = {"samples": empty_latent} + return (positive, negative, out_latent) + + NODE_CLASS_MAPPINGS = { "IAMCCS_WanImageMotion": IAMCCS_WanImageMotion, + "WanImageMotionPro": WanImageMotionPro, + "IAMCCS_WanImageMotionPro": WanImageMotionPro, } NODE_DISPLAY_NAME_MAPPINGS = { "IAMCCS_WanImageMotion": "IAMCCS WanImageMotion", + "WanImageMotionPro": "WanImageMotionPro", + "IAMCCS_WanImageMotionPro": "WanImageMotionPro", } diff --git a/version.json b/version.json index e4e344e..f646a95 100644 --- a/version.json +++ b/version.json @@ -1,6 +1,6 @@ { "name": "iamccs-nodes", - "version": "1.3.4", + "version": "1.3.5", "author": "Carmine Cristallo Scalzi (IAMCCS)", "description": "IAMCCS nodes for ComfyUI: IAMCCS echosystem for ComfyUI, nodes 4 LoRA, WAN 2.2, WAN 2.1 and LTX-2 pipelines." } diff --git a/web/iamccs_autolink_converter.js b/web/iamccs_autolink_converter.js index 1679fc8..86bb81e 100644 --- a/web/iamccs_autolink_converter.js +++ b/web/iamccs_autolink_converter.js @@ -153,6 +153,20 @@ function normalizeAutolinkIOSlots(graph, node, { wantInputs = 0, wantOutputs = 0 if (!node.inputs) node.inputs = []; if (!node.outputs) node.outputs = []; + // Remove ALL inputs when none are wanted (fixes Get nodes that were incorrectly + // serialized with stale input slots from old buggy workflows or the queue patch). + if (wantInputs === 0 && node.inputs.length > 0) { + try { + for (let i = node.inputs.length - 1; i >= 0; i--) { + try { if (typeof node.disconnectInput === "function") node.disconnectInput(i); } catch {} + try { + if (typeof node.removeInput === "function") node.removeInput(i); + else node.inputs.splice(i, 1); + } catch {} + } + } catch {} + } + // Ensure at least one slot exists when requested if (wantInputs > 0 && node.inputs.length === 0 && typeof node.addInput === "function") { node.addInput("*", "*"); @@ -1350,6 +1364,10 @@ app.registerExtension({ if (nodeData?.name === SET_TYPE) { const onNodeCreated = nodeType.prototype.onNodeCreated; nodeType.prototype.onNodeCreated = function() { + // --- MUST be set BEFORE anything else so ComfyUI skips this node + // during graphToPrompt serialization and uses getInputLink chain --- + this.isVirtualNode = true; + const result = onNodeCreated?.apply(this, arguments); const node = this; @@ -1411,38 +1429,67 @@ app.registerExtension({ // Input/output: normalizza workflow vecchi (che possono avere input duplicati) normalizeAutolinkIOSlots(app.graph, node, { wantInputs: 1, wantOutputs: 1 }); - - // Callback quando si collega - this.onConnectionsChange = function(slotType, slot, isConnect, link_info) { - if (slotType === 1 && isConnect && link_info) { - const fromNode = app.graph.getNodeById(link_info.origin_id); - if (fromNode && fromNode.outputs && fromNode.outputs[link_info.origin_slot]) { - const outputType = fromNode.outputs[link_info.origin_slot].type; - // Usa sempre un nome leggibile e stabile (slot-name), non il tipo puro. - // Questo produce base tipo "model"/"image" e poi lo rendiamo unico: model_0, image_2, ... - let suggestedBase = getSlotName(fromNode, link_info.origin_slot, true); - if (!isValidAutolinkKey(suggestedBase)) { - if (outputType && outputType !== "*") suggestedBase = String(outputType).trim().toLowerCase(); - else suggestedBase = `output_${link_info.origin_slot}`; + // --- KJ-style update: propagate type changes to all matching Get nodes --- + this._iamccsUpdateGetters = function() { + try { + const key = getAutolinkKey(node); + if (!key) return; + const curType = node.inputs?.[0]?.type || "*"; + const gets = _iamccsGraphNodes(app.graph).filter( + n => n?.type === GET_TYPE && getAutolinkKey(n) === key + ); + for (const g of gets) { + if (g.outputs?.[0]) { + g.outputs[0].type = curType; + g.outputs[0].name = curType; } - - // Imposta tipo - node.inputs[0].type = outputType; - node.outputs[0].type = outputType; + // Validate and remove type-incompatible links from each Get + try { g.validateLinks?.(); } catch {} + } + } catch {} + }; - // Se il converter ha giΓ  impostato un nome unico, NON sovrascriverlo qui. - // Auto-fill solo quando il widget Γ¨ vuoto. - const currentKey = getAutolinkKey(node); - const desiredKey = isValidAutolinkKey(currentKey) - ? currentKey - : makeUniqueAutolinkSetName(app.graph, suggestedBase); + // Callback quando si collega / scollega + this.onConnectionsChange = function(slotType, slot, isConnect, link_info) { + if (slotType === 1) { // input slot changed + if (isConnect && link_info) { + const fromNode = app.graph.getNodeById + ? app.graph.getNodeById(link_info.origin_id) + : getNodeById(app.graph, link_info.origin_id); + if (fromNode?.outputs?.[link_info.origin_slot]) { + const outputType = fromNode.outputs[link_info.origin_slot].type; - // Imposta chiave + UI + porta coerenti - setAutolinkKeyAndTitle(node, desiredKey); + // Stabilize a name using the slot label (not the raw type) + let suggestedBase = getSlotName(fromNode, link_info.origin_slot, true); + if (!isValidAutolinkKey(suggestedBase)) { + if (outputType && outputType !== "*") + suggestedBase = String(outputType).trim().toLowerCase(); + else + suggestedBase = `output_${link_info.origin_slot}`; + } - // Applica colore testo titolo (se configurato) - applyNodeTitleTextColor(node, getCurrentColorTitlesMode()); + // Update type on both slots + if (node.inputs?.[0]) node.inputs[0].type = outputType; + if (node.outputs?.[0]) node.outputs[0].type = outputType; + + // Auto-fill name only when the widget is still empty + const currentKey = getAutolinkKey(node); + const desiredKey = isValidAutolinkKey(currentKey) + ? currentKey + : makeUniqueAutolinkSetName(app.graph, suggestedBase); + + setAutolinkKeyAndTitle(node, desiredKey); + applyNodeTitleTextColor(node, getCurrentColorTitlesMode()); + + // Propagate new type to all Get nodes sharing our key + node._iamccsUpdateGetters?.(); + } + } else if (!isConnect) { + // On disconnect: reset type to wildcard + if (node.inputs?.[0]) { node.inputs[0].type = "*"; node.inputs[0].name = "*"; } + if (node.outputs?.[0]) { node.outputs[0].type = "*"; node.outputs[0].name = "*"; } + node._iamccsUpdateGetters?.(); } } }; @@ -1469,23 +1516,26 @@ app.registerExtension({ if (String(rawName ?? "").trim() === "*") setWidgetValue(node, "name", ""); if (String(node.title ?? "").trim() === "*") node.title = "Set AutoLink"; } catch {} - - // Nodo virtuale - non serializza per il prompt + + // isVirtualNode already set at top – keep here for safety (serialization guard) this.isVirtualNode = true; - + return result; }; } - - // Get node - come KJ GetNode + + // Get node - 1:1 KJ GetNode pattern with IAMCCS styling if (nodeData?.name === GET_TYPE) { const onNodeCreated = nodeType.prototype.onNodeCreated; nodeType.prototype.onNodeCreated = function() { + // --- MUST be set BEFORE anything else --- + this.isVirtualNode = true; + const result = onNodeCreated?.apply(this, arguments); const node = this; - - // Combo dinamico con lista Set disponibili - this.addWidget("combo", "name", "", (value) => { + + // Combo dinamico con lista Set disponibili (identical to KJ "Constant" combo) + this.addWidget("combo", "name", "", () => { node.onRename(); }, { values: () => { @@ -1494,22 +1544,64 @@ app.registerExtension({ } }); - // Normalizza output duplicati su workflow vecchi + // Normalizza output/input duplicati su workflow vecchi + // wantInputs: 0 β†’ any stale inputs will be stripped by normalizeAutolinkIOSlots normalizeAutolinkIOSlots(app.graph, node, { wantInputs: 0, wantOutputs: 1 }); - + + // --- KJ-style: remove links whose type no longer matches our output --- + this.validateLinks = function() { + try { + if (!node.outputs?.[0]) return; + const outType = node.outputs[0].type; + if (!outType || outType === "*") return; + const links = node.outputs[0].links; + if (!Array.isArray(links) || links.length === 0) return; + for (const linkId of [...links]) { + const link = _iamccsGetLink(app.graph, linkId); + if (!link) continue; + const lt = link.type; + if (lt && lt !== "*" && lt !== outType && + !lt.split(",").includes(outType) && + !outType.split(",").includes(lt)) { + try { app.graph.removeLink(linkId); } catch {} + } + } + } catch {} + }; + + this.setType = function(type) { + if (!node.outputs?.[0]) return; + node.outputs[0].name = type; + node.outputs[0].type = type; + node.validateLinks(); + }; + + // KJ-style setName: updates widget and triggers onRename + this.setName = function(name) { + setWidgetValue(node, "name", name); + node.onRename(); + }; + this.onRename = function() { const setterName = getAutolinkKey(node); - const setter = _iamccsGraphNodes(app.graph).find(n => + const setter = _iamccsGraphNodes(app.graph).find(n => n.type === SET_TYPE && getAutolinkKey(n) === setterName ); - - if (setter) { - const linkType = setter.outputs[0].type; - node.outputs[0].type = linkType; - node.outputs[0].name = linkType; - node.title = setterName; + if (setter) { + const linkType = setter.inputs?.[0]?.type || "*"; + node.setType(linkType); + node.title = setterName; applyNodeTitleTextColor(node, getCurrentColorTitlesMode()); + } else { + node.setType("*"); + } + }; + + // On output connection change, validate link types (KJ pattern) + this.onConnectionsChange = function(slotType /*1=input,2=output*/, slot, isConnect) { + if (slotType === 2) { + node.validateLinks(); } }; @@ -1517,31 +1609,34 @@ app.registerExtension({ try { if (String(node.title ?? "").trim() === "*") node.title = "Get AutoLink"; } catch {} - // (removed) per-node overlay title drawing; native title is colored via canvas hook - - // Override getInputLink per prendere da Set + + // getInputLink: called by ComfyUI graphToPrompt to resolve the real source link. + // ComfyUI calls this with `slot` = the OUTPUT slot index of this GetNode (always 0). + // We look up our paired SetNode and return the link on its input slot 0, + // which has origin_id = the real upstream node (not virtual). this.getInputLink = function(slot) { - const setterName = getAutolinkKey(node); - const setter = _iamccsGraphNodes(app.graph).find(n => - n.type === SET_TYPE && getAutolinkKey(n) === setterName - ); - - if (setter) { - const slotInfo = setter.inputs[slot]; - if (slotInfo) { - const linkId = slotInfo.link != null - ? slotInfo.link - : (Array.isArray(slotInfo.links) ? slotInfo.links[0] : null); - const link = _iamccsGetLink(app.graph, linkId); - return link || null; - } + try { + const setterName = getAutolinkKey(node); + if (!setterName) return null; + const setter = _iamccsGraphNodes(app.graph).find(n => + n.type === SET_TYPE && getAutolinkKey(n) === setterName + ); + if (!setter?.inputs?.length) return null; + // slot index maps to the Set's input (both nodes use slot 0) + const slotInfo = setter.inputs[0]; + if (!slotInfo) return null; + const linkId = slotInfo.link != null + ? slotInfo.link + : (Array.isArray(slotInfo.links) ? slotInfo.links[0] : null); + return _iamccsGetLink(app.graph, linkId) || null; + } catch { + return null; } - return null; }; - - // Nodo virtuale - non serializza per il prompt + + // isVirtualNode redundant here (set at top) but kept as an insurance belt this.isVirtualNode = true; - + return result; }; } @@ -2538,7 +2633,7 @@ function convertAllLinks( const occupiedSetPositions = new Set(); const occupiedGetPositions = new Set(); const createdSets = new Map(); - + // Traccia i nomi dei Set giΓ  esistenti/creati per evitare duplicati // - Preserva nomi numerati come "model_0" (non li riduce a "model") const usedExactSetNames = new Set(); @@ -2550,6 +2645,18 @@ function convertAllLinks( if (name && String(name).trim()) usedExactSetNames.add(String(name).trim()); } + // Build a map of (srcNodeId, originSlot) β†’ existing SetNode so we can REUSE + // a Set that already consumes from that output instead of creating a duplicate. + // This prevents the "double Set for same source slot" bug when Convert is called + // on a partially-converted graph. + const existingSetBySourceKey = new Map(); + for (const es of existingSets) { + const inLink = es?.inputs?.[0]?.link != null ? _iamccsGetLink(graph, es.inputs[0].link) : null; + if (!inLink) continue; + const sk = `${inLink.origin_id}_${inLink.origin_slot}`; + if (!existingSetBySourceKey.has(sk)) existingSetBySourceKey.set(sk, es); + } + function makeUniqueSetName(desiredName) { const desired = String(desiredName ?? "").trim(); if (!desired) { @@ -2582,83 +2689,96 @@ function convertAllLinks( return candidate; } - console.log(`[IAMCCS AutoLink] Creating ${linksByOrigin.size} Set nodes...`); - - // Crea Set nodes (uno per origine) + console.log(`[IAMCCS AutoLink] Creating/reusing ${linksByOrigin.size} Set nodes...`); + + // Crea Set nodes (uno per origine), RIUTILIZZANDO Set giΓ  esistenti per la stessa sorgente. + // Questo evita il bug "doppio Set per lo stesso slot" quando Convert viene chiamato + // su un grafo parzialmente convertito. for (const [key, originData] of linksByOrigin) { const { srcNode, originSlot, outputName, destinations } = originData; - - const setPos = findFreePosition( - graph, - srcNode.pos[0] + (srcNode.size?.[0] || 200), - (alignMode === "Proportional" ? getAnchorY(srcNode, originSlot, true) : srcNode.pos[1]), - 20, - occupiedSetPositions, - alignMode - ); - - const setNode = createNode(graph, SET_TYPE, setPos[0], setPos[1]); - if (!setNode) continue; - // If the source is inside a hidden/disabled group, the created AutoLink must follow. - _iamccsApplyGroupStateToNode( - graph, - setNode, - srcNode, - { x: setPos[0] + 75, y: setPos[1] + 13 } - ); - - setNode.properties = setNode.properties || {}; - setNode.properties.autolink_color_name = colorSet; - if (setNode.properties.autolink_color_locked === undefined) setNode.properties.autolink_color_locked = false; - applyNodeColors(setNode, getAutolinkColorPreset(colorSet, 'set', separateCol, colorGet)); - applyNodeTitleTextColor(setNode, colorTitles); - - // Ottieni il tipo dall'output del nodo sorgente + // Tipo dell'output sorgente const outputType = srcNode.outputs?.[originSlot]?.type || "*"; let outputSlotName = getSlotName(srcNode, originSlot, true); - - // Ulteriore fix: non permettere mai "*" come chiave if (!outputSlotName || String(outputSlotName).trim() === "*") { if (outputType && outputType !== "*") outputSlotName = String(outputType).trim().toLowerCase(); else outputSlotName = `output_${originSlot}`; } - + console.log(`[IAMCCS AutoLink] Processing: ${srcNode.title || srcNode.type}[${originSlot}] with name "${outputSlotName}"`); - - // Genera nome unico se esiste giΓ  un Set con questo nome. - // Importante: NON ridurre mai "model_0" a "model". - const uniqueName = makeUniqueSetName(outputSlotName); - - // Imposta tipo e nome correttamente - if (setNode.inputs && setNode.inputs[0]) { - setNode.inputs[0].type = outputType; - setNode.inputs[0].name = uniqueName; + + // --- CHECK: is there already a Set node consuming this exact source slot? --- + const sourceKey = `${srcNode.id}_${originSlot}`; + const reuseSet = existingSetBySourceKey.get(sourceKey); + + let setNode; + let uniqueName; + + if (reuseSet) { + // Reuse the existing Set node β€” just add new Get nodes for the new destinations. + setNode = reuseSet; + uniqueName = getAutolinkKey(reuseSet) || makeUniqueSetName(outputSlotName); + console.log(`[IAMCCS AutoLink] β†Ί Reusing existing Set node: "${uniqueName}" for ${srcNode.title || srcNode.type}[${originSlot}]`); + // Make sure the type name widget is still consistent + if (setNode.inputs?.[0]) setNode.inputs[0].type = outputType; + if (setNode.outputs?.[0]) setNode.outputs[0].type = outputType; + } else { + // Create a brand-new Set node + const setPos = findFreePosition( + graph, + srcNode.pos[0] + (srcNode.size?.[0] || 200), + (alignMode === "Proportional" ? getAnchorY(srcNode, originSlot, true) : srcNode.pos[1]), + 20, + occupiedSetPositions, + alignMode + ); + + setNode = createNode(graph, SET_TYPE, setPos[0], setPos[1]); + if (!setNode) continue; + + _iamccsApplyGroupStateToNode( + graph, + setNode, + srcNode, + { x: setPos[0] + 75, y: setPos[1] + 13 } + ); + + setNode.properties = setNode.properties || {}; + setNode.properties.autolink_color_name = colorSet; + if (setNode.properties.autolink_color_locked === undefined) setNode.properties.autolink_color_locked = false; + applyNodeColors(setNode, getAutolinkColorPreset(colorSet, 'set', separateCol, colorGet)); + applyNodeTitleTextColor(setNode, colorTitles); + + // Genera nome unico β€” NON ridurre mai "model_0" a "model" + uniqueName = makeUniqueSetName(outputSlotName); + + if (setNode.inputs?.[0]) { setNode.inputs[0].type = outputType; setNode.inputs[0].name = uniqueName; } + if (setNode.outputs?.[0]) { setNode.outputs[0].type = outputType; setNode.outputs[0].name = uniqueName; } + + setWidgetValue(setNode, "name", uniqueName); + setNode.title = uniqueName; + + console.log(`[IAMCCS AutoLink] βœ“ Created Set node: "${uniqueName}" (from ${srcNode.title || srcNode.type})`); + + const nameWidget = getWidget(setNode, "name"); + if (nameWidget) nameWidget.lastValue = uniqueName; + + // Collapse immediately + setTimeout(() => { + if (setNode.collapse) setNode.collapse(); + setNode.size = [150, 26]; + }, 0); + + // Connect source β†’ Set; any existing sourceβ†’somewhere link on this slot is preserved + // (LiteGraph allows multiple outgoing links; old dstNode links will be replaced when + // we create the Get nodes below and call _iamccsRemoveOtherLinksToTarget). + srcNode.connect(originSlot, setNode, 0); + + // Register for future reuse in this same convertAllLinks call + existingSetBySourceKey.set(sourceKey, setNode); } - if (setNode.outputs && setNode.outputs[0]) { - setNode.outputs[0].type = outputType; - setNode.outputs[0].name = uniqueName; - } - - setWidgetValue(setNode, "name", uniqueName); - setNode.title = `${uniqueName}`; - - console.log(`[IAMCCS AutoLink] βœ“ Created Set node: "${uniqueName}" (from ${srcNode.title || srcNode.type})`); - - // Inizializza lastValue per il tracking delle modifiche - const nameWidget = getWidget(setNode, "name"); - if (nameWidget) nameWidget.lastValue = uniqueName; - - // Collassa - setTimeout(() => { - if (setNode.collapse) setNode.collapse(); - setNode.size = [150, 26]; - }, 0); - - // Collega il Set al nodo sorgente - srcNode.connect(originSlot, setNode, 0); - - // Salva il Set creato per creare i Get dopo + + // Salva il Set (nuovo o riutilizzato) per creare i Get dopo createdSets.set(key, { setNode, outputName: uniqueName, @@ -2733,12 +2853,17 @@ function convertAllLinks( getNode.size = [150, 26]; }, 0); + const ts = Number(targetSlot); + const targetInputName = Number.isFinite(ts) ? (dstNode?.inputs?.[ts]?.name ?? null) : null; + // Salva metadata per restore const metadata = { iamccs_autolink: true, output_name: outputName, origin: { id: srcNode.id, slot: originSlot }, - target: { id: dstNode.id, slot: targetSlot } + target: { id: dstNode.id, slot: targetSlot }, + // Used by Restore Direct Links to avoid slot drift + target_input_name: targetInputName, }; setNode.properties = setNode.properties || {}; @@ -2749,7 +2874,6 @@ function convertAllLinks( getNode.properties.metadata = metadata; // Safe rewire: preserve the previous direct link if any. - const ts = Number(targetSlot); if (!Number.isFinite(ts)) { try { graph.remove(getNode); } catch {} continue; @@ -3059,6 +3183,13 @@ function restoreDirectLinks(graph, options = {}) { const ts = Number(targetSlot); if (!Number.isFinite(os) || !Number.isFinite(ts)) return null; + // Safety: never connect to an out-of-range target slot. + // On many LiteGraph/ComfyUI builds this triggers dynamic input creation, + // which shows up as many inactive/empty inputs after Restore. + if (Array.isArray(dstNode?.inputs)) { + if (ts < 0 || ts >= dstNode.inputs.length) return null; + } + try { srcNode.connect(os, dstNode, ts); const ok = _iamccsDidConnect(srcNode, os, dstNode, ts); @@ -3090,6 +3221,26 @@ function restoreDirectLinks(graph, options = {}) { return null; }; + const resolveTargetSlot = (dstNode, md) => { + if (!dstNode) return null; + const inputs = dstNode.inputs; + + // Prefer restoring by input name when available (slot indices can drift). + const targetName = md?.target_input_name; + if (targetName && Array.isArray(inputs)) { + const idx = inputs.findIndex((i) => i?.name === targetName); + if (idx >= 0) return idx; + } + + const slot = md?.target?.slot; + const ts = Number(slot); + if (!Number.isFinite(ts)) return null; + if (Array.isArray(inputs)) { + if (ts < 0 || ts >= inputs.length) return null; + } + return ts; + }; + const isTargetCurrentlyFromGetNode = (dstNode, targetSlot, getNodeId) => { try { const ts = Number(targetSlot); @@ -3136,7 +3287,7 @@ function restoreDirectLinks(graph, options = {}) { const srcNode = getNodeById(graph, origin.id); const dstNode = getNodeById(graph, target.id); const originSlot = origin.slot; - const targetSlot = target.slot; + const ts = resolveTargetSlot(dstNode, md); if (!srcNode || !dstNode) { failed++; @@ -3144,8 +3295,7 @@ function restoreDirectLinks(graph, options = {}) { continue; } - const ts = Number(targetSlot); - if (!Number.isFinite(ts)) { + if (ts == null) { failed++; if (key) keysWithFailures.add(key); continue; @@ -3232,9 +3382,8 @@ function restoreDirectLinks(graph, options = {}) { const dstNode = getNodeById(graph, outLink.target_id); if (!dstNode) continue; - const targetSlot = outLink.target_slot; - const ts = Number(targetSlot); - if (!Number.isFinite(ts)) continue; + const ts = resolveTargetSlot(dstNode, { target: { slot: outLink.target_slot } }); + if (ts == null) continue; // Already correct? Mark restored so we can remove the AutoLink nodes. try { @@ -3369,65 +3518,12 @@ function restoreDirectLinks(graph, options = {}) { console.log("[IAMCCS AutoLink] Extension loaded"); -// ---- Runtime patch: ensure AutoLink graphs can execute ---- -// AutoLink Set/Get nodes are frontend helpers; the backend nodes are no-op. -// To run a workflow, we temporarily restore direct links before queueing, -// then reload the original graph so the user keeps AutoLink nodes. -function _iamccsPatchQueuePromptForAutolink() { - try { - if (app.__iamccs_autolink_queue_patch_installed) return; - if (typeof app?.queuePrompt !== "function") return; - - const originalQueuePrompt = app.queuePrompt; - app.queuePrompt = async function (...args) { - const graph = app?.graph; - const hasAutoLink = _iamccsGraphNodes(graph).some(n => n?.type === SET_TYPE || n?.type === GET_TYPE); - if (!hasAutoLink) { - return await originalQueuePrompt.apply(this, args); - } - - let snapshot = null; - try { - snapshot = typeof graph.serialize === "function" ? graph.serialize() : null; - } catch (e) { - snapshot = null; - } - - try { - // Convert AutoLink nodes into direct links (and remove them) for execution. - // This mutates the live graph, so we restore from snapshot in finally. - try { - restoreDirectLinks(graph, { removeNodes: false, asyncRemove: false, pruneTargetDuplicates: false }); - } catch (e) { - console.warn("[IAMCCS AutoLink] restoreDirectLinks failed before queuePrompt", e); - } - - try { _iamccsFixLinkIntegrity(graph); } catch {} - - return await originalQueuePrompt.apply(this, args); - } finally { - // Restore the original AutoLink graph for the UI. - if (snapshot) { - try { - if (typeof app.loadGraphData === "function") { - await app.loadGraphData(snapshot); - } else if (typeof graph?.configure === "function") { - graph.configure(snapshot); - graph.setDirtyCanvas?.(true, true); - } - } catch (e) { - console.warn("[IAMCCS AutoLink] Failed to restore graph snapshot after queuePrompt", e); - } - } - } - }; - - app.__iamccs_autolink_queue_patch_installed = true; - console.log("[IAMCCS AutoLink] Patched app.queuePrompt (temporary restore for execution)"); - } catch (e) { - console.warn("[IAMCCS AutoLink] Failed to patch queuePrompt", e); - } -} - -_iamccsPatchQueuePromptForAutolink(); +// Queue patch REMOVED. +// AutoLink Set/Get nodes use isVirtualNode = true + getInputLink(), which is the same +// mechanism as KJ SetNode/GetNode. ComfyUI's graphToPrompt() traces through virtual +// nodes transparently, so no live-graph mutation is needed before queueing. +// The old patch was destructive: it called restoreDirectLinks() (mutating the graph), +// then reloaded the graph via loadGraphData() on every queue, causing duplicated inputs +// on Get nodes and broken link chains. +console.log("[IAMCCS AutoLink] Queue execution via isVirtualNode/getInputLink (no patch needed)");