The memory analysis function, `analyze_safetensor_distorch`, has been improved to provide a more accurate and detailed report.
Instead of estimating the number of swapped blocks based on an average size, the function now receives the actual list of blocks being swapped. It generates a per-block table detailing each block's ID, type, size, and its final assignment (COMPUTE or SWAP).
This provides users with a precise breakdown of the memory offload, reflecting the actual state of the model rather than a theoretical calculation.
Additionally, the unused `log_memory_usage` helper function has been removed.
This commit completely rewrites the block swapping implementation for improved stability, correctness, and code structure.
Key changes:
- Replaces the fragile monkey-patching of the `forward` method with the standard PyTorch `register_forward_pre_hook`.
- Introduces a `BlockSwapManager` class to encapsulate all swapping logic, separating it from the ComfyUI node.
- Implements a "Sequential Swapping" strategy: the previously active block is offloaded before the current block is loaded, ensuring only one block is on the active device at a time.
- Adds a `cleanup` method to properly remove hooks after execution, preventing state leakage between runs.
- Fixes a critical bug where the hook signature was incorrect.
- Adds a memory logging utility for easier debugging.
This commit removes the unused `reserved_swap_gb` parameter from the `analyze_safetensor_distorch` function and its call sites. This simplifies the function's signature and cleans up the analysis output by removing the "Reserve" metric, which was always zero.
Additionally, the `.gitignore` file is simplified by removing entries for the `binaries/` directory, which are no longer needed.
This commit introduces a major update, "DisTorch v2", which integrates the new `BlockSwap` system for more efficient and dynamic memory management across multiple GPUs.
Key changes:
- **BlockSwap Integration:** GGUF model loading is completely refactored to use `BlockSwap`, enabling more intelligent VRAM allocation based on tensor analysis.
- **Node Renaming:** All SafeTensor loader nodes are renamed from `...DisTorchMultiGPU` to `...DisTorch2MultiGPU` to clearly distinguish the new implementation from the old one.
- **Legacy Support:** The previous GGUF loader is preserved as a legacy option for backward compatibility.
- **Improved Memory Calculation:** A more accurate memory calculation function (`get_total_memory_v2`) is implemented and used by the new system.
This commit refactors the codebase by extracting major components from the main `__init__.py` file into their own dedicated modules. This improves code organization, readability, and maintainability.
- **`distorch.py`**: New file containing the `DisTorch` class, which manages multi-GPU device patching and distribution logic.
- **`block_swap.py`**: New file containing the generic `BlockSwap` class for UNet block swapping to manage VRAM.
- **`wanvideo.py`**: New file containing the `WanVideoBlockSwap` class, a specialized implementation for WanVideo models.
- **`__init__.py`**: Simplified to handle node registration and imports from the new modules.
- Added comprehensive file logging with timestamps
- Log file created in logs/distorch_TIMESTAMP.log
- Debug logging for all major operations:
- Model structure analysis
- Memory calculations
- Block partitioning strategy
- Device movements and swaps
- Hook installations
- Performance metrics
- Console output for important INFO messages
- Full exception tracebacks captured
- Simplified node naming from DisTorchBlockSwap to DisTorch
- Cleaned up accidentally added main_branch directory
- Updated all references in __init__.py and core/blockswap.py
- Created comprehensive architecture documentation (ARCHITECTURE_V2.0.0.md)
- Added DOE optimization planning document (DOE_OPTIMIZATION.md)
- Implemented DisTorchBlockSwap node for safetensor models
- Created core/blockswap.py with BlockSwapManager
- Unified VirtualVRAM interface design
- Based on analysis of WanVideo's block swap methodology
Populate 'type' options by sourcing from core nodes to avoid drift:\n- CLIPLoaderGGUF now derives 'type' from nodes.CLIPLoader.INPUT_TYPES()\n- DualCLIPLoaderGGUF now derives 'type' from nodes.DualCLIPLoader.INPUT_TYPES()\nThis fixes missing or outdated 'type' options in GGUF Single and Dual CLIP loaders.\n\nchore: bump version to 1.8.2
Add guarded Intel XPU support alongside CUDA:
- get_device_list now includes xpu:N when available
- device selection (model/text encoder) considers CUDA or XPU and validates devices
- DisTorch donor/offload selection includes xpu devices
Also: remove unused MergeFluxLoRAs node and mapping; delete tools/ and precompiled_binaries/; bump project version to 1.8.1.
Problem: WanVideoWrapper caches device at module load time, causing timesteps
and tensors to be created on wrong device when looping between models on
different GPUs.
Solution: WanVideoSamplerMultiGPU wrapper updates module-level device variable
to match current model's device before sampling.
Changes:
- Added comprehensive logging to trace device allocation through pipeline
- Identified module-level device caching as root cause
- Simplified WanVideoSamplerMultiGPU to only update device variable
- Verified fix works for multi-model workflows with looping
- Created custom implementations for all WanVideo nodes with explicit device selection
- Added WanVideoBlockSwap with dual device control (swap_device and model_offload_device)
- Created WanVideoModelLoader_TWO for multi-model workflows to avoid race conditions
- Discovered core ComfyUI bug: safetensors loader ignores device index (uses device.type instead of str(device))
- All wrapper nodes use runtime module patching to override WanVideoWrapper's cached device variables
- Extensive logging added for debugging device assignments