Problem: WanVideoWrapper caches device at module load time, causing timesteps
and tensors to be created on wrong device when looping between models on
different GPUs.
Solution: WanVideoSamplerMultiGPU wrapper updates module-level device variable
to match current model's device before sampling.
Changes:
- Added comprehensive logging to trace device allocation through pipeline
- Identified module-level device caching as root cause
- Simplified WanVideoSamplerMultiGPU to only update device variable
- Verified fix works for multi-model workflows with looping
- Created custom implementations for all WanVideo nodes with explicit device selection
- Added WanVideoBlockSwap with dual device control (swap_device and model_offload_device)
- Created WanVideoModelLoader_TWO for multi-model workflows to avoid race conditions
- Discovered core ComfyUI bug: safetensors loader ignores device index (uses device.type instead of str(device))
- All wrapper nodes use runtime module patching to override WanVideoWrapper's cached device variables
- Extensive logging added for debugging device assignments
If you were using nodes with DiffSynth in their name (like ...DiffSynthMultiGPU), please switch to the standard MultiGPU versions for now (e.g., ...MultiGPU). This change eliminates a device management issue that was affecting some Windows users.
See: https://github.com/pollockjj/ComfyUI-MultiGPU/issues/13
Most users won't be affected as this only impacts the DiffSynth variants of nodes.
If you need help modifying your workflows, please open an issue. Will revisit DiffSynth functionality once issue can be contained or worked-around
Update comfyregistry to 1.5.0 to reflect major change in functionality, in this case, a reduction.
These binaries are built from https://github.com/pollockjj/llama.cpp/tree/flux-quant-b3600
Linux binary (SHA256: 846c0bab3c7f7c6729f22b6229a2db5da2da63a4549ef8bf76546092f56fec15):
- Ubuntu 22.04
- gcc 13.3.0
- Debug build with flags: --config Debug -j10 --target llama-quantize
Windows binary (SHA256: 121c02184fcf30dc4ff3dcf4a69177e44aced9d7ba8c77f49cd138c2d73f7e6b):
- Windows 11
- MSVC 19.42.34436.0
- Debug build with flags: -G "Visual Studio 17 2022" -A x64 -DBUILD_SHARED_LIBS=OFF
Binaries are verified and released at:
https://github.com/pollockjj/llama.cpp/releases/tag/1.0.0
In Windows, module detection was failing because the method couldn't find the hard-coded custom_nodes/ folder in os.join.path.
We switch to using folder_paths which will return a correct path regardless of the platform.