John Pollock
657fdac13a
Fix WanVideo multi-GPU device mismatch issue
...
Problem: WanVideoWrapper caches device at module load time, causing timesteps
and tensors to be created on wrong device when looping between models on
different GPUs.
Solution: WanVideoSamplerMultiGPU wrapper updates module-level device variable
to match current model's device before sampling.
Changes:
- Added comprehensive logging to trace device allocation through pipeline
- Identified module-level device caching as root cause
- Simplified WanVideoSamplerMultiGPU to only update device variable
- Verified fix works for multi-model workflows with looping
2025-08-06 04:30:03 -05:00
John Pollock
582ca6a247
WanVideoWrapper MultiGPU integration - custom wrapper nodes
...
- Created custom implementations for all WanVideo nodes with explicit device selection
- Added WanVideoBlockSwap with dual device control (swap_device and model_offload_device)
- Created WanVideoModelLoader_TWO for multi-model workflows to avoid race conditions
- Discovered core ComfyUI bug: safetensors loader ignores device index (uses device.type instead of str(device))
- All wrapper nodes use runtime module patching to override WanVideoWrapper's cached device variables
- Extensive logging added for debugging device assignments
2025-08-05 18:59:16 -05:00
John Pollock
a05823ff0a
feat: add CLIPVisionLoaderMultiGPU support and update version to 1.7.3
2025-04-17 18:43:01 -05:00
John Pollock
4ff9b80286
feat: add QuadrupleCLIPLoader / QuadrupleCLIPLoaderGGUF support and update version to 1.7.2
2025-04-17 17:00:20 -05:00
John Pollock
2d81ef0a21
Support for kijai's ComfyUI-WanVideoWrapper
2025-03-23 13:40:05 -05:00
John Pollock
a2093a4fc9
feat: add text encoder device handling, whereas CLIP can sometimes default to CPU, whereas using a DisTorch CLIP load you can load the layes on CPU buy use CUDA for processing. Especially helpful llava-llama
2025-02-12 11:51:58 -06:00
John Pollock
9bd984b420
Update default value for virtual VRAM GB to 4.0 in override_class_with_distorch
2025-02-07 18:27:14 -06:00
John Pollock
de2219c974
Refactor virtual VRAM allocation logic and improve logging format
2025-02-07 18:24:28 -06:00
John Pollock
c98a535435
Refactor logging in DisTorch analysis and update allocation handling for virtual VRAM
2025-02-07 16:09:35 -06:00
John Pollock
3a4c6d50c8
Virtual VRAM "automatic" mode for DisTorch, WIP but working
2025-02-07 15:05:08 -06:00
John Pollock
5a403e638c
MergeFluxLoRAsQuantizeAndLoad, WIP
2025-02-07 04:43:45 -06:00
John Pollock
4a8d70a0d4
refactored to move stable wrapper nodes into nodes.py and remainder in init.py
2025-02-03 09:15:05 -06:00
John Pollock
3e130e3dfb
Remove log_comfy_states function - no longer needed
2025-01-31 06:45:43 -06:00
John Pollock
005b5b1882
This release includes an embeddings adapter for the IP2V part of kijai's CLIP loader for HunyuanVideo. See examples. Bump version to 1.4.3 and update category for HunyuanVideoEmbeddingsAdapter to multigpu; enhance README with new workflow examples for HunyuanVideo GGUF-quantized models.
2025-01-29 11:55:16 -06:00
John Pollock
3260b7e38e
Add HunyuanVideoEmbeddingsAdapter class for using kijai's IP2V conditioning video embeddings in the standard sampler, allowing it to be used with GGUF/DisTorch methods.
2025-01-29 09:25:19 -06:00
3dluvr
379ecce687
Fix check_module_exists() to use folder_paths
...
In Windows, module detection was failing because the method couldn't find the hard-coded custom_nodes/ folder in os.join.path.
We switch to using folder_paths which will return a correct path regardless of the platform.
2025-01-28 19:59:44 -05:00
pollock
c07a345c45
Fix case sensitivity in module check for HunyuanVideoWrapper in __init__.py
2025-01-27 18:32:20 -05:00
John Pollock
7ccea97c52
Update device allocation format and enhance module check for case insensitivity in __init__.py
2025-01-27 17:24:15 -06:00
John Pollock
3dbfcc7135
Chasing down bug causing incorrect patched device with distorch code. Re-integrated distorch into __init__.py as one of the consequences.
2025-01-25 17:49:12 -06:00
John Pollock
624c893942
Refactored into init.py and distorch.py
2025-01-23 13:16:18 -06:00
pollock
43d7d1582d
Refactor UnetLoaderGGUF registration to support MultiGPU and DisTorch versions
2025-01-23 12:00:24 -05:00
pollock
b8f314921c
Refactor imports and logging messages for clarity and consistency
2025-01-23 08:50:48 -05:00
John Pollock
d9899d4df1
Refactor GGUF model patcher and analysis functions to improve device handling and logging
2025-01-21 06:50:10 -06:00
John Pollock
5eb03a220a
Add .vscode/settings.json to .gitignore to exclude IDE-specific settings, DisTorch work in progress. Clip also loaded and distributed. Much WIP.
2025-01-21 02:50:10 -06:00
John Pollock
d8d122f397
Refactor MultiGPU module registration and improve device handling logic towards releasing on :main:
2025-01-20 12:11:12 -06:00
John Pollock
1a3cc9d151
Refactor MultiGPU device handling and improve logging for better traceability
2025-01-20 11:34:14 -06:00
John Pollock
455a4ef3f8
Refactor code structure for improved readability and maintainability
2025-01-20 07:40:33 -06:00
John Pollock
a7c424f238
Continued clean-up of DisTorch code.
2025-01-19 21:04:16 -06:00
John Pollock
018aec49b0
Add detailed logging for MultiGPU device memory allocation analysis
2025-01-19 17:00:26 -06:00
John Pollock
f119787bee
ggml learning
2025-01-19 14:24:11 -06:00
John Pollock
e73c8bd5e3
Add experimental DiffSynth block-swapping support via new GPU offload device
...
- Adds HyVideoModelLoaderDiffSynthMultiGPU node implementing DiffSynth block-swapping
- Introduces offload_device selection for secondary GPU utilization
- Updates documentation with known behaviors and expected OOM patterns
- Adds example workflow demonstrating higher resolution/longer duration video generation
- Maintains backwards compatibility with existing MultiGPU workflows
2025-01-15 06:53:59 -06:00
John Pollock
ba24f572ee
feat: Add DeviceSelectorMultiGPU node and device selection functionality - allowing the linking of one or more MultiGPU nodes to the same cuda device in cases where this would prevent accidental errors if should there be a device mismatch futher along in the pipeline due to non-loader nodes performing device-specific tasks or logic.
2025-01-11 10:03:40 -06:00
John Pollock
5c3c1a7f3b
feat: Implement MultiGPU support for Hunyuan models and add module existence checks
...
Nodes work, investating how determinisitically we MultiGPU can play nice with these nodes.
2025-01-07 12:18:03 -06:00
John Pollock
3ce3598a5d
feat: Add initial MultiGPU support for Pulid model, added module existence checks before creating MultiGPU node variant
2025-01-02 22:11:28 -06:00
John Pollock
edcd5cc612
Merge branch 'main' of https://github.com/pollockjj/ComfyUI-MultiGPU
2025-01-02 21:15:08 -06:00
John Pollock
4aec967384
fix: hard coded all supported nodes. If the parent custom_node is installed then it will inherit the needed functionality at run-time, fully eliminating any load depenencies. No outside custom_nodes need to be pre-loaded and no python is inspected. New MultiGPU variants are created in a self-contained manner.
2025-01-02 21:09:39 -06:00
John Pollock
1d020adbdb
feat: Add hard-coded registration for LTX and Florence2 nodes with module existence checks
2024-12-30 22:22:52 -06:00
John Pollock
1e553cfe22
Add utility function to check module existence before registration for hard-coded MultiGPU nodes
2024-12-30 21:00:48 -06:00
John Pollock
198365dfbe
Add hard-coded registration for LTX and Florence2 nodes in MultiGPU setup for debug purposes.
...
Actual nodes pick up the underlying structure at runtime now that the global NODE_CLASS_MAPPINGS has been updated with their information, I pull it directly from there.
A work-around for the loading sequencing problems, but hopefully one that requrires little upkeep as any changes to the underlying structure is picked-up at runtime.
2024-12-30 20:48:41 -06:00
John Pollock
9464cf6cc1
fix: Enhance error handling during module execution in MultiGPU registration
2024-12-30 15:13:12 -06:00
John Pollock
4ac8d33270
fix: Use local module dictionary instead of global for custom node registration
...
- Added a local_map_name ("NODE_CLASS_MAPPINGS") and retrieve it from the module
immediately after loading.
- If the local dictionary exists, wrap the target nodes from there, rather than
relying on the global dictionary.
- Removed references to GLOBAL_NODE_CLASS_MAPPINGS for custom nodes and replaced
them with the local module mapping lookup.
2024-12-30 11:17:31 -06:00
John Pollock
fcf054006d
Mew method merged in with old code base so less of a shock.
2024-12-30 11:03:55 -06:00
John Pollock
14758b1985
changed methodology (again) to run their code and then scan their LOCAL dict, not the global dict. This will work and be robust I strongly believe
...
Still too much new code, but we'll slim it down later
2024-12-30 10:12:28 -06:00
John Pollock
caee8716b6
Refactor - intermediate step of Implementing MultiGPU node registration and class definition retrieval for custom nodes
2024-12-30 07:54:57 -06:00
John Pollock
58cb0ab59a
Officially adding CheckpointLoaderNF4 support from ComfyUI_bitsandbytes_NF4
2024-12-29 23:14:11 -06:00
John Pollock
2c6fb487d3
editing for clarity
2024-12-28 14:49:02 -06:00
John Pollock
0b3ba045c4
Change loading to a determinsitc process by querying each node. This relies upon each node's idempotency that will need to get checked.
2024-12-28 13:39:23 -06:00
John Pollock
e4b570de43
contunued NF4, working towards general solution.
2024-12-28 12:13:21 -06:00
John Pollock
d9f7ab23e5
debugging NF4 loader issue and sequence loading in general.
...
Have a method that works, broke it, so re-assembling it slowly.
This is the initial commit where N4 is importint correctly.
2024-12-28 12:03:07 -06:00
John Pollock
d099ecd497
Enhance device selection in get_torch_device_patched and override_class to include 'cpu' as a valid option for multi-GPU support
2024-12-27 03:06:50 -06:00