Commit Graph
50 Commits
Author SHA1 Message Date
John Pollock de2219c974 Refactor virtual VRAM allocation logic and improve logging format 2025-02-07 18:24:28 -06:00
John Pollock c98a535435 Refactor logging in DisTorch analysis and update allocation handling for virtual VRAM 2025-02-07 16:09:35 -06:00
John Pollock 3a4c6d50c8 Virtual VRAM "automatic" mode for DisTorch, WIP but working 2025-02-07 15:05:08 -06:00
John Pollock 5a403e638c MergeFluxLoRAsQuantizeAndLoad, WIP 2025-02-07 04:43:45 -06:00
John Pollock 4a8d70a0d4 refactored to move stable wrapper nodes into nodes.py and remainder in init.py 2025-02-03 09:15:05 -06:00
John Pollock 3e130e3dfb Remove log_comfy_states function - no longer needed 2025-01-31 06:45:43 -06:00
John Pollock 005b5b1882 This release includes an embeddings adapter for the IP2V part of kijai's CLIP loader for HunyuanVideo. See examples. Bump version to 1.4.3 and update category for HunyuanVideoEmbeddingsAdapter to multigpu; enhance README with new workflow examples for HunyuanVideo GGUF-quantized models. 2025-01-29 11:55:16 -06:00
John Pollock 3260b7e38e Add HunyuanVideoEmbeddingsAdapter class for using kijai's IP2V conditioning video embeddings in the standard sampler, allowing it to be used with GGUF/DisTorch methods. 2025-01-29 09:25:19 -06:00
3dluvr 379ecce687 Fix check_module_exists() to use folder_paths
In Windows, module detection was failing because the method couldn't find the hard-coded custom_nodes/ folder in os.join.path.

We switch to using folder_paths which will return a correct path regardless of the platform.
2025-01-28 19:59:44 -05:00
pollock c07a345c45 Fix case sensitivity in module check for HunyuanVideoWrapper in __init__.py 2025-01-27 18:32:20 -05:00
John Pollock 7ccea97c52 Update device allocation format and enhance module check for case insensitivity in __init__.py 2025-01-27 17:24:15 -06:00
John Pollock 3dbfcc7135 Chasing down bug causing incorrect patched device with distorch code. Re-integrated distorch into __init__.py as one of the consequences. 2025-01-25 17:49:12 -06:00
John Pollock 624c893942 Refactored into init.py and distorch.py 2025-01-23 13:16:18 -06:00
pollock 43d7d1582d Refactor UnetLoaderGGUF registration to support MultiGPU and DisTorch versions 2025-01-23 12:00:24 -05:00
pollock b8f314921c Refactor imports and logging messages for clarity and consistency 2025-01-23 08:50:48 -05:00
John Pollock d9899d4df1 Refactor GGUF model patcher and analysis functions to improve device handling and logging 2025-01-21 06:50:10 -06:00
John Pollock 5eb03a220a Add .vscode/settings.json to .gitignore to exclude IDE-specific settings, DisTorch work in progress. Clip also loaded and distributed. Much WIP. 2025-01-21 02:50:10 -06:00
John Pollock d8d122f397 Refactor MultiGPU module registration and improve device handling logic towards releasing on :main: 2025-01-20 12:11:12 -06:00
John Pollock 1a3cc9d151 Refactor MultiGPU device handling and improve logging for better traceability 2025-01-20 11:34:14 -06:00
John Pollock 455a4ef3f8 Refactor code structure for improved readability and maintainability 2025-01-20 07:40:33 -06:00
John Pollock a7c424f238 Continued clean-up of DisTorch code. 2025-01-19 21:04:16 -06:00
John Pollock 018aec49b0 Add detailed logging for MultiGPU device memory allocation analysis 2025-01-19 17:00:26 -06:00
John Pollock f119787bee ggml learning 2025-01-19 14:24:11 -06:00
John Pollock e73c8bd5e3 Add experimental DiffSynth block-swapping support via new GPU offload device
- Adds HyVideoModelLoaderDiffSynthMultiGPU node implementing DiffSynth block-swapping
- Introduces offload_device selection for secondary GPU utilization
- Updates documentation with known behaviors and expected OOM patterns
- Adds example workflow demonstrating higher resolution/longer duration video generation
- Maintains backwards compatibility with existing MultiGPU workflows
2025-01-15 06:53:59 -06:00
John Pollock ba24f572ee feat: Add DeviceSelectorMultiGPU node and device selection functionality - allowing the linking of one or more MultiGPU nodes to the same cuda device in cases where this would prevent accidental errors if should there be a device mismatch futher along in the pipeline due to non-loader nodes performing device-specific tasks or logic. 2025-01-11 10:03:40 -06:00
John Pollock 5c3c1a7f3b feat: Implement MultiGPU support for Hunyuan models and add module existence checks
Nodes work, investating how determinisitically we MultiGPU can play nice with these nodes.
2025-01-07 12:18:03 -06:00
John Pollock 3ce3598a5d feat: Add initial MultiGPU support for Pulid model, added module existence checks before creating MultiGPU node variant 2025-01-02 22:11:28 -06:00
John Pollock edcd5cc612 Merge branch 'main' of https://github.com/pollockjj/ComfyUI-MultiGPU 2025-01-02 21:15:08 -06:00
John Pollock 4aec967384 fix: hard coded all supported nodes. If the parent custom_node is installed then it will inherit the needed functionality at run-time, fully eliminating any load depenencies. No outside custom_nodes need to be pre-loaded and no python is inspected. New MultiGPU variants are created in a self-contained manner. 2025-01-02 21:09:39 -06:00
John Pollock 1d020adbdb feat: Add hard-coded registration for LTX and Florence2 nodes with module existence checks 2024-12-30 22:22:52 -06:00
John Pollock 1e553cfe22 Add utility function to check module existence before registration for hard-coded MultiGPU nodes 2024-12-30 21:00:48 -06:00
John Pollock 198365dfbe Add hard-coded registration for LTX and Florence2 nodes in MultiGPU setup for debug purposes.
Actual nodes pick up the underlying structure at runtime now that the global NODE_CLASS_MAPPINGS has been updated with their information, I pull it directly from there.

A work-around for the loading sequencing problems, but hopefully one that requrires little upkeep as any changes to the underlying structure is picked-up at runtime.
2024-12-30 20:48:41 -06:00
John Pollock 9464cf6cc1 fix: Enhance error handling during module execution in MultiGPU registration 2024-12-30 15:13:12 -06:00
John Pollock 4ac8d33270 fix: Use local module dictionary instead of global for custom node registration
- Added a local_map_name ("NODE_CLASS_MAPPINGS") and retrieve it from the module
  immediately after loading.
- If the local dictionary exists, wrap the target nodes from there, rather than
  relying on the global dictionary.
- Removed references to GLOBAL_NODE_CLASS_MAPPINGS for custom nodes and replaced
  them with the local module mapping lookup.
2024-12-30 11:17:31 -06:00
John Pollock fcf054006d Mew method merged in with old code base so less of a shock. 2024-12-30 11:03:55 -06:00
John Pollock 14758b1985 changed methodology (again) to run their code and then scan their LOCAL dict, not the global dict. This will work and be robust I strongly believe
Still too much new code, but we'll slim it down later
2024-12-30 10:12:28 -06:00
John Pollock caee8716b6 Refactor - intermediate step of Implementing MultiGPU node registration and class definition retrieval for custom nodes 2024-12-30 07:54:57 -06:00
John Pollock 58cb0ab59a Officially adding CheckpointLoaderNF4 support from ComfyUI_bitsandbytes_NF4 2024-12-29 23:14:11 -06:00
John Pollock 2c6fb487d3 editing for clarity 2024-12-28 14:49:02 -06:00
John Pollock 0b3ba045c4 Change loading to a determinsitc process by querying each node. This relies upon each node's idempotency that will need to get checked. 2024-12-28 13:39:23 -06:00
John Pollock e4b570de43 contunued NF4, working towards general solution. 2024-12-28 12:13:21 -06:00
John Pollock d9f7ab23e5 debugging NF4 loader issue and sequence loading in general.
Have a method that works, broke it, so re-assembling it slowly.

This is the initial commit where N4 is importint correctly.
2024-12-28 12:03:07 -06:00
John Pollock d099ecd497 Enhance device selection in get_torch_device_patched and override_class to include 'cpu' as a valid option for multi-GPU support 2024-12-27 03:06:50 -06:00
John Pollock dead358227 Add MMAudioSampler to experimental audio model loaders in __init__.py so it stays in sync with the Model and FeatureUtils Loaders. I suspect this is because it is querying the cuda device directly. This would give the wrong answer half the time depending on how the other loaders were ran via the sequencing logic.
This will sync up those cuda device queries with the patched version from MultiGPU.
2024-12-26 14:59:16 -06:00
John Pollock 94cf7d9779 Update project description and version in pyproject.toml; add experimental audio model loaders in __init__.py 2024-12-26 08:30:27 -06:00
John Pollock 19c0abd5ec Updated example workflows, added additional experimental workflows, added coverage for LTX loader 2024-12-23 15:39:23 -06:00
John Pollock 56e2e61e52 refactor: streamline MultiGPU node initialization
Code Changes (__init__.py):
- Switches from deepcopy to standard copy for better efficiency
- Initializes current_device from model_management instead of hardcoded "cuda:0"
- Simplifies node class mapping by removing intermediate TARGET_NODE_CLASS_MAPPINGS
- Updates node naming convention: removes underscore from MultiGPU suffix
- Adds new supported nodes: "CheckpointLoaderSimple", "ControlNetLoader", "LoadFluxControlNet"
- Improves code organization with better comment clarity

COMPATIBILITY NOTE: This version restores backward compatibility with workflows
using the previous node naming scheme. Both old and new node names will work.

Documentation Changes (README.md):
- Updates node list to reflect automatic detection system
- Adds proper attribution links for required dependencies (ComfyUI-GGUF, x-flux-comfy)
- Links to example quantized models like flux1-dev-gguf
- Reorganizes loader sections with clear dependency requirements
- Updates support links to new maintainer
- Removes business/commercial references
- Updates credits section to reflect current project status
2024-12-22 22:36:59 -06:00
John Pollock a652a7f264 refactor: adopt City96's streamlined MultiGPU implementation
BREAKING CHANGE: This replaces the existing MultiGPU implementation with a new system
designed by City96 (https://v100s.net/). Users with existing workflows using the older
MultiGPU nodes will need to update their workflows.

The new implementation by City96 significantly improves the codebase:
- Replaces manual class definitions with a dynamic class override system
- Reduces code from 400+ lines to ~50 through smart use of inheritance
- Handles all loader types via TARGET_NODE_NAMES configuration
- Provides consistent behavior across standard and GGUF loaders
- Creates a unified "multigpu" category for better organization

The system creates MultiGPU versions of nodes by wrapping original classes through
an override mechanism, maintaining functionality while adding device selection.

Migration:
- Replace old MultiGPU nodes with new versions (same names with "_MultiGPU" suffix)
- Node functionality remains the same, only the implementation has changed
2024-12-22 08:28:35 -06:00
John Pollock 186723f34d Add mochi and ltxv support to CLIPLoaderMultiGPU 2024-12-18 17:01:01 -06:00
Alexander Dzhoganov b5462877f3 Initial commit 2024-08-04 19:36:59 +03:00