docs: Add documentation for standard, gguf, and DisTorch nodes/wrappers
This commit is contained in:
@@ -0,0 +1,52 @@
|
||||
# CLIPLoaderDisTorch2MultiGPU
|
||||
|
||||
The `CLIPLoaderDisTorch2MultiGPU` node is used to load standard CLIP text encoder models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger text encoding models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name` | `STRING` | The name of the CLIP model to load. |
|
||||
| `type` | `STRING` | The type of CLIP model (e.g., 'stable_diffusion', 'stable_diffusion_xl'). |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded CLIP text encoder model with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,52 @@
|
||||
# CLIPLoaderGGUFDisTorch2MultiGPU
|
||||
|
||||
The `CLIPLoaderGGUFDisTorch2MultiGPU` node is used to load GGUF format CLIP text encoder models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger text encoding models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` and `ComfyUI/models/clip_gguf` folders, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name` | `STRING` | The name of the CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `type` | `STRING` | The type of CLIP model (e.g., 'stable_diffusion', 'stable_diffusion_xl'). |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded CLIP text encoder model with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,19 @@
|
||||
# CLIPLoaderGGUFMultiGPU
|
||||
|
||||
The `CLIPLoaderGGUFMultiGPU` node is used to load GGUF format CLIP text encoder models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` and `ComfyUI/models/clip_gguf` folders, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name` | `STRING` | The name of the CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `type` | `STRING` | The type of CLIP model (e.g., 'stable_diffusion', 'stable_diffusion_xl'). |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded CLIP text encoder model. |
|
||||
@@ -0,0 +1,19 @@
|
||||
# CLIPLoaderMultiGPU
|
||||
|
||||
The `CLIPLoaderMultiGPU` node is used to load CLIP text encoder models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name` | `STRING` | The name of the CLIP model to load. |
|
||||
| `type` | `STRING` | The type of CLIP model (e.g., 'stable_diffusion', 'stable_diffusion_xl'). |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded CLIP text encoder model. |
|
||||
@@ -0,0 +1,51 @@
|
||||
# CLIPVisionLoaderDisTorch2MultiGPU
|
||||
|
||||
The `CLIPVisionLoaderDisTorch2MultiGPU` node is used to load CLIP Vision models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger vision encoder models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip_vision` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_vision` | `STRING` | The name of the CLIP Vision model to load. |
|
||||
| `device` | `STRING` | Target device for vision encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP_VISION` | `CLIP_VISION` | The loaded CLIP Vision model with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,18 @@
|
||||
# CLIPVisionLoaderMultiGPU
|
||||
|
||||
The `CLIPVisionLoaderMultiGPU` node is used to load CLIP Vision models with device selection capability, enabling users to specify which GPU or device should be used for vision encoder execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip_vision` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_vision` | `STRING` | The name of the CLIP Vision model to load. |
|
||||
| `device` | `STRING` | Target device for vision encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP_VISION` | `CLIP_VISION` | The loaded CLIP Vision model. |
|
||||
@@ -0,0 +1,53 @@
|
||||
# CheckpointLoaderSimpleDisTorch2MultiGPU
|
||||
|
||||
The `CheckpointLoaderSimpleDisTorch2MultiGPU` node is used to load checkpoint models (complete diffusion models containing UNet, CLIP, and VAE components) with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/checkpoints` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `ckpt_name` | `STRING` | The name of the checkpoint model to load. |
|
||||
| `compute_device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded UNet diffusion model with DisTorch2 distributed allocation applied. |
|
||||
| `CLIP` | `CLIP` | The loaded CLIP text encoder model. |
|
||||
| `VAE` | `VAE` | The loaded VAE decoder/encoder model. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `compute_device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,20 @@
|
||||
# CheckpointLoaderSimpleMultiGPU
|
||||
|
||||
The `CheckpointLoaderSimpleMultiGPU` node is used to load checkpoint models (complete diffusion models containing UNet, CLIP, and VAE components) with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/checkpoints` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `ckpt_name` | `STRING` | The name of the checkpoint model to load. |
|
||||
| `device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded UNet diffusion model. |
|
||||
| `CLIP` | `CLIP` | The loaded CLIP text encoder model. |
|
||||
| `VAE` | `VAE` | The loaded VAE decoder/encoder model. |
|
||||
@@ -0,0 +1,51 @@
|
||||
# ControlNetLoaderDisTorch2MultiGPU
|
||||
|
||||
The `ControlNetLoaderDisTorch2MultiGPU` node is used to load ControlNet models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger conditional generation models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/controlnet` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `control_net_name` | `STRING` | The name of the ControlNet model to load. |
|
||||
| `compute_device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CONTROL_NET` | `CONTROL_NET` | The loaded ControlNet model with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `compute_device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,18 @@
|
||||
# ControlNetLoaderMultiGPU
|
||||
|
||||
The `ControlNetLoaderMultiGPU` node is used to load ControlNet models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/controlnet` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `control_net_name` | `STRING` | The name of the ControlNet model to load. |
|
||||
| `device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CONTROL_NET` | `CONTROL_NET` | The loaded ControlNet model. |
|
||||
@@ -0,0 +1,51 @@
|
||||
# DiffControlNetLoaderDisTorch2MultiGPU
|
||||
|
||||
The `DiffControlNetLoaderDisTorch2MultiGPU` node is used to load Diffusers ControlNet models (HuggingFace Hub repositories) with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger conditional generation models across multiple GPUs.
|
||||
|
||||
This node loads ControlNet models directly from HuggingFace model repositories by specifying the repository ID (e.g., "diffusers/controlnet-canny-sdxl-1.0").
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `model_path` | `STRING` | The HuggingFace repository ID or local path of the diffusers ControlNet model to load. |
|
||||
| `compute_device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CONTROL_NET` | `CONTROL_NET` | The loaded diffusers ControlNet model with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `compute_device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,18 @@
|
||||
# DiffControlNetLoaderMultiGPU
|
||||
|
||||
The `DiffControlNetLoaderMultiGPU` node is used to load Diffusers ControlNet models (HuggingFace Hub repositories) with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node loads ControlNet models directly from HuggingFace model repositories by specifying the repository ID (e.g., "diffusers/controlnet-canny-sdxl-1.0").
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `model_path` | `STRING` | The HuggingFace repository ID or local path of the diffusers ControlNet model to load. |
|
||||
| `device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CONTROL_NET` | `CONTROL_NET` | The loaded diffusers ControlNet model. |
|
||||
@@ -0,0 +1,51 @@
|
||||
# DiffusersLoaderDisTorch2MultiGPU
|
||||
|
||||
The `DiffusersLoaderDisTorch2MultiGPU` node is used to load Diffusers models (HuggingFace Hub repositories) with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger diffusion models across multiple GPUs.
|
||||
|
||||
This node loads models directly from HuggingFace model repositories by specifying the repository ID (e.g., "stabilityai/stable-diffusion-xl-base-1.0").
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `model_path` | `STRING` | The HuggingFace repository ID or local path of the diffusers model to load (e.g., 'stabilityai/stable-diffusion-xl-base-1.0'). |
|
||||
| `compute_device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded diffusers model with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `compute_device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,18 @@
|
||||
# DiffusersLoaderMultiGPU
|
||||
|
||||
The `DiffusersLoaderMultiGPU` node is used to load Diffusers models (HuggingFace Hub repositories) with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node loads models directly from HuggingFace model repositories by specifying the repository ID (e.g., "stabilityai/stable-diffusion-xl-base-1.0").
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `model_path` | `STRING` | The HuggingFace repository ID or local path of the diffusers model to load (e.g., 'stabilityai/stable-diffusion-xl-base-1.0'). |
|
||||
| `device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded diffusers model. |
|
||||
@@ -0,0 +1,53 @@
|
||||
# DualCLIPLoaderDisTorch2MultiGPU
|
||||
|
||||
The `DualCLIPLoaderDisTorch2MultiGPU` node is used to load dual standard CLIP text encoder models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger text encoding models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load. |
|
||||
| `type` | `STRING` | The type of CLIP model configuration for dual loading. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded dual CLIP text encoder models with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,53 @@
|
||||
# DualCLIPLoaderGGUFDisTorch2MultiGPU
|
||||
|
||||
The `DualCLIPLoaderGGUFDisTorch2MultiGPU` node is used to load dual GGUF format CLIP text encoder models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger text encoding models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` and `ComfyUI/models/clip_gguf` folders, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `type` | `STRING` | The type of CLIP model configuration for dual loading. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded dual CLIP text encoder models with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,20 @@
|
||||
# DualCLIPLoaderGGUFMultiGPU
|
||||
|
||||
The `DualCLIPLoaderGGUFMultiGPU` node is used to load dual GGUF format CLIP text encoder models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` and `ComfyUI/models/clip_gguf` folders, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `type` | `STRING` | The type of CLIP model configuration for dual loading. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded dual CLIP text encoder models. |
|
||||
@@ -0,0 +1,20 @@
|
||||
# DualCLIPLoaderMultiGPU
|
||||
|
||||
The `DualCLIPLoaderMultiGPU` node is used to load dual CLIP text encoder models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load. |
|
||||
| `type` | `STRING` | The type of CLIP model configuration for dual loading. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded dual CLIP text encoder models. |
|
||||
@@ -0,0 +1,54 @@
|
||||
# QuadrupleCLIPLoaderDisTorch2MultiGPU
|
||||
|
||||
The `QuadrupleCLIPLoaderDisTorch2MultiGPU` node is used to load quadruple standard CLIP text encoder models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger text encoding models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load. |
|
||||
| `clip_name3` | `STRING` | The name of the third CLIP model to load. |
|
||||
| `clip_name4` | `STRING` | The name of the fourth CLIP model to load. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded quadruple CLIP text encoder models with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,54 @@
|
||||
# QuadrupleCLIPLoaderGGUFDisTorch2MultiGPU
|
||||
|
||||
The `QuadrupleCLIPLoaderGGUFDisTorch2MultiGPU` node is used to load quadruple GGUF format CLIP text encoder models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger text encoding models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` and `ComfyUI/models/clip_gguf` folders, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name3` | `STRING` | The name of the third CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name4` | `STRING` | The name of the fourth CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded quadruple CLIP text encoder models with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,21 @@
|
||||
# QuadrupleCLIPLoaderGGUFMultiGPU
|
||||
|
||||
The `QuadrupleCLIPLoaderGGUFMultiGPU` node is used to load quadruple GGUF format CLIP text encoder models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` and `ComfyUI/models/clip_gguf` folders, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name3` | `STRING` | The name of the third CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name4` | `STRING` | The name of the fourth CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded quadruple CLIP text encoder models. |
|
||||
@@ -0,0 +1,21 @@
|
||||
# QuadrupleCLIPLoaderMultiGPU
|
||||
|
||||
The `QuadrupleCLIPLoaderMultiGPU` node is used to load quadruple CLIP text encoder models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load. |
|
||||
| `clip_name3` | `STRING` | The name of the third CLIP model to load. |
|
||||
| `clip_name4` | `STRING` | The name of the fourth CLIP model to load. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded quadruple CLIP text encoder models. |
|
||||
@@ -0,0 +1,53 @@
|
||||
# TripleCLIPLoaderDisTorch2MultiGPU
|
||||
|
||||
The `TripleCLIPLoaderDisTorch2MultiGPU` node is used to load triple standard CLIP text encoder models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger text encoding models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load. |
|
||||
| `clip_name3` | `STRING` | The name of the third CLIP model to load. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded triple CLIP text encoder models configured for SD3 with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,53 @@
|
||||
# TripleCLIPLoaderGGUFDisTorch2MultiGPU
|
||||
|
||||
The `TripleCLIPLoaderGGUFDisTorch2MultiGPU` node is used to load triple GGUF format CLIP text encoder models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger text encoding models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` and `ComfyUI/models/clip_gguf` folders, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name3` | `STRING` | The name of the third CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded triple CLIP text encoder models configured for SD3 with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,20 @@
|
||||
# TripleCLIPLoaderGGUFMultiGPU
|
||||
|
||||
The `TripleCLIPLoaderGGUFMultiGPU` node is used to load triple GGUF format CLIP text encoder models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` and `ComfyUI/models/clip_gguf` folders, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `clip_name3` | `STRING` | The name of the third CLIP model to load from combined clip and clip_gguf folders. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded triple CLIP text encoder models configured for SD3. |
|
||||
@@ -0,0 +1,20 @@
|
||||
# TripleCLIPLoaderMultiGPU
|
||||
|
||||
The `TripleCLIPLoaderMultiGPU` node is used to load triple CLIP text encoder models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/clip` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `clip_name1` | `STRING` | The name of the first CLIP model to load. |
|
||||
| `clip_name2` | `STRING` | The name of the second CLIP model to load. |
|
||||
| `clip_name3` | `STRING` | The name of the third CLIP model to load. |
|
||||
| `device` | `STRING` | Target device for text encoder compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `CLIP` | `CLIP` | The loaded triple CLIP text encoder models configured for SD3. |
|
||||
@@ -0,0 +1,51 @@
|
||||
# UNETLoaderDisTorch2MultiGPU
|
||||
|
||||
The `UNETLoaderDisTorch2MultiGPU` node is used to load UNet diffusion models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/unet` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `unet_name` | `STRING` | The name of the UNet model to load. |
|
||||
| `compute_device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded UNet model with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `compute_device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,18 @@
|
||||
# UNETLoaderMultiGPU
|
||||
|
||||
The `UNETLoaderMultiGPU` node is used to load diffusion model UNet components with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/unet` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `unet_name` | `STRING` | The name of the UNet model to load. |
|
||||
| `device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded UNet diffusion model. |
|
||||
@@ -0,0 +1,54 @@
|
||||
# UnetLoaderGGUFAdvancedDisTorch2MultiGPU
|
||||
|
||||
The `UnetLoaderGGUFAdvancedDisTorch2MultiGPU` node is used to load GGUF format UNet models with advanced quantization options and DisTorch2 distributed tensor allocation, enabling sophisticated multi-device VRAM management across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/unet_gguf` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `unet_name` | `STRING` | The name of the GGUF format UNet model to load. |
|
||||
| `dequant_dtype` | `STRING` | Target data type for model dequantization during loading (options: 'default', 'target', 'float32', 'float16', 'bfloat16'). |
|
||||
| `patch_dtype` | `STRING` | Data type for LoRA patches applied to the model (options: 'default', 'target', 'float32', 'float16', 'bfloat16'). |
|
||||
| `patch_on_device` | `BOOLEAN` | Whether to apply LoRA patches directly on the target device. |
|
||||
| `compute_device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded GGUF format UNet model with advanced quantization settings and DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `compute_device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,21 @@
|
||||
# UnetLoaderGGUFAdvancedMultiGPU
|
||||
|
||||
The `UnetLoaderGGUFAdvancedMultiGPU` node is used to load GGUF format UNet models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/unet_gguf` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `unet_name` | `STRING` | The name of the GGUF format UNet model to load. |
|
||||
| `dequant_dtype` | `STRING` | Target data type for model dequantization during loading (options: 'default', 'target', 'float32', 'float16', 'bfloat16'). |
|
||||
| `patch_dtype` | `STRING` | Data type for LoRA patches applied to the model (options: 'default', 'target', 'float32', 'float16', 'bfloat16'). |
|
||||
| `patch_on_device` | `BOOLEAN` | Whether to apply LoRA patches directly on the target device. |
|
||||
| `device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded GGUF format UNet model with advanced quantization settings. |
|
||||
@@ -0,0 +1,51 @@
|
||||
# UnetLoaderGGUFDisTorch2MultiGPU
|
||||
|
||||
The `UnetLoaderGGUFDisTorch2MultiGPU` node is used to load GGUF format UNet models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/unet_gguf` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `unet_name` | `STRING` | The name of the GGUF format UNet model to load. |
|
||||
| `compute_device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded GGUF format UNet model with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `compute_device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,18 @@
|
||||
# UnetLoaderGGUFMultiGPU
|
||||
|
||||
The `UnetLoaderGGUFMultiGPU` node is used to load GGUF format UNet models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/unet_gguf` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `unet_name` | `STRING` | The name of the GGUF format UNet model to load. |
|
||||
| `device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `MODEL` | `MODEL` | The loaded GGUF format UNet model. |
|
||||
@@ -0,0 +1,51 @@
|
||||
# VAELoaderDisTorch2MultiGPU
|
||||
|
||||
The `VAELoaderDisTorch2MultiGPU` node is used to load VAE (Variational Autoencoder) models with DisTorch2 distributed tensor allocation, enabling advanced multi-device VRAM management to handle larger models across multiple GPUs.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/vae` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `vae_name` | `STRING` | The name of the VAE model to load. |
|
||||
| `compute_device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
| `virtual_vram_gb` | `FLOAT` | Amount of virtual VRAM in gigabytes to allocate for distributed tensor management (default: 4.0, range: 0.0-128.0). |
|
||||
| `donor_device` | `STRING` | Device to donate VRAM from when allocating virtual memory (default: 'cpu'). |
|
||||
| `expert_mode_allocations` | `STRING` | Advanced allocation string for expert users to manually specify device/ratio distributions (e.g., 'cuda:0,50%;cpu,*'). |
|
||||
| `keep_loaded` | `BOOLEAN` | Whether to keep the model loaded when triggering memory cleanup operations (default: true). |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `VAE` | `VAE` | The loaded VAE decoder/encoder with DisTorch2 distributed allocation applied. |
|
||||
|
||||
## DisTorch2 Distributed Loading
|
||||
|
||||
DisTorch2 is an advanced memory management system that enables loading and running large diffusion models across multiple GPUs by intelligently distributing tensor allocations. Instead of loading an entire model on a single device, DisTorch2 splits the model's layers across available devices while maintaining computational efficiency.
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**Virtual VRAM Allocation**: Artificially increases the available VRAM on the compute device by borrowing memory capacity from donor devices through intelligent tensor distribution.
|
||||
|
||||
**Expert Mode Allocations**: Advanced users can manually specify exactly how much of the model should be placed on each device using ratio or byte-based allocation strings.
|
||||
|
||||
### Allocation Examples
|
||||
|
||||
**Basic Virtual VRAM Mode**:
|
||||
- `compute_device`: `cuda:0`
|
||||
- `virtual_vram_gb`: `8.0`
|
||||
- `donor_device`: `cuda:1`
|
||||
- Result: Loads model as if cuda:0 had 8GB more VRAM available, using cuda:1 as memory donor.
|
||||
|
||||
**Expert Ratio Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,60%;cuda:1,30%;cpu,10%`
|
||||
- Distributes model layers with 60% on GPU 0, 30% on GPU 1, and 10% on CPU.
|
||||
|
||||
**Expert Byte Allocation**:
|
||||
- `expert_mode_allocations`: `cuda:0,4gb;cuda:1,2gb;cpu,*`
|
||||
- Allocates exactly 4GB to cuda:0, 2GB to cuda:1, and remaining to CPU.
|
||||
|
||||
**Mixed Mode**:
|
||||
Combines virtual VRAM with expert allocations for complex multi-device scenarios.
|
||||
@@ -0,0 +1,18 @@
|
||||
# VAELoaderMultiGPU
|
||||
|
||||
The `VAELoaderMultiGPU` node is used to load VAE (Variational Autoencoder) models with device selection capability, enabling users to specify which GPU or device should be used for model execution.
|
||||
|
||||
This node automatically detects models located in the `ComfyUI/models/vae` folder, and it will also read models from additional paths configured in the `extra_model_paths.yaml` file. Sometimes, you may need to **refresh the ComfyUI interface** to allow it to read the model files from the corresponding folder.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `vae_name` | `STRING` | The name of the VAE model to load. |
|
||||
| `device` | `STRING` | Target device for compute operations (e.g., 'cuda:0', 'cuda:1', 'cpu'). Selected from available devices on your system. |
|
||||
|
||||
## Outputs
|
||||
|
||||
| Output Name | Data Type | Description |
|
||||
| --- | --- | --- |
|
||||
| `VAE` | `VAE` | The loaded VAE decoder/encoder model. |
|
||||
Reference in New Issue
Block a user