13 KiB
System Architecture & Patterns (Updated 2025-09-29)
Core Architecture
Dynamic Class Override System
Foundation Pattern: City96's elegant inheritance-based approach (Dec 2024 revolution)
def override_class(original_class, device_param="device"):
class MultiGPUClass(original_class):
@classmethod
def INPUT_TYPES(cls):
inputs = original_class.INPUT_TYPES()
inputs["required"][device_param] = (get_device_list(),)
return inputs
def override(self, *args, **kwargs):
device = kwargs.pop(device_param, None)
mm.text_encoder_device = device
return original_class.FUNCTION(self, *args, **kwargs)
return MultiGPUClass
Key Benefits:
- 50 lines vs 400+: Eliminated manual class definitions
- Universal Support: Works with any ComfyUI loader node
- Maintenance: Auto-adapts to ComfyCore changes
- Consistency: Unified behavior across all MultiGPU nodes
Load-Patch-Distribute (LPD) Method
DisTorch2 Core Process:
# 1. LOAD - Always on compute device first
tensor = load_tensor_on_compute_device(tensor_name)
# 2. PATCH - Apply all LoRAs at full precision
if lora_patches:
tensor = apply_lora_patches(tensor, lora_patches, precision=torch.float16)
# 3. DISTRIBUTE - Move to target device after patching
final_tensor = tensor.to(target_device)
Design Principles:
- Quality First: No precision loss during LoRA application
- Deterministic: Same allocation every time
- ComfyUI Native: Works with existing ComfyCore patterns
Memory Management Architecture
Virtual VRAM System
Concept: Make CPU/secondary GPU memory appear as extended VRAM
class VirtualVRAM:
def __init__(self, compute_device, donor_device, virtual_gb):
self.compute_device = compute_device # e.g., "cuda:0"
self.donor_device = donor_device # e.g., "cpu" or "cuda:1"
self.virtual_gb = virtual_gb # Extended memory pool
def allocate_layers(self, model_layers, allocation_string):
# "cuda:0,2.5gb;cpu,*" -> assign layers based on cumulative memory
Expert Allocation Modes
Bytes Mode (Recommended):
# "cuda:0,2.5gb;cuda:1,3.0g;cpu,*"
def parse_bytes_allocation(allocation_string):
devices = []
for device_spec in allocation_string.split(';'):
device_name, memory_spec = device_spec.split(',')
if memory_spec == '*':
memory_bytes = float('inf') # Overflow device
else:
memory_bytes = parse_memory_string(memory_spec) # 2.5gb -> bytes
devices.append((device_name, memory_bytes))
return devices
Ratio Mode (llama.cpp style):
# "cuda:0,25%;cpu,75%" -> 1:3 split
def parse_ratio_allocation(allocation_string):
total_ratio = sum(float(spec.split(',')[1].rstrip('%')) for spec in allocation_string.split(';'))
device_ratios = []
for device_spec in allocation_string.split(';'):
device_name, ratio_spec = device_spec.split(',')
ratio = float(ratio_spec.rstrip('%')) / total_ratio
device_ratios.append((device_name, ratio))
return device_ratios
Selective Ejection Pipeline (Current)
Updated to reflect current code (Phase 3 implemented without global sentinel):
- Load-time flagging (per-model transient):
- In each DisTorch2 override, after the real loader returns:
out[0].model._mgpu_unload_distorch_model = (keep_loaded == False)
- Purpose: mark this specific DisTorch model for ejection only when the user unchecked “keep loaded”.
- In each DisTorch2 override, after the real loader returns:
- Manager-parity cleanup trigger:
force_full_system_cleanup(reason, force=True)sets:unload_models=True,free_memory=Trueon PromptQueue (exactly what Manager’s “Free model and node cache” does).
- Selective unloading:
mm.unload_all_modelsis patched (_mgpu_patched_unload_all_modelsinmodel_management_mgpu.py):- Splits
mm.current_loaded_modelsintomodels_to_unload(flag==True) andkept_models(flag==False). - If any are flagged, unloads only those and resets
mm.current_loaded_models = kept_models. - Note: If no models are flagged, the current code delegates to the original
unload_all_models()(this is under review; see “Hardened Rule” below).
- Splits
- Multi-device VRAM cache + CPU reset:
mm.soft_empty_cacheis patched tosoft_empty_cache_distorch2_patched:- Detects DisTorch2-active state and clears allocator caches on all devices via
soft_empty_cache_multigpu() - Adaptive CPU memory reset (threshold-based), and optional forced
PromptExecutor.reset()whenforce=Truefor Manager parity.
- Detects DisTorch2-active state and clears allocator caches on all devices via
Hardened Rule (target behavior to restore):
- If
models_to_unloadis empty,unload_all_modelsshould be a strict no-op (do not delegate to the original). Retained models must never be ejected when no flags are set. This will be re-applied during the rediscovery step.
Device Detection & Management
Multi-Device Enumeration:
def get_device_list():
devices = ["cpu"] # Always available
if torch.cuda.is_available():
devices.extend([f"cuda:{i}" for i in range(torch.cuda.device_count())])
# XPU/NPU/MLU/MPS/DirectML/CoreX detection...
return devices
Device Bandwidth Intelligence (from benchmarking):
- NVLINK (~50.8 GB/s)
- PCIe 4.0 x16 (~27.2 GB/s)
- PCIe 3.0 x8 (~6.8 GB/s)
- PCIe 3.0 x4 (~2.1 GB/s)
Integration Patterns
ComfyCore Alignment
Philosophy: Work WITH ComfyUI, not against it
# GOOD: Use ComfyCore's device management
current_device = mm.get_torch_device()
mm.text_encoder_device = target_device
# AVOID: Direct PyTorch device manipulation
torch.cuda.set_device(device_id) # Bypasses ComfyCore
Node Registration System
# Dynamic registration based on available dependencies
if "ComfyUI-GGUF" in installed_modules:
NODE_CLASS_MAPPINGS["UnetLoaderGGUFDisTorch2MultiGPU"] = create_gguf_distorch_node()
Dependency Detection
def check_module_availability(module_paths):
for path in module_paths:
if os.path.exists(os.path.join(custom_nodes_dir, path)):
return True
return False
Performance Optimization Patterns
Layer Transfer Optimization
def optimized_layer_transfer(layer, source_device, target_device):
if source_device == target_device:
return layer
non_blocking = "cuda" in source_device and "cuda" in target_device
if source_device == "cpu" and "cuda" in target_device:
layer = layer.pin_memory()
return layer.to(target_device, non_blocking=non_blocking)
Memory Pressure Management
def should_auto_offload(model_size_gb, vram_available_gb, threshold=0.9):
return model_size_gb > (vram_available_gb * threshold)
def calculate_offload_amount(model_size_gb, target_vram_usage_gb):
return max(0, model_size_gb - target_vram_usage_gb)
Error Handling Philosophy
Fail Loudly Pattern
# GOOD: Let ComfyCore changes surface immediately
def load_model(self, model_name, device):
return original_loader.load_unet(model_name, device)
# AVOID: Defensive coding that masks issues
try:
return original_loader.load_unet(model_name, device)
except AttributeError:
return fallback_method()
Integration Validation
def validate_comfycore_integration():
required_attrs = ['FUNCTION', 'INPUT_TYPES', 'RETURN_TYPES']
for attr in required_attrs:
if not hasattr(target_class, attr):
raise AttributeError(f"ComfyCore node missing {attr} - API changed")
Code Style Patterns
Self-Documenting Code
def override_class_with_device_selection(original_class, device_param_name="device"):
compute_device = kwargs.get(device_param_name, mm.get_torch_device())
Minimal Comments Philosophy
Prefer structure and naming to convey intent; use comments for non-obvious constraints/assumptions.
Architectural Decision Records
Why Dynamic Class Override vs Manual Definitions
Decision: Use inheritance-based class override (City96 approach) Rationale:
- Reduces code from 400+ lines to ~50 lines
- Auto-adapts to ComfyCore changes
- Eliminates maintenance burden of manual node definitions
- Provides consistent behavior across all node types
Why Load-Patch-Distribute vs Direct Distribution
Decision: Always load on compute device first, then distribute Rationale:
- Ensures LoRA patches applied at full precision
- Maintains quality parity with single-GPU workflows
- Predictable behavior regardless of target device
- Works with ComfyCore’s existing patching mechanisms
Why Expert Modes vs Automatic Only
Decision: Provide both automatic and expert allocation modes Rationale:
- Automatic mode enables low-VRAM users immediately
- Expert modes allow optimization for specific hardware
- Performance depends on bandwidth topology; experts need control
Why Universal Device Support vs CUDA-Only
Decision: Support CPU, XPU, NPU, MLU, MPS, DirectML alongside CUDA Rationale:
- ComfyUI’s user base spans diverse hardware
- Future-proof for emerging accelerators
- Hardware democracy principle
Why Per-Model Flag vs Global Sentinel (Updated)
Decision: Use per-model _mgpu_unload_distorch_model instead of a global “DISTORCH2_UNLOAD_MODEL” sentinel
Rationale:
- Surgical precision at model granularity
- No persistent or cross-workflow state
- Cleaner semantics under ComfyUI’s queue/flag model
Hardened unloading rule (target to re-apply):
- If no models are flagged for ejection,
mm.unload_all_modelsmust be a strict no-op to preserve retained models across the full Manager-parity flow.
Testing & Validation Patterns
Hardware Configuration Testing
HARDWARE_CONFIGS = [
{"compute": "cuda:0", "donor": "cpu", "connection": "PCIe 4.0 x16"},
{"compute": "cuda:0", "donor": "cuda:1", "connection": "NVLink"},
{"compute": "cuda:0", "donor": "cuda:1", "connection": "PCIe 3.0 x8"},
{"compute": "cuda:0", "donor": "cuda:1", "connection": "PCIe 3.0 x4"},
]
Model Compatibility Validation
TEST_MODELS = [
{"name": "FLUX.1-dev", "format": ".safetensors", "size_gb": 23.8},
{"name": "WAN 2.2", "format": ".safetensors", "size_gb": 14.0},
{"name": "FLUX-GGUF", "format": ".gguf", "size_gb": 11.8},
{"name": "QWEN Image", "format": ".safetensors", "size_gb": 38.0},
]
Performance Regression Testing
def benchmark_allocation_performance(model, hardware_config, allocation_configs):
baseline_time = benchmark_single_gpu(model)
for allocation in allocation_configs:
distributed_time = benchmark_distributed(model, hardware_config, allocation)
performance_ratio = distributed_time / baseline_time
assert performance_ratio < expected_slowdown_threshold(hardware_config)
Module Architecture (Post-Refactoring)
Core Module Separation
Problem Solved: Eliminated circular import device_utils.py ↔ distorch_2.py
Solution: model_management_mgpu.py as central model lifecycle hub
Module Responsibilities
device_utils.py (Base Layer):
- Device enumeration and detection
- VRAM cache management (
soft_empty_cache_multigpu) - Pure hardware abstraction – no model tracking
model_management_mgpu.py (Core Layer):
- Model lifecycle tracking and memory logging
- Cleanup orchestration (
force_full_system_cleanup,trigger_executor_cache_reset,check_cpu_memory_threshold) - Patched unload path (selective ejection)
distorch_2.py/distorch.py (Feature Layer):
- DisTorch distribution algorithms and allocation analysis
- Per-model flagging (
_mgpu_unload_distorch_model) during DisTorch loads - Imports FROM Core/Base only
UI Layer: nodes.py, checkpoint_multigpu.py
- Device-aware user interfaces and node definitions
Assembly: __init__.py
- Final integration/patch registration (
mm.soft_empty_cachepatch, node maps)
Import Flow Architecture
┌─────────────────┐
│ __init__.py │ ← Assembly Layer
└─────────────────┘
↑
┌─────────────────┐
│ UI Layer │ ← nodes.py, checkpoint_multigpu.py
└─────────────────┘
↑
┌─────────────────┐
│ Feature Layer │ ← distorch_2.py, distorch.py
└─────────────────┘
↑
┌─────────────────┐
│ Core Layer │ ← model_management_mgpu.py
└─────────────────┘
↑
┌─────────────────┐
│ Base Layer │ ← device_utils.py
└─────────────────┘
Architectural Validation
Rule: Dependencies only flow UPWARD. Violations create circular imports.
Prevention: Before any import, verify it respects the layer hierarchy.
Function Migration Record
Moved from device_utils.py to model_management_mgpu.py:
multigpu_memory_log– memory state loggingtrigger_executor_cache_reset– CPU memory managementcheck_cpu_memory_threshold– adaptive cleanup triggersforce_full_system_cleanup– Manager-parity free flow
Rationale: These belong to model lifecycle/cleanup, not hardware enumeration.