docs(memory-bank): sync with current code state for selective ejection\n\n- Document Phase 3 implemented without global sentinel (_mgpu_unload_distorch_model per-model flag)\n- Describe patched unload_all_models selective behavior and current all-kept delegation caveat\n- Outline rediscovery plan and strict no-op target when no models are flagged\n- Update active context, system patterns, code references, progress, tech context, and lineage
This commit is contained in:
+105
-186
@@ -1,4 +1,4 @@
|
||||
# Active Context: Current Development Focus
|
||||
# Active Context: Current Development Focus (Updated 2025-09-29)
|
||||
|
||||
## Current Work Focus
|
||||
|
||||
@@ -8,223 +8,142 @@
|
||||
**Community**: Active user base with consistent feedback
|
||||
**Performance**: Benchmarked and validated across hardware configurations
|
||||
|
||||
### Recent Major Achievements (Last 6 Months)
|
||||
### Recent Major Achievements (Last 6–12 Months)
|
||||
|
||||
#### DisTorch V2.0 Release (August 2025)
|
||||
- **Universal SafeTensor Support**: Extended beyond GGUF to all model formats
|
||||
- **10% Performance Improvement**: Optimized memory transfer patterns
|
||||
- **Load-Patch-Distribute Pipeline**: Ensures quality parity with single-GPU
|
||||
- **Expert Allocation Modes**: Bytes, ratios, fraction-based distribution
|
||||
- Universal SafeTensor support (beyond GGUF)
|
||||
- ~10% performance improvement over DisTorch V1
|
||||
- Load-Patch-Distribute (LPD) pipeline: load on compute → patch LoRAs at full precision → distribute
|
||||
- Expert allocation modes: bytes, ratios, fractions
|
||||
|
||||
#### City96 Architecture Integration (December 2024 - Ongoing)
|
||||
- **Code Reduction**: 400+ lines → 50 lines via inheritance
|
||||
- **Dynamic Class Override**: Automatic node creation from existing loaders
|
||||
- **Maintenance Simplification**: Auto-adapts to ComfyCore API changes
|
||||
- **Universal Support**: Works with any ComfyUI loader pattern
|
||||
#### City96 Architecture Integration (Dec 2024 – Ongoing)
|
||||
- Code reduction: ~400 lines → ~50 lines via inheritance-based dynamic override
|
||||
- Automatic node creation from existing loaders
|
||||
- Maintenance simplification (fail-loudly alignment with ComfyCore API)
|
||||
- Universal support for loader patterns
|
||||
|
||||
#### Comprehensive Hardware Validation
|
||||
- **6 Hardware Configurations**: NVLink to PCIe 3.0 x4 coverage
|
||||
- **5 Model Architectures**: FLUX, WAN, QWEN, HunyuanVideo tested
|
||||
- **Performance Benchmarking**: Quantified bandwidth vs. performance relationships
|
||||
- **Strategic Recommendations**: Clear guidance for different use cases
|
||||
- 6 hardware configurations (NVLink to PCIe 3.0 x4)
|
||||
- 5 model families validated (FLUX, WAN, QWEN, HunyuanVideo, Florence2)
|
||||
- Clear bandwidth vs performance characterization and recommendations
|
||||
|
||||
## Current Development Priorities
|
||||
|
||||
### 1. CPU Memory Leak Resolution (FINAL SOLUTION CONCEPTUALIZED)
|
||||
**Goal**: eliminate CPU DRAM memory leaks through 3-flag surgical ejection system
|
||||
### 1) CPU Memory Leak Resolution: Status and What’s Left
|
||||
Current code state (verified in repo):
|
||||
- Selective ejection (Phase 3) is implemented without the Phase 1 global sentinel.
|
||||
- During load in DisTorch2 wrappers (UNET/CLIP/VAE), we set a per-model transient flag:
|
||||
- `_mgpu_unload_distorch_model = (keep_loaded == False)`
|
||||
- End-of-workflow “free” path mirrors Manager parity by setting:
|
||||
- `unload_models=True`, `free_memory=True`
|
||||
- Patches in place:
|
||||
- `mm.unload_all_models` → selectively unloads only models with `_mgpu_unload_distorch_model == True` and rebuilds `mm.current_loaded_models` from kept models
|
||||
- `mm.soft_empty_cache` → `soft_empty_cache_distorch2_patched` (multi-device VRAM clear + adaptive CPU reset, and forceable executor reset for parity)
|
||||
|
||||
**Finalized Conceptual Solution**:
|
||||
- **keep_loaded Boolean Engineering**: Drives preservation, trigger, and selective destruction ✅
|
||||
- **3-Transient-Flags Architecture**: Execution-scoped flags with complete isolation ✅
|
||||
- **Surgical Ejection Logic**: Design for processing only models with ejection flag set ✅
|
||||
- **Elimination Design**: Memory leaks designed for clinical resolution through distributed cleanup ✅
|
||||
Outstanding defect:
|
||||
- In some flows, retained (keep_loaded=True) models are still being ejected downstream.
|
||||
- Two likely culprits:
|
||||
1) “All-kept delegation” in our patched unload: when no models are flagged, delegation to the original `unload_all_models()` unloads everything.
|
||||
2) Post-unload follow-on flows (e.g., `PromptExecutor.reset()`, GC, `soft_empty_cache()`, or a core `free_memory(...)` path) may cause unintended detaches for retained models.
|
||||
|
||||
**Status**: Complete conceptual solution designed and documented. Requires implementation and testing to eliminate memory leaks.
|
||||
Immediate Actions:
|
||||
- Documentation sync (this update) and commit
|
||||
- Rediscover the previously working selective retention variant from branch history and reinstate it
|
||||
- Harden no-op path in `unload_all_models`:
|
||||
- If no models are flagged for ejection, do nothing (strict no-op), never delegate to the original
|
||||
- Add temporary instrumentation:
|
||||
- Memory/log snapshots at: pre-unload → post-unload → post-reset → post-gc/soft_empty
|
||||
- ERROR if any kept model is missing after the full `/free` flow
|
||||
|
||||
### 2. Ecosystem Expansion (High Priority)
|
||||
**Goal**: Support emerging model formats and custom nodes
|
||||
Verification Matrix:
|
||||
- Minimal retention: A(keep=false), B(true), C(true) → A ejected, B/C retained after complete free flow
|
||||
- All-kept: D(true), E(true) → no ejection, only allocator/cache cleanups
|
||||
|
||||
**Active Integrations**:
|
||||
- **ComfyUI-GGUF**: 6 DisTorch-enabled GGUF nodes (complete)
|
||||
- **WanVideoWrapper**: 8 MultiGPU video nodes (complete)
|
||||
- **Florence2**: Vision model support (complete)
|
||||
- **HunyuanVideoWrapper**: Native VAE + device selection (active development)
|
||||
Rediscovery Plan:
|
||||
- Search recent commits where logs indicate successful retention after free
|
||||
- Diff `_mgpu_patched_unload_all_models` vs current to recover exact guard/flow
|
||||
- Confirm Manager parity (`/free` flags) still routes through patched unload and retains kept models across reset/GC
|
||||
|
||||
**Next Targets**:
|
||||
- **LTX Video**: New video architecture support
|
||||
- **Mochi**: Performance-optimized video models
|
||||
- **Community Requests**: Issue-driven integration priorities
|
||||
### 2) Ecosystem Expansion (High Priority)
|
||||
Goal: Support emerging model formats and custom nodes
|
||||
|
||||
### 2. User Experience Optimization (Medium Priority)
|
||||
**Goal**: Reduce complexity for new users while maintaining expert capabilities
|
||||
Active Integrations:
|
||||
- ComfyUI-GGUF: DisTorch-enabled GGUF nodes (complete)
|
||||
- WanVideoWrapper: MultiGPU video nodes (complete)
|
||||
- Florence2: Vision model support (complete)
|
||||
- HunyuanVideoWrapper: Native VAE + device selection (in progress)
|
||||
|
||||
**Recent Improvements**:
|
||||
- **Automatic Mode**: Intelligent offloading based on VRAM availability
|
||||
- **Error Messages**: Clear guidance when allocation fails
|
||||
- **Example Workflows**: 20+ example JSON files covering major use cases
|
||||
Next Targets:
|
||||
- LTX Video
|
||||
- Mochi
|
||||
- Issue-driven community requests
|
||||
|
||||
**Ongoing Work**:
|
||||
- **Configuration Validation**: Prevent invalid allocation strings
|
||||
- **Performance Prediction**: Estimate slowdown before execution
|
||||
- **Documentation**: User-friendly guides for different hardware scenarios
|
||||
### 3) User Experience Optimization (Medium Priority)
|
||||
Goal: Reduce complexity while preserving expert control
|
||||
|
||||
### 3. Advanced Features (Low Priority)
|
||||
**Goal**: Push boundaries of multi-device inference
|
||||
Recent Improvements:
|
||||
- Automatic Mode: Intelligent offloading based on VRAM availability
|
||||
- Error messages: Clearer guidance for allocation failures
|
||||
- Documentation: 20+ example JSON workflows
|
||||
|
||||
**Research Areas**:
|
||||
- **Model Parallelism**: Split individual layers across multiple devices
|
||||
- **Pipeline Parallelism**: Concurrent execution of different workflow stages
|
||||
- **Memory Compression**: Runtime compression of stored model layers
|
||||
- **Quality Metrics**: Quantitative measurement of output quality preservation
|
||||
Ongoing:
|
||||
- Configuration validation and performance prediction
|
||||
- “First-run” guides for low-VRAM and multi-GPU users
|
||||
|
||||
### 4) Advanced Features (Low Priority)
|
||||
Research Areas:
|
||||
- Model parallelism and pipeline parallelism
|
||||
- Memory compression, fragmentation handling
|
||||
- Quality metrics and deterministic parity checks
|
||||
|
||||
## Active Technical Decisions
|
||||
|
||||
### Memory Management Philosophy
|
||||
**Current Approach**: Conservative with user control
|
||||
- **Default Behavior**: Minimal offloading unless user specifies
|
||||
- **Safety First**: Automatic fallbacks when allocations fail
|
||||
- **Transparency**: Clear logging of memory operations
|
||||
- **User Choice**: Expert modes for power users
|
||||
|
||||
**Alternative Considered**: Aggressive automatic optimization
|
||||
- **Rejected**: Too unpredictable, quality concerns with LoRAs
|
||||
- **Lesson**: Users prefer control over convenience
|
||||
- Conservative by default with explicit user control
|
||||
- Preserve quality: Patch LoRAs before distributing
|
||||
- Transparency: Verbose and structured memory logging
|
||||
- Fail-loudly alignment with ComfyCore
|
||||
|
||||
### Integration Strategy
|
||||
**Current Approach**: Inheritance-based class override
|
||||
- **City96 Pattern**: Dynamic class creation at runtime
|
||||
- **Minimal API Surface**: Reduces maintenance burden
|
||||
- **ComfyCore Alignment**: Works with existing patterns
|
||||
|
||||
**Alternative Considered**: Direct node registration
|
||||
- **Rejected**: Maintenance nightmare, API fragility
|
||||
- **Lesson**: Elegant code reduces long-term costs
|
||||
- Inheritance-based node override (City96 pattern)
|
||||
- Minimal patch surface area with explicit patch points:
|
||||
- `mm.get_torch_device`/`mm.text_encoder_device` override for device selection
|
||||
- `mm.soft_empty_cache` override for multi-device cache clear + CPU reset
|
||||
- `mm.unload_all_models` selective unload path
|
||||
|
||||
### Hardware Support Priority
|
||||
**Current Approach**: Universal device support with quality tiers
|
||||
- **Tier 1**: CUDA (primary development and testing)
|
||||
- **Tier 2**: CPU, MPS (community validated)
|
||||
- **Tier 3**: XPU, NPU, DirectML (experimental support)
|
||||
|
||||
**Rationale**: ComfyUI's diverse hardware ecosystem demands inclusivity
|
||||
- Tier 1: CUDA
|
||||
- Tier 2: CPU, MPS
|
||||
- Tier 3: XPU, NPU, MLU, DirectML (experimental footprint grows with community validation)
|
||||
|
||||
## User Behavior Patterns (Observed)
|
||||
|
||||
### Common Usage Scenarios
|
||||
1. **Low-VRAM Image Generation** (40% of users)
|
||||
- Single GPU systems (RTX 4070, RTX 3080)
|
||||
- Running FLUX.1-dev, QWEN models
|
||||
- Primary strategy: CPU offloading
|
||||
|
||||
2. **Multi-GPU Video Generation** (30% of users)
|
||||
- Dual-GPU setups (mixed architectures common)
|
||||
- WAN, HunyuanVideo workflows
|
||||
- Primary strategy: GPU-to-GPU distribution
|
||||
|
||||
3. **Professional Workflows** (20% of users)
|
||||
- High-end hardware (3090s, 4090s)
|
||||
- Batch processing, high resolutions
|
||||
- Primary strategy: Optimization for throughput
|
||||
|
||||
4. **Enthusiast Experimentation** (10% of users)
|
||||
- Cutting-edge models, extreme configurations
|
||||
- Custom allocation strings, performance tweaking
|
||||
- Primary strategy: Push hardware limits
|
||||
|
||||
### Support Request Patterns
|
||||
1. **"Only cuda:0 visible"** - Device detection issues (25%)
|
||||
2. **"Out of memory errors"** - Allocation configuration (20%)
|
||||
3. **"Slower than expected"** - Hardware optimization (15%)
|
||||
4. **"Node missing after install"** - Dependency conflicts (15%)
|
||||
5. **"Quality differences"** - LoRA/quantization concerns (10%)
|
||||
6. **"Integration requests"** - New model support (15%)
|
||||
|
||||
### Configuration Preferences
|
||||
- **Bytes Mode**: 60% adoption (preferred for precision)
|
||||
- **Fraction Mode**: 25% adoption (simple but limited)
|
||||
- **Ratio Mode**: 15% adoption (llama.cpp familiarity)
|
||||
|
||||
**Automatic vs Expert**: 70% start automatic, 40% graduate to expert modes
|
||||
|
||||
## Project Learnings & Insights
|
||||
|
||||
### What Works Well
|
||||
1. **Inheritance Pattern**: City96's architecture scales beautifully
|
||||
2. **Load-Patch-Distribute**: Maintains quality while enabling distribution
|
||||
3. **Comprehensive Testing**: Hardware benchmarking prevents regression
|
||||
4. **Conservative Defaults**: Users prefer working slowly to not working
|
||||
5. **Clear Documentation**: Example workflows accelerate adoption
|
||||
|
||||
### What We've Learned to Avoid
|
||||
1. **Defensive Programming**: Masks ComfyCore API changes, creates maintenance debt
|
||||
2. **Automatic LoRA Offloading**: Quality concerns outweigh convenience
|
||||
3. **Over-Optimization**: Complex algorithms often perform worse than simple ones
|
||||
4. **API Abstraction**: Users want direct control over model placement
|
||||
5. **Hardware Assumptions**: Every configuration is someone's primary system
|
||||
|
||||
### Development Philosophy Evolution
|
||||
**Early**: "Make it work on as many systems as possible"
|
||||
**Current**: "Make it work reliably, then optimize for common cases"
|
||||
**Future**: "Provide the tools, let users choose their tradeoffs"
|
||||
- Low-VRAM image gen, multi-GPU video gen, professional pipelines, enthusiast experiments
|
||||
- Support requests: device detection, OOM, performance expectations, missing nodes, quality concerns, integration requests
|
||||
- Allocation preferences: bytes (most common), fraction, ratio
|
||||
|
||||
## Next Steps & Immediate Actions
|
||||
Short-term (2–4 weeks):
|
||||
- Commit Memory Bank updates (this change)
|
||||
- Rediscover and reinstate the selective retention behavior that worked
|
||||
- Harden no-op branch in unload patch and add retention instrumentation
|
||||
- Run verification matrix and update docs with results
|
||||
- Triage top GitHub issues
|
||||
|
||||
### Short-term (Next 2-4 weeks)
|
||||
1. **Issue Triage**: Address 5 highest-priority GitHub issues
|
||||
2. **HunyuanVideo Integration**: Complete native VAE support
|
||||
3. **Documentation Update**: Refresh README with current capabilities
|
||||
4. **Example Refresh**: Update workflow examples for new features
|
||||
Medium-term (2–3 months):
|
||||
- LTX Video integration
|
||||
- Performance dashboard and quality measurement runs
|
||||
- Tutorials and doc refresh based on latest capabilities
|
||||
|
||||
### Medium-term (Next 2-3 months)
|
||||
1. **LTX Video Support**: Integrate new video model architecture
|
||||
2. **Performance Dashboard**: Web-based hardware configuration guide
|
||||
3. **Quality Validation**: Systematic output quality measurement
|
||||
4. **Community Outreach**: Tutorial videos, blog posts
|
||||
|
||||
### Long-term (6-12 months)
|
||||
1. **Model Parallelism**: Research splitting individual layers
|
||||
2. **Streaming Inference**: Real-time video generation support
|
||||
3. **Cloud Integration**: Multi-node distributed inference
|
||||
4. **Professional Tools**: Batch processing, API server modes
|
||||
|
||||
## Knowledge Gaps & Research Areas
|
||||
|
||||
### Technical Uncertainties
|
||||
1. **Future ComfyUI Changes**: Core API evolution risk
|
||||
2. **Next-Gen Hardware**: PCIe 5.0, NVLink 5.0 optimization opportunities
|
||||
3. **Model Architecture Evolution**: MoE, multimodal impact on distribution
|
||||
4. **PyTorch Updates**: Memory management changes in newer versions
|
||||
|
||||
### Community Questions
|
||||
1. **Adoption Barriers**: What prevents users from trying MultiGPU?
|
||||
2. **Quality Perception**: Do users trust distributed inference quality?
|
||||
3. **Hardware Investment**: Will users buy hardware based on MultiGPU support?
|
||||
4. **Professional Use**: What features do commercial users need?
|
||||
|
||||
### Performance Mysteries
|
||||
1. **Transfer Prediction**: Can we accurately predict slowdown before execution?
|
||||
2. **Memory Fragmentation**: How do repeated loads/unloads affect performance?
|
||||
3. **Thermal Behavior**: Does extended use show different performance patterns?
|
||||
4. **OS Differences**: Are there meaningful Windows vs Linux performance gaps?
|
||||
Long-term (6–12 months):
|
||||
- Model/pipeline parallelism experiments
|
||||
- Streaming inference for video
|
||||
- Multi-node/cloud integration and orchestration
|
||||
|
||||
## Current Environment State
|
||||
- IDE: VSCode
|
||||
- Version Control: Git with conventional commits
|
||||
- Testing: Manual validation on available hardware + community contributions
|
||||
- Primary Dev HW: RTX 3090 + mixed secondaries
|
||||
- Known Limitation: Limited access to newest GPUs (e.g., RTX 5090)
|
||||
|
||||
### Development Tools
|
||||
- **Primary IDE**: VSCode with Python extensions
|
||||
- **Version Control**: Git with conventional commits
|
||||
- **Testing**: Manual validation across available hardware
|
||||
- **Documentation**: Markdown files, example JSON workflows
|
||||
|
||||
### Hardware Access
|
||||
- **Primary Development**: RTX 3090 with various secondary GPUs
|
||||
- **Testing Network**: Community contributors with diverse configurations
|
||||
- **Benchmarking**: Systematic testing across 6 hardware configurations
|
||||
- **Limitations**: Limited access to newest hardware (RTX 5090, etc.)
|
||||
|
||||
### Community Engagement
|
||||
- **GitHub Issues**: Active monitoring and response
|
||||
- **Discord**: ComfyUI community support channel participation
|
||||
- **Documentation**: Comprehensive README and example workflows
|
||||
- **Support**: Personal responses to complex issues
|
||||
|
||||
This Memory Bank serves as my only link to previous work. Each reset, I depend entirely on these files to understand the project state and continue development effectively.
|
||||
This Active Context reflects the current codebase reality: Phase 3 selective ejection is in place (per-model flags + selective unload patch), but a retention defect remains when no models are flagged and/or after the free path completes. The immediate roadmap is to commit these updates, then locate and reinstate the previously working selective retention behavior and add guards to ensure robust “keep_loaded=True” semantics across the full Manager parity flow.
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
|
||||
Purpose
|
||||
- Provide an end-to-end, fully verified lineage of the ComfyUI Manager “Free model and node cache” button through to the exact consumption of flags in ComfyUI core, with exact file paths and code excerpts captured from the current snapshot in this workspace.
|
||||
- Document MultiGPU patch integration points that participate in the free/unload flow, including selective unload behavior and current caveats.
|
||||
|
||||
End‑to‑End Flow (Current Snapshot)
|
||||
1) UI Button (Manager) → 2) JS helper free_models(...) → 3) POST /free (Comfy core) → 4) main.py prompt_worker thread polls flags and performs:
|
||||
@@ -144,3 +145,87 @@ Verification Status
|
||||
- Manager JS files under ../ComfyUI-Manager/js/
|
||||
- ComfyUI server and main under ../../server.py and ../../main.py
|
||||
- Consumption site conclusively identified in ../../main.py prompt_worker via q.get_flags → unload_all_models + PromptExecutor.reset
|
||||
|
||||
---
|
||||
|
||||
MultiGPU Integration Points (This Repository)
|
||||
|
||||
Overview
|
||||
- In addition to the core /free flow, MultiGPU patches (in this repository) alter both the unload and soft-empty behaviors to enable selective ejection of DisTorch-managed models and multi-device cache clearing.
|
||||
|
||||
1) Per-model transient flag (DisTorch2 nodes)
|
||||
- File: memory-bank reference → implemented in code at: ./distorch_2.py
|
||||
- Where:
|
||||
- In DisTorch2 wrappers (UNET/CLIP/VAE) inside `override(...)`, after calling the original node:
|
||||
- `out[0].model._mgpu_unload_distorch_model = (not keep_loaded)`
|
||||
- Purpose:
|
||||
- Mark models for ejection only when the user disables “keep_loaded”.
|
||||
- This supplants the previously planned global sentinel; the implemented design is purely per-model.
|
||||
|
||||
2) Selective unloading (patched unload_all_models)
|
||||
- File: ./model_management_mgpu.py
|
||||
- Patch site notes:
|
||||
- At import time, we patch `mm.unload_all_models` with `_mgpu_patched_unload_all_models`.
|
||||
- Behavior:
|
||||
- Iterate `mm.current_loaded_models` into:
|
||||
- `models_to_unload`: those with `_mgpu_unload_distorch_model == True`
|
||||
- `kept_models`: the rest
|
||||
- If any are flagged, unload only `models_to_unload` and rebuild `mm.current_loaded_models = kept_models`.
|
||||
- If none are flagged (all kept), current code delegates to original `unload_all_models()` (known caveat; see below).
|
||||
- Known caveat (to be fixed next):
|
||||
- The “all kept” branch currently delegates to the original unload, which unloads everything. Target behavior is strict no-op when no models are flagged.
|
||||
|
||||
3) Multi-device VRAM cache and CPU reset (patched soft_empty_cache)
|
||||
- File: ./__init__.py
|
||||
- Patch site notes:
|
||||
- `mm.soft_empty_cache` → `soft_empty_cache_distorch2_patched`
|
||||
- Behavior:
|
||||
- Detect DisTorch2 active state; clear allocator caches on ALL devices via `soft_empty_cache_multigpu()` from `device_utils.py`
|
||||
- Adaptive CPU memory reset with optional force to emulate Manager “free_memory”.
|
||||
- This ensures cache clearing covers all devices in MultiGPU environments beyond a single `mm.get_torch_device()`.
|
||||
|
||||
4) Manager parity helper
|
||||
- File: ./model_management_mgpu.py
|
||||
- Function: `force_full_system_cleanup(reason="manual", force=True)`
|
||||
- Sets both flags (`unload_models=True`, `free_memory=True`) on PromptQueue, identical to Manager’s “Free model and node cache”.
|
||||
- Useful for testing and ensuring parity from MultiGPU paths.
|
||||
|
||||
Behavioral Summary
|
||||
- End-to-end Manager parity:
|
||||
- Manager “Free model and node cache” → POST /free sets flags → Comfy’s prompt_worker calls our patched `unload_all_models` (selective) → `PromptExecutor.reset()` → our patched `soft_empty_cache` (multi-device) → GC.
|
||||
- Selectiveness guarantee (intended):
|
||||
- Only DisTorch2 models flagged with `_mgpu_unload_distorch_model=True` are ejected.
|
||||
- Unflagged models (keep_loaded=True) remain in `mm.current_loaded_models` after the entire flow.
|
||||
- Current discrepancy:
|
||||
- When no models are flagged, our patch currently delegates to the original unload (unloads everything). Target fix is to convert this branch to a strict no-op.
|
||||
|
||||
Validation & Logging Hooks
|
||||
- Memory snapshots:
|
||||
- Use `multigpu_memory_log(identifier, tag)` in `model_management_mgpu.py` for timestamped CPU/VRAM snapshot lines.
|
||||
- VRAM cache clearing:
|
||||
- `soft_empty_cache_multigpu()` logs per-device clearing events (pre/post) in `device_utils.py`.
|
||||
- Unload path tracing:
|
||||
- `_mgpu_patched_unload_all_models` logs the counts of kept/unloaded models and updates to `mm.current_loaded_models`.
|
||||
|
||||
Practical Test Recipes
|
||||
1) Minimal retention test
|
||||
- Load A(keep=false), B(keep=true), C(keep=true)
|
||||
- POST /free payload: {"unload_models": true, "free_memory": true}
|
||||
- Expected:
|
||||
- Only A is ejected; B and C remain in `mm.current_loaded_models` post-flow.
|
||||
- CPU RAM drops; VRAM caches clear on all devices.
|
||||
|
||||
2) All-kept test
|
||||
- Load D(keep=true), E(keep=true)
|
||||
- POST /free payload: {"unload_models": true, "free_memory": true}
|
||||
- Expected target behavior:
|
||||
- No models are ejected (strict no-op in unload step), allocator/cache cleaning only.
|
||||
- Current behavior (caveat):
|
||||
- Delegates to original unload → all models may be ejected. This is the next change to reinstate strict no-op.
|
||||
|
||||
References (paths in this repo)
|
||||
- Per-model flagging: ./distorch_2.py
|
||||
- Selective unload patch: ./model_management_mgpu.py
|
||||
- Patched soft empty: ./__init__.py (soft_empty_cache_distorch2_patched)
|
||||
- Multi-device cache clear: ./device_utils.py
|
||||
- Manager parity helper: ./model_management_mgpu.py (force_full_system_cleanup)
|
||||
|
||||
+85
-241
@@ -1,264 +1,108 @@
|
||||
# ComfyUI Core Lineage & Integration Analysis
|
||||
# ComfyUI Core Lineage & Integration Analysis (Updated 2025-09-29)
|
||||
|
||||
## Overview
|
||||
|
||||
After analyzing `comfy/model_management.py`, the lineage of ComfyUI-MultiGPU becomes clear: **MultiGPU extends and enhances ComfyUI's existing memory management rather than replacing it**. This explains the project's "fail loudly" philosophy and deep integration patterns.
|
||||
ComfyUI‑MultiGPU extends (does not replace) ComfyUI core. Principles:
|
||||
- Extend, not replace: patch specific core functions and inherit existing nodes
|
||||
- Fail loudly: small, explicit patch points so core API changes surface quickly
|
||||
- User agency: device placement is explicit and honored
|
||||
- Multi‑device native: treat all devices as first‑class
|
||||
|
||||
## ComfyUI Core Foundation
|
||||
Current code reality:
|
||||
- Phase 3 “Selective Ejection” is implemented via a per‑model flag (no global sentinel).
|
||||
- Outstanding caveat: when no models are flagged, the current unload path delegates to the original core unload (unloads everything). Target is strict no‑op in this branch.
|
||||
|
||||
### Memory Management Architecture
|
||||
ComfyUI already provides sophisticated memory management through:
|
||||
## ComfyUI Core Foundation (Reference)
|
||||
|
||||
```python
|
||||
# Core VRAM state management
|
||||
class VRAMState(Enum):
|
||||
DISABLED = 0 # No vram present
|
||||
NO_VRAM = 1 # Very low vram: enable all options to save vram
|
||||
LOW_VRAM = 2
|
||||
NORMAL_VRAM = 3
|
||||
HIGH_VRAM = 4
|
||||
SHARED = 5 # Memory shared between CPU and GPU
|
||||
Key concepts implemented by ComfyUI core (see memory-bank/comfy_core.py snapshot):
|
||||
- Global list: `current_loaded_models`
|
||||
- Model wrapper: `LoadedModel` with methods like `model_load`, `model_unload`, `model_memory_required`
|
||||
- Memory utilities: `soft_empty_cache()`, `get_free_memory()`, etc.
|
||||
- Prompt execution:
|
||||
- `/free` endpoint sets queue flags: `unload_models`, `free_memory` (server.py)
|
||||
- `main.py` prompt worker consumes flags:
|
||||
- If `unload_models` (or `free_memory`): `comfy.model_management.unload_all_models()`
|
||||
- If `free_memory`: `PromptExecutor.reset()`
|
||||
- Then GC + `comfy.model_management.soft_empty_cache()`
|
||||
|
||||
# Device state tracking
|
||||
class CPUState(Enum):
|
||||
GPU = 0
|
||||
CPU = 1
|
||||
MPS = 2
|
||||
```
|
||||
|
||||
### Universal Device Detection (ComfyUI Core)
|
||||
ComfyUI already detects multiple device types:
|
||||
- **CUDA**: `torch.cuda.is_available()`
|
||||
- **DirectML**: `torch_directml` integration
|
||||
- **XPU**: Intel GPU support via `intel_extension_for_pytorch`
|
||||
- **NPU**: Ascend NPUs via `torch_npu`
|
||||
- **MLU**: Cambricon MLUs via `torch_mlu`
|
||||
- **MPS**: Apple Silicon via `torch.backends.mps`
|
||||
- **IXUCA**: CoreX accelerators
|
||||
|
||||
### LoadedModel Management System
|
||||
```python
|
||||
class LoadedModel:
|
||||
def __init__(self, model):
|
||||
self._model = weakref.ref(model)
|
||||
self.device = model.load_device
|
||||
self.currently_used = True
|
||||
|
||||
def model_load(self, lowvram_model_memory=0, force_patch_weights=False):
|
||||
# Core loading logic that MultiGPU patches
|
||||
|
||||
def model_unload(self, memory_to_free=None, unpatch_weights=True):
|
||||
# Unloading logic that MultiGPU extends
|
||||
|
||||
current_loaded_models = [] # Global list MultiGPU works with
|
||||
```
|
||||
This is the canonical “Manager button” path for model + execution cache cleanup.
|
||||
|
||||
## How MultiGPU Extends ComfyUI Core
|
||||
|
||||
### 1. Device Detection Enhancement
|
||||
**ComfyUI Core**:
|
||||
```python
|
||||
def get_torch_device():
|
||||
if directml_enabled:
|
||||
return directml_device
|
||||
if cpu_state == CPUState.MPS:
|
||||
return torch.device("mps")
|
||||
# ... single device selection logic
|
||||
```
|
||||
MultiGPU adds small patches and inherits nodes to enable multi‑device behavior while preserving ComfyUI’s flow.
|
||||
|
||||
**MultiGPU Extension**:
|
||||
```python
|
||||
def get_device_list():
|
||||
# Returns ALL available devices, not just primary
|
||||
devices = ["cpu"]
|
||||
if torch.cuda.is_available():
|
||||
devices.extend([f"cuda:{i}" for i in range(torch.cuda.device_count())])
|
||||
# Extended detection for ALL instances of each type
|
||||
```
|
||||
### 1) Device selection alignment
|
||||
- File: `__init__.py`
|
||||
- Patches:
|
||||
- `mm.get_torch_device = get_torch_device_patched`
|
||||
- `mm.text_encoder_device = text_encoder_device_patched`
|
||||
- Purpose: Respect user‑selected devices supplied by MultiGPU wrappers while staying coherent with ComfyUI’s device model.
|
||||
|
||||
### 2. Memory Management Patching
|
||||
**ComfyUI Core**:
|
||||
```python
|
||||
def get_torch_device():
|
||||
# Returns single primary device
|
||||
|
||||
def soft_empty_cache():
|
||||
# Clears cache on single device
|
||||
```
|
||||
### 2) Multi‑device VRAM cache + CPU reset
|
||||
- File: `__init__.py`
|
||||
- Patch:
|
||||
- `mm.soft_empty_cache = soft_empty_cache_distorch2_patched`
|
||||
- Behavior:
|
||||
- Detects if any DisTorch2 model is active and clears allocator caches on ALL devices via `soft_empty_cache_multigpu()` (from `device_utils.py`)
|
||||
- Integrates adaptive CPU memory reset; can force `PromptExecutor.reset()` on `force=True` for Manager parity
|
||||
|
||||
**MultiGPU Patches**:
|
||||
```python
|
||||
# Patch the core functions to be MultiGPU-aware
|
||||
mm.get_torch_device = get_torch_device_patched
|
||||
mm.soft_empty_cache = soft_empty_cache_distorch2_patched
|
||||
### 3) Selective ejection (patched unload)
|
||||
- File: `model_management_mgpu.py`
|
||||
- Patch:
|
||||
- `mm.unload_all_models = _mgpu_patched_unload_all_models`
|
||||
- Behavior:
|
||||
- Iterate `mm.current_loaded_models` and split into:
|
||||
- `models_to_unload`: models with per‑model flag `_mgpu_unload_distorch_model == True`
|
||||
- `kept_models`: all others
|
||||
- If any flagged: unload only the flagged models and set `mm.current_loaded_models = kept_models`
|
||||
- Current caveat: If none are flagged (all kept), code delegates to original core unload, which unloads everything (target: strict no‑op for this branch)
|
||||
|
||||
def soft_empty_cache_multigpu():
|
||||
# Clear cache on ALL devices
|
||||
for device_str in get_device_list():
|
||||
# Clear each device type appropriately
|
||||
```
|
||||
### 4) Per‑model flag is set at load time (no global sentinel)
|
||||
- File: `distorch_2.py`
|
||||
- Where:
|
||||
- In DisTorch2 wrappers (UNET/CLIP/VAE) inside `override(...)`, after calling the original loader:
|
||||
- `out[0].model._mgpu_unload_distorch_model = (keep_loaded == False)`
|
||||
- Rationale:
|
||||
- Surgical precision at model granularity and no persistent global state
|
||||
|
||||
### 3. Model Loading Enhancement
|
||||
**ComfyUI Core**:
|
||||
```python
|
||||
def load_models_gpu(models, memory_required=0, force_patch_weights=False):
|
||||
# Load models to single GPU with CPU offloading
|
||||
```
|
||||
### 5) Manager parity helper for tests/flows
|
||||
- File: `model_management_mgpu.py`
|
||||
- Function:
|
||||
- `force_full_system_cleanup(reason="manual", force=True)`
|
||||
- Behavior:
|
||||
- Sets both `unload_models=True` and `free_memory=True` on the PromptQueue, just like the Manager “Free model and node cache” button
|
||||
|
||||
**MultiGPU Proactive Enhancement**:
|
||||
```python
|
||||
# Patch load_models_gpu for DisTorch2 awareness
|
||||
original_load_models_gpu = mm.load_models_gpu
|
||||
## End‑to‑End Free Flow (Now)
|
||||
|
||||
def patched_load_models_gpu(models, memory_required=0, ...):
|
||||
# Detect DisTorch2 models
|
||||
# Proactively unload on multiple devices
|
||||
# Call original with enhanced context
|
||||
```
|
||||
“Manager button” or parity helper triggers the same core actions:
|
||||
|
||||
## Integration Patterns
|
||||
1) POST /free with `{"unload_models": true, "free_memory": true}`
|
||||
2) `main.py` prompt worker consumes flags:
|
||||
- Calls `comfy.model_management.unload_all_models()`
|
||||
- MultiGPU patched unload runs:
|
||||
- If any models flagged via `_mgpu_unload_distorch_model=True`: unload only those and retain others
|
||||
- If none are flagged: current code delegates to original unload (unloads everything) — under review
|
||||
- Calls `PromptExecutor.reset()`
|
||||
- GC + `comfy.model_management.soft_empty_cache()`
|
||||
- MultiGPU patched soft empty runs:
|
||||
- Multi‑device allocator cache clear (CUDA/MPS/XPU/NPU/MLU/DirectML/CoreX as available)
|
||||
- Optional CPU reset behavior when forced
|
||||
|
||||
### 1. Inheritance-Based Override (City96 Pattern)
|
||||
Instead of rewriting ComfyUI nodes, MultiGPU dynamically inherits and extends:
|
||||
Intended invariant (target):
|
||||
- Only flagged DisTorch2 models are ejected; unflagged (keep_loaded=True) models remain live after the full flow.
|
||||
|
||||
```python
|
||||
def override_class(cls):
|
||||
class MultiGPUClass(cls):
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
inputs = cls.INPUT_TYPES() # Get original inputs
|
||||
inputs["optional"]["device"] = (get_device_list(),) # Add device selection
|
||||
return inputs
|
||||
|
||||
def override(self, *args, **kwargs):
|
||||
# Set device context, call original, restore context
|
||||
return super().FUNCTION(*args, **kwargs)
|
||||
```
|
||||
## Behavior Notes & Next Step
|
||||
|
||||
### 2. Core Function Patching
|
||||
MultiGPU patches specific ComfyUI functions rather than replacing entire modules:
|
||||
- Implemented:
|
||||
- Per‑model selective ejection (Phase 3) without global sentinel
|
||||
- Multi‑device allocator clearing and Manager parity semantics
|
||||
- Caveat:
|
||||
- If no models are flagged, current patched unload delegates to original unload (unloads everything)
|
||||
- This can defeat selectiveness when all models are intended to be retained
|
||||
- Next step (hardening):
|
||||
- Reinstate “strict no‑op” in the all‑kept branch of `_mgpu_patched_unload_all_models` (never delegate to original unload if nothing is flagged)
|
||||
- Add instrumentation around pre/post unload, post reset, post soft‑empty to ensure retained models remain alive
|
||||
|
||||
```python
|
||||
# Patch specific functions while preserving ecosystem
|
||||
mm.get_torch_device = get_torch_device_patched
|
||||
mm.text_encoder_device = text_encoder_device_patched
|
||||
comfy.model_patcher.ModelPatcher.partially_load = new_partially_load
|
||||
```
|
||||
## Sequence Summary
|
||||
|
||||
### 3. Integration with LoadedModel System
|
||||
MultiGPU works with ComfyUI's existing model tracking:
|
||||
|
||||
```python
|
||||
# Use existing current_loaded_models list
|
||||
for lm in mm.current_loaded_models:
|
||||
mp = lm.model # Work with existing ModelPatcher
|
||||
if is_distorch_model(mp):
|
||||
apply_multidevice_logic(mp)
|
||||
```
|
||||
|
||||
## Why This Architecture Works
|
||||
|
||||
### 1. Minimal API Surface
|
||||
By extending rather than replacing, MultiGPU:
|
||||
- Maintains compatibility with ComfyUI updates
|
||||
- Preserves existing workflow compatibility
|
||||
- Reduces maintenance burden
|
||||
- Enables gradual adoption
|
||||
|
||||
### 2. Fail-Loudly Benefits
|
||||
When ComfyUI core changes:
|
||||
- MultiGPU patches break immediately (desired behavior)
|
||||
- No silent failures or degraded performance
|
||||
- Clear indication of needed updates
|
||||
- Prevents hidden incompatibilities
|
||||
|
||||
### 3. Ecosystem Harmony
|
||||
MultiGPU's approach allows:
|
||||
- Other custom nodes to work unchanged
|
||||
- ComfyUI core development to continue
|
||||
- Users to mix MultiGPU and standard nodes
|
||||
- Gradual migration rather than replacement
|
||||
|
||||
## Code Lineage Examples
|
||||
|
||||
### Memory Query Functions
|
||||
**ComfyUI Core**:
|
||||
```python
|
||||
def get_free_memory(dev=None, torch_free_too=False):
|
||||
# Single device memory query with device-specific logic
|
||||
if hasattr(dev, 'type') and (dev.type == 'cpu' or dev.type == 'mps'):
|
||||
mem_free_total = psutil.virtual_memory().available
|
||||
elif is_intel_xpu():
|
||||
stats = torch.xpu.memory_stats(dev)
|
||||
# ... XPU-specific logic
|
||||
```
|
||||
|
||||
**MultiGPU Usage**:
|
||||
```python
|
||||
def comfyui_memory_load(tag: str) -> str:
|
||||
# Use ComfyUI's functions for each device
|
||||
for dev_str in devices:
|
||||
device = torch.device(dev_str)
|
||||
total = mm.get_total_memory(device) # Use ComfyUI function
|
||||
free_info = mm.get_free_memory(device, torch_free_too=True) # Use ComfyUI function
|
||||
```
|
||||
|
||||
### Device Selection Logic
|
||||
**ComfyUI Core**:
|
||||
```python
|
||||
def text_encoder_device():
|
||||
if args.gpu_only:
|
||||
return get_torch_device()
|
||||
elif vram_state == VRAMState.HIGH_VRAM or vram_state == VRAMState.NORMAL_VRAM:
|
||||
return get_torch_device()
|
||||
else:
|
||||
return torch.device("cpu")
|
||||
```
|
||||
|
||||
**MultiGPU Override**:
|
||||
```python
|
||||
def text_encoder_device_patched():
|
||||
# Respect user's explicit device choice
|
||||
devs = set(get_device_list())
|
||||
device = torch.device(current_text_encoder_device) if str(current_text_encoder_device) in devs else torch.device("cpu")
|
||||
return device
|
||||
```
|
||||
|
||||
## Architectural Insights
|
||||
|
||||
### 1. ComfyUI's Memory Philosophy
|
||||
- **Conservative by default**: Prefers CPU offloading over OOM
|
||||
- **State-driven**: Uses VRAM state to guide decisions
|
||||
- **Single-device focused**: Optimized for primary GPU + CPU paradigm
|
||||
|
||||
### 2. MultiGPU's Enhancement Philosophy
|
||||
- **User agency**: Let users specify device placement explicitly
|
||||
- **Multi-device native**: Treat all devices as equal citizens
|
||||
- **Distributed intelligence**: Spread models across available hardware
|
||||
|
||||
### 3. Symbiotic Relationship
|
||||
- ComfyUI provides the foundation and compatibility
|
||||
- MultiGPU provides the multi-device extensions
|
||||
- Both evolve independently while maintaining integration
|
||||
- Users benefit from both developments
|
||||
|
||||
## Evolution Path
|
||||
|
||||
This lineage explains MultiGPU's evolution:
|
||||
|
||||
1. **Phase 1**: Simple device selection (override device choice)
|
||||
2. **Phase 2**: Memory management extensions (multi-device cache clearing)
|
||||
3. **Phase 3**: Model distribution (DisTorch distributed loading)
|
||||
4. **Phase 4**: Production integration (proactive unloading, comprehensive patching)
|
||||
|
||||
Each phase built upon ComfyUI's existing capabilities rather than replacing them, leading to the elegant and maintainable architecture we see today.
|
||||
|
||||
## Future Considerations
|
||||
|
||||
Understanding this lineage suggests future development should:
|
||||
- Continue the extension pattern rather than replacement
|
||||
- Monitor ComfyUI core changes for integration opportunities
|
||||
- Contribute improvements back to ComfyUI core where appropriate
|
||||
- Maintain the fail-loudly approach for API changes
|
||||
|
||||
The symbiotic relationship between ComfyUI core and MultiGPU represents a model for how complex extensions can enhance rather than fragment open-source ecosystems.
|
||||
A) Vanilla ComfyUI Manager “Free
|
||||
|
||||
@@ -1,20 +1,121 @@
|
||||
No. It is clear that you do not given multiple failed implementations past this point. So, lets do this in phases.
|
||||
# CPU Memory Leak Fix Plan (Updated to Current Code State)
|
||||
|
||||
Phase 1: Implement DISTORCH2_UNLOAD_MODEL Global correctly. It should be set to True when it sees a keep_loaded=false and should be reset at the end of our patched unload_all_models. No other code changes. Document with device snapshot and memory datalog each time a new operation is done - so when it is set and unset so it can been seen in the datalog.
|
||||
Last updated: 2025-09-29
|
||||
|
||||
Phase 2: In Distorch_2.py, implement `_mgpu_unload` flag to any DisTorch model when keep_loaded=false and at the same time as setting DISTORCH2_UNLOAD_MODEL=True. In our patched unload_all_models() we create a simple evaluatioon loop with my pseudocode:
|
||||
Executive summary
|
||||
- Phase 3 (Selective Ejection) is implemented in code without the Phase 1 global sentinel.
|
||||
- Current mechanism:
|
||||
- During load, DisTorch2 nodes set a per-model transient flag: `_mgpu_unload_distorch_model = (keep_loaded == False)`.
|
||||
- End-of-workflow cleanup uses ComfyUI’s standard flags (unload_models/free_memory), which route through our patched code:
|
||||
- `mm.unload_all_models` is patched to selectively unload only models where `_mgpu_unload_distorch_model == True` and retain others (rebuilds `mm.current_loaded_models` with `kept_models`).
|
||||
- `mm.soft_empty_cache` is patched to `soft_empty_cache_distorch2_patched` for multi-device VRAM clear + adaptive CPU reset, and forced `PromptExecutor.reset()` when `force=True` (Manager parity).
|
||||
- `force_full_system_cleanup()` sets both flags exactly like Manager’s “Free model and node cache”.
|
||||
- Remaining defect (to fix next): In some flows, retained models are still ejected downstream. We had the selectiveness working earlier on this branch, so the next action is to rediscover and reinstate the exact working variant.
|
||||
|
||||
if hasattr(getattr(model, 'model', None), '_mgpu_unload'):
|
||||
multigpu_memory_log(model_hash, "_mgpu_unload=true")
|
||||
logger.mgpu_mm_log(f"[THREE_FLAG_DEBUG] model has `_mpgu_unload` flag")
|
||||
else:
|
||||
logger.mgpu_mm_log(f"[THREE_FLAG_DEBUG] model does not have _mpgu_unload flag")
|
||||
Current implementation snapshot
|
||||
|
||||
At the end of the loop no matter what calls it, DISTORCH2_UNLOAD_MODEL = FALSE with an appropriate log:
|
||||
logger.mgpu_mm_log("Setting DISTORCH2_UNLOAD_MODEL=False")
|
||||
- Per-model transient flag (set at load time)
|
||||
- File: `distorch_2.py`
|
||||
- Where: In each DisTorch2 override (UNET/CLIP/VAE), after calling the real loader:
|
||||
- `out[0].model._mgpu_unload_distorch_model = (not keep_loaded)`
|
||||
- Purpose: Mark this model for selective ejection at unload time only if the user asked not to keep it loaded.
|
||||
|
||||
Phase 3: Replace existing faulty retention or ejection logic with the loop from Phase 2:
|
||||
- Selective unload (end-of-workflow)
|
||||
- File: `model_management_mgpu.py`
|
||||
- Patch: `mm.unload_all_models` → `_mgpu_patched_unload_all_models`
|
||||
- Behavior:
|
||||
- Iterate `mm.current_loaded_models` and split into:
|
||||
- `models_to_unload`: those with `_mgpu_unload_distorch_model == True`
|
||||
- `kept_models`: everything else
|
||||
- If all models are kept (no flags set), it delegates to the original `mm.unload_all_models()`.
|
||||
- Else it unloads only `models_to_unload`, and then sets `mm.current_loaded_models = kept_models`.
|
||||
|
||||
1. At the beginning of our patched unload_all_models, check DISTORCH2_UNLOAD_MODEL
|
||||
If FALSE: run _original_unload_all_models()
|
||||
IF TRUE: Using the loop from Phase 2, apply only the unload_all_models routine to the models with `_mpgu_unload` flag set, else do nothing to other models, exactly like Else loop from Phase 2.
|
||||
- Manager parity (trigger path)
|
||||
- File: `model_management_mgpu.py`
|
||||
- `force_full_system_cleanup(reason, force=True)` sets both flags on the queue:
|
||||
- `"unload_models": True`
|
||||
- `"free_memory": True`
|
||||
- ComfyUI worker thread consumes these flags:
|
||||
- Calls `comfy.model_management.unload_all_models()` (our patched version runs)
|
||||
- Calls `PromptExecutor.reset()` when `free_memory=True`
|
||||
- Performs GC and `mm.soft_empty_cache()` (our patched version runs)
|
||||
|
||||
- Multi-device cache and CPU reset
|
||||
- File: `__init__.py`
|
||||
- Patch: `mm.soft_empty_cache` → `soft_empty_cache_distorch2_patched`
|
||||
- Detects if DisTorch2 is active
|
||||
- Clears VRAM on all devices via `soft_empty_cache_multigpu()`
|
||||
- Checks CPU pressure and optionally triggers executor reset (when forced)
|
||||
|
||||
What is not used (vs. earlier plan)
|
||||
- No global executing sentinel (e.g., `DISTORCH2_UNLOAD_MODEL`). The selective logic is driven entirely by per-model `_mgpu_unload_distorch_model` flags plus the patched unload path and standard ComfyUI flags.
|
||||
|
||||
Observed defect (to fix next)
|
||||
- In some flows (e.g., after selective unload completes), retained models get ejected anyway. Evidence points to two hotspots:
|
||||
1) The “all kept” branch in `_mgpu_patched_unload_all_models`:
|
||||
- Current code:
|
||||
- If `len(kept_models) == len(mm.current_loaded_models)`, it delegates to original `mm.unload_all_models()`.
|
||||
- That call will unload everything, defeating the selective policy when no models were flagged.
|
||||
2) Downstream actions after returning from our unload:
|
||||
- `PromptExecutor.reset()`, GC, and `soft_empty_cache()` shouldn’t unload models, but other core flows (e.g., a subsequent `free_memory()` call or a clone swap) might detach/evict retained models if not guarded.
|
||||
|
||||
Hypotheses to validate
|
||||
1) “All-kept delegation” wipes retained models
|
||||
- When no models are flagged (`models_to_unload` empty), our patch delegates to the original unload which unloads everything.
|
||||
- Fix approach: If there are zero models to unload, do nothing (no-op) — do not delegate to original unload.
|
||||
|
||||
2) Post-unload follow-on flows eject retained models
|
||||
- After selective unload, `PromptExecutor.reset()` and GC execute. These should not trigger unloads for retained models, but there may be a core call path that drives unload/evict regardless.
|
||||
- Fix approach: Instrumenting and asserting retained references across the entire `/free` flow to pinpoint where the undesired eviction occurs.
|
||||
|
||||
Rediscovery plan (the next step after committing this Memory Bank update)
|
||||
|
||||
1) Locate previously working selective retention commit(s)
|
||||
- Search this branch history for commits that logged successful retention:
|
||||
- Look for “[UNLOAD_DEBUG] Updated mm.current_loaded_models…” followed by a subsequent flow where retained models remained alive.
|
||||
- Diff the unload patch in those commits against the current `_mgpu_patched_unload_all_models` implementation.
|
||||
|
||||
2) Reinstate the proven selective no-op guard
|
||||
- Ensure this rule:
|
||||
- If `models_to_unload` is empty, return immediately (no-op). Do not delegate to original.
|
||||
- If `models_to_unload` is non-empty, unload only those and rebuild `mm.current_loaded_models = kept_models`.
|
||||
|
||||
3) Add hardening logs and assertions
|
||||
- Around unload:
|
||||
- “pre-unload snapshot”, “post-unload snapshot”, “post-reset snapshot”, “post-gc/soft_empty snapshot”.
|
||||
- If any object in `kept_models` is missing/evicted after the full free flow, log an ERROR with class name/hash.
|
||||
- Keep these until regression is confidently resolved, then demote to DEBUG if too noisy.
|
||||
|
||||
Verification matrix
|
||||
|
||||
- Minimal retention test
|
||||
- Load models: A (keep=false), B (keep=true), C (keep=true).
|
||||
- Trigger Manager-parity cleanup: unload_models=true, free_memory=true.
|
||||
- Expectation:
|
||||
- `A` is ejected. `B` and `C` remain in `mm.current_loaded_models`.
|
||||
- Memory snapshots show CPU memory decreases; VRAM caches cleared; retained models still live after the whole free flow.
|
||||
|
||||
- All kept test
|
||||
- Load models: D (keep=true), E (keep=true).
|
||||
- Trigger Manager-parity cleanup.
|
||||
- Expectation:
|
||||
- No models are ejected (strict no-op on unload when none are flagged).
|
||||
- Snapshots reflect cache cleaning only (allocator/torch caches), not model unloads.
|
||||
|
||||
Acceptance criteria
|
||||
|
||||
- After cleanup:
|
||||
- Only models flagged with `_mgpu_unload_distorch_model=True` are ejected.
|
||||
- Models with `_mgpu_unload_distorch_model=False` remain referenced by `mm.current_loaded_models` and alive after `PromptExecutor.reset()`, GC, and `soft_empty_cache()`.
|
||||
|
||||
Next steps (after this doc commit)
|
||||
- Run git history to identify the prior working selective retention commit(s).
|
||||
- Reinstate the working no-op behavior for the “all-kept” branch.
|
||||
- Add targeted logging to confirm no retained models are ejected downstream.
|
||||
- Re-run verification matrix and keep the Memory Bank synchronized.
|
||||
|
||||
Appendix: Relevant code touch points (as of today)
|
||||
- Per-model flag: `distorch_2.py` (DisTorch2 overrides)
|
||||
- Patched unload: `model_management_mgpu.py` (`mm.unload_all_models` → `_mgpu_patched_unload_all_models`)
|
||||
- Patched soft empty: `__init__.py` (`mm.soft_empty_cache` → `soft_empty_cache_distorch2_patched`)
|
||||
- Manager parity: `model_management_mgpu.py` (`force_full_system_cleanup` sets both queue flags)
|
||||
|
||||
+125
-225
@@ -1,273 +1,173 @@
|
||||
# Project Progress & Status
|
||||
# Project Progress & Status (Updated 2025-09-29)
|
||||
|
||||
## What Works (Production Ready)
|
||||
|
||||
### Core MultiGPU Infrastructure ✅
|
||||
- **Dynamic Class Override System**: City96's inheritance pattern enables automatic node creation
|
||||
- **Device Detection**: Universal support for CUDA, CPU, MPS, XPU, NPU, DirectML
|
||||
- **Memory Management**: ComfyUI-compatible device allocation and management
|
||||
- **Node Registration**: Automatic registration based on available dependencies
|
||||
- Dynamic Class Override System (City96): inheritance-based node wrapping, auto-adapts to ComfyCore
|
||||
- Device Detection: CPU, CUDA, MPS, XPU, NPU, MLU, DirectML, CoreX
|
||||
- VRAM Management: Multi-device cache clearing via `soft_empty_cache_multigpu`
|
||||
- Node Registration: Automatic node creation based on available dependencies
|
||||
|
||||
### DisTorch2 Distributed Loading ✅
|
||||
- **Universal Model Support**: .safetensors, .gguf, .bin format compatibility
|
||||
- **Load-Patch-Distribute Pipeline**: Quality-preserving LoRA application
|
||||
- **Expert Allocation Modes**: Bytes, ratios, and fraction-based distribution
|
||||
- **Performance Optimization**: 10% improvement over DisTorch V1
|
||||
- **Memory Safety**: Automatic fallbacks and error handling
|
||||
- Universal SafeTensor support (beyond GGUF)
|
||||
- Load-Patch-Distribute pipeline (quality-preserving LoRA patching on compute device)
|
||||
- Expert allocation modes (bytes, ratios, fractions)
|
||||
- ~10% performance improvement over DisTorch V1
|
||||
|
||||
### Selective Unloading (Implemented) ✅
|
||||
- Per-model transient flag is set by DisTorch2 loader wrappers:
|
||||
- `_mgpu_unload_distorch_model = (keep_loaded == False)`
|
||||
- Patched unload path:
|
||||
- `mm.unload_all_models` → selectively unloads models with `_mgpu_unload_distorch_model=True` and rebuilds `mm.current_loaded_models` with retained models
|
||||
- Patched soft empty:
|
||||
- `mm.soft_empty_cache` → `soft_empty_cache_distorch2_patched`: multi-device allocator cache clearing + adaptive CPU reset; can force executor reset for Manager parity
|
||||
- Manager parity helper:
|
||||
- `force_full_system_cleanup` sets `unload_models` and `free_memory` flags to mirror the “Free model and node cache” button
|
||||
|
||||
### Hardware Configuration Support ✅
|
||||
- **NVLink Optimization**: Near-native performance (5-7% slowdown)
|
||||
- **PCIe 4.0 CPU Offloading**: Excellent performance (40-50% slowdown)
|
||||
- **Legacy Hardware**: PCIe 3.0 support with acceptable performance
|
||||
- **Mixed Architectures**: Old + new GPU combinations work seamlessly
|
||||
- **Bandwidth Intelligence**: Performance predictions based on connection speed
|
||||
- NVLink: near-native performance
|
||||
- PCIe 4.0 CPU offloading: excellent performance
|
||||
- Legacy hardware: PCIe 3.0 coverage with documented trade-offs
|
||||
- Mixed architectures: supported
|
||||
|
||||
### External Integrations ✅
|
||||
- **ComfyUI-GGUF**: 6 DisTorch-enabled quantized model nodes
|
||||
- **WanVideoWrapper**: 8 MultiGPU video generation nodes
|
||||
- **Florence2**: Vision model multi-device support
|
||||
- **HunyuanVideoWrapper**: Native VAE + device selection support
|
||||
- **Dynamic Discovery**: Automatic node creation based on installed extensions
|
||||
- ComfyUI-GGUF: DisTorch-enabled quantized model nodes
|
||||
- WanVideoWrapper: MultiGPU video nodes
|
||||
- Florence2: Vision model support
|
||||
- HunyuanVideoWrapper: Native VAE + device selection (active)
|
||||
|
||||
### Documentation & Examples ✅
|
||||
- **Comprehensive README**: Installation, configuration, troubleshooting
|
||||
- **20+ Example Workflows**: Covering major model architectures and use cases
|
||||
- **Performance Benchmarks**: Quantified performance across hardware configurations
|
||||
- **Strategic Recommendations**: Clear guidance for different user scenarios
|
||||
- Comprehensive README
|
||||
- 20+ example workflows
|
||||
- Performance benchmarks and configuration recommendations
|
||||
|
||||
## What's Left to Build (Development Roadmap)
|
||||
## What’s Left to Build (Development Roadmap)
|
||||
|
||||
### Short-term Enhancements (Next 2-4 weeks)
|
||||
### Short-term Enhancements (Next 2–4 weeks)
|
||||
|
||||
#### Selective Retention Hardening (Top Priority) 🔄
|
||||
- Current state:
|
||||
- Phase 3 selective ejection implemented without global sentinel
|
||||
- In some flows, retained models (keep_loaded=True) are still ejected downstream
|
||||
- Likely culprits:
|
||||
1) “All-kept delegation” in patched unload: when no models are flagged, current code delegates to original unload which unloads everything
|
||||
2) Post-unload follow-on flows (PromptExecutor.reset/GC/soft_empty/free_memory path) may detach retained models
|
||||
- Action plan:
|
||||
- Rediscover prior commit(s) where selectiveness worked end-to-end
|
||||
- Reinstate strict no-op when `models_to_unload` is empty (do not delegate to original)
|
||||
- Add instrumentation: pre/post unload → post reset → post GC/soft_empty snapshots; ERROR if any kept model disappears
|
||||
- Re-run verification matrix (A=false, B/C=true; D/E all kept)
|
||||
|
||||
#### User Experience Improvements 🔄
|
||||
- **Configuration Validation**: Prevent invalid allocation strings before execution
|
||||
- **Performance Prediction**: Show estimated slowdown before model loading
|
||||
- **Better Error Messages**: Context-aware troubleshooting guidance
|
||||
- **Auto-Configuration**: Intelligent defaults based on hardware detection
|
||||
- Configuration validation and performance prediction
|
||||
- Refined error messaging for allocation/placement issues
|
||||
- Documentation refresh for current state (this update)
|
||||
|
||||
#### Integration Expansion 🔄
|
||||
- **LTX Video Support**: Next-generation video model architecture
|
||||
- **Mochi Integration**: Performance-optimized video models
|
||||
- **Community Requests**: Issue-driven custom node support
|
||||
- **Dependency Robustness**: Better handling of missing/incompatible extensions
|
||||
- LTX Video support
|
||||
- Mochi integration
|
||||
- Issue-driven community requests
|
||||
|
||||
### Medium-term Goals (2-3 months)
|
||||
### Medium-term Goals (2–3 months)
|
||||
|
||||
#### Advanced Memory Management 📋
|
||||
- **3-Flag Surgical Ejection System**: Conceptual transient flags design for CPU memory leak elimination ✅
|
||||
- **keep_loaded Boolean Engineering**: Conceptual triple-duty design for preservation, triggers, and destruction ✅
|
||||
- **Transient Flag Architecture**: Conceptual execution-scoped flags with complete external isolation ✅
|
||||
- **Smart Offloading**: Machine learning-based allocation optimization
|
||||
- **Memory Compression**: Runtime compression of stored model layers
|
||||
- **Fragmentation Handling**: Better memory pool management
|
||||
- **Pressure Monitoring**: Proactive memory pressure detection
|
||||
- Memory compression / fragmentation handling research
|
||||
- Enhanced retention/eviction policies under pressure
|
||||
- Robust regression tests for retention across `/free` flow
|
||||
|
||||
#### Professional Features 📋
|
||||
- **Batch Processing**: Multi-image/video queue optimization
|
||||
- **API Server Mode**: RESTful interface for workflow automation
|
||||
- **Quality Metrics**: Quantitative output quality measurement
|
||||
- **Performance Dashboard**: Web-based configuration and monitoring
|
||||
- Batch processing tooling
|
||||
- API server modes for automation
|
||||
- Quality metrics and reproducibility checks
|
||||
- Performance dashboard
|
||||
|
||||
#### Community Tools 📋
|
||||
- **Configuration Generator**: GUI tool for allocation string creation
|
||||
- **Hardware Profiler**: Automated bandwidth and VRAM testing
|
||||
- **Model Compatibility Database**: Community-maintained model support matrix
|
||||
- **Tutorial Content**: Video guides, blog posts, documentation expansion
|
||||
- Allocation string generator w/ validation
|
||||
- Hardware profiler (bandwidth/VRAM/latency)
|
||||
- Compatibility matrix (community-maintained)
|
||||
- Tutorials and video guides
|
||||
|
||||
### Long-term Research (6-12 months)
|
||||
### Long-term Research (6–12 months)
|
||||
|
||||
#### Next-Generation Features 🔬
|
||||
- **Model Parallelism**: Split individual layers across multiple devices
|
||||
- **Pipeline Parallelism**: Concurrent execution of workflow stages
|
||||
- **Streaming Inference**: Real-time video generation support
|
||||
- **Quality Preservation**: Mathematically proven output equivalence
|
||||
|
||||
#### Distributed Computing 🔬
|
||||
- **Multi-Node Support**: Network-distributed model inference
|
||||
- **Cloud Integration**: AWS, GCP, Azure multi-GPU instances
|
||||
- **Container Orchestration**: Kubernetes-based scaling
|
||||
- **Edge Computing**: Mobile/embedded device support
|
||||
|
||||
#### Hardware Evolution 🔬
|
||||
- **PCIe 5.0 Optimization**: Next-generation bandwidth utilization
|
||||
- **NVLink 5.0 Support**: Advanced interconnect technologies
|
||||
- **Emerging Architectures**: ARM64, RISC-V, custom AI chips
|
||||
- **Memory Technologies**: CXL, DDR6, high-bandwidth memory
|
||||
- Model parallelism and pipeline parallelism
|
||||
- Streaming inference for video
|
||||
- Multi-node/cloud distributed inference
|
||||
- Deterministic output equivalence verification
|
||||
|
||||
## Current Status Assessment
|
||||
|
||||
### Stability Rating: **Production Grade** (8/10)
|
||||
- **Memory Leaks**: CPU leaks still present - final solution conceptualized but not implemented
|
||||
- **Crash Rate**: <0.1% based on community feedback
|
||||
- **API Compatibility**: Stable across ComfyUI versions
|
||||
- **Hardware Compatibility**: 95%+ success rate across configurations
|
||||
### Stability: Production Grade (8/10)
|
||||
- CPU memory leak: Phase 3 implemented, retention bug remains in some flows
|
||||
- Crash rate: Low based on community feedback
|
||||
- API compatibility: Stable with ComfyCore
|
||||
- Hardware coverage: Broad and documented
|
||||
|
||||
### Performance Rating: **Optimized** (8/10)
|
||||
- **NVLink Performance**: Near-native (5-7% slowdown)
|
||||
- **CPU Offloading**: Excellent on modern systems (40-50% slowdown)
|
||||
- **Memory Efficiency**: Minimal overhead beyond base model requirements
|
||||
- **Transfer Optimization**: Bandwidth-optimized with predictable scaling
|
||||
### Performance: Optimized (8/10)
|
||||
- NVLink: 5–7% slowdown vs native in typical cases
|
||||
- PCIe 4.0 CPU offloading: ~40–50% slowdown with excellent price/perf
|
||||
- Predictable tradeoffs based on bandwidth hierarchy
|
||||
|
||||
### Feature Completeness: **Comprehensive** (8.5/10)
|
||||
- **Core Functionality**: All essential features implemented
|
||||
- **Model Support**: Major architectures covered (FLUX, WAN, QWEN, etc.)
|
||||
- **Hardware Support**: Universal device compatibility
|
||||
- **User Experience**: Good documentation, examples, error handling
|
||||
### Feature Completeness: Comprehensive (8.5/10)
|
||||
- Core functionality: Implemented
|
||||
- Model support: Major families (FLUX, WAN, QWEN, etc.)
|
||||
- Hardware support: Universal
|
||||
- UX: Good docs/examples; ongoing improvement
|
||||
|
||||
### Community Adoption: **Growing** (7/10)
|
||||
- **GitHub Stars**: Steady growth in community interest
|
||||
- **Issue Resolution**: 90+ issues resolved, active maintenance
|
||||
- **User Feedback**: Positive reception, feature requests indicate engagement
|
||||
- **Ecosystem Integration**: Multiple custom node dependencies
|
||||
### Community Adoption: Growing (7/10)
|
||||
- Active stars/issues/discussions
|
||||
- Integration requests from other node ecosystems
|
||||
- Positive feedback with actionable feature requests
|
||||
|
||||
## Known Issues & Limitations
|
||||
|
||||
### Technical Limitations 🐛
|
||||
### Selective Retention Bug 🐛
|
||||
- Symptom: Retained models (keep_loaded=True) sometimes ejected during `/free`
|
||||
- Cause suspects:
|
||||
- All-kept delegation to original unload
|
||||
- Post-unload flows (reset/GC/soft_empty/free_memory)
|
||||
- Status: High priority; rediscovery and hardening planned
|
||||
|
||||
#### ComfyUI API Dependencies
|
||||
- **Breaking Changes**: ComfyCore evolution can break integrations
|
||||
- **Mitigation**: Fail-loudly pattern exposes issues immediately
|
||||
- **Status**: Monitoring required, no current blocking issues
|
||||
### ComfyUI API Dependencies
|
||||
- Core changes can impact patch points
|
||||
- Fail-loudly approach surfaces issues quickly
|
||||
- Ongoing monitoring required
|
||||
|
||||
#### Hardware Edge Cases
|
||||
- **Unusual Configurations**: Some exotic hardware combinations untested
|
||||
- **Memory Allocation**: Occasional allocation failures with complex setups
|
||||
- **Status**: Community-reported, investigated on case-by-case basis
|
||||
### Hardware Edge Cases
|
||||
- Exotic configurations may need targeted validation
|
||||
- System RAM bandwidth can impact offloading performance
|
||||
|
||||
#### Performance Bottlenecks
|
||||
- **PCIe 3.0 x4**: Severe performance penalty for image generation
|
||||
- **System RAM Speed**: DDR4-2400 shows measurable slowdowns
|
||||
- **Status**: Documented limitations, not blocking for intended use cases
|
||||
### Documentation Gaps
|
||||
- Hardware selection and configuration recipes (ongoing)
|
||||
- Edge-case troubleshooting
|
||||
|
||||
### User Experience Issues 🔧
|
||||
## Evolution of Project Decisions (Highlights)
|
||||
|
||||
#### Configuration Complexity
|
||||
- **Expert Modes**: Allocation strings require technical knowledge
|
||||
- **Error Messages**: Sometimes cryptic for allocation failures
|
||||
- **Status**: Planned improvements in UX roadmap
|
||||
|
||||
#### Documentation Gaps
|
||||
- **Hardware Selection**: Users struggle with optimal hardware choices
|
||||
- **Troubleshooting**: Some edge case scenarios poorly documented
|
||||
- **Status**: Active documentation improvement effort
|
||||
|
||||
### Ecosystem Dependencies 🔗
|
||||
|
||||
#### External Custom Nodes
|
||||
- **Version Compatibility**: Breaking changes in dependencies affect integration
|
||||
- **Installation Order**: Some configurations require specific installation sequences
|
||||
- **Status**: Dependency management improvements planned
|
||||
|
||||
#### Model Format Evolution
|
||||
- **New Formats**: FP4, INT8, block-wise quantization not yet supported
|
||||
- **Architecture Changes**: New model architectures require integration updates
|
||||
- **Status**: Research ongoing, implementations follow community demand
|
||||
|
||||
## Evolution of Project Decisions
|
||||
|
||||
### Architecture Evolution Timeline
|
||||
|
||||
#### Phase 1: Basic Multi-Device (Aug 2024)
|
||||
**Decision**: Simple device selection for model loaders
|
||||
**Outcome**: Enabled multi-GPU setups but limited functionality
|
||||
**Learning**: Users wanted more than just device selection
|
||||
|
||||
#### Phase 2: Manual Node Definitions (Sep-Nov 2024)
|
||||
**Decision**: Create explicit MultiGPU versions of every loader
|
||||
**Outcome**: 400+ lines of code, maintenance nightmare
|
||||
**Learning**: Manual approaches don't scale
|
||||
|
||||
#### Phase 3: City96 Revolution (Dec 2024)
|
||||
**Decision**: Adopt inheritance-based dynamic class override
|
||||
**Outcome**: 400+ lines → 50 lines, universal compatibility
|
||||
**Learning**: Elegant architecture scales beautifully
|
||||
|
||||
#### Phase 4: DisTorch V1 (Jan-Jul 2025)
|
||||
**Decision**: GGUF-specific distributed loading
|
||||
**Outcome**: Enabled large model usage on limited VRAM
|
||||
**Learning**: Model-specific solutions don't generalize
|
||||
|
||||
#### Phase 5: DisTorch V2.0 (Aug 2025)
|
||||
**Decision**: Universal SafeTensor support with Load-Patch-Distribute
|
||||
**Outcome**: Quality parity with single-GPU, 10% performance improvement
|
||||
**Learning**: Quality preservation must be engineered, not assumed
|
||||
|
||||
#### Phase 6: Production Hardening (Sep 2025 - Current)
|
||||
**Decision**: Comprehensive benchmarking and documentation
|
||||
**Outcome**: Production-grade stability, clear performance expectations
|
||||
**Learning**: Reliability requires systematic validation
|
||||
|
||||
### Key Decision Reversals
|
||||
|
||||
#### Defensive Programming → Fail Loudly
|
||||
**Original Approach**: Try to handle all possible ComfyCore changes gracefully
|
||||
**Problem**: Masked API changes, created maintenance debt
|
||||
**New Approach**: Fail immediately when ComfyCore changes break compatibility
|
||||
**Result**: Earlier problem detection, faster fixes
|
||||
|
||||
#### Automatic Optimization → User Control
|
||||
**Original Approach**: Smart automatic allocation based on model analysis
|
||||
**Problem**: Unpredictable behavior, quality concerns with LoRA handling
|
||||
**New Approach**: Conservative defaults with expert override options
|
||||
**Result**: Predictable behavior, user trust
|
||||
|
||||
#### Custom API → ComfyUI Native
|
||||
**Original Approach**: Create abstraction layer over ComfyUI device management
|
||||
**Problem**: Broke existing workflows, fought ComfyUI patterns
|
||||
**New Approach**: Work within ComfyUI's existing device management system
|
||||
**Result**: Seamless integration, compatibility
|
||||
- Dynamic class override over manual node duplication
|
||||
- Load-Patch-Distribute over direct distribution
|
||||
- Per-model unload flag over global sentinel
|
||||
- Fail-loudly over defensive abstraction
|
||||
|
||||
## Success Metrics & Validation
|
||||
|
||||
### Technical Success Indicators
|
||||
- **Zero Crash Reports**: No memory corruption or system instability reports
|
||||
- **Quality Parity**: Bit-identical outputs vs single-GPU (with proper configuration)
|
||||
- **Performance Predictability**: Measured performance matches theoretical calculations
|
||||
- **Hardware Compatibility**: 95%+ success rate across diverse configurations
|
||||
### Technical
|
||||
- Zero regressions in selective retention tests
|
||||
- Predictable performance across bandwidth tiers
|
||||
- Quality parity with single-GPU baselines
|
||||
|
||||
### User Success Indicators
|
||||
- **Workflow Enablement**: Users running previously impossible model combinations
|
||||
- **Hardware Utilization**: Old GPUs finding new life in MultiGPU setups
|
||||
- **Community Growth**: Increasing GitHub stars, issue engagement, feature requests
|
||||
- **Professional Adoption**: Commercial users deploying in production workflows
|
||||
### User
|
||||
- Previously impossible workflows now run reliably
|
||||
- Clear guidance for low-VRAM and multi-GPU users
|
||||
- Reduced support load for common issues
|
||||
|
||||
### Ecosystem Success Indicators
|
||||
- **Integration Requests**: Other custom nodes requesting MultiGPU versions
|
||||
- **Developer Recognition**: ComfyUI core team awareness and acknowledgment
|
||||
- **Hardware Vendor Interest**: GPU manufacturers citing MultiGPU in optimization discussions
|
||||
- **Educational Impact**: Universities and courses teaching multi-GPU AI techniques
|
||||
### Ecosystem
|
||||
- Broader adoption in custom node projects
|
||||
- Recognition in optimization discussions
|
||||
- Community contributions to validation
|
||||
|
||||
## Lessons for Future Development
|
||||
|
||||
### What Scales Well
|
||||
1. **Inheritance Patterns**: Dynamic class override adapts to ecosystem evolution
|
||||
2. **Conservative Defaults**: Users prefer reliable slow over unreliable fast
|
||||
3. **Comprehensive Testing**: Systematic validation prevents regression issues
|
||||
4. **Clear Documentation**: Examples accelerate adoption more than features
|
||||
5. **Community Engagement**: User feedback drives meaningful improvements
|
||||
|
||||
### What Doesn't Scale
|
||||
1. **Manual Node Definitions**: Maintenance burden grows exponentially
|
||||
2. **Over-Engineering**: Complex solutions often perform worse than simple ones
|
||||
3. **API Abstraction**: Fighting the host framework creates ongoing conflicts
|
||||
4. **Defensive Programming**: Masking problems creates technical debt
|
||||
5. **Feature Creep**: Adding features without validation reduces quality
|
||||
|
||||
### Principles for Future Work
|
||||
1. **Work WITH ComfyUI**: Leverage existing patterns, don't fight core architecture
|
||||
2. **Validate Systematically**: Every feature needs benchmarking and testing
|
||||
3. **Document Thoroughly**: Code structure should tell the story
|
||||
4. **Engage Community**: Users know their needs better than developers assume
|
||||
5. **Fail Fast**: Early problem detection beats graceful degradation
|
||||
|
||||
## Current State Summary
|
||||
|
||||
**Production Status**: ✅ Ready for professional use
|
||||
**Performance**: ✅ Benchmarked and optimized
|
||||
**Compatibility**: ✅ Universal hardware support
|
||||
**Documentation**: ✅ Comprehensive guides and examples
|
||||
**Community**: ✅ Active user base with positive feedback
|
||||
|
||||
**Next Phase Focus**: User experience refinement and ecosystem expansion
|
||||
|
||||
The ComfyUI-MultiGPU project has evolved from a simple device selector to a comprehensive multi-device AI inference platform. Through systematic development, community feedback, and technical innovation, it now enables previously impossible AI workflows across diverse hardware configurations while maintaining production-grade reliability.
|
||||
## Next Steps (Actionable)
|
||||
- Commit Memory Bank sync (this change)
|
||||
- Git archeology to recover working selective retention diff
|
||||
- Implement strict no-op for all-kept branch in unload
|
||||
- Add temporary instrumentation; run verification matrix
|
||||
- Update docs with results and remove extra logs after stabilization
|
||||
|
||||
+97
-124
@@ -1,9 +1,9 @@
|
||||
# System Architecture & Patterns
|
||||
# System Architecture & Patterns (Updated 2025-09-29)
|
||||
|
||||
## Core Architecture
|
||||
|
||||
### Dynamic Class Override System
|
||||
**Foundation Pattern**: City96's elegant inheritance-based approach (Dec 2024 revolution)
|
||||
Foundation Pattern: City96's elegant inheritance-based approach (Dec 2024 revolution)
|
||||
|
||||
```python
|
||||
def override_class(original_class, device_param="device"):
|
||||
@@ -22,14 +22,14 @@ def override_class(original_class, device_param="device"):
|
||||
return MultiGPUClass
|
||||
```
|
||||
|
||||
**Key Benefits**:
|
||||
- **50 lines vs 400+**: Eliminated manual class definitions
|
||||
- **Universal Support**: Works with any ComfyUI loader node
|
||||
- **Maintenance**: Auto-adapts to ComfyCore changes
|
||||
- **Consistency**: Unified behavior across all MultiGPU nodes
|
||||
Key Benefits:
|
||||
- 50 lines vs 400+: Eliminated manual class definitions
|
||||
- Universal Support: Works with any ComfyUI loader node
|
||||
- Maintenance: Auto-adapts to ComfyCore changes
|
||||
- Consistency: Unified behavior across all MultiGPU nodes
|
||||
|
||||
### Load-Patch-Distribute (LPD) Method
|
||||
**DisTorch2 Core Process**:
|
||||
DisTorch2 Core Process:
|
||||
|
||||
```python
|
||||
# 1. LOAD - Always on compute device first
|
||||
@@ -43,15 +43,15 @@ if lora_patches:
|
||||
final_tensor = tensor.to(target_device)
|
||||
```
|
||||
|
||||
**Design Principles**:
|
||||
- **Quality First**: No precision loss during LoRA application
|
||||
- **Deterministic**: Same allocation every time
|
||||
- **ComfyUI Native**: Works with existing ComfyCore patterns
|
||||
Design Principles:
|
||||
- Quality First: No precision loss during LoRA application
|
||||
- Deterministic: Same allocation every time
|
||||
- ComfyUI Native: Works with existing ComfyCore patterns
|
||||
|
||||
## Memory Management Architecture
|
||||
|
||||
### Virtual VRAM System
|
||||
**Concept**: Make CPU/secondary GPU memory appear as extended VRAM
|
||||
Concept: Make CPU/secondary GPU memory appear as extended VRAM
|
||||
|
||||
```python
|
||||
class VirtualVRAM:
|
||||
@@ -61,13 +61,12 @@ class VirtualVRAM:
|
||||
self.virtual_gb = virtual_gb # Extended memory pool
|
||||
|
||||
def allocate_layers(self, model_layers, allocation_string):
|
||||
# Parse: "cuda:0,2.5gb;cpu,*"
|
||||
# Assign layers based on cumulative memory requirements
|
||||
# "cuda:0,2.5gb;cpu,*" -> assign layers based on cumulative memory
|
||||
```
|
||||
|
||||
### Expert Allocation Modes
|
||||
|
||||
**Bytes Mode** (Recommended):
|
||||
Bytes Mode (Recommended):
|
||||
```python
|
||||
# "cuda:0,2.5gb;cuda:1,3.0g;cpu,*"
|
||||
def parse_bytes_allocation(allocation_string):
|
||||
@@ -82,7 +81,7 @@ def parse_bytes_allocation(allocation_string):
|
||||
return devices
|
||||
```
|
||||
|
||||
**Ratio Mode** (llama.cpp style):
|
||||
Ratio Mode (llama.cpp style):
|
||||
```python
|
||||
# "cuda:0,25%;cpu,75%" -> 1:3 split
|
||||
def parse_ratio_allocation(allocation_string):
|
||||
@@ -95,36 +94,50 @@ def parse_ratio_allocation(allocation_string):
|
||||
return device_ratios
|
||||
```
|
||||
|
||||
## Device Detection & Management
|
||||
### Selective Ejection Pipeline (Current)
|
||||
Updated to reflect current code (Phase 3 implemented without global sentinel):
|
||||
- Load-time flagging (per-model transient):
|
||||
- In each DisTorch2 override, after the real loader returns:
|
||||
- `out[0].model._mgpu_unload_distorch_model = (keep_loaded == False)`
|
||||
- Purpose: mark this specific DisTorch model for ejection only when the user unchecked “keep loaded”.
|
||||
- Manager-parity cleanup trigger:
|
||||
- `force_full_system_cleanup(reason, force=True)` sets:
|
||||
- `unload_models=True`, `free_memory=True` on PromptQueue (exactly what Manager’s “Free model and node cache” does).
|
||||
- Selective unloading:
|
||||
- `mm.unload_all_models` is patched (`_mgpu_patched_unload_all_models` in `model_management_mgpu.py`):
|
||||
- Splits `mm.current_loaded_models` into `models_to_unload` (flag==True) and `kept_models` (flag==False).
|
||||
- If any are flagged, unloads only those and resets `mm.current_loaded_models = kept_models`.
|
||||
- Note: If no models are flagged, the current code delegates to the original `unload_all_models()` (this is under review; see “Hardened Rule” below).
|
||||
- Multi-device VRAM cache + CPU reset:
|
||||
- `mm.soft_empty_cache` is patched to `soft_empty_cache_distorch2_patched`:
|
||||
- Detects DisTorch2-active state and clears allocator caches on all devices via `soft_empty_cache_multigpu()`
|
||||
- Adaptive CPU memory reset (threshold-based), and optional forced `PromptExecutor.reset()` when `force=True` for Manager parity.
|
||||
|
||||
### Multi-Device Enumeration
|
||||
Hardened Rule (target behavior to restore):
|
||||
- If `models_to_unload` is empty, `unload_all_models` should be a strict no-op (do not delegate to the original). Retained models must never be ejected when no flags are set. This will be re-applied during the rediscovery step.
|
||||
|
||||
### Device Detection & Management
|
||||
|
||||
Multi-Device Enumeration:
|
||||
```python
|
||||
def get_device_list():
|
||||
devices = ["cpu"] # Always available
|
||||
|
||||
# CUDA detection
|
||||
if torch.cuda.is_available():
|
||||
devices.extend([f"cuda:{i}" for i in range(torch.cuda.device_count())])
|
||||
|
||||
# Extended device support
|
||||
for device_type in ["xpu", "npu", "mlu", "mps"]:
|
||||
if device_available(device_type):
|
||||
devices.append(device_type)
|
||||
|
||||
# XPU/NPU/MLU/MPS/DirectML/CoreX detection...
|
||||
return devices
|
||||
```
|
||||
|
||||
### Device Bandwidth Intelligence
|
||||
**Hierarchy** (from benchmarking data):
|
||||
1. **NVLINK**: ~50.8 GB/s (near-native performance)
|
||||
2. **PCIe 4.0 x16**: ~27.2 GB/s (excellent CPU offloading)
|
||||
3. **PCIe 3.0 x8**: ~6.8 GB/s (acceptable for video models)
|
||||
4. **PCIe 3.0 x4**: ~2.1 GB/s (slow but viable for capacity)
|
||||
Device Bandwidth Intelligence (from benchmarking):
|
||||
1. NVLINK (~50.8 GB/s)
|
||||
2. PCIe 4.0 x16 (~27.2 GB/s)
|
||||
3. PCIe 3.0 x8 (~6.8 GB/s)
|
||||
4. PCIe 3.0 x4 (~2.1 GB/s)
|
||||
|
||||
## Integration Patterns
|
||||
|
||||
### ComfyCore Alignment
|
||||
**Philosophy**: Work WITH ComfyUI, not against it
|
||||
Philosophy: Work WITH ComfyUI, not against it
|
||||
|
||||
```python
|
||||
# GOOD: Use ComfyCore's device management
|
||||
@@ -140,9 +153,6 @@ torch.cuda.set_device(device_id) # Bypasses ComfyCore
|
||||
# Dynamic registration based on available dependencies
|
||||
if "ComfyUI-GGUF" in installed_modules:
|
||||
NODE_CLASS_MAPPINGS["UnetLoaderGGUFDisTorch2MultiGPU"] = create_gguf_distorch_node()
|
||||
|
||||
if "ComfyUI-WanVideoWrapper" in installed_modules:
|
||||
NODE_CLASS_MAPPINGS["WanVideoModelLoaderMultiGPU"] = create_wanvideo_node()
|
||||
```
|
||||
|
||||
### Dependency Detection
|
||||
@@ -152,10 +162,6 @@ def check_module_availability(module_paths):
|
||||
if os.path.exists(os.path.join(custom_nodes_dir, path)):
|
||||
return True
|
||||
return False
|
||||
|
||||
# Example: Check for multiple possible names
|
||||
GGUF_PATHS = ["ComfyUI-GGUF", "comfyui-gguf", "ComfyUI_GGUF"]
|
||||
has_gguf = check_module_availability(GGUF_PATHS)
|
||||
```
|
||||
|
||||
## Performance Optimization Patterns
|
||||
@@ -163,28 +169,20 @@ has_gguf = check_module_availability(GGUF_PATHS)
|
||||
### Layer Transfer Optimization
|
||||
```python
|
||||
def optimized_layer_transfer(layer, source_device, target_device):
|
||||
"""Optimized tensor transfer with memory management"""
|
||||
if source_device == target_device:
|
||||
return layer
|
||||
|
||||
# Use non_blocking for CUDA->CUDA transfers
|
||||
non_blocking = "cuda" in source_device and "cuda" in target_device
|
||||
|
||||
# Pin memory for CPU->GPU transfers
|
||||
if source_device == "cpu" and "cuda" in target_device:
|
||||
layer = layer.pin_memory()
|
||||
|
||||
return layer.to(target_device, non_blocking=non_blocking)
|
||||
```
|
||||
|
||||
### Memory Pressure Management
|
||||
```python
|
||||
def should_auto_offload(model_size_gb, vram_available_gb, threshold=0.9):
|
||||
"""Automatic offloading when model exceeds 90% of available VRAM"""
|
||||
return model_size_gb > (vram_available_gb * threshold)
|
||||
|
||||
def calculate_offload_amount(model_size_gb, target_vram_usage_gb):
|
||||
"""Calculate exact amount to offload for target VRAM usage"""
|
||||
return max(0, model_size_gb - target_vram_usage_gb)
|
||||
```
|
||||
|
||||
@@ -194,20 +192,17 @@ def calculate_offload_amount(model_size_gb, target_vram_usage_gb):
|
||||
```python
|
||||
# GOOD: Let ComfyCore changes surface immediately
|
||||
def load_model(self, model_name, device):
|
||||
# No try/except - we want to know if ComfyCore changes break us
|
||||
return original_loader.load_unet(model_name, device)
|
||||
|
||||
# AVOID: Defensive coding that masks issues
|
||||
try:
|
||||
return original_loader.load_unet(model_name, device)
|
||||
except AttributeError:
|
||||
# This hides when ComfyCore API changes
|
||||
return fallback_method()
|
||||
```
|
||||
|
||||
### Integration Validation
|
||||
```python
|
||||
# Validate ComfyCore compatibility at startup
|
||||
def validate_comfycore_integration():
|
||||
required_attrs = ['FUNCTION', 'INPUT_TYPES', 'RETURN_TYPES']
|
||||
for attr in required_attrs:
|
||||
@@ -219,79 +214,59 @@ def validate_comfycore_integration():
|
||||
|
||||
### Self-Documenting Code
|
||||
```python
|
||||
# GOOD: Names explain purpose
|
||||
def override_class_with_device_selection(original_class, device_param_name="device"):
|
||||
compute_device = kwargs.get(device_param_name, mm.get_torch_device())
|
||||
|
||||
# AVOID: Cryptic naming requiring comments
|
||||
def oc_wds(oc, dpn="device"): # override class with device selection
|
||||
cd = kwargs.get(dpn, mm.gtd()) # compute device = get torch device
|
||||
```
|
||||
|
||||
### Minimal Comments Philosophy
|
||||
```python
|
||||
# GOOD: Code structure tells the story
|
||||
class DisTorchLoader:
|
||||
def __init__(self, compute_device, donor_device, virtual_vram_gb):
|
||||
self.compute_device = compute_device
|
||||
self.donor_device = donor_device
|
||||
self.virtual_vram_gb = virtual_vram_gb
|
||||
|
||||
def load_model_with_distribution(self, model_path, allocation_string):
|
||||
model = self.load_on_compute_device(model_path)
|
||||
distributed_model = self.distribute_layers(model, allocation_string)
|
||||
return distributed_model
|
||||
|
||||
# AVOID: Over-commenting obvious code
|
||||
class DisTorchLoader:
|
||||
def __init__(self, compute_device, donor_device, virtual_vram_gb):
|
||||
# Set the compute device for processing
|
||||
self.compute_device = compute_device
|
||||
# Set the donor device for storage
|
||||
self.donor_device = donor_device
|
||||
# Set the virtual VRAM amount in gigabytes
|
||||
self.virtual_vram_gb = virtual_vram_gb
|
||||
```
|
||||
Prefer structure and naming to convey intent; use comments for non-obvious constraints/assumptions.
|
||||
|
||||
## Architectural Decision Records
|
||||
|
||||
### Why Dynamic Class Override vs Manual Definitions
|
||||
**Decision**: Use inheritance-based class override (City96 approach)
|
||||
**Rationale**:
|
||||
Decision: Use inheritance-based class override (City96 approach)
|
||||
Rationale:
|
||||
- Reduces code from 400+ lines to ~50 lines
|
||||
- Auto-adapts to ComfyCore changes
|
||||
- Eliminates maintenance burden of manual node definitions
|
||||
- Provides consistent behavior across all node types
|
||||
|
||||
### Why Load-Patch-Distribute vs Direct Distribution
|
||||
**Decision**: Always load on compute device first, then distribute
|
||||
**Rationale**:
|
||||
Decision: Always load on compute device first, then distribute
|
||||
Rationale:
|
||||
- Ensures LoRA patches applied at full precision
|
||||
- Maintains quality parity with single-GPU workflows
|
||||
- Predictable behavior regardless of target device
|
||||
- Works with ComfyCore's existing patching mechanisms
|
||||
- Works with ComfyCore’s existing patching mechanisms
|
||||
|
||||
### Why Expert Modes vs Automatic Only
|
||||
**Decision**: Provide both automatic and expert allocation modes
|
||||
**Rationale**:
|
||||
Decision: Provide both automatic and expert allocation modes
|
||||
Rationale:
|
||||
- Automatic mode enables low-VRAM users immediately
|
||||
- Expert modes allow optimization for specific hardware
|
||||
- Benchmarking shows performance depends on hardware configuration
|
||||
- Power users need fine-grained control
|
||||
- Performance depends on bandwidth topology; experts need control
|
||||
|
||||
### Why Universal Device Support vs CUDA-Only
|
||||
**Decision**: Support CPU, XPU, NPU, MLU, MPS, DirectML alongside CUDA
|
||||
**Rationale**:
|
||||
- ComfyUI runs on diverse hardware platforms
|
||||
- Apple Silicon (MPS) and Intel hardware (XPU) growing user bases
|
||||
- Future-proofing for emerging compute devices
|
||||
- Principle of hardware democracy
|
||||
Decision: Support CPU, XPU, NPU, MLU, MPS, DirectML alongside CUDA
|
||||
Rationale:
|
||||
- ComfyUI’s user base spans diverse hardware
|
||||
- Future-proof for emerging accelerators
|
||||
- Hardware democracy principle
|
||||
|
||||
### Why Per-Model Flag vs Global Sentinel (Updated)
|
||||
Decision: Use per-model `_mgpu_unload_distorch_model` instead of a global “DISTORCH2_UNLOAD_MODEL” sentinel
|
||||
Rationale:
|
||||
- Surgical precision at model granularity
|
||||
- No persistent or cross-workflow state
|
||||
- Cleaner semantics under ComfyUI’s queue/flag model
|
||||
|
||||
Hardened unloading rule (target to re-apply):
|
||||
- If no models are flagged for ejection, `mm.unload_all_models` must be a strict no-op to preserve retained models across the full Manager-parity flow.
|
||||
|
||||
## Testing & Validation Patterns
|
||||
|
||||
### Hardware Configuration Testing
|
||||
```python
|
||||
# Test matrix for different hardware combinations
|
||||
HARDWARE_CONFIGS = [
|
||||
{"compute": "cuda:0", "donor": "cpu", "connection": "PCIe 4.0 x16"},
|
||||
{"compute": "cuda:0", "donor": "cuda:1", "connection": "NVLink"},
|
||||
@@ -302,7 +277,6 @@ HARDWARE_CONFIGS = [
|
||||
|
||||
### Model Compatibility Validation
|
||||
```python
|
||||
# Test different model architectures and formats
|
||||
TEST_MODELS = [
|
||||
{"name": "FLUX.1-dev", "format": ".safetensors", "size_gb": 23.8},
|
||||
{"name": "WAN 2.2", "format": ".safetensors", "size_gb": 14.0},
|
||||
@@ -314,9 +288,7 @@ TEST_MODELS = [
|
||||
### Performance Regression Testing
|
||||
```python
|
||||
def benchmark_allocation_performance(model, hardware_config, allocation_configs):
|
||||
"""Ensure performance doesn't regress with updates"""
|
||||
baseline_time = benchmark_single_gpu(model)
|
||||
|
||||
for allocation in allocation_configs:
|
||||
distributed_time = benchmark_distributed(model, hardware_config, allocation)
|
||||
performance_ratio = distributed_time / baseline_time
|
||||
@@ -326,26 +298,32 @@ def benchmark_allocation_performance(model, hardware_config, allocation_configs)
|
||||
## Module Architecture (Post-Refactoring)
|
||||
|
||||
### Core Module Separation
|
||||
**Problem Solved**: Eliminated circular import `device_utils.py` ↔ `distorch_2.py`
|
||||
Problem Solved: Eliminated circular import `device_utils.py` ↔ `distorch_2.py`
|
||||
|
||||
**Solution**: Created `model_management_mgpu.py` as central model lifecycle hub
|
||||
Solution: `model_management_mgpu.py` as central model lifecycle hub
|
||||
|
||||
### Module Responsibilities
|
||||
|
||||
**device_utils.py** (Base Layer):
|
||||
device_utils.py (Base Layer):
|
||||
- Device enumeration and detection
|
||||
- VRAM cache management (`soft_empty_cache_multigpu`)
|
||||
- Pure hardware abstraction - NO model tracking
|
||||
- Pure hardware abstraction – no model tracking
|
||||
|
||||
**model_management_mgpu.py** (Core Layer):
|
||||
- Model lifecycle tracking (`track_modelpatcher`)
|
||||
- Memory logging (`multigpu_memory_log`)
|
||||
- System cleanup (`force_full_system_cleanup`, `trigger_executor_cache_reset`)
|
||||
model_management_mgpu.py (Core Layer):
|
||||
- Model lifecycle tracking and memory logging
|
||||
- Cleanup orchestration (`force_full_system_cleanup`, `trigger_executor_cache_reset`, `check_cpu_memory_threshold`)
|
||||
- Patched unload path (selective ejection)
|
||||
|
||||
**distorch_2.py/distorch.py** (Feature Layer):
|
||||
- DisTorch distribution algorithms
|
||||
- SafeTensor/GGUF specific logic
|
||||
- Imports FROM core/base layers ONLY
|
||||
distorch_2.py/distorch.py (Feature Layer):
|
||||
- DisTorch distribution algorithms and allocation analysis
|
||||
- Per-model flagging (`_mgpu_unload_distorch_model`) during DisTorch loads
|
||||
- Imports FROM Core/Base only
|
||||
|
||||
UI Layer: nodes.py, checkpoint_multigpu.py
|
||||
- Device-aware user interfaces and node definitions
|
||||
|
||||
Assembly: __init__.py
|
||||
- Final integration/patch registration (`mm.soft_empty_cache` patch, node maps)
|
||||
|
||||
### Import Flow Architecture
|
||||
```
|
||||
@@ -355,36 +333,31 @@ def benchmark_allocation_performance(model, hardware_config, allocation_configs)
|
||||
↑
|
||||
┌─────────────────┐
|
||||
│ UI Layer │ ← nodes.py, checkpoint_multigpu.py
|
||||
│ (User Interface)│
|
||||
└─────────────────┘
|
||||
↑
|
||||
┌─────────────────┐
|
||||
│ Feature Layer │ ← distorch_2.py, distorch.py
|
||||
│ (DisTorch Logic)│
|
||||
└─────────────────┘
|
||||
↑
|
||||
┌─────────────────┐
|
||||
│ Core Layer │ ← model_management_mgpu.py
|
||||
│ (Model Lifecycle)│
|
||||
└─────────────────┘
|
||||
↑
|
||||
┌─────────────────┐
|
||||
│ Base Layer │ ← device_utils.py
|
||||
│ (Hardware) │
|
||||
└─────────────────┘
|
||||
```
|
||||
|
||||
### Architectural Validation
|
||||
**Rule**: Dependencies only flow UPWARD. Violations create circular imports.
|
||||
Rule: Dependencies only flow UPWARD. Violations create circular imports.
|
||||
|
||||
**Prevention**: Before any import, ask "Does this violate the layer hierarchy?"
|
||||
Prevention: Before any import, verify it respects the layer hierarchy.
|
||||
|
||||
### Function Migration Record
|
||||
**Moved from device_utils.py to model_management_mgpu.py:**
|
||||
- `multigpu_memory_log` - Memory state logging
|
||||
- `track_modelpatcher` - ModelPatcher lifecycle tracking
|
||||
- `trigger_executor_cache_reset` - CPU memory management
|
||||
- `check_cpu_memory_threshold` - Adaptive cleanup triggers
|
||||
- `force_full_system_cleanup` - Full system reset
|
||||
Moved from device_utils.py to model_management_mgpu.py:
|
||||
- `multigpu_memory_log` – memory state logging
|
||||
- `trigger_executor_cache_reset` – CPU memory management
|
||||
- `check_cpu_memory_threshold` – adaptive cleanup triggers
|
||||
- `force_full_system_cleanup` – Manager-parity free flow
|
||||
|
||||
**Rationale**: These functions manage model lifecycle and memory state, not hardware detection. Separation prevents circular dependencies while maintaining clean responsibilities.
|
||||
Rationale: These belong to model lifecycle/cleanup, not hardware enumeration.
|
||||
|
||||
+118
-225
@@ -1,25 +1,25 @@
|
||||
# Technical Context & Dependencies
|
||||
# Technical Context & Dependencies (Updated 2025-09-29)
|
||||
|
||||
## Core Technology Stack
|
||||
|
||||
### Python Environment
|
||||
**Requirements**:
|
||||
- **Python 3.8+**: ComfyUI minimum requirement
|
||||
- **PyTorch 2.0+**: Core tensor operations and device management
|
||||
- **CUDA 11.8+/12.x**: GPU compute support (when available)
|
||||
- **ComfyUI**: Host framework (dynamic dependency)
|
||||
Requirements:
|
||||
- Python 3.10+ recommended
|
||||
- PyTorch 2.x (CUDA/HIP/XPU backends as available)
|
||||
- ComfyUI as host framework
|
||||
|
||||
### Framework Dependencies
|
||||
|
||||
#### Required (ComfyUI Core)
|
||||
Required (ComfyUI Core)
|
||||
```python
|
||||
import torch
|
||||
import comfy.model_management as mm
|
||||
import comfy.model_patcher
|
||||
import comfy.utils
|
||||
import folder_paths
|
||||
```
|
||||
|
||||
#### Optional (External Custom Nodes)
|
||||
Optional (External Custom Nodes)
|
||||
```python
|
||||
# ComfyUI-GGUF Integration
|
||||
try:
|
||||
@@ -38,245 +38,138 @@ except ImportError:
|
||||
|
||||
## Device Support Matrix
|
||||
|
||||
### Primary Support (Tested)
|
||||
- **CUDA**: GeForce RTX series, Professional/Quadro cards
|
||||
- **CPU**: x86_64 systems with sufficient RAM (16GB+ recommended)
|
||||
- **MPS**: Apple Silicon (M1/M2/M3) via Metal Performance Shaders
|
||||
Primary Support (tested)
|
||||
- CUDA (NVIDIA)
|
||||
- CPU
|
||||
- MPS (Apple Metal)
|
||||
|
||||
### Extended Support (Community Validated)
|
||||
- **XPU**: Intel Arc GPUs, integrated graphics
|
||||
- **NPU**: Intel NPU for Core 7 processors
|
||||
- **HIP/ROCm**: AMD GPUs on Linux (community contributed)
|
||||
- **DirectML**: Windows ML acceleration layer
|
||||
Extended/Community
|
||||
- XPU (Intel)
|
||||
- NPU (Ascend)
|
||||
- MLU (Cambricon)
|
||||
- DirectML (Windows)
|
||||
- CoreX/IXUCA
|
||||
|
||||
### Hardware Constraints
|
||||
## Integration Architecture (Current Patch Points)
|
||||
|
||||
#### Memory Requirements
|
||||
- **Minimum RAM**: 16GB system memory
|
||||
- **Recommended RAM**: 32GB+ for large model offloading
|
||||
- **VRAM**: No minimum (CPU-only operation supported)
|
||||
- **Storage**: NVMe SSD recommended for model loading speed
|
||||
This project extends ComfyUI through carefully scoped patches and runtime overrides. The current core integration points are:
|
||||
|
||||
#### Connection Bandwidth Hierarchy
|
||||
1. **NVLINK 2x3090**: 50.8 GB/s (optimal)
|
||||
2. **PCIe 5.0 x16**: ~63 GB/s theoretical (future GPUs)
|
||||
3. **PCIe 4.0 x16**: ~27.2 GB/s measured
|
||||
4. **PCIe 3.0 x16**: ~15.8 GB/s theoretical
|
||||
5. **PCIe 3.0 x8**: ~6.8 GB/s measured
|
||||
6. **PCIe 3.0 x4**: ~2.1 GB/s measured
|
||||
1) get_torch_device/text_encoder_device override (device selection)
|
||||
- File: `__init__.py`
|
||||
- Patch:
|
||||
- `mm.get_torch_device = get_torch_device_patched`
|
||||
- `mm.text_encoder_device = text_encoder_device_patched`
|
||||
- Purpose: Respect user-selected devices handoff by MultiGPU wrappers and maintain ComfyUI alignment.
|
||||
|
||||
2) soft_empty_cache (multi-device + CPU reset)
|
||||
- File: `__init__.py`
|
||||
- Patch:
|
||||
- `mm.soft_empty_cache = soft_empty_cache_distorch2_patched`
|
||||
- Behavior:
|
||||
- Detects DisTorch2 activity, clears allocator caches across ALL devices via `soft_empty_cache_multigpu()` (from `device_utils.py`)
|
||||
- Adaptive CPU memory reset (threshold-based), and optional forced `PromptExecutor.reset()` when `force=True` (Manager parity)
|
||||
|
||||
3) unload_all_models (selective ejection)
|
||||
- File: `model_management_mgpu.py`
|
||||
- Patch:
|
||||
- `mm.unload_all_models = _mgpu_patched_unload_all_models`
|
||||
- Behavior:
|
||||
- Splits `mm.current_loaded_models` into:
|
||||
- `models_to_unload` where per-model `_mgpu_unload_distorch_model == True`
|
||||
- `kept_models` for all others
|
||||
- If flagged models exist: unload them only, then set `mm.current_loaded_models = kept_models`
|
||||
- Current caveat: When none are flagged, the code delegates to the original unload (target is strict no-op; see System Patterns and Fix Plan)
|
||||
|
||||
4) DisTorch2 load-time model flagging (per-model transient)
|
||||
- File: `distorch_2.py`
|
||||
- Where:
|
||||
- In DisTorch2 wrappers (UNET/CLIP/VAE) within `override(...)` after original call:
|
||||
- `out[0].model._mgpu_unload_distorch_model = (keep_loaded == False)`
|
||||
- Rationale:
|
||||
- Surgical per-model control enables selective ejection in patched unload without any global sentinel
|
||||
|
||||
5) Manager parity helper
|
||||
- File: `model_management_mgpu.py`
|
||||
- Function:
|
||||
- `force_full_system_cleanup(reason="manual", force=True)`
|
||||
- Behavior:
|
||||
- Sets both `unload_models=True` and `free_memory=True` on PromptQueue, matching Manager’s “Free model and node cache” button behavior
|
||||
|
||||
## Selective Ejection Flow (Technical Overview)
|
||||
|
||||
- Load time (DisTorch2 wrappers):
|
||||
- Mark models for ejection if keep_loaded=False
|
||||
- Free flow (Manager or programmatic parity):
|
||||
- /free → prompt_worker picks flags → calls `mm.unload_all_models()` (selective) → `PromptExecutor.reset()` → GC → `mm.soft_empty_cache()` (multi-device)
|
||||
- Intended properties:
|
||||
- Models flagged for ejection are destroyed
|
||||
- Retained models remain live after full flow (including reset/GC/soft_empty)
|
||||
|
||||
Current caveat (to fix next):
|
||||
- When no models are flagged, the patched unload delegates to the original unload, which unloads everything. The target is strict no-op in this branch.
|
||||
|
||||
## Development Environment
|
||||
|
||||
### Supported Operating Systems
|
||||
- **Linux**: Primary development platform (Ubuntu 20.04+, others)
|
||||
- **Windows 10/11**: Full support with CUDA/DirectML
|
||||
- **macOS**: MPS support for Apple Silicon
|
||||
Supported OS
|
||||
- Linux (primary)
|
||||
- Windows 10/11
|
||||
- macOS (Apple Silicon via MPS)
|
||||
|
||||
### Development Tools
|
||||
- **IDE**: VSCode with Python extensions
|
||||
- **Version Control**: Git with conventional commits
|
||||
- **Testing**: Manual validation across hardware configurations
|
||||
- **Performance**: Built-in benchmarking tools
|
||||
|
||||
### Build System
|
||||
```toml
|
||||
# pyproject.toml
|
||||
[build-system]
|
||||
requires = ["setuptools", "wheel"]
|
||||
|
||||
[project]
|
||||
name = "comfyui-multigpu"
|
||||
version = "2.4.7"
|
||||
dependencies = [] # All dependencies via ComfyUI
|
||||
```
|
||||
|
||||
## Integration Architecture
|
||||
|
||||
### ComfyUI Core Integration Points
|
||||
|
||||
#### Model Management Hooks
|
||||
```python
|
||||
# Patch ComfyUI's device management
|
||||
original_get_torch_device = mm.get_torch_device
|
||||
original_text_encoder_device = mm.text_encoder_device
|
||||
|
||||
def get_torch_device_patched():
|
||||
return current_multigpu_device or original_get_torch_device()
|
||||
```
|
||||
|
||||
#### Node Registration System
|
||||
```python
|
||||
# Dynamic node creation based on available dependencies
|
||||
NODE_CLASS_MAPPINGS = {}
|
||||
|
||||
# Core MultiGPU nodes (always available)
|
||||
for node_name in ["UNETLoader", "VAELoader", "CLIPLoader"]:
|
||||
if node_name in GLOBAL_NODE_CLASS_MAPPINGS:
|
||||
NODE_CLASS_MAPPINGS[f"{node_name}MultiGPU"] = override_class(
|
||||
GLOBAL_NODE_CLASS_MAPPINGS[node_name]
|
||||
)
|
||||
|
||||
# Conditional nodes based on extensions
|
||||
if GGUF_AVAILABLE:
|
||||
NODE_CLASS_MAPPINGS["UnetLoaderGGUFDisTorch2MultiGPU"] = create_gguf_distorch_node()
|
||||
```
|
||||
|
||||
### External Custom Node Integrations
|
||||
|
||||
#### ComfyUI-GGUF
|
||||
- **Purpose**: GGUF quantized model support
|
||||
- **Integration**: DisTorch for layer-wise distribution
|
||||
- **Requirements**: city96/ComfyUI-GGUF installed
|
||||
- **Nodes Created**: 6 GGUF-specific MultiGPU nodes
|
||||
|
||||
#### ComfyUI-WanVideoWrapper
|
||||
- **Purpose**: Kijai's optimized video model support
|
||||
- **Integration**: BlockSwap + MultiGPU device selection
|
||||
- **Requirements**: kijai/ComfyUI-WanVideoWrapper installed
|
||||
- **Nodes Created**: 8 WanVideo-specific MultiGPU nodes
|
||||
|
||||
#### ComfyUI-Florence2
|
||||
- **Purpose**: Microsoft Florence2 vision model support
|
||||
- **Integration**: Standard MultiGPU device override
|
||||
- **Requirements**: kijai/ComfyUI-Florence2 installed
|
||||
- **Nodes Created**: 2 Florence2-specific MultiGPU nodes
|
||||
Tools
|
||||
- IDE: VSCode
|
||||
- VCS: Git (conventional commits encouraged)
|
||||
- Testing: Manual validation across available hardware + community testing
|
||||
|
||||
## Performance Characteristics
|
||||
|
||||
### Memory Transfer Patterns
|
||||
Bandwidth hierarchy
|
||||
1. NVLink (~50.8 GB/s) – near-native performance
|
||||
2. PCIe 4.0 x16 (~27.2 GB/s) – excellent offloading
|
||||
3. PCIe 3.0 x8 (~6.8 GB/s)
|
||||
4. PCIe 3.0 x4 (~2.1 GB/s)
|
||||
|
||||
#### Optimal Configurations
|
||||
```python
|
||||
OPTIMAL_CONFIGS = {
|
||||
"image_generation": {
|
||||
"priority": "bandwidth",
|
||||
"recommended": ["nvlink", "pcie_4_0_x16_cpu"],
|
||||
"acceptable": ["pcie_3_0_x16_cpu"],
|
||||
"avoid": ["pcie_3_0_x8_gpu", "pcie_3_0_x4_gpu"]
|
||||
},
|
||||
"video_generation": {
|
||||
"priority": "capacity",
|
||||
"recommended": ["any_available"],
|
||||
"acceptable": ["pcie_3_0_x4_gpu", "slow_cpu"],
|
||||
"avoid": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Transfer Optimization
|
||||
- **Pinned Memory**: CPU→GPU transfers use pinned memory allocation
|
||||
- **Non-blocking Transfers**: GPU→GPU uses asynchronous copying
|
||||
- **Batch Transfers**: Multiple small layers combined into single transfer
|
||||
- **Memory Pressure**: Automatic garbage collection during heavy usage
|
||||
|
||||
### Model-Specific Behaviors
|
||||
|
||||
#### GGUF Models (DisTorch V1/V2)
|
||||
- **Quantization**: Q8_0, Q6_K, Q4_K_M supported
|
||||
- **Layer Granularity**: Individual GGML tensor distribution
|
||||
- **Performance**: 10% speed improvement in DisTorch V2
|
||||
- **Memory**: Native quantized storage, no dequantization overhead
|
||||
|
||||
#### SafeTensor Models (DisTorch V2)
|
||||
- **Precision**: FP16, BF16, FP8 native support
|
||||
- **LoRA Compatibility**: Full-precision patching on compute device
|
||||
- **Layer Distribution**: Based on tensor memory footprint
|
||||
- **Quality**: No quality loss vs single-GPU operation
|
||||
Load-Patch-Distribute (LPD)
|
||||
- Always load on compute device first
|
||||
- Apply LoRAs at full precision
|
||||
- Distribute blocks to assigned devices for final placement
|
||||
- Ensures quality preservation and deterministic behavior
|
||||
|
||||
## Configuration Management
|
||||
|
||||
### Expert Allocation String Formats
|
||||
|
||||
#### Bytes Mode (Recommended)
|
||||
```python
|
||||
# Format: "device1,amount1;device2,amount2;overflow_device,*"
|
||||
BYTES_EXAMPLES = [
|
||||
"cuda:0,2.5gb;cpu,*", # Simple CPU offload
|
||||
"cuda:0,500mb;cuda:1,3.0g;cpu,*", # Multi-GPU distribution
|
||||
"cuda:0,1024mb;cuda:1,2048mb;cpu,*" # Exact memory control
|
||||
]
|
||||
```
|
||||
|
||||
#### Ratio Mode (llama.cpp style)
|
||||
```python
|
||||
# Format: "device1,percentage1%;device2,percentage2%"
|
||||
RATIO_EXAMPLES = [
|
||||
"cuda:0,25%;cpu,75%", # 1:3 split
|
||||
"cuda:0,40%;cuda:1,60%", # GPU-only distribution
|
||||
"cuda:0,10%;cuda:1,10%;cpu,80%" # Multi-device split
|
||||
]
|
||||
```
|
||||
|
||||
#### Legacy Fraction Mode
|
||||
```python
|
||||
# Format: fraction of device VRAM to use
|
||||
FRACTION_EXAMPLES = [
|
||||
0.8, # Use 80% of available VRAM
|
||||
0.5, # Use 50% of available VRAM
|
||||
0.95 # Use 95% of available VRAM
|
||||
]
|
||||
```
|
||||
|
||||
## Development Constraints
|
||||
|
||||
### ComfyUI API Stability
|
||||
- **Challenge**: ComfyUI core evolves rapidly
|
||||
- **Strategy**: Minimal API surface area, fail-loudly on changes
|
||||
- **Pattern**: Use inheritance to adapt to API evolution
|
||||
- **Testing**: Validate against multiple ComfyUI versions
|
||||
|
||||
### Hardware Diversity
|
||||
- **Challenge**: Thousands of possible hardware combinations
|
||||
- **Strategy**: Focus on most common configurations
|
||||
- **Community**: User-contributed validation for edge cases
|
||||
- **Benchmarking**: Systematic performance characterization
|
||||
|
||||
### Memory Management Complexity
|
||||
- **Challenge**: PyTorch + CUDA memory semantics
|
||||
- **Strategy**: Leverage ComfyUI's existing memory management
|
||||
- **Safety**: Automatic fallbacks for allocation failures
|
||||
- **Monitoring**: Built-in memory pressure detection
|
||||
Expert allocation strings
|
||||
- Bytes mode (recommended):
|
||||
- `"cuda:0,2.5gb;cuda:1,3.0g;cpu,*"`
|
||||
- Ratio mode:
|
||||
- `"cuda:0,25%;cpu,75%"`
|
||||
- Fraction mode (legacy):
|
||||
- `0.8`, `0.5`, `0.95`
|
||||
|
||||
## Debugging & Monitoring
|
||||
|
||||
### Logging Infrastructure
|
||||
```python
|
||||
import logging
|
||||
logger = logging.getLogger("MultiGPU")
|
||||
Logging
|
||||
- `logger.mgpu_mm_log(...)` for structured memory/system logs
|
||||
- `multigpu_memory_log(identifier, tag)` for timestamped CPU/VRAM snapshots
|
||||
|
||||
# Structured logging for performance analysis
|
||||
logger.info(f"[DisTorch2] Model {model_id} allocated: {allocation_summary}")
|
||||
logger.debug(f"Layer {layer_name} transferred {source} -> {target} in {transfer_time}ms")
|
||||
```
|
||||
Inspection
|
||||
- `device_utils.comfyui_memory_load(tag)` for one-line current memory snapshot
|
||||
- VRAM cache clearing logs around `soft_empty_cache_multigpu()`
|
||||
|
||||
### Performance Telemetry
|
||||
- **Transfer Times**: Track layer transfer latencies
|
||||
- **Memory Usage**: Monitor VRAM/RAM utilization per device
|
||||
- **Model Loading**: Time model initialization phases
|
||||
- **Inference Impact**: Measure per-step slowdown vs baseline
|
||||
## Architectural Rationale (Updated)
|
||||
|
||||
### Error Categories
|
||||
1. **Device Detection**: Missing GPUs, driver issues
|
||||
2. **Memory Allocation**: OOM, fragmentation problems
|
||||
3. **Model Loading**: Corrupt files, missing dependencies
|
||||
4. **Integration**: ComfyUI API changes, extension conflicts
|
||||
Per-model flag over global sentinel
|
||||
- Granular control, no persistent global state
|
||||
- Isolated to each loaded model, matches ComfyUI lifecycle
|
||||
|
||||
## Future Technology Considerations
|
||||
Patched unload behavior (selective)
|
||||
- Maintain `kept_models` across the full free path
|
||||
- Only eject DisTorch2 models when explicitly requested via keep_loaded=False
|
||||
|
||||
### Next-Generation Hardware
|
||||
- **PCIe 5.0**: 63 GB/s bandwidth capability
|
||||
- **NVLink 4.0**: 112.5 GB/s for future GPUs
|
||||
- **DDR5**: Higher memory bandwidth for CPU offloading
|
||||
- **CXL Memory**: Unified memory pool architectures
|
||||
Patched soft empty (multi-device)
|
||||
- Ensure cache clearing is not limited to the single `mm.get_torch_device()` device
|
||||
- CPU memory behavior integrated with PromptExecutor.reset() semantics
|
||||
|
||||
### Emerging Platforms
|
||||
- **Intel Arc**: XPU support expanding
|
||||
- **AMD RDNA**: HIP/ROCm improvements
|
||||
- **ARM64**: Apple Silicon and server adoption
|
||||
- **Distributed**: Multi-node inference possibilities
|
||||
## Known Technical Work (Next)
|
||||
|
||||
### Model Architecture Evolution
|
||||
- **Mixture of Experts**: Sparse model support
|
||||
- **Multimodal**: Vision+Language combined models
|
||||
- **Streaming**: Real-time model serving requirements
|
||||
- **Quantization**: Advanced formats (FP4, INT8, block-wise)
|
||||
- Reinstate strict no-op in `_mgpu_patched_unload_all_models` when `models_to_unload` is empty (no delegation to original unload)
|
||||
- Add instrumentation and assertions to guarantee no unintended ejection of retained models after `/free` flow
|
||||
- Re-run verification matrix and capture logs in Memory Bank
|
||||
|
||||
Reference in New Issue
Block a user