Files
pollockjj-ComfyUI-MultiGPU/memory-bank/code-references.md
T

10 KiB
Raw Blame History

Code References (Definitive): ComfyUI Manager “Free model and node cache”

Purpose

  • Provide an end-to-end, fully verified lineage of the ComfyUI Manager “Free model and node cache” button through to the exact consumption of flags in ComfyUI core, with exact file paths and code excerpts captured from the current snapshot in this workspace.
  • Document MultiGPU patch integration points that participate in the free/unload flow, including selective unload behavior and current caveats.

End‑to‑End Flow (Current Snapshot)

  1. UI Button (Manager) → 2) JS helper free_models(...) → 3) POST /free (Comfy core) → 4) main.py prompt_worker thread polls flags and performs:
    • unload_models: comfy.model_management.unload_all_models()
    • free_memory: PromptExecutor.reset()
    • Additionally triggers GC and comfy.model_management.soft_empty_cache()

A) Frontend UI trigger (ComfyUI Manager)

  • File: ../ComfyUI-Manager/js/comfyui-manager.js
  • Location: app.registerExtension({ name: "Comfy.ManagerMenu", ... }) → setup() → ComfyButtonGroup
new(await import("../../scripts/ui/components/button.js")).ComfyButton({
  icon: "vacuum-outline",
  action: () => {
    free_models();
  },
  tooltip: "Unload Models"
}).element,
new(await import("../../scripts/ui/components/button.js")).ComfyButton({
  icon: "vacuum",
  action: () => {
    free_models(true);
  },
  tooltip: "Free model and node cache"
}).element,

Semantics:

  • “Unload Models” → free_models() (models only)
  • “Free model and node cache” → free_models(true) (models + execution cache)

B) Frontend request construction (ComfyUI Manager)

  • File: ../ComfyUI-Manager/js/common.js
  • Function: export async function free_models(free_execution_cache)
export async function free_models(free_execution_cache) {
  try {
    let mode = "";
    if (free_execution_cache) {
      mode = '{"unload_models": true, "free_memory": true}';
    } else {
      mode = '{"unload_models": true}';
    }

    console.log(`[ManagerFreePath] POST /free payload: ${mode}`);
    let res = await api.fetchApi(`/free`, {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: mode
    });
    console.log(`[ManagerFreePath] /free status: ${res.status}`);

    if (res.status == 200) {
      if (free_execution_cache) {
        showToast("'Models' and 'Execution Cache' have been cleared.", 3000);
      } else {
        showToast("Models' have been unloaded.", 3000);
      }
    } else {
      showToast('Unloading of models failed. Installed ComfyUI may be an outdated version.', 5000);
    }
  } catch (error) {
    console.error('[ManagerFreePath] /free error:', error);
    showToast('An error occurred while trying to unload models.', 5000);
  }
}

Semantics:

  • free_models(true) → POST /free with {"unload_models": true, "free_memory": true}
  • free_models() → POST /free with {"unload_models": true}

C) Core server endpoint (flags are set on the queue)

  • File: ../../server.py
  • Route: @routes.post("/free")
@routes.post("/free")
async def post_free(request):
    json_data = await request.json()
    unload_models = json_data.get("unload_models", False)
    free_memory = json_data.get("free_memory", False)
    if unload_models:
        self.prompt_queue.set_flag("unload_models", unload_models)
    if free_memory:
        self.prompt_queue.set_flag("free_memory", free_memory)
    return web.Response(status=200)

Semantics:

  • The HTTP endpoint itself does not unload/reset; instead it sets flags on PromptServer.prompt_queue for the background worker to consume.

D) Flag consumption and execution (definitive mechanism)

  • File: ../../main.py
  • Function: prompt_worker(q, server_instance)
  • Excerpt (poll and handle flags, then clean up):
        flags = q.get_flags()
        free_memory = flags.get("free_memory", False)

        if flags.get("unload_models", free_memory):
            comfy.model_management.unload_all_models()
            need_gc = True
            last_gc_collect = 0

        if free_memory:
            e.reset()
            need_gc = True
            last_gc_collect = 0

        if need_gc:
            current_time = time.perf_counter()
            if (current_time - last_gc_collect) > gc_collect_interval:
                gc.collect()
                comfy.model_management.soft_empty_cache()
                last_gc_collect = current_time
                need_gc = False
                hook_breaker_ac10a0.restore_functions()

Context:

  • e is a PromptExecutor (created earlier in prompt_worker): e = execution.PromptExecutor(server_instance, ...)
  • The worker thread is started in start_comfyui():
threading.Thread(target=prompt_worker, daemon=True, args=(prompt_server.prompt_queue, prompt_server,)).start()

Interpretation (What the Manager button actually does)

  • “Free model and node cache” sets unload_models: true and free_memory: true via POST /free.
  • The background prompt_worker then:
    • Calls comfy.model_management.unload_all_models()
    • Calls e.reset() on the PromptExecutor to drop execution caches
    • Performs gc.collect() and comfy.model_management.soft_empty_cache()
  • This matches the “benchmark button” behavior required for CPU memory reclamation (models fully unloaded + executor reset + allocator/cache cleanup).

Implications for MultiGPU P1 (force_full_system_cleanup)

  • To 100% replicate the benchmark button behavior from within MultiGPU code paths:
    • Call comfy.model_management.unload_all_models()
    • Trigger PromptExecutor.reset() on the active executor
    • Follow up with gc.collect() and comfy.model_management.soft_empty_cache()
  • Or, trigger the core behavior indirectly by POST /free with both flags set, relying on ComfyUI’s running prompt worker.

Verification Status

  • All file paths and snippets above were extracted from this workspace:
    • Manager JS files under ../ComfyUI-Manager/js/
    • ComfyUI server and main under ../../server.py and ../../main.py
  • Consumption site conclusively identified in ../../main.py prompt_worker via q.get_flags → unload_all_models + PromptExecutor.reset

MultiGPU Integration Points (This Repository)

Overview

  • In addition to the core /free flow, MultiGPU patches (in this repository) alter both the unload and soft-empty behaviors to enable selective ejection of DisTorch-managed models and multi-device cache clearing.
  1. Per-model transient flag (DisTorch2 nodes)
  • File: memory-bank reference → implemented in code at: ./distorch_2.py
  • Where:
    • In DisTorch2 wrappers (UNET/CLIP/VAE) inside override(...), after calling the original node:
      • out[0].model._mgpu_unload_distorch_model = (not keep_loaded)
  • Purpose:
    • Mark models for ejection only when the user disables “keep_loaded”.
    • This supplants the previously planned global sentinel; the implemented design is purely per-model.
  1. Selective unloading (patched unload_all_models)
  • File: ./model_management_mgpu.py
  • Patch site notes:
    • At import time, we patch mm.unload_all_models with _mgpu_patched_unload_all_models.
    • Behavior:
      • Iterate mm.current_loaded_models into:
        • models_to_unload: those with _mgpu_unload_distorch_model == True
        • kept_models: the rest
      • If any are flagged, unload only models_to_unload and rebuild mm.current_loaded_models = kept_models.
      • If none are flagged (all kept), current code delegates to original unload_all_models() (known caveat; see below).
  • Known caveat (to be fixed next):
    • The “all kept” branch currently delegates to the original unload, which unloads everything. Target behavior is strict no-op when no models are flagged.
  1. Multi-device VRAM cache and CPU reset (patched soft_empty_cache)
  • File: ./__init__.py
  • Patch site notes:
    • mm.soft_empty_cache → soft_empty_cache_distorch2_patched
    • Behavior:
      • Detect DisTorch2 active state; clear allocator caches on ALL devices via soft_empty_cache_multigpu() from device_utils.py
      • Adaptive CPU memory reset with optional force to emulate Manager “free_memory”.
    • This ensures cache clearing covers all devices in MultiGPU environments beyond a single mm.get_torch_device().
  1. Manager parity helper
  • File: ./model_management_mgpu.py
  • Function: force_full_system_cleanup(reason="manual", force=True)
    • Sets both flags (unload_models=True, free_memory=True) on PromptQueue, identical to Manager’s “Free model and node cache”.
    • Useful for testing and ensuring parity from MultiGPU paths.

Behavioral Summary

  • End-to-end Manager parity:
    • Manager “Free model and node cache” → POST /free sets flags → Comfy’s prompt_worker calls our patched unload_all_models (selective) → PromptExecutor.reset() → our patched soft_empty_cache (multi-device) → GC.
  • Selectiveness guarantee (intended):
    • Only DisTorch2 models flagged with _mgpu_unload_distorch_model=True are ejected.
    • Unflagged models (keep_loaded=True) remain in mm.current_loaded_models after the entire flow.
  • Current discrepancy:
    • When no models are flagged, our patch currently delegates to the original unload (unloads everything). Target fix is to convert this branch to a strict no-op.

Validation & Logging Hooks

  • Memory snapshots:
    • Use multigpu_memory_log(identifier, tag) in model_management_mgpu.py for timestamped CPU/VRAM snapshot lines.
  • VRAM cache clearing:
    • soft_empty_cache_multigpu() logs per-device clearing events (pre/post) in device_utils.py.
  • Unload path tracing:
    • _mgpu_patched_unload_all_models logs the counts of kept/unloaded models and updates to mm.current_loaded_models.

Practical Test Recipes

  1. Minimal retention test
  • Load A(keep=false), B(keep=true), C(keep=true)
  • POST /free payload: {"unload_models": true, "free_memory": true}
  • Expected:
    • Only A is ejected; B and C remain in mm.current_loaded_models post-flow.
    • CPU RAM drops; VRAM caches clear on all devices.
  1. All-kept test
  • Load D(keep=true), E(keep=true)
  • POST /free payload: {"unload_models": true, "free_memory": true}
  • Expected target behavior:
    • No models are ejected (strict no-op in unload step), allocator/cache cleaning only.
  • Current behavior (caveat):
    • Delegates to original unload → all models may be ejected. This is the next change to reinstate strict no-op.

References (paths in this repo)

  • Per-model flagging: ./distorch_2.py
  • Selective unload patch: ./model_management_mgpu.py
  • Patched soft empty: ./__init__.py (soft_empty_cache_distorch2_patched)
  • Multi-device cache clear: ./device_utils.py
  • Manager parity helper: ./model_management_mgpu.py (force_full_system_cleanup)