Files
xmarre-ComfyUI-GPU-Resident…/README.md

16 KiB
Raw Permalink Blame History

ComfyUI GPU Resident Loader

A ComfyUI custom-node pack for faster time-to-VRAM, selective safetensors loading, and sticky GPU residency control.

This repo does three related jobs:

  1. Installs startup-time monkey patches before any workflow nodes run.
  2. Ships KJ-style resident loader nodes for diffusion models and checkpoints.
  3. Maintains a live residency registry with preload / pin / evict / report controls for tracked objects.

It is not just a “clean RAM” addon. The main target is the path from model file -> tensors -> live ComfyUI object -> VRAM retention.

Why this exists

ComfyUI’s default behavior mixes together two separate concerns:

  • ingest path — where tensors are first materialized while a model is being read, and
  • residency policy — where the finished model tends to live afterwards.

Those are not the same problem.

This repo focuses on both:

  • For .safetensors, it tries to keep eligible loads on the narrowest, most GPU-friendly path it can.
  • For resident diffusion and checkpoint-model loads, it avoids broad checkpoint materialization by selecting only the detected UNet keys where possible.
  • For runtime VRAM pressure, it adds a sticky-priority registry and teaches ComfyUI’s unload path to protect higher-value resident entries until enough VRAM must be reclaimed.
  • For manual control, it exposes nodes that let you preload, pin, evict, and inspect tracked models, CLIPs, and VAEs.

What changes at startup

__init__.py calls startup.install_patches(), which applies the monkey patches exactly once when the custom node is imported.

Patched functions / methods

Current patch surface:

  • comfy.utils.load_torch_file
  • comfy.clip_vision.load_torch_file (redirected to the patched comfy.utils.load_torch_file when present)
  • comfy.model_management.free_memory
  • comfy.model_management.load_models_gpu
  • comfy.model_management.unet_offload_device
  • comfy.model_management.text_encoder_offload_device
  • comfy.model_management.vae_offload_device
  • comfy.model_management.text_encoder_device
  • comfy.model_management.vae_device
  • comfy.model_management.unet_inital_load_device
  • comfy.model_management.LoadedModel.model_unload
  • comfy.model_patcher.ModelPatcher.detach
  • comfy.sd.load_checkpoint_guess_config
  • comfy.sd.load_diffusion_model
  • comfy.sd.load_clip
  • comfy.sd.VAE.encode
  • comfy.sd.VAE.decode
  • comfy.clip_vision.load
  • comfy.controlnet.load_controlnet
  • comfy.diffusers_load.load_diffusers

What those patches do

1) load_torch_file becomes residency-aware

The patched loader:

  • detects the active load context (model, clip, vae, checkpoint, etc.),
  • picks an explicit GPU target device when the active policy wants GPU ingest,
  • attempts direct safetensors reads on the requested device,
  • falls back to CPU read + tensor-by-tensor copy if direct GPU safetensors loading fails,
  • still uses CPU-first torch.load() for pickle formats (.ckpt, .pt, .pth, .bin),
  • records the actual load method in the residency registry.

For safetensors loads happening inside a model / clip / vae context, it can select only the detected component keys from the file header instead of pulling the full file into memory first.

2) Loader contexts are attached to stock ComfyUI load paths

These stock paths are wrapped with registry context and output binding:

  • checkpoint loads,
  • diffusion-model loads,
  • CLIP loads,
  • CLIP Vision loads,
  • ControlNet loads,
  • diffusers loads.

That means the registry is not limited to the custom resident nodes. Stock ComfyUI loaders that pass through these paths are also tracked.

3) Device/offload policy is overridden

Depending on the active policy, the patcher can steer:

  • initial UNet load device,
  • CLIP/Text Encoder device,
  • VAE device,
  • offload devices for UNet / CLIP / VAE.

This is how prefer_gpu and sticky_gpu keep more of the hot path on the GPU side than stock ComfyUI would.

4) free_memory() becomes sticky-aware

Under sticky_gpu, comfy.model_management.free_memory() is patched so that:

  • sticky tracked wrappers are considered first,
  • higher-priority sticky entries are protected first,
  • lower-priority or older sticky entries yield first when VRAM must be reclaimed,
  • a transient protection floor is applied so ComfyUI does not immediately tear down high-value resident entries for small requests.

5) Clone replacement is hardened

load_models_gpu() is patched to fully unload clone-conflict wrappers before replacement instead of relying on a shallow detach path that can leave base weights patched.

6) Unload / detach is redirected to CPU when needed

LoadedModel.model_unload() and ModelPatcher.detach() are patched so that unloads which would otherwise not reclaim VRAM are redirected through a CPU offload target first.

7) VAE encode/decode gets a sticky-safe path

Under sticky_gpu, patched VAE.encode() and VAE.decode():

  • cap the working batch count when necessary to preserve transient VRAM headroom, and
  • retry with tiled VAE encode/decode on OOM.

That behavior is not a general performance feature toggle. It exists to reduce avoidable VRAM spikes while sticky residency is active.

Policies

The registry exposes four global policies:

legacy

Stay closest to stock ComfyUI behavior. Registry tracking still exists, but the patcher does not aggressively steer ingest/offload toward the GPU path.

balanced

Keep registry tracking and diagnostics without aggressive GPU residency behavior.

prefer_gpu

Prefer GPU ingest for tracked model-like loads and keep the faster side of the device/offload policy for:

  • diffusion models,
  • CLIP / text encoders,
  • ControlNets.

This policy does not auto-pin tracked objects.

sticky_gpu

Builds on prefer_gpu and additionally:

  • auto-pins newly bound models and CLIPs,
  • keeps VAE offload on the GPU side as well,
  • patches free_memory() to protect sticky tracked wrappers by priority,
  • uses the sticky-safe VAE encode/decode behavior.

Default policy selection

Selection order is:

  1. COMFYUI_GPU_RESIDENT_POLICY, if set to a supported value.
  2. sticky_gpu when ComfyUI is started with --gpu_only.
  3. sticky_gpu when ComfyUI is started with --highvram.
  4. otherwise prefer_gpu.

Supported values are:

  • legacy
  • balanced
  • prefer_gpu
  • sticky_gpu

Included nodes

All nodes live under the GPU Resident Loader category.

Loader nodes

Diffusion Model Selector Resident

Returns an absolute path string for a selected diffusion model.

Notes:

  • resolves from diffusion_models, and
  • also exposes text_encoders entries whose filename contains connector.

Diffusion Model Loader Resident

KJ-style diffusion-model loader with these controls:

  • weight_dtype
  • compute_dtype
  • patch_cublaslinear
  • sage_attention
  • enable_fp16_accumulation
  • optional extra_state_dict
  • optional policy_override

Behavior:

  • for .safetensors, it loads only the detected UNet portion of the file,
  • if extra_state_dict is provided, only matching UNet keys are merged,
  • repeated loads reuse a live equivalent model when the source path and loader-relevant options still match,
  • before GPU-bound loads, it estimates the upcoming footprint and trims only enough lower-priority residency to cover the request plus adaptive headroom.

Checkpoint Loader Resident

Full checkpoint loader that returns:

  • MODEL
  • CLIP
  • VAE

Behavior:

  • shares the same tuning knobs as the resident diffusion-model loader for the model component,
  • reuses already-live equivalent components when possible,
  • composes the final output from model / clip / vae component loaders instead of always rebuilding the whole checkpoint path from scratch.

Checkpoint Model Loader Resident

Model-only checkpoint loader.

Behavior:

  • takes the same selective safetensors UNet fast path as the diffusion-model loader,
  • reuses a live equivalent model when available,
  • uses the same dtype / attention / cublas / fp16-accumulation knobs as the full checkpoint loader.

Checkpoint Clip Loader Resident

CLIP-only checkpoint loader.

Behavior:

  • can reuse a live equivalent CLIP object,
  • avoids rebuilding the diffusion model and VAE outputs when only CLIP is needed.

Checkpoint VAE Loader Resident

VAE-only checkpoint loader.

Behavior:

  • can reuse a live equivalent VAE object,
  • avoids rebuilding the diffusion model and CLIP outputs when only VAE is needed.

Residency nodes

Set Global Residency Policy

Sets the active global policy and returns it as a STRING.

The loader nodes also expose an optional policy_override string input for one-off loads.

Registry Snapshot

Returns the whole registry as formatted JSON.

Pin Model Residency / Pin CLIP Residency / Pin VAE Residency

Marks a tracked object as sticky or non-sticky and optionally changes its priority.

Preload Model To GPU / Preload CLIP To GPU / Preload VAE To GPU

Calls load_models_gpu(..., force_full_load=True) for the selected object, then updates sticky state / priority in the registry.

Evict Model From GPU / Evict CLIP From GPU / Evict VAE From GPU

Attempts to unload the selected object from the current loaded-model set.

unpatch_weights=True performs a full unload path. When eviction succeeds, the node returns evicted; otherwise not_loaded.

Report Model Residency / Report CLIP Residency / Report VAE Residency

Returns a JSON report for a single tracked object.

If the object is not currently bound in the registry, the node returns a JSON payload with tracked: false.

Adaptive trimming before resident loads

The resident loaders now do load-scoped VRAM trimming themselves.

Before a GPU-bound resident load, the loader estimates required bytes from:

  • the safetensors header when possible,
  • the detected checkpoint component subset when possible,
  • otherwise the source file size as a fallback.

It then requests enough free VRAM for:

  • the estimated load size, plus
  • adaptive headroom.

Current adaptive headroom policy:

  • ratio: 12.5% of the estimated load,
  • floor: 256 MiB,
  • ceiling: 1 GiB.

The trim path prefers to:

  • unload non-sticky entries first,
  • then lower-priority sticky entries,
  • preserve explicitly kept models,
  • use partial unload where available.

This logic lives in the resident loader path. You do not need a separate “target free VRAM” node for it.

Registry and observability

The registry tracks residency metadata for bound objects.

Typical per-entry fields include:

  • entry_id
  • kind
  • source_path
  • basename
  • sticky
  • priority
  • created_at
  • last_touched
  • loaded_bytes
  • total_bytes
  • load_device
  • offload_device
  • current_device
  • last_method
  • last_report
  • loader_key
  • notes
  • alive

The last_method / last_report fields let you see whether a load actually used:

  • direct safetensors GPU ingest,
  • safetensors CPU -> CUDA fallback,
  • safetensors component-only load,
  • CPU-first torch.load() compatibility path,
  • or a recorded load failure.

What gets tracked

Tracked/bound paths include:

  • resident node loads from this repo,
  • stock checkpoint loads,
  • stock diffusion-model loads,
  • stock CLIP loads,
  • stock CLIP Vision loads,
  • stock diffusers loads.

ControlNet loads also participate in the patched load context and device-policy path, but this repo does not currently expose dedicated ControlNet residency nodes.

Important limits and non-goals

Best path is still .safetensors

The narrow fast path is built around .safetensors.

That is where this repo can:

  • inspect headers cheaply,
  • select only model / clip / vae subsets,
  • estimate component bytes more accurately,
  • attempt direct device-targeted reads.

.ckpt / .pt / pickle formats are still CPU-first

For pickle-based formats, PyTorch still goes through torch.load() on CPU first.

The repo can still:

  • track those loads,
  • keep the resulting live objects resident,
  • reuse equivalent live objects later.

It does not claim direct-to-GPU ingest for those formats.

Cross-process persistence is out of scope

This repo does not keep VRAM allocations alive after ComfyUI, Python, or WSL exits.

CUDA memory lifetime is process/context scoped. True persistence across process shutdown would need a separate long-lived keeper process or service that owns the CUDA context.

It does not automatically capture arbitrary custom loader implementations

The registry only sees objects that pass through the patched ComfyUI load paths or through this repo’s resident nodes.

If another custom node loads models through its own private code path and bypasses those patched entry points, that object may never become a tracked registry entry. In that case, the preload / pin / evict / report nodes from this repo cannot manage it until that external loader is integrated or patched.

Installation

Clone into custom_nodes:

git clone https://github.com/xmarre/ComfyUI-GPU-Resident-Loader ComfyUI/custom_nodes/ComfyUI-GPU-Resident-Loader

Install dependencies into the same Python environment ComfyUI uses:

pip install -r ComfyUI/custom_nodes/ComfyUI-GPU-Resident-Loader/requirements.txt

Requirements declared by the repo:

  • Python >=3.10
  • safetensors>=0.4.3

Optional SageAttention dependencies are not installed by default. Install those separately if you plan to use a SageAttention mode in the resident loaders.

Basic usage patterns

1) Large-VRAM, mostly resident workflow

Recommended baseline:

  • start ComfyUI with --highvram or set policy manually to sticky_gpu,
  • prefer .safetensors for hot models,
  • load diffusion models through Diffusion Model Loader Resident,
  • use Preload ... To GPU for models you know you will reuse,
  • inspect with Report ... Residency or Registry Snapshot.

2) Full checkpoint workflow

Use Checkpoint Loader Resident when you want MODEL + CLIP + VAE together.

That path can reuse already-live components instead of always rebuilding all three outputs.

3) Staged checkpoint workflow

Use component loaders when the graph does not need the whole checkpoint at once:

  • Checkpoint Model Loader Resident for diffusion model only,
  • Checkpoint Clip Loader Resident for CLIP only,
  • Checkpoint VAE Loader Resident for VAE only.

4) Manual residency control

Use:

  • Pin ... Residency to mark a tracked entry sticky / non-sticky,
  • Preload ... To GPU to force a full live load now,
  • Evict ... From GPU to unload it from the current loaded-model set.

Notes on compatibility and migration

Legacy wiring: extra_state_dict used as a policy string

The resident diffusion-model loader contains a compatibility shim for older graphs:

  • if extra_state_dict receives one of the known policy names,
  • and that value is not an existing file path,
  • it is interpreted as policy_override instead.

New graphs should connect policy strings to policy_override, not to extra_state_dict.

Convert hot pickle checkpoints to safetensors

scripts/convert_checkpoint_to_safetensors.py is included for one-time conversion of hot .ckpt / .pt / .pth style checkpoints.

Example:

python ComfyUI/custom_nodes/ComfyUI-GPU-Resident-Loader/scripts/convert_checkpoint_to_safetensors.py \
  --input /path/to/model.ckpt \
  --output /path/to/model.safetensors

Optional flags:

  • --state-dict-key <key> to extract a different top-level dict key
  • --allow-non-tensor-values to skip non-tensor entries instead of failing

License

GPL-3.0-or-later.

This repo stays GPL-compatible because it adapts behavior from GPL-licensed ComfyUI and mirrors relevant loader behavior from GPL-3.0-licensed KJNodes.