Commit Graph

  • fc2a732419 prepare for final release candidate John Pollock 2025-09-30 09:15:32 -05:00
  • 9a526e2546 prepare for final release candidate John Pollock 2025-09-30 09:10:37 -05:00
  • 23ed34df1b prepare for final release candidate John Pollock 2025-09-30 09:08:15 -05:00
  • 429be7c912 docs: update activeContext for v2.5.0 release with refactoring summary John Pollock 2025-09-30 09:07:05 -05:00
  • e7d8113a86 refactored in to one analyze_safetensor_loading John Pollock 2025-09-30 08:44:16 -05:00
  • 8b8a16e982 Major architectural refactor: Consolidate wrappers, fix CheckpointLoader bug, improve separation of concerns (-531 lines) John Pollock 2025-09-30 08:13:21 -05:00
  • 07b429f3f9 fix(distorch): Add GC anchor protection for selective model retention John Pollock 2025-09-29 16:01:24 -05:00
  • bde51c6236 docs(distorch): document GC anchor solution for selective unload John Pollock 2025-09-29 15:45:46 -05:00
  • c23dc083d3 WIP John Pollock 2025-09-29 14:04:05 -05:00
  • 6e1f9671f2 docs(memory): update Phase 3 implementation status and clarify bug John Pollock 2025-09-29 13:21:57 -05:00
  • 01df082651 docs(memory-bank): sync with current code state for selective ejection\n\n- Document Phase 3 implemented without global sentinel (_mgpu_unload_distorch_model per-model flag)\n- Describe patched unload_all_models selective behavior and current all-kept delegation caveat\n- Outline rediscovery plan and strict no-op target when no models are flagged\n- Update active context, system patterns, code references, progress, tech context, and lineage John Pollock 2025-09-29 07:57:35 -05:00
  • 1ca3daf0d8 At least now the logs reflect it is now trying to do what I know we have figured out how to do in the past in one of these commits. . . John Pollock 2025-09-29 07:08:03 -05:00
  • d61ca7b06f back to setting full reset flag if there is a distorch unload pending. John Pollock 2025-09-29 05:27:53 -05:00
  • ede0957f65 feat: add caching and logging to IS_CHANGED methods in safetensor overrides John Pollock 2025-09-29 04:46:07 -05:00
  • c4ae5e9e08 extensive clean-up, WIP John Pollock 2025-09-29 03:53:21 -05:00
  • 23d2abe237 fixed incorrect info John Pollock 2025-09-28 19:19:20 -05:00
  • 8591063a3c incremental progress (I think, hard to tell) John Pollock 2025-09-28 19:14:47 -05:00
  • 18493f5277 refactor: simplify model retention logic in multi-GPU unload John Pollock 2025-09-28 13:22:19 -05:00
  • 0d056141c0 an interesting experiment that produces wrong behavior but no OOM. Looks like we are circling it and I don't want to lose this intermediate step. John Pollock 2025-09-28 12:15:52 -05:00
  • ae8bb7cf2c feat: Refine model retention logic in multi-GPU unloading John Pollock 2025-09-28 11:14:34 -05:00
  • fda5d6ed00 commiting this steaming pile of hot garbage for future dissection to see if I want any organs from this terminally ill branch John Pollock 2025-09-25 14:36:13 -05:00
  • bd672479fa refactor: eliminate circular import by separating model management functions John Pollock 2025-09-24 17:38:15 -05:00
  • ff6efb4217 Scorched Earth, but it works. John Pollock 2025-09-24 13:55:31 -05:00
  • 7b319544e0 feat: implement comprehensive memory management and OOM prevention John Pollock 2025-09-23 22:52:20 -05:00
  • cd7a536645 docs: add development rules and project context in .clinerules John Pollock 2025-09-23 20:58:03 -05:00
  • 3121b2f70c feat(mgpu): scoped MM logger; parse compute device/VRAM plan John Pollock 2025-09-23 04:41:44 -05:00
  • a0fe72e290 Additonal refinements to DisTorch2 cache/unload to avoid OOM. Needs at least one more clean-up pass. John Pollock 2025-09-21 09:30:03 -05:00
  • 55a0d22b01 refactor: simplify memory logging in checkpoint loading John Pollock 2025-09-21 06:12:20 -05:00
  • 63ff1a4064 committing so we don't lose verbose logging. John Pollock 2025-09-20 11:58:29 -05:00
  • 8e4c7fed14 Potential improvement - committing for additional testing John Pollock 2025-09-20 07:08:45 -05:00
  • 57c7d3da8e Merge pull request #109 from pollockjj/low_vram_clip John Pollock 2025-09-15 13:22:40 -05:00
  • f7942dca93 Fix for (#104): drop text_encoder_initial_device patch and state - these were part of an attempt to solve a CLIP compute issue that was recently solved another way (Commit edc8a4d) John Pollock 2025-09-15 12:46:02 -05:00
  • 80f8a14dea Fix non-deterministic behavior of CLIP compute device when ~100% offloading. John Pollock 2025-09-14 14:59:01 -05:00
  • aa00a682d0 feat: add model inspection utilities for tracking and analysis John Pollock 2025-09-14 00:22:28 -05:00
  • edc8a4dd2b Identified a long-standing bug where fully-allocated CLIP (for example 99G of VirtualVRAM = 100% of major blocks no matter the model) proceeded to execute on the donor device (e.g. cpu) instead of the indicated compute device. Turns out, it only happens when *all* blocks are identified to go onto the donor card. In the case of the donor being the cpu this was irritatingly slow. John Pollock 2025-09-14 00:07:47 -05:00
  • d34a32f097 Fix for Triple/Quad Clip Loaders (#99) John Pollock 2025-09-12 23:44:20 -05:00
  • afafc8042d Add no-device variants for multi-GPU CLIP loaders John Pollock 2025-09-12 22:39:50 -05:00
  • 5bb7add514 Remove unused 'device' parameter from CLIP loader methods John Pollock 2025-09-12 22:18:26 -05:00
  • fabbc9e7be Merge branch 'lora_reapply_fix' John Pollock 2025-09-10 19:44:08 -05:00
  • 5b62671f0c roll back aggresive memory management John Pollock 2025-09-10 19:17:05 -05:00
  • e9fb4a8c2f Hot FixL: Revert aggresive memory management until a more targeted approach can be developed. This was causing OOMs on models that should load normally using the normal loader. John Pollock 2025-09-10 18:24:12 -05:00
  • be9cc21d4d preliminary changes John Pollock 2025-09-10 14:55:02 -05:00
  • e1635e9996 Improve memory handling for safetensor models in corner cases John Pollock 2025-09-09 12:00:56 -05:00
  • 803cf542d9 MultiGPU garbage collection/cache clearing and DisTorch2 Clip device node bug John Pollock 2025-09-08 23:10:54 -05:00
  • c63b539f1e Additional garbage/cache collection (#101) addressed DisTorch2 Device issue for CLIP hopefully closing (#99,#104) John Pollock 2025-09-08 23:06:21 -05:00
  • 0adf219f60 Hot fix for (https://github.com/pollockjj/ComfyUI-MultiGPU/issues/99). It might not be 100% but will prevent error and I will revisit to ensure John Pollock 2025-09-02 07:50:46 -05:00
  • 54b7c5b0e6 Advanced Checkpoint and Advanced DisTorch2 Checkpoint loaders (https://github.com/pollockjj/ComfyUI-MultiGPU/issues/95), Fix CLiP loading device when MultGPU or DisTorch2 is invoked. John Pollock 2025-08-31 01:45:13 -05:00
  • 0b1511edee refactor: Simplify checkpoint loading and fix text encoder device John Pollock 2025-08-31 01:00:53 -05:00
  • 9e14e4622c fix: Overhaul checkpoint loader for proper device handling John Pollock 2025-08-30 19:51:15 -05:00
  • f07c2d2b89 feat: Add advanced checkpoint loaders for MultiGPU and DisTorch2 John Pollock 2025-08-30 19:19:10 -05:00
  • 4d0d4a673f fix for issue https://github.com/pollockjj/ComfyUI-MultiGPU/issues/87: ComfyU-MultiGPU not supporting all device types currently supported by Comfy Core. John Pollock 2025-08-30 07:39:26 -05:00
  • 06bc2c3ac8 Fix for https://github.com/pollockjj/ComfyUI-MultiGPU/issues/96 John Pollock 2025-08-29 23:18:14 -05:00
  • 5127e1807e sync changes to kijai's nodes John Pollock 2025-08-29 21:11:11 -05:00
  • 336e236105 Fix for wanvideo bug and add new example John Pollock 2025-08-29 20:50:55 -05:00
  • 0ca771fe68 Fix for : https://github.com/pollockjj/ComfyUI-MultiGPU/issues/93 John Pollock 2025-08-28 20:47:12 -05:00
  • 2481b52064 Refactor: Improve block allocation and expert string parsing John Pollock 2025-08-26 16:43:04 -05:00
  • a5ff7fd878 Fixed one of two glaring bugs introduced by recent "improvements" John Pollock 2025-08-26 15:21:35 -05:00
  • 6f5c4aa901 Preparing for 2.2.0 release (byte and ratio model allocation schemes) John Pollock 2025-08-26 14:42:25 -05:00
  • baa31a1961 refactor(distorch): Default unassigned tensor blocks to CPU John Pollock 2025-08-26 14:08:30 -05:00
  • 40cccdf01d Refactor: Improve DisTorch2 allocation logic and robustness John Pollock 2025-08-26 11:13:38 -05:00
  • 6c2a3d5b15 feat(distorch): Improve device discovery and add CPU fallback John Pollock 2025-08-26 09:01:03 -05:00
  • 5643a616e5 docs: Clarify CPU is default wildcard in Expert Mode John Pollock 2025-08-26 07:36:47 -05:00
  • 4fd2456ed5 docs: Add 'bytes' and 'ratio' expert modes to README John Pollock 2025-08-26 07:28:58 -05:00
  • 56b8dd233e feat: Add byte-based model allocation mode John Pollock 2025-08-26 06:45:55 -05:00
  • c58ffaeb05 introduces calculate_fraction_from_ratio_expert_string to correctly handle the 'ratio' allocation mode. This function translates a user-provided model-split ratio (e.g., '75% on GPU, 25% on CPU') into the device VRAM fractions required by the internal allocation system, aligning the feature's behavior with user expectations. John Pollock 2025-08-26 01:29:24 -05:00
  • b351c5dbc3 feat: Detect and log VRAM allocation mode John Pollock 2025-08-26 00:14:37 -05:00
  • 07df43b863 reverting disaster commit adding back in as an altenate file for reference to at least attempt salvage of what I was attempting to build before it getting butchered by incapable assistants. John Pollock 2025-08-25 23:26:19 -05:00
  • f88a2fca8d parking this total piece of garbage. John Pollock 2025-08-25 22:23:20 -05:00
  • f656195653 feat: Add expert implementation of distorch_2 John Pollock 2025-08-25 21:03:07 -05:00
  • b6d3403b71 Add memory parsing and flexible allocation support for safetensor loading John Pollock 2025-08-25 19:57:55 -05:00
  • 47ed1bed69 Reference file no longer needed. John Pollock 2025-08-24 06:05:18 -05:00
  • de00faaa3d Fixes for DisTorch V2 LoRA loading as well as sticky allocations when using standard loader John Pollock 2025-08-24 06:01:07 -05:00
  • 842ee650ed Optimize DisTorchV2 loader and FP8 casting logic John Pollock 2025-08-24 05:58:07 -05:00
  • afd8fecd94 Refactor DisTorch model patching logic for improved device assignment and FP8 casting John Pollock 2025-08-24 05:37:04 -05:00
  • 4367c892d8 Eliminate unused safetensor loading analysis method and update example configurations, adding one with LoRAs as one of the tested configurations to avoid the issue seen during initial release. John Pollock 2025-08-24 04:29:32 -05:00
  • e28b040cda Refactor DisTorchV2 loader to support both on-device (to avoid tensor mis-match on some models, but much slower patching) and on-compute (faster, highest fidelity for the combination of [fp8 model/LoRAs/store-on-CPU]) John Pollock 2025-08-24 04:13:01 -05:00
  • dfe6612880 Refactor safetensor loading logic and standardize logging John Pollock 2025-08-24 01:20:13 -05:00
  • 2331710c50 Enhance partially_load with fallback and reduced logging John Pollock 2025-08-23 14:31:32 -05:00
  • 543a0dc1eb patching logic from model_patcher load John Pollock 2025-08-23 10:06:54 -05:00
  • 956bd3bfa0 Enhance partially_load with weight unpatching and static assignments John Pollock 2025-08-23 09:09:31 -05:00
  • 240acae8c5 Simplify safetensor loading analysis and device assignment (from lowvram branch) John Pollock 2025-08-23 01:48:02 -05:00
  • 6195ed24c6 Refactor memory analysis to use ComfyUI's _load_list method (from lowvram_fix branch) John Pollock 2025-08-23 01:13:06 -05:00
  • 299c087a84 Pulling in deciding block allocation based on Comfy's own model_patcher._load_list().sort(reverse=True) for offload suitibility John Pollock 2025-08-22 21:57:07 -05:00
  • db697f1ccb Updating allocation logic based on exact placement and not the DistorchV1 methodology of CPU overrun. From lowvram_fix branch. John Pollock 2025-08-22 21:50:09 -05:00
  • d205f4da9e Adding improvements/updates to override_class_with_distorch_safetensor_v2 from previous partially_load development branch John Pollock 2025-08-22 21:45:35 -05:00
  • d0c4cd26fb Sync with main from last branch John Pollock 2025-08-22 21:34:34 -05:00
  • 24510c34ef Pulling in the work on new_load as reference for partially_load implementation John Pollock 2025-08-22 21:24:08 -05:00
  • 6e4181a7bb Refactor: Remove debugging and memory audit utilities John Pollock 2025-08-15 08:25:18 -05:00
  • ddd159ef23 docs: Clarify GGUF performance gain comparison in README John Pollock 2025-08-15 06:13:34 -05:00
  • 9838f2cc04 Fixed some confusing text John Pollock 2025-08-15 05:19:15 -05:00
  • 148f74c503 Update documentation to reflect DisTorch V2 John Pollock 2025-08-15 05:07:00 -05:00
  • 291a4a4572 feat: Add support for Apple MPS devices John Pollock 2025-08-14 12:58:34 -05:00
  • 545da7f741 Refactor: Reorganize and update example workflows John Pollock 2025-08-14 12:45:15 -05:00
  • d1c88a7cdb feat(distorch): Add universal .safetensors support & memory-based distribution John Pollock 2025-08-14 08:17:15 -05:00
  • fb6e2e6ffa refactor(distorch): Implement IS_CHANGED for robust model reloading John Pollock 2025-08-13 16:27:02 -05:00
  • e288152dae refactor: Introduce DisTorch V2 architecture John Pollock 2025-08-13 13:37:23 -05:00
  • d5dc678c04 Add FLUX support with new safetensor v2 implementation John Pollock 2025-08-12 20:16:29 -05:00
  • 298b4b829b Parking code. A new tack is needed. John Pollock 2025-08-12 12:06:19 -05:00
  • 235cd267bf feat(swap): Add shell-based block swapping for WanVideo models John Pollock 2025-08-11 23:41:51 -05:00
  • 6cb71ac51c feat(swap): Add block swap support for Qwen models John Pollock 2025-08-11 20:45:25 -05:00