6e1f9671f2
docs(memory): update Phase 3 implementation status and clarify bug
John Pollock2025-09-29 13:21:57 -05:00
01df082651
docs(memory-bank): sync with current code state for selective ejection\n\n- Document Phase 3 implemented without global sentinel (_mgpu_unload_distorch_model per-model flag)\n- Describe patched unload_all_models selective behavior and current all-kept delegation caveat\n- Outline rediscovery plan and strict no-op target when no models are flagged\n- Update active context, system patterns, code references, progress, tech context, and lineage
John Pollock2025-09-29 07:57:35 -05:00
1ca3daf0d8
At least now the logs reflect it is now trying to do what I know we have figured out how to do in the past in one of these commits. . .
John Pollock2025-09-29 07:08:03 -05:00
d61ca7b06f
back to setting full reset flag if there is a distorch unload pending.
John Pollock2025-09-29 05:27:53 -05:00
ede0957f65
feat: add caching and logging to IS_CHANGED methods in safetensor overrides
John Pollock2025-09-29 04:46:07 -05:00
8591063a3c
incremental progress (I think, hard to tell)
John Pollock2025-09-28 19:14:47 -05:00
18493f5277
refactor: simplify model retention logic in multi-GPU unload
John Pollock2025-09-28 13:22:19 -05:00
0d056141c0
an interesting experiment that produces wrong behavior but no OOM. Looks like we are circling it and I don't want to lose this intermediate step.
John Pollock2025-09-28 12:15:52 -05:00
ae8bb7cf2c
feat: Refine model retention logic in multi-GPU unloading
John Pollock2025-09-28 11:14:34 -05:00
fda5d6ed00
commiting this steaming pile of hot garbage for future dissection to see if I want any organs from this terminally ill branch
John Pollock2025-09-25 14:36:13 -05:00
bd672479fa
refactor: eliminate circular import by separating model management functions
John Pollock2025-09-24 17:38:15 -05:00
f7942dca93
Fix for (#104): drop text_encoder_initial_device patch and state - these were part of an attempt to solve a CLIP compute issue that was recently solved another way (Commit edc8a4d)
John Pollock2025-09-15 12:46:02 -05:00
80f8a14dea
Fix non-deterministic behavior of CLIP compute device when ~100% offloading.
John Pollock2025-09-14 14:59:01 -05:00
aa00a682d0
feat: add model inspection utilities for tracking and analysis
John Pollock2025-09-14 00:22:28 -05:00
edc8a4dd2b
Identified a long-standing bug where fully-allocated CLIP (for example 99G of VirtualVRAM = 100% of major blocks no matter the model) proceeded to execute on the donor device (e.g. cpu) instead of the indicated compute device. Turns out, it only happens when *all* blocks are identified to go onto the donor card. In the case of the donor being the cpu this was irritatingly slow.
John Pollock2025-09-14 00:07:47 -05:00
e9fb4a8c2f
Hot FixL: Revert aggresive memory management until a more targeted approach can be developed. This was causing OOMs on models that should load normally using the normal loader.
John Pollock2025-09-10 18:24:12 -05:00
2481b52064
Refactor: Improve block allocation and expert string parsing
John Pollock2025-08-26 16:43:04 -05:00
a5ff7fd878
Fixed one of two glaring bugs introduced by recent "improvements"
John Pollock2025-08-26 15:21:35 -05:00
6f5c4aa901
Preparing for 2.2.0 release (byte and ratio model allocation schemes)
John Pollock2025-08-26 14:42:25 -05:00
baa31a1961
refactor(distorch): Default unassigned tensor blocks to CPU
John Pollock2025-08-26 14:08:30 -05:00
40cccdf01d
Refactor: Improve DisTorch2 allocation logic and robustness
John Pollock2025-08-26 11:13:38 -05:00
6c2a3d5b15
feat(distorch): Improve device discovery and add CPU fallback
John Pollock2025-08-26 09:01:03 -05:00
5643a616e5
docs: Clarify CPU is default wildcard in Expert Mode
John Pollock2025-08-26 07:36:47 -05:00
4fd2456ed5
docs: Add 'bytes' and 'ratio' expert modes to README
John Pollock2025-08-26 07:28:58 -05:00
56b8dd233e
feat: Add byte-based model allocation mode
John Pollock2025-08-26 06:45:55 -05:00
c58ffaeb05
introduces calculate_fraction_from_ratio_expert_string to correctly handle the 'ratio' allocation mode. This function translates a user-provided model-split ratio (e.g., '75% on GPU, 25% on CPU') into the device VRAM fractions required by the internal allocation system, aligning the feature's behavior with user expectations.
John Pollock2025-08-26 01:29:24 -05:00
b351c5dbc3
feat: Detect and log VRAM allocation mode
John Pollock2025-08-26 00:14:37 -05:00
07df43b863
reverting disaster commit adding back in as an altenate file for reference to at least attempt salvage of what I was attempting to build before it getting butchered by incapable assistants.
John Pollock2025-08-25 23:26:19 -05:00
de00faaa3d
Fixes for DisTorch V2 LoRA loading as well as sticky allocations when using standard loader
John Pollock2025-08-24 06:01:07 -05:00
842ee650ed
Optimize DisTorchV2 loader and FP8 casting logic
John Pollock2025-08-24 05:58:07 -05:00
afd8fecd94
Refactor DisTorch model patching logic for improved device assignment and FP8 casting
John Pollock2025-08-24 05:37:04 -05:00
4367c892d8
Eliminate unused safetensor loading analysis method and update example configurations, adding one with LoRAs as one of the tested configurations to avoid the issue seen during initial release.
John Pollock2025-08-24 04:29:32 -05:00
e28b040cda
Refactor DisTorchV2 loader to support both on-device (to avoid tensor mis-match on some models, but much slower patching) and on-compute (faster, highest fidelity for the combination of [fp8 model/LoRAs/store-on-CPU])
John Pollock2025-08-24 04:13:01 -05:00
dfe6612880
Refactor safetensor loading logic and standardize logging
John Pollock2025-08-24 01:20:13 -05:00
2331710c50
Enhance partially_load with fallback and reduced logging
John Pollock2025-08-23 14:31:32 -05:00
956bd3bfa0
Enhance partially_load with weight unpatching and static assignments
John Pollock2025-08-23 09:09:31 -05:00
240acae8c5
Simplify safetensor loading analysis and device assignment (from lowvram branch)
John Pollock2025-08-23 01:48:02 -05:00
6195ed24c6
Refactor memory analysis to use ComfyUI's _load_list method (from lowvram_fix branch)
John Pollock2025-08-23 01:13:06 -05:00
299c087a84
Pulling in deciding block allocation based on Comfy's own model_patcher._load_list().sort(reverse=True) for offload suitibility
John Pollock2025-08-22 21:57:07 -05:00
db697f1ccb
Updating allocation logic based on exact placement and not the DistorchV1 methodology of CPU overrun. From lowvram_fix branch.
John Pollock2025-08-22 21:50:09 -05:00
d205f4da9e
Adding improvements/updates to override_class_with_distorch_safetensor_v2 from previous partially_load development branch
John Pollock2025-08-22 21:45:35 -05:00