14 Commits
Author SHA1 Message Date
John Pollock 0b438f7a8b Refactor code for improved readability and performance
- Cleaned up unnecessary whitespace and comments in model_management_mgpu.py, nodes.py, wanvideo.py, and wrappers.py for better code clarity.
- Replaced list comprehensions with direct list conversions in nodes.py for efficiency.
- Updated memory logging format in model_management_mgpu.py to streamline data capture.
- Enhanced device management in wanvideo.py by ensuring consistent device setting and loading.
- Added linting configurations in pyproject.toml to enforce code quality standards.
- Removed unused imports and optimized existing ones across multiple files.
2026-03-06 05:59:12 -06:00
John Pollock 1951513e92 refactor: migrate DisTorch2 allocation tracking to per-model metadata
- Replace global safetensor_allocation_store/safetensor_settings_store and create_safetensor_model_hash
  with a per-model annotation (_distorch_v2_meta) stored directly on the inner model object.
- Update distorch_2 to remove global stores and hash creation; parse and consume allocation strings
  from inner_model._distorch_v2_meta during model registration and loading.
- Update wrappers, checkpoint_multigpu, device_utils, and __init__ to set and read the new metadata
  instead of writing/reading global stores.
- Simplify detection of DisTorch-managed models (check inner_model._distorch_v2_meta) and adjust
  logging to surface inner model ids and allocation info.
- Clean up related imports and dead code paths.

Files changed: distorch_2.py, wrappers.py, checkpoint_multigpu.py, device_utils.py, model_management_mgpu.py, __init__.py
2025-10-13 19:00:59 -05:00
John Pollock e6d19951d7 Paranoia before cleanup 2025-10-04 12:32:37 -05:00
John Pollock 62752d1bbf Standardize doc strings and make PEP 257 compliant 2025-09-30 09:34:41 -05:00
John Pollock 429be7c912 docs: update activeContext for v2.5.0 release with refactoring summary
Update active context documentation to reflect v2.5.0 release candidate status.
Major achievements documented:

- DisTorch2 allocation refactoring (-179 lines, unified UNET/CLIP logic)
- Production cleanup removing debug instrumentation (-40 lines)
- Verified selective unload system working with production logs
- Architecture status showing all core files production-ready
- Updated memory management pipeline with verification details

Reorganized content to prioritize recent session achievements (2025-09-30)
and production readiness status. Total code reduction: 219 lines through
consolidation and cleanup while maintaining full functionality.
2025-09-30 09:07:05 -05:00
John Pollock 07b429f3f9 fix(distorch): Add GC anchor protection for selective model retention
Problem: Models correctly categorized as "keep loaded" during selective
unload were disappearing before the next cleanup cycle. After reassigning
mm.current_loaded_models = kept_models, Python's garbage collector would
clear the models because the list was their only remaining strong reference.

Solution: Implement GC anchor system using a global set to hold strong
references to ModelPatcher objects that must survive cleanup cycles.

Changes:
- Add _MGPU_RETENTION_ANCHORS global set and helper functions
- Add early delegation check: if no DisTorch models want unload, clear
  anchors and delegate to original unload_all_models
- Add retention anchor when categorizing kept models
- Clear anchors before delegating to allow normal cleanup

Result: Self-contained, reversible protection mechanism. Models with
keep_loaded=True survive automatic cleanup but can be cleared with
explicit "Clear All Models" button. Tested on both keep_loaded=True
and keep_loaded=False scenarios.

Refs: memory-bank/distorch_selective_unload_solution.md
2025-09-29 16:01:24 -05:00
John Pollock c23dc083d3 WIP 2025-09-29 14:04:05 -05:00
John Pollock 1ca3daf0d8 At least now the logs reflect it is now trying to do what I know we have figured out how to do in the past in one of these commits. . . 2025-09-29 07:08:03 -05:00
John Pollock c4ae5e9e08 extensive clean-up, WIP 2025-09-29 03:53:21 -05:00
John Pollock 8591063a3c incremental progress (I think, hard to tell) 2025-09-28 19:14:47 -05:00
John Pollock 18493f5277 refactor: simplify model retention logic in multi-GPU unload
- Renamed `keep_loaded` variable to `should_retain` for improved clarity
- Simplified assignment by directly retrieving `_mgpu_keep_loaded` attribute with default False
- Updated logging accordingly; may alter behavior for non-DisTorch models to no longer retain automatically
2025-09-28 13:22:19 -05:00
John Pollock ae8bb7cf2c feat: Refine model retention logic in multi-GPU unloading
- Modified condition to retain models lacking `_mgpu_keep_loaded` attribute or with `keep_loaded=True`
- Improves reliability of unloading by distinguishing DisTorch and non-DisTorch models
- Addresses potential premature unloading of intended persistent models in multi-GPU setups
2025-09-28 11:14:34 -05:00
John Pollock fda5d6ed00 commiting this steaming pile of hot garbage for future dissection to see if I want any organs from this terminally ill branch 2025-09-25 14:36:13 -05:00
John Pollock bd672479fa refactor: eliminate circular import by separating model management functions
- Create model_management_mgpu.py for centralized model lifecycle tracking
- Move memory management functions from device_utils.py to new module:
  * multigpu_memory_log, track_modelpatcher, trigger_executor_cache_reset
  * check_cpu_memory_threshold, prune_distorch_stores, try_malloc_trim
  * force_full_system_cleanup
- Update imports across codebase (distorch_2.py, distorch.py, __init__.py,
  nodes.py, checkpoint_multigpu.py)
- Resolves device_utils.py ↔ distorch_2.py circular dependency
- Follows established clean coding patterns with fail-fast error handling

Addresses critical CPU memory leak investigation infrastructure by ensuring
proper module separation for comprehensive memory management utilities.
2025-09-24 17:38:15 -05:00