Commit Graph
203 Commits
Author SHA1 Message Date
Adrien Toupet 4a3efc7883 feat: add CPU offloading for intermediate data to reduce VRAM usage
- Store latents and transformed videos on CPU between processing phases reducing VRAM usage to enable larger batch processing
- Make transformed video storage conditional (only when color_correction != none)
- Add all_ori_lengths tracking for consistent trimming
- Use non_blocking=False for CPU-GPU transfers to avoid pinned memory issues
Fixes https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/pull/163#issuecomment-3333686784
2025-09-25 11:30:32 -04:00
Adrien Toupet 547ccc749b Merge pull request #163 from AInVFX/nightly
WIP - 3 phases / GGUF / Relative Imports / Extra parms ...
2025-09-25 00:34:54 -04:00
Adrien Toupet 9b4a7dfa3e fix: unify model registry for CLI and improve GGUF implementation
Model Registry Integration:
- Replace hardcoded CLI model choices with dynamic get_available_models()
- Use centralized DEFAULT_MODEL constant from model_registry
- Achieve single source of truth between CLI and ComfyUI interfaces

GGUF Optimizations:
- Create tensors directly on target device to avoid CPU->GPU copy overhead
- Remove excessive memory cleanup during tensor loading for better performance
- Always use meta init for memory-efficient model creation
- Implement precision-optimized dequantization path (GGUF → FP16 → compute dtype)

GGUF Precision Handling:
- Add GGUFQuantizedLinear/Conv2d layers with get_dequantized_weight_for_compute() to preserve precision
- Track and report quantization types (Q4_K_M, Q5_K_M, etc.) in model
- Fix __torch_function__ as classmethod to resolve deprecation warning

Code Quality:
- Restructure _load_model_weights() with modular helper functions to reduce duplication
- Improve separation of concerns between standard and GGUF weight loading
- Enhance logging to always display WARNING/ERROR messages
- Add comprehensive GGUF architecture validation
- Remove debug-only log_memory_state calls
- Fix type function definitions
2025-09-25 00:09:56 -04:00
Adrien Toupet 1fa603899f Fix GGUF implementation: centralize warning suppression, improve type hints, enhance error handling
- Add suppress_tensor_warnings() to constants.py to centralize tensor/numpy warning handling
- Fix type hints and List/Tuple imports
- Improve GGUFTensor.__torch_function__ to prevent recursion, add debug null checks
- Fix GGUFTensor.to() to properly preserve tensor_shape attribute
- Enhance GGUF error handling with proper exceptions instead of just warnings
- Simplify gguf_ops dequantize_weight by removing redundant fallback paths
- Add memory logging and proper CUDA sync cleanup in GGUF loading
- Translate French comments to English in euler.py
- Remove unnecessary GGUF special handling in memory_manager
- Fix _propagate_debug_to_modules to handle None debug instance
2025-09-24 15:47:18 -04:00
Adrien Toupet 0b0c87ed4a Add GGUF quantized model support (based on PR #121 from @cmeka / @lihaoyun6)
- Implement GGUF model loading with Q3_K_M through Q8_K_M quantization support
- Add GGUFTensor wrapper to preserve quantization and enable on-demand dequantization
- Maintain tensors in quantized format to reduce VRAM usage
- Add GGUF dequantization operations for inference
- Update model registry to include GGUF variants for 3B/7B models
- Fix wavelet blur radius limit to prevent OOM at high resolutions (max 1/8 of image dimension)
- Add safety clamp [-1,1] for SDR color range to prevent numerical errors
- This is a WIP commit as some additional cleaning/testing is needed
- Add type hints throughout for better code maintainability
2025-09-24 11:49:57 -04:00
Adrien Toupet 953524b13b feat: Add support for extra_model_paths.yaml configuration
- Implement model discovery across multiple paths via ComfyUI's folder_paths API
- Add centralized model file scanning in get_all_model_files() to avoid duplication
- Update model loading to search all registered paths before downloading
- Prevent redundant downloads when files exist in alternate paths
- Users can now specify custom model directories using 'seedvr2' key in extra_model_paths.yaml

Addresses request from https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/issues/128#issuecomment-3312407641
2025-09-23 00:56:26 -04:00
Adrien Toupet 0967bb153c fix: update dependencies to include all required packages for CLI and ComfyUI 2025-09-23 00:01:02 -04:00
Adrien Toupet efa851009f Fix corrupted model downloads with hash validation and resume support
- Add SHA256 hash validation for all models and VAE
- Implement download resume capability with progress bars
- Add validation cache to avoid redundant checks
- Replace unreliable torchvision downloads with urllib
- Fix safetensors header corruption detection
- Move SeedVR2 ASCII logo with credits before download operations for clearer logs

Addresses #128, #151
2025-09-22 23:45:43 -04:00
Adrien Toupet 2b00252a49 Simplify __init__.py files + fix dependencies + add pyproject.toml
- Clean up all __init__.py files to remove commented code and unnecessary imports
- Root __init__.py now only exports NODE_CLASS_MAPPINGS and NODE_DISPLAY_NAME_MAPPINGS
- Remove unused dependencies: pytorch-extension (replaced by custom implementations) and peft
- Fix diffusers version syntax (remove 'v' prefix)
- Add pyproject.toml for ComfyUI registry compliance
2025-09-22 15:47:04 -04:00
Adrien Toupet 8f46898073 Fix OOM when loading models without preserve_vram by using meta device initialization
- Always use meta device initialization to avoid double memory allocation
- Load weights directly to target device with assign=True for in-place materialization
- Initialize non-persistent buffers (RoPE frequencies) after materialization
- Use strict=False universally in load_state_dict to handle missing buffers
2025-09-22 14:48:15 -04:00
Adrien Toupet 9927f45f02 Fix BlockSwap config detection and progress bar cleanup
- Fix BlockSwap config changes not being detected when cache_model=True
  * Use tuple comparison for configs: (blocks, offload_io)
  * Handle all transitions: enabled↔disabled, config changes, disconnection
  * Show readable format in logs: blocks=14, offload_io=True

- Fix progress bar stuck after errors by clearing it in cleanup()

- Fix BlockSwap cleanup by checking any state existence, not just _blockswap_active
  * Properly cleans up deactivated BlockSwap with remaining config
2025-09-22 13:46:50 -04:00
Adrien Toupet 22f540a7d6 feat: Add input noise parameter and rename latent noise for clarity
- Add input_noise_scale (0-1) to reduce artifacts at high resolutions
  - Applies subtle noise before VAE encoding with progressive blend (0-50%)
  - Based on GitHub issue #64 community findings

- Rename cond_noise_scale to latent_noise_scale for clarity
  - Better distinguishes between input (pixel) and latent (diffusion) noise
  - Maintains same functionality, just clearer naming

- Update both ComfyUI and CLI interfaces with new parameters
2025-09-22 11:47:49 -04:00
Adrien Toupet d0d24a18cb Refactor: Replace dynamic relative imports with model class registry for cross-platform compatibility. Minor: Adjust ASCII logo width for better screen fit 2025-09-20 01:53:25 -04:00
Adrien Toupet e89a8b00a5 feat: promote cond_noise_scale to UI for controlling conditioning noise level helping mitigate noise artifacts in high-res upscaling 2025-09-18 23:19:35 -04:00
Adrien Toupet e6da542b7b VAE tiling defaults + Fix BlockSwap + preserve_vram compatibility and device management
- Changed VAE tiling default to False to restore original behavior
- Fixed BlockSwap bypass mode during cache cleanup to allow proper CPU offloading
- Added runner parameter to all manage_model_device calls for BlockSwap detection
- Removed conditional movement flags, always offload when preserve_vram=True
- Fixed device comparison to use full device strings (cuda:0) not just types
- Added debug logging for BlockSwap device skip scenarios
2025-09-17 23:39:03 -04:00
Adrien Toupet ee95a3a673 feat: Add configurable color correction methods and improve memory management
- Add color_correction parameter (wavelet/adain/none) to both ComfyUI node and CLI interface to choose between wavelet reconstruction (frequency-based), AdaIN (statistical matching), or no color correction
- Improve memory management after upscaling batches with targeted clear_memory calls to reduce VRAM pressure during VAE decoding
2025-09-16 22:21:45 -04:00
Adrien Toupet 1518ecc1a5 refactor: Fix ComfyUI node conflicts via relative imports and clearer model structure
- Renamed model directories for clarity: dit -> dit_7b, dit_v2 -> dit_3b
- Converted all absolute imports to relative imports throughout codebase
- Removed sys.path.append() manipulations that caused namespace conflicts
- Updated YAML configs to reference renamed model directories
- Simplified model variant detection logic using new directory names
- Standardized function calls with named arguments for better clarity
- Fixed generation context initialization and interrupt handling

This resolves import conflicts with other ComfyUI nodes (e.g., Basic data handling)
that use sys.path manipulation, making the module properly isolated and compatible.

Fixes #29, #114, #136
2025-09-16 16:45:31 -04:00
Adrien Toupet b41998c5e1 docs: Clean up README formatting 2025-09-16 15:34:48 -04:00
Adrien Toupet 2c570105b1 docs: Update project documentation and credits
- Fix project names and URLs in CONTRIBUTING.md
- Add dual branch workflow (main/nightly) guidelines
- Update LICENSE with correct copyright holders
- Expand README credits to acknowledge all contributors
- Add proper contact information for maintainers
2025-09-16 15:29:39 -04:00
Adrien Toupet 477f57fd5a Refactor: Three-phase batch processing pipeline for improved performance
Major architectural change to minimize model swapping overhead by processing
all batches in three distinct phases instead of sequential per-batch processing:
- Phase 1: Encode all batches with VAE
- Phase 2: Upscale all latents with DiT
- Phase 3: Decode all latents with VAE

Core changes:
- Split monolithic generation_loop into modular functions:
  - prepare_generation_context(): Shared state management
  - setup_device_environment(): Device configuration
  - prepare_runner(): Model loading with cache support
  - encode_all_batches(): Batch VAE encoding
  - upscale_all_batches(): Batch DiT upscaling
  - decode_all_batches(): Batch VAE decoding
- Removed generation_step function (logic integrated into upscale phase)
- Added lazy precision initialization to avoid redundant setup

Performance improvements:
- Pre-allocated lists for memory efficiency
- Better cleanup of intermediate storage between phases
- Added unique timer names to clear_memory() to avoid naming conflicts
- Improved model state management with change detection and caching

UI/UX enhancements:
- Switched to ComfyUI's native ProgressBar with weighted phase progress
- Changed from per-batch FPS to overall average FPS (always visible)
- Improved log clarity with clear phase separators
- Added ASCII art logo to clearly identify SeedVR2 process start
- Better progress tracking with weighted percentages across three phases

Code cleanup:
- Removed deprecated timer_context from Debug class
- Removed unused time imports across multiple files
- Fixed LOCAL_RANK environment variable to handle string conversion properly
- Improved error handling with try/except/finally blocks in all phases
2025-09-16 14:38:57 -04:00
Adrien Toupet d504eeafa6 fix(regression): Restore MPS compatibility for VAE models
- Set VAE to BFloat16 on MPS to prevent dtype mismatch errors
- Add override_dtype parameter to model loading pipeline
- Convert VAE weights to target dtype during load for efficiency

This fixes https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/issues/73#issuecomment-3237331971
2025-08-29 18:15:56 +02:00
Adrien Toupet cf274a990b fix(regression): Restore MPS compatibility for VAE models
- Set VAE to BFloat16 on MPS to prevent dtype mismatch errors
- Add override_dtype parameter to model loading pipeline
- Convert VAE weights to target dtype during load for efficiency
2025-08-29 12:14:23 -04:00
Adrien Toupet 36f4d80e6f Merge pull request #137 from AInVFX/nightly - disable detailed tensor inspection
fix: disable detailed tensor inspection to avoid deprecation warnings and indirect dependency calls
2025-08-29 15:07:37 +02:00
Adrien Toupet fc27097045 fix: disable detailed tensor inspection to avoid deprecation warnings and indirect dependency calls 2025-08-29 09:05:29 -04:00
Adrien Toupet fe61f1866e Memory efficiency, performance improvements,and architecture refactoring
### 🚀 Performance Optimizations

#### Model Initialization & Loading
- **Faster model initialization** using meta device to avoid unnecessary memory allocation during DiT and VAE creation
- **Memory-mapped VAE loading** with `assign=True` to reference tensors directly instead of copying, matching DiT's loading strategy
- **Optimized state dict management** with explicit deletion after loading to free memory immediately
- **Removed slow preinitialize_rope_cache()** bottleneck during model preparation

#### Generation Pipeline  
- **Streamlined batch processing** by merging redundant log steps and unifying "Model Configuration" and "Input Preparation" into single "Generation Setup" step
- **Optimized memory usage** by passing compute_dtype through entire diffusion pipeline, ensuring timesteps use bfloat16 instead of float32
- **Code deduplication** by reusing existing functions: `prepare_video_transforms()`, `load_text_embeddings()`, `calculate_optimal_batch_params()`
- **Improved batch recommendations** with `calculate_optimal_batch_params()` now prioritizing temporal stability by suggesting largest valid 4n+1 batch size

#### VAE Operations
- **Optimized tile blending** with pre-computed cosine ramps instead of recalculating per-tile
- **In-place operations** using `mul_` and `div_` to reduce memory allocations
- **Reduced redundant transfers** by skipping device moves when tensors are already on correct device
- **Improved slicing operations** with consistent memory cleanup for both encode and decode paths

### 💾 Memory Management

#### Critical Fixes
- **Fixed major VRAM leak** in VAE decoder by clearing InflatedCausalConv3d memory buffers when moving VAE to CPU
- **Fixed RoPE OOM issues** by computing only required dimensions instead of maximum (1024x128x128), enabling higher resolution upscaling
- **Added OOM retry mechanism** for VAE operations (Upsample3D, ResnetBlock3D, InflatedCausalConv3d, GroupNorm) with single retry on failure
- **Added cuBLAS workspace clearing** for better GPU memory management

#### BlockSwap Integration
- **Full preserve_vram compatibility** with bypass mechanism to temporarily disable BlockSwap protection during offload
- **Transparent handling** in `manage_model_device()` for BlockSwap models
- **Proper state restoration** of blocks to correct devices (partial GPU config) on reload
- **Simplified preserve_vram logic** by removing BlockSwap checks from infer.py

### 🏗️ Architecture Refactoring

#### Code Organization
- **Unified model configuration** by merging `configure_dit_model_inference` and `configure_vae_model_inference` into single `configure_model_inference` function  
- **Focused helper functions** extracted from `configure_model_inference()` 
- **Removed redundant code**: gradient checkpointing (training-only), model_weight variables, platform-specific VAE dtype logic
- **Simplified FP8 logic** by removing unnecessary conversion code from `load_quantized_state_dict`

#### Dtype Handling
- **Unified dtype flow**: VAE now uses configured dtype consistently, removing redundant vae_dtype variable
- **Fixed FP16 black frames** by using BFloat16 compute/autocast for FP16 models
- **Natural dtype propagation** from generation_loop → configure_diffusion → sampling, eliminating manual conversions
- **Removed nested autocast contexts** relying on outer context from generation_step

### 🐛 Bug Fixes & Improvements

#### Debug & Logging
- **Optimized debug performance** by making tensor counting opt-in (`show_tensors=False` by default)
- **Conditional logging** with `reset_vram_peak()` only logging when debug is enabled
- **Early return pattern** in log() method to avoid string operations when disabled
- **Standardized formatting**: consistent .2f float precision, uppercase device names, unified category names
- **Improved clarity**: better frame padding logs, clearer FP8 RoPE conversion messages
2025-08-29 06:18:44 +02:00
Adrien Toupet 66bcf436ab perf: optimize debug logging to reduce performance overhead
- Change log_memory_state() to default show_tensors=False (opt-in tensor counting), skipping expensive tensor counting during intermediate operations
- Enable detailed tensor analysis only at final cleanup
- Add early return pattern in log() method to avoid string operations when disabled
- Make reset_vram_peak() logging conditional on debug.enabled state
- Remove unused vram_info call in model_manager.py
2025-08-28 23:50:49 -04:00
Adrien Toupet 22ffc86af7 perf: Optimize VAE operations and add missing memory cleanup to decode
- Add InflatedCausalConv3d memory cleanup to slicing_decode to match slicing_encode
- Skip redundant device transfers when tensors are already on correct device
- Pre-compute cosine ramps for tile blending instead of recalculating per-tile
- Use in-place operations (mul_, div_) to reduce memory allocations
- Reduce debug logging frequency and clarify tile progress messages
- Filter modules with memory before clearing to avoid unnecessary iterations

Minor performance improvements with no memory increase.
2025-08-28 16:18:12 -04:00
Adrien Toupet 5e2fb76414 fix: optimize RoPE frequency computation and remove initialization bottleneck
- Fix NaMMRotaryEmbedding3d to compute only required dimensions instead of maximum (1024x128x128), enabling much higher resolution upscaling with 3B model without running OOM
- Remove slow preinitialize_rope_cache() which was causing bottleneck during model preparation - no longer needed with NaMMRotaryEmbedding3d fixed
- Clean up unnecessary memory clearing calls and improve debug logging clarity
2025-08-28 14:20:41 -04:00
Adrien Toupet feabb94e5a perf: optimize memory usage and streamline generation pipeline
Dtype Consistency:
- Pass compute_dtype through entire diffusion pipeline to ensure timesteps use bfloat16 instead of float32, reducing memory usage and eliminating dtype conversion overhead
- Propagate dtype from generation_loop → configure_diffusion → sampling timesteps creation

Code Deduplication:
- Removed duplicated code and made use of existing reusable functions: prepare_video_transforms(), load_text_embeddings(), calculate_optimal_batch_params()
- Improve calculate_optimal_batch_params() to prioritize temporal stability by recommending largest valid 4n+1 batch size

Memory Optimizations:
- Remove unnecessary GPU↔CPU offloading for text embeddings and timesteps during preserve_vram mode (added overhead without meaningful memory benefits)
- Simplify text embeddings cleanup to work directly with dictionary structure
- Add cuBLAS workspace clearing for better GPU memory management

Code Cleanup:
- Remove redundant self.runner = None in _internal_execute (handled by cleanup())
- Enhance user messaging for batch padding waste with clearer explanations
2025-08-27 14:15:41 -04:00
Adrien Toupet 267bc099e6 Fix BlockSwap compatibility with preserve_vram mode
- Add bypass mechanism to temporarily disable BlockSwap protection during offload
- Refactor manage_model_device() to handle BlockSwap models transparently
- Restore blocks to correct devices (partial GPU config) on reload
- Simplify preserve_vram logic in infer.py by removing BlockSwap checks

This allows BlockSwap to work correctly with preserve_vram, offloading the entire model to CPU after inference and restoring only necessary blocks to GPU before next inference.
2025-08-26 17:31:04 -04:00
Adrien Toupet 89bde29cbe perf: optimize model initialization with meta device for 90% speedup
- Use meta device initialization to avoid unnecessary memory allocation during model creation
  - Reduces DiT and VAE initialization time by ~90% when loading to CPU
  - Explicitly delete state dicts after loading to free memory immediately
- Refactor configure_model_inference() into focused helper functions for DRY code
- Improve debug logging with timestamps and cleaner section separators for readability
2025-08-26 13:23:30 -04:00
Adrien Toupet 1ee1ee0fb5 fix: add OOM retry mechanism and fix FP16 black frames regression
- Implement retry_on_oom() helper with single retry for all VAE operations (Upsample3D, ResnetBlock3D, InflatedCausalConv3d, GroupNorm)
- Fix regression: use BFloat16 compute/autocast for FP16 models to prevent black frames
- Improve frame padding logs to clarify model constraint (frames % 4 == 1)
2025-08-26 00:03:02 -04:00
Adrien Toupet 0358198804 refactor: Remove redundant dtype conversions and autocast contexts
- Remove hardcoded target_dtype and manual tensor conversions in inference()
- Remove redundant inner autocast context - rely on outer context from generation_step
- Align compute_dtype and autocast_dtype for FP16 models (both now float16)
- Remove unused block_swap_config parameter
- Fix device logging to properly display device names

The dtype is now determined once in generation_loop based on model weights and flows naturally through the pipeline via the outer autocast context, eliminating unnecessary conversions and improving performance.
2025-08-25 20:37:52 -04:00
Adrien Toupet 5595d58597 refactor: streamline generation pipeline and improve dtype handling
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
2025-08-25 17:11:29 -04:00
Adrien Toupet 47552c0ed5 refactor: unify model configuration and improve logging precision
- Merge configure_dit_model_inference and configure_vae_model_inference into single configure_model_inference function
- Simplify load_quantized_state_dict by removing unnecessary FP8 conversion logic
- Remove redundant gradient checkpointing call (training-only feature)
- Remove redundant model_weight variables
- Standardize float formatting to .2f across all debug logs for consistency
- Eliminate redundant device movement operations already handled during model creation
2025-08-25 13:45:13 -04:00
Adrien Toupet 53f41770c1 fix: use memory-mapped loading for VAE via assign=True
- Match DiT's load_state_dict(assign=True) for VAE to reference memory-mapped tensors directly instead of copying, reducing peak RAM
- Move post-inference memory logging after offloading for accurate metrics
2025-08-25 11:03:35 -04:00
Adrien Toupet 8c48aa59ea Fix VRAM memory leak in VAE decoder
Updated manage_model_device() to clear InflatedCausalConv3d memory buffers
when moving VAE to CPU. These buffers were not automatically released with
.to('cpu'), causing intermediate tensors to remain in VRAM after batch processing.
2025-08-25 10:12:40 -04:00
Adrien Toupet 7e7853df54 Blockswap regression with cache model
Preserve BlockSwap's GPU/CPU distribution when cache_model=True to avoid unnecessary reconfiguration on subsequent runs. Also fixes regression error where tensors were not on the correct device after caching.
2025-08-24 18:15:31 +02:00
Adrien Toupet 627f215b57 Fix BlockSwap memory configuration being lost during model caching
Preserve BlockSwap's GPU/CPU distribution when cache_model=True to avoid unnecessary reconfiguration on subsequent runs. Also fixes regression error where tensors were not on the correct device after caching.
2025-08-24 11:38:02 -04:00
Adrien Toupet a2b546c0e7 Fix BlockSwap regression: preserve block positions when caching
Regression from previous commit - cleanup_blockswap was removing forward wrappers/RoPE patches
but leaving blocks in position, then re-applying BlockSwap duplicated memory allocations
causing OOM on second generation.

Now when caching with same config:
- Blocks remain in their GPU/CPU positions for next inference
- Forward wrappers and RoPE patches stay intact
- Only toggles _blockswap_active flag for safety
- Reuses exact same VRAM allocation as previous run

Full cleanup only happens when cache_model=False or config changes.

Fixes OOM errors and eliminates redundant GPU transfers between cached generations.
2025-08-23 22:58:53 -04:00
Adrien Toupet 5bf439ee3d refactor: Memory management overhaul with enhanced debugging and performance optimizations
#### Memory Management & Leak Fixes
- **Fixed BlockSwap pinned memory leak**: Removed `non_blocking` parameter that was preventing pinned memory release
- **Eliminated logging-induced memory leaks**: Modified `log_memory_state()` to avoid holding tensor references and added checkpoint history limits
- **Complete model cleanup**: Ensured proper deletion and tensor storage release when `cache_model=False`
- **Fixed lingering GPU tensors**: Resolved issues with scalar tensors and text embeddings remaining on GPU between batches with `preserve_vram`

#### Performance Optimizations
- **Direct-to-CPU model loading**: DiT/VAE weights load directly to CPU when `preserve_vram` or BlockSwap is active, avoiding VRAM spikes
- **Reactive OOM handling**: Replaced proactive memory clearing during VAE decode with reactive OOM recovery, fixing performance regression (removed the need for the keep_vae_loaded_during_decode flag)
- **Optimized cleanup functions**: Consolidated all cleanup into core functions (`manage_*`, `release_*`, `clear_*`) 

#### Enhanced Debugging & Monitoring
- **Unified debug system**: Improved logging categories, icons, and hierarchical timers
- **Comprehensive memory tracking**: Per-interval VRAM peak tracking, accurate process-specific RAM monitoring
- **Debug propagation**: Debug instance now properly propagates to all submodules (sampler, VAE, etc.)
- **Improved log clarity**: Standardized message formatting, removed redundant logs, added critical operation snapshots

#### Code Architecture Improvements
- **Type hints **: Added  type hints to all memory management functions
- **Centralized device management**: New `manage_model_device()` function for consistent device movement logic
- **Error handling**: Added try/except blocks to all critical memory and device operations
- **Removed private APIs**: Eliminated `torch._C._clear_cache()` usage that was triggering warning

#### Bug Fixes
- Resolved duplicate dtype detection in generation loop
- Corrected CPU offloading logic for text embeddings and sampling timesteps
- Fixed timer namespace collisions in batch processing
- Removed redundant RoPE cache clearing operations

#### Various
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove unused imports and legacy code/comments
2025-08-23 08:25:20 +02:00
Adrien Toupet 8333fb856e refactor(WIP): complete memory management overhaul with proper error handling
Memory Management:
- Add type hints to all memory functions for better IDE support and maintainability
- Replace silent exception handling with debug logging across all operations
- Removed unnecessary CPU transfers for GPU memory release
- Introduce unified manage_model_device() for consistent device management
- Remove torch._C._clear_cache() private API usage (incompatible across PyTorch versions)

Performance & Debugging:
- Add debug timers to critical operations (clear_memory, clear_runtime_caches, etc.)
- Consolidate all cleanup code into core functions: manage_*, release_*, clear_*, complete_cleanup
- Improved timer log messages for clarity

Code Quality:
- Remove unused imports
- Remove redundant RoPE cache clearing (clear_runtime_caches handles it)
- Simplify configure_runner() by eliminating duplicate code paths
- Update infer.py to use generic device management functions
- Add release_text_embeddings() helper to deduplicate embedding cleanup
2025-08-23 01:13:45 -04:00
Adrien Toupet faea9397d9 Fix: Remove non_blocking from BlockSwap to prevent CPU-GPU pinned memory leaks 2025-08-22 16:57:48 -04:00
Adrien Toupet f51b0b3991 Fix debug logging: propagate debug instance to sampler and VAE submodules 2025-08-22 16:29:44 -04:00
Adrien Toupet 2053a80f38 refactor(WIP): Improve debug logging consistency and reduce redundancy
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state()  formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
2025-08-22 15:38:00 -04:00
Adrien Toupet 9b7c113681 refactor(WIP): Memory Management Overhaul
This is a work-in-progress commit that consolidates a series of changes to fix memory leaks, optimize VRAM/RAM usage, improve performance, and enhance code maintainability.

*   **BlockSwap Pinned Memory:** Disabled `use_non_blocking=True` for CPU-to-GPU transfers to resolve a memory leak where pinned memory was not being released.
*   **Logging-Induced Leaks:** Modified `log_memory_state()` to avoid holding references to tensors during analysis and added a history limit to the checkpoint system to prevent unbounded memory growth.
*   **Incomplete Model Cleanup:** Ensured models are completely deleted and their tensor storage is released when `cache_model=False`.
*   **Lingering Tensors:** Fixed an issue where a scalar tensor from sampling timesteps and text embeddings remained on the GPU between batches when `preserve_vram` is active.

*   **Centralized Cleanup Functions:** Introduced `clear_memory()` to replace `clear_vram_cache()` and all manual `torch.cuda.empty_cache()` calls, providing consistent VRAM/RAM cleanup logic. The function features a `full` parameter to distinguish between a fast, GPU-only cache clear (~1-5ms) for frequent operations and a full cleanup with garbage collection (~10-50ms) for critical stages.
*   **Direct-to-CPU Model Loading:** Modified DiT/VAE weight loading to load directly onto the CPU when `preserve_vram` or `BlockSwap` is active, avoiding unnecessary VRAM spikes during model preparation.
*   **VAE Device Management:** Created the `manage_vae_device()` helper function to centralize the logic for moving the VAE between the CPU and GPU, reducing code duplication. This also fixed a bug that incorrectly kept the VAE on the GPU when `preserve_vram` was active.
*   **CPU Offloading:** Implemented logic to move text embeddings and sampling timesteps to the CPU after each batch when `preserve_vram` is active, reducing idle VRAM usage.

*   **VAE Decode Performance:** Replaced proactive, frequent memory clearing during VAE decode with a reactive Out-of-Memory (OOM) handling system. This fixed a significant performance regression and eliminated the need for the `keep_vae_loaded_during_decode` flag.
*   **Reduced Overhead:** Removed redundant `gc.collect()` calls from multiple locations to decrease unnecessary processing overhead.

*   **Interval-Based VRAM Tracking:** Modified the logging system to reset peak VRAM statistics after each `log_memory_state()` call, enabling accurate tracking of peak memory usage for specific processing intervals (e.g., encode, inference, decode).
*   **Accurate RAM Monitoring:** Added the `get_ram_usage()` function for correct process-specific RAM tracking.
*   **Efficient Log Refactoring:** Refactored `log_memory_state()` into modular helper methods, optimizing tensor analysis into a single-pass `gc` iteration to improve both performance and maintainability.
*   **Log Clarity:** Refined memory state and debug logging to remove redundant snapshots and add new ones for critical operations like model loading, weight loading, VAE encoding, and decoding. Standardized log message conventions.
*   **Per-Batch Timers:** Implemented timer namespacing to ensure that performance timers for each batch are logged correctly without overwriting one another.

*   **Error Handling:** Added `try/except` blocks to key memory and device management functions to handle edge cases and improve robustness.
*   **Code Cleanup:** Removed deprecated code and outdated comments throughout the related modules.
*   **Documentation:** Updated comments and function docstrings to reflect the new memory management architecture.
2025-08-22 09:20:41 -04:00
Adrien Toupet 0373a81d2b Merge pull request #126 from benjaminherb/fix-prepend-frames
Fix for the --prepend_frames logic
2025-08-21 14:17:58 +02:00
Benjamin Herb 22aa85abb7 fix: Updates --prepend_frames logic 2025-08-21 08:12:07 +02:00
Adrien Toupet 1a12465588 Merge pull request #125 from JohnAlcatraz/nightly
Add keep_vae_loaded_during_decode setting for speeding up VAE decode more than 2x with preserve_vram=true
2025-08-20 22:53:06 +02:00
JohnAlcatraz 2eaebf7709 Update comfyui_node.py 2025-08-20 22:07:37 +02:00