Commit Graph
27 Commits
Author SHA1 Message Date
Adrien Toupet 477f57fd5a Refactor: Three-phase batch processing pipeline for improved performance
Major architectural change to minimize model swapping overhead by processing
all batches in three distinct phases instead of sequential per-batch processing:
- Phase 1: Encode all batches with VAE
- Phase 2: Upscale all latents with DiT
- Phase 3: Decode all latents with VAE

Core changes:
- Split monolithic generation_loop into modular functions:
  - prepare_generation_context(): Shared state management
  - setup_device_environment(): Device configuration
  - prepare_runner(): Model loading with cache support
  - encode_all_batches(): Batch VAE encoding
  - upscale_all_batches(): Batch DiT upscaling
  - decode_all_batches(): Batch VAE decoding
- Removed generation_step function (logic integrated into upscale phase)
- Added lazy precision initialization to avoid redundant setup

Performance improvements:
- Pre-allocated lists for memory efficiency
- Better cleanup of intermediate storage between phases
- Added unique timer names to clear_memory() to avoid naming conflicts
- Improved model state management with change detection and caching

UI/UX enhancements:
- Switched to ComfyUI's native ProgressBar with weighted phase progress
- Changed from per-batch FPS to overall average FPS (always visible)
- Improved log clarity with clear phase separators
- Added ASCII art logo to clearly identify SeedVR2 process start
- Better progress tracking with weighted percentages across three phases

Code cleanup:
- Removed deprecated timer_context from Debug class
- Removed unused time imports across multiple files
- Fixed LOCAL_RANK environment variable to handle string conversion properly
- Improved error handling with try/except/finally blocks in all phases
2025-09-16 14:38:57 -04:00
Adrien Toupet 66bcf436ab perf: optimize debug logging to reduce performance overhead
- Change log_memory_state() to default show_tensors=False (opt-in tensor counting), skipping expensive tensor counting during intermediate operations
- Enable detailed tensor analysis only at final cleanup
- Add early return pattern in log() method to avoid string operations when disabled
- Make reset_vram_peak() logging conditional on debug.enabled state
- Remove unused vram_info call in model_manager.py
2025-08-28 23:50:49 -04:00
Adrien Toupet 5e2fb76414 fix: optimize RoPE frequency computation and remove initialization bottleneck
- Fix NaMMRotaryEmbedding3d to compute only required dimensions instead of maximum (1024x128x128), enabling much higher resolution upscaling with 3B model without running OOM
- Remove slow preinitialize_rope_cache() which was causing bottleneck during model preparation - no longer needed with NaMMRotaryEmbedding3d fixed
- Clean up unnecessary memory clearing calls and improve debug logging clarity
2025-08-28 14:20:41 -04:00
Adrien Toupet feabb94e5a perf: optimize memory usage and streamline generation pipeline
Dtype Consistency:
- Pass compute_dtype through entire diffusion pipeline to ensure timesteps use bfloat16 instead of float32, reducing memory usage and eliminating dtype conversion overhead
- Propagate dtype from generation_loop → configure_diffusion → sampling timesteps creation

Code Deduplication:
- Removed duplicated code and made use of existing reusable functions: prepare_video_transforms(), load_text_embeddings(), calculate_optimal_batch_params()
- Improve calculate_optimal_batch_params() to prioritize temporal stability by recommending largest valid 4n+1 batch size

Memory Optimizations:
- Remove unnecessary GPU↔CPU offloading for text embeddings and timesteps during preserve_vram mode (added overhead without meaningful memory benefits)
- Simplify text embeddings cleanup to work directly with dictionary structure
- Add cuBLAS workspace clearing for better GPU memory management

Code Cleanup:
- Remove redundant self.runner = None in _internal_execute (handled by cleanup())
- Enhance user messaging for batch padding waste with clearer explanations
2025-08-27 14:15:41 -04:00
Adrien Toupet 267bc099e6 Fix BlockSwap compatibility with preserve_vram mode
- Add bypass mechanism to temporarily disable BlockSwap protection during offload
- Refactor manage_model_device() to handle BlockSwap models transparently
- Restore blocks to correct devices (partial GPU config) on reload
- Simplify preserve_vram logic in infer.py by removing BlockSwap checks

This allows BlockSwap to work correctly with preserve_vram, offloading the entire model to CPU after inference and restoring only necessary blocks to GPU before next inference.
2025-08-26 17:31:04 -04:00
Adrien Toupet 1ee1ee0fb5 fix: add OOM retry mechanism and fix FP16 black frames regression
- Implement retry_on_oom() helper with single retry for all VAE operations (Upsample3D, ResnetBlock3D, InflatedCausalConv3d, GroupNorm)
- Fix regression: use BFloat16 compute/autocast for FP16 models to prevent black frames
- Improve frame padding logs to clarify model constraint (frames % 4 == 1)
2025-08-26 00:03:02 -04:00
Adrien Toupet 5595d58597 refactor: streamline generation pipeline and improve dtype handling
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
2025-08-25 17:11:29 -04:00
Adrien Toupet 47552c0ed5 refactor: unify model configuration and improve logging precision
- Merge configure_dit_model_inference and configure_vae_model_inference into single configure_model_inference function
- Simplify load_quantized_state_dict by removing unnecessary FP8 conversion logic
- Remove redundant gradient checkpointing call (training-only feature)
- Remove redundant model_weight variables
- Standardize float formatting to .2f across all debug logs for consistency
- Eliminate redundant device movement operations already handled during model creation
2025-08-25 13:45:13 -04:00
Adrien Toupet 8c48aa59ea Fix VRAM memory leak in VAE decoder
Updated manage_model_device() to clear InflatedCausalConv3d memory buffers
when moving VAE to CPU. These buffers were not automatically released with
.to('cpu'), causing intermediate tensors to remain in VRAM after batch processing.
2025-08-25 10:12:40 -04:00
Adrien Toupet 627f215b57 Fix BlockSwap memory configuration being lost during model caching
Preserve BlockSwap's GPU/CPU distribution when cache_model=True to avoid unnecessary reconfiguration on subsequent runs. Also fixes regression error where tensors were not on the correct device after caching.
2025-08-24 11:38:02 -04:00
Adrien Toupet 8333fb856e refactor(WIP): complete memory management overhaul with proper error handling
Memory Management:
- Add type hints to all memory functions for better IDE support and maintainability
- Replace silent exception handling with debug logging across all operations
- Removed unnecessary CPU transfers for GPU memory release
- Introduce unified manage_model_device() for consistent device management
- Remove torch._C._clear_cache() private API usage (incompatible across PyTorch versions)

Performance & Debugging:
- Add debug timers to critical operations (clear_memory, clear_runtime_caches, etc.)
- Consolidate all cleanup code into core functions: manage_*, release_*, clear_*, complete_cleanup
- Improved timer log messages for clarity

Code Quality:
- Remove unused imports
- Remove redundant RoPE cache clearing (clear_runtime_caches handles it)
- Simplify configure_runner() by eliminating duplicate code paths
- Update infer.py to use generic device management functions
- Add release_text_embeddings() helper to deduplicate embedding cleanup
2025-08-23 01:13:45 -04:00
Adrien Toupet 2053a80f38 refactor(WIP): Improve debug logging consistency and reduce redundancy
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state()  formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
2025-08-22 15:38:00 -04:00
Adrien Toupet 9b7c113681 refactor(WIP): Memory Management Overhaul
This is a work-in-progress commit that consolidates a series of changes to fix memory leaks, optimize VRAM/RAM usage, improve performance, and enhance code maintainability.

*   **BlockSwap Pinned Memory:** Disabled `use_non_blocking=True` for CPU-to-GPU transfers to resolve a memory leak where pinned memory was not being released.
*   **Logging-Induced Leaks:** Modified `log_memory_state()` to avoid holding references to tensors during analysis and added a history limit to the checkpoint system to prevent unbounded memory growth.
*   **Incomplete Model Cleanup:** Ensured models are completely deleted and their tensor storage is released when `cache_model=False`.
*   **Lingering Tensors:** Fixed an issue where a scalar tensor from sampling timesteps and text embeddings remained on the GPU between batches when `preserve_vram` is active.

*   **Centralized Cleanup Functions:** Introduced `clear_memory()` to replace `clear_vram_cache()` and all manual `torch.cuda.empty_cache()` calls, providing consistent VRAM/RAM cleanup logic. The function features a `full` parameter to distinguish between a fast, GPU-only cache clear (~1-5ms) for frequent operations and a full cleanup with garbage collection (~10-50ms) for critical stages.
*   **Direct-to-CPU Model Loading:** Modified DiT/VAE weight loading to load directly onto the CPU when `preserve_vram` or `BlockSwap` is active, avoiding unnecessary VRAM spikes during model preparation.
*   **VAE Device Management:** Created the `manage_vae_device()` helper function to centralize the logic for moving the VAE between the CPU and GPU, reducing code duplication. This also fixed a bug that incorrectly kept the VAE on the GPU when `preserve_vram` was active.
*   **CPU Offloading:** Implemented logic to move text embeddings and sampling timesteps to the CPU after each batch when `preserve_vram` is active, reducing idle VRAM usage.

*   **VAE Decode Performance:** Replaced proactive, frequent memory clearing during VAE decode with a reactive Out-of-Memory (OOM) handling system. This fixed a significant performance regression and eliminated the need for the `keep_vae_loaded_during_decode` flag.
*   **Reduced Overhead:** Removed redundant `gc.collect()` calls from multiple locations to decrease unnecessary processing overhead.

*   **Interval-Based VRAM Tracking:** Modified the logging system to reset peak VRAM statistics after each `log_memory_state()` call, enabling accurate tracking of peak memory usage for specific processing intervals (e.g., encode, inference, decode).
*   **Accurate RAM Monitoring:** Added the `get_ram_usage()` function for correct process-specific RAM tracking.
*   **Efficient Log Refactoring:** Refactored `log_memory_state()` into modular helper methods, optimizing tensor analysis into a single-pass `gc` iteration to improve both performance and maintainability.
*   **Log Clarity:** Refined memory state and debug logging to remove redundant snapshots and add new ones for critical operations like model loading, weight loading, VAE encoding, and decoding. Standardized log message conventions.
*   **Per-Batch Timers:** Implemented timer namespacing to ensure that performance timers for each batch are logged correctly without overwriting one another.

*   **Error Handling:** Added `try/except` blocks to key memory and device management functions to handle edge cases and improve robustness.
*   **Code Cleanup:** Removed deprecated code and outdated comments throughout the related modules.
*   **Documentation:** Updated comments and function docstrings to reflect the new memory management architecture.
2025-08-22 09:20:41 -04:00
lihaoyun6 1cfde3d1c7 Added multi-gpu support for ComfyUI nodes 2025-08-15 21:06:16 +08:00
lihaoyun6 176562f829 Restored compatibility with FP8 safetensors for MPS backend; Changed the way to detect MPS device 2025-08-13 00:53:05 +08:00
lihaoyun6 3b0699ead9 Merge remote-tracking branch 'upstream/nightly' 2025-08-12 22:28:36 +08:00
Adrien Toupet 4716d1820e Removed ComfyUI dependency + improved VRAM clearing between batch 2025-08-08 00:53:48 +02:00
Adrien Toupet 6d8c278a4e Fix QC bugs post updates 2025-08-08 00:53:19 +02:00
Adrien Toupet 1da2cdaf12 Fix VRAM memory leak when using cache_model, added log_memory_state detailed_tensors debug output for troubleshooting 2025-08-08 00:51:02 +02:00
Adrien Toupet d075dde799 feat: Extract model caching to main node + add hierarchical debug system
- Move cache_model and enable_debug options from BlockSwap to main SeedVR2 node

- cache_model: Keep models in RAM between runs (skip reload for faster iterations)

- enable_debug: Show detailed memory/timing tracking with hierarchical display

- Replace hardcoded debug statements with unified Debug class

- Add visual icons for log categorization

- Implement parent-child timing relationships for operation breakdown

- Add VRAM/RAM usage tracking with before/after comparisons

Model caching now available for all workflows, not just BlockSwap users.

Debug system provides granular performance insights when enabled.
2025-08-08 00:49:46 +02:00
lihaoyun6 3b530dc983 Added MPS backend support (for running on macOS) 2025-08-06 18:52:33 +08:00
NumZ 9a129bfacf Update memory_manager.py for stand alone 2025-07-29 17:19:38 +02:00
Adrien Toupet 68f7e913df Avoid debug logging/operations with debug false 2025-07-10 13:43:34 -04:00
Adrien Toupet c88c9e61c1 Blockswap with better memory management 2025-07-09 03:53:24 -04:00
Adrien Toupet 50dfb771e6 Add BlockSwap support 2025-07-04 21:28:47 -04:00
NumZ eb05b8ad26 fix import 2025-07-01 13:33:50 +02:00
NumZ 2786c05faa Speed Up 30 to 50%, fix Memory Leak, refacto) 2025-06-30 12:46:41 +02:00