- Store latents and transformed videos on CPU between processing phases reducing VRAM usage to enable larger batch processing
- Make transformed video storage conditional (only when color_correction != none)
- Add all_ori_lengths tracking for consistent trimming
- Use non_blocking=False for CPU-GPU transfers to avoid pinned memory issues
Fixes https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/pull/163#issuecomment-3333686784
Model Registry Integration:
- Replace hardcoded CLI model choices with dynamic get_available_models()
- Use centralized DEFAULT_MODEL constant from model_registry
- Achieve single source of truth between CLI and ComfyUI interfaces
GGUF Optimizations:
- Create tensors directly on target device to avoid CPU->GPU copy overhead
- Remove excessive memory cleanup during tensor loading for better performance
- Always use meta init for memory-efficient model creation
- Implement precision-optimized dequantization path (GGUF → FP16 → compute dtype)
GGUF Precision Handling:
- Add GGUFQuantizedLinear/Conv2d layers with get_dequantized_weight_for_compute() to preserve precision
- Track and report quantization types (Q4_K_M, Q5_K_M, etc.) in model
- Fix __torch_function__ as classmethod to resolve deprecation warning
Code Quality:
- Restructure _load_model_weights() with modular helper functions to reduce duplication
- Improve separation of concerns between standard and GGUF weight loading
- Enhance logging to always display WARNING/ERROR messages
- Add comprehensive GGUF architecture validation
- Remove debug-only log_memory_state calls
- Fix type function definitions
- Add suppress_tensor_warnings() to constants.py to centralize tensor/numpy warning handling
- Fix type hints and List/Tuple imports
- Improve GGUFTensor.__torch_function__ to prevent recursion, add debug null checks
- Fix GGUFTensor.to() to properly preserve tensor_shape attribute
- Enhance GGUF error handling with proper exceptions instead of just warnings
- Simplify gguf_ops dequantize_weight by removing redundant fallback paths
- Add memory logging and proper CUDA sync cleanup in GGUF loading
- Translate French comments to English in euler.py
- Remove unnecessary GGUF special handling in memory_manager
- Fix _propagate_debug_to_modules to handle None debug instance
- Implement GGUF model loading with Q3_K_M through Q8_K_M quantization support
- Add GGUFTensor wrapper to preserve quantization and enable on-demand dequantization
- Maintain tensors in quantized format to reduce VRAM usage
- Add GGUF dequantization operations for inference
- Update model registry to include GGUF variants for 3B/7B models
- Fix wavelet blur radius limit to prevent OOM at high resolutions (max 1/8 of image dimension)
- Add safety clamp [-1,1] for SDR color range to prevent numerical errors
- This is a WIP commit as some additional cleaning/testing is needed
- Add type hints throughout for better code maintainability
- Implement model discovery across multiple paths via ComfyUI's folder_paths API
- Add centralized model file scanning in get_all_model_files() to avoid duplication
- Update model loading to search all registered paths before downloading
- Prevent redundant downloads when files exist in alternate paths
- Users can now specify custom model directories using 'seedvr2' key in extra_model_paths.yaml
Addresses request from https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/issues/128#issuecomment-3312407641
- Clean up all __init__.py files to remove commented code and unnecessary imports
- Root __init__.py now only exports NODE_CLASS_MAPPINGS and NODE_DISPLAY_NAME_MAPPINGS
- Remove unused dependencies: pytorch-extension (replaced by custom implementations) and peft
- Fix diffusers version syntax (remove 'v' prefix)
- Add pyproject.toml for ComfyUI registry compliance
- Always use meta device initialization to avoid double memory allocation
- Load weights directly to target device with assign=True for in-place materialization
- Initialize non-persistent buffers (RoPE frequencies) after materialization
- Use strict=False universally in load_state_dict to handle missing buffers
- Fix BlockSwap config changes not being detected when cache_model=True
* Use tuple comparison for configs: (blocks, offload_io)
* Handle all transitions: enabled↔disabled, config changes, disconnection
* Show readable format in logs: blocks=14, offload_io=True
- Fix progress bar stuck after errors by clearing it in cleanup()
- Fix BlockSwap cleanup by checking any state existence, not just _blockswap_active
* Properly cleans up deactivated BlockSwap with remaining config
- Add input_noise_scale (0-1) to reduce artifacts at high resolutions
- Applies subtle noise before VAE encoding with progressive blend (0-50%)
- Based on GitHub issue #64 community findings
- Rename cond_noise_scale to latent_noise_scale for clarity
- Better distinguishes between input (pixel) and latent (diffusion) noise
- Maintains same functionality, just clearer naming
- Update both ComfyUI and CLI interfaces with new parameters
- Changed VAE tiling default to False to restore original behavior
- Fixed BlockSwap bypass mode during cache cleanup to allow proper CPU offloading
- Added runner parameter to all manage_model_device calls for BlockSwap detection
- Removed conditional movement flags, always offload when preserve_vram=True
- Fixed device comparison to use full device strings (cuda:0) not just types
- Added debug logging for BlockSwap device skip scenarios
- Add color_correction parameter (wavelet/adain/none) to both ComfyUI node and CLI interface to choose between wavelet reconstruction (frequency-based), AdaIN (statistical matching), or no color correction
- Improve memory management after upscaling batches with targeted clear_memory calls to reduce VRAM pressure during VAE decoding
- Renamed model directories for clarity: dit -> dit_7b, dit_v2 -> dit_3b
- Converted all absolute imports to relative imports throughout codebase
- Removed sys.path.append() manipulations that caused namespace conflicts
- Updated YAML configs to reference renamed model directories
- Simplified model variant detection logic using new directory names
- Standardized function calls with named arguments for better clarity
- Fixed generation context initialization and interrupt handling
This resolves import conflicts with other ComfyUI nodes (e.g., Basic data handling)
that use sys.path manipulation, making the module properly isolated and compatible.
Fixes#29, #114, #136
Major architectural change to minimize model swapping overhead by processing
all batches in three distinct phases instead of sequential per-batch processing:
- Phase 1: Encode all batches with VAE
- Phase 2: Upscale all latents with DiT
- Phase 3: Decode all latents with VAE
Core changes:
- Split monolithic generation_loop into modular functions:
- prepare_generation_context(): Shared state management
- setup_device_environment(): Device configuration
- prepare_runner(): Model loading with cache support
- encode_all_batches(): Batch VAE encoding
- upscale_all_batches(): Batch DiT upscaling
- decode_all_batches(): Batch VAE decoding
- Removed generation_step function (logic integrated into upscale phase)
- Added lazy precision initialization to avoid redundant setup
Performance improvements:
- Pre-allocated lists for memory efficiency
- Better cleanup of intermediate storage between phases
- Added unique timer names to clear_memory() to avoid naming conflicts
- Improved model state management with change detection and caching
UI/UX enhancements:
- Switched to ComfyUI's native ProgressBar with weighted phase progress
- Changed from per-batch FPS to overall average FPS (always visible)
- Improved log clarity with clear phase separators
- Added ASCII art logo to clearly identify SeedVR2 process start
- Better progress tracking with weighted percentages across three phases
Code cleanup:
- Removed deprecated timer_context from Debug class
- Removed unused time imports across multiple files
- Fixed LOCAL_RANK environment variable to handle string conversion properly
- Improved error handling with try/except/finally blocks in all phases
- Set VAE to BFloat16 on MPS to prevent dtype mismatch errors
- Add override_dtype parameter to model loading pipeline
- Convert VAE weights to target dtype during load for efficiency
### 🚀 Performance Optimizations
#### Model Initialization & Loading
- **Faster model initialization** using meta device to avoid unnecessary memory allocation during DiT and VAE creation
- **Memory-mapped VAE loading** with `assign=True` to reference tensors directly instead of copying, matching DiT's loading strategy
- **Optimized state dict management** with explicit deletion after loading to free memory immediately
- **Removed slow preinitialize_rope_cache()** bottleneck during model preparation
#### Generation Pipeline
- **Streamlined batch processing** by merging redundant log steps and unifying "Model Configuration" and "Input Preparation" into single "Generation Setup" step
- **Optimized memory usage** by passing compute_dtype through entire diffusion pipeline, ensuring timesteps use bfloat16 instead of float32
- **Code deduplication** by reusing existing functions: `prepare_video_transforms()`, `load_text_embeddings()`, `calculate_optimal_batch_params()`
- **Improved batch recommendations** with `calculate_optimal_batch_params()` now prioritizing temporal stability by suggesting largest valid 4n+1 batch size
#### VAE Operations
- **Optimized tile blending** with pre-computed cosine ramps instead of recalculating per-tile
- **In-place operations** using `mul_` and `div_` to reduce memory allocations
- **Reduced redundant transfers** by skipping device moves when tensors are already on correct device
- **Improved slicing operations** with consistent memory cleanup for both encode and decode paths
### 💾 Memory Management
#### Critical Fixes
- **Fixed major VRAM leak** in VAE decoder by clearing InflatedCausalConv3d memory buffers when moving VAE to CPU
- **Fixed RoPE OOM issues** by computing only required dimensions instead of maximum (1024x128x128), enabling higher resolution upscaling
- **Added OOM retry mechanism** for VAE operations (Upsample3D, ResnetBlock3D, InflatedCausalConv3d, GroupNorm) with single retry on failure
- **Added cuBLAS workspace clearing** for better GPU memory management
#### BlockSwap Integration
- **Full preserve_vram compatibility** with bypass mechanism to temporarily disable BlockSwap protection during offload
- **Transparent handling** in `manage_model_device()` for BlockSwap models
- **Proper state restoration** of blocks to correct devices (partial GPU config) on reload
- **Simplified preserve_vram logic** by removing BlockSwap checks from infer.py
### 🏗️ Architecture Refactoring
#### Code Organization
- **Unified model configuration** by merging `configure_dit_model_inference` and `configure_vae_model_inference` into single `configure_model_inference` function
- **Focused helper functions** extracted from `configure_model_inference()`
- **Removed redundant code**: gradient checkpointing (training-only), model_weight variables, platform-specific VAE dtype logic
- **Simplified FP8 logic** by removing unnecessary conversion code from `load_quantized_state_dict`
#### Dtype Handling
- **Unified dtype flow**: VAE now uses configured dtype consistently, removing redundant vae_dtype variable
- **Fixed FP16 black frames** by using BFloat16 compute/autocast for FP16 models
- **Natural dtype propagation** from generation_loop → configure_diffusion → sampling, eliminating manual conversions
- **Removed nested autocast contexts** relying on outer context from generation_step
### 🐛 Bug Fixes & Improvements
#### Debug & Logging
- **Optimized debug performance** by making tensor counting opt-in (`show_tensors=False` by default)
- **Conditional logging** with `reset_vram_peak()` only logging when debug is enabled
- **Early return pattern** in log() method to avoid string operations when disabled
- **Standardized formatting**: consistent .2f float precision, uppercase device names, unified category names
- **Improved clarity**: better frame padding logs, clearer FP8 RoPE conversion messages
- Change log_memory_state() to default show_tensors=False (opt-in tensor counting), skipping expensive tensor counting during intermediate operations
- Enable detailed tensor analysis only at final cleanup
- Add early return pattern in log() method to avoid string operations when disabled
- Make reset_vram_peak() logging conditional on debug.enabled state
- Remove unused vram_info call in model_manager.py
- Add InflatedCausalConv3d memory cleanup to slicing_decode to match slicing_encode
- Skip redundant device transfers when tensors are already on correct device
- Pre-compute cosine ramps for tile blending instead of recalculating per-tile
- Use in-place operations (mul_, div_) to reduce memory allocations
- Reduce debug logging frequency and clarify tile progress messages
- Filter modules with memory before clearing to avoid unnecessary iterations
Minor performance improvements with no memory increase.
- Fix NaMMRotaryEmbedding3d to compute only required dimensions instead of maximum (1024x128x128), enabling much higher resolution upscaling with 3B model without running OOM
- Remove slow preinitialize_rope_cache() which was causing bottleneck during model preparation - no longer needed with NaMMRotaryEmbedding3d fixed
- Clean up unnecessary memory clearing calls and improve debug logging clarity
Dtype Consistency:
- Pass compute_dtype through entire diffusion pipeline to ensure timesteps use bfloat16 instead of float32, reducing memory usage and eliminating dtype conversion overhead
- Propagate dtype from generation_loop → configure_diffusion → sampling timesteps creation
Code Deduplication:
- Removed duplicated code and made use of existing reusable functions: prepare_video_transforms(), load_text_embeddings(), calculate_optimal_batch_params()
- Improve calculate_optimal_batch_params() to prioritize temporal stability by recommending largest valid 4n+1 batch size
Memory Optimizations:
- Remove unnecessary GPU↔CPU offloading for text embeddings and timesteps during preserve_vram mode (added overhead without meaningful memory benefits)
- Simplify text embeddings cleanup to work directly with dictionary structure
- Add cuBLAS workspace clearing for better GPU memory management
Code Cleanup:
- Remove redundant self.runner = None in _internal_execute (handled by cleanup())
- Enhance user messaging for batch padding waste with clearer explanations
- Add bypass mechanism to temporarily disable BlockSwap protection during offload
- Refactor manage_model_device() to handle BlockSwap models transparently
- Restore blocks to correct devices (partial GPU config) on reload
- Simplify preserve_vram logic in infer.py by removing BlockSwap checks
This allows BlockSwap to work correctly with preserve_vram, offloading the entire model to CPU after inference and restoring only necessary blocks to GPU before next inference.
- Use meta device initialization to avoid unnecessary memory allocation during model creation
- Reduces DiT and VAE initialization time by ~90% when loading to CPU
- Explicitly delete state dicts after loading to free memory immediately
- Refactor configure_model_inference() into focused helper functions for DRY code
- Improve debug logging with timestamps and cleaner section separators for readability
- Implement retry_on_oom() helper with single retry for all VAE operations (Upsample3D, ResnetBlock3D, InflatedCausalConv3d, GroupNorm)
- Fix regression: use BFloat16 compute/autocast for FP16 models to prevent black frames
- Improve frame padding logs to clarify model constraint (frames % 4 == 1)
- Remove hardcoded target_dtype and manual tensor conversions in inference()
- Remove redundant inner autocast context - rely on outer context from generation_step
- Align compute_dtype and autocast_dtype for FP16 models (both now float16)
- Remove unused block_swap_config parameter
- Fix device logging to properly display device names
The dtype is now determined once in generation_loop based on model weights and flows naturally through the pipeline via the outer autocast context, eliminating unnecessary conversions and improving performance.
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
- Merge configure_dit_model_inference and configure_vae_model_inference into single configure_model_inference function
- Simplify load_quantized_state_dict by removing unnecessary FP8 conversion logic
- Remove redundant gradient checkpointing call (training-only feature)
- Remove redundant model_weight variables
- Standardize float formatting to .2f across all debug logs for consistency
- Eliminate redundant device movement operations already handled during model creation
- Match DiT's load_state_dict(assign=True) for VAE to reference memory-mapped tensors directly instead of copying, reducing peak RAM
- Move post-inference memory logging after offloading for accurate metrics
Updated manage_model_device() to clear InflatedCausalConv3d memory buffers
when moving VAE to CPU. These buffers were not automatically released with
.to('cpu'), causing intermediate tensors to remain in VRAM after batch processing.
Preserve BlockSwap's GPU/CPU distribution when cache_model=True to avoid unnecessary reconfiguration on subsequent runs. Also fixes regression error where tensors were not on the correct device after caching.
Preserve BlockSwap's GPU/CPU distribution when cache_model=True to avoid unnecessary reconfiguration on subsequent runs. Also fixes regression error where tensors were not on the correct device after caching.
Regression from previous commit - cleanup_blockswap was removing forward wrappers/RoPE patches
but leaving blocks in position, then re-applying BlockSwap duplicated memory allocations
causing OOM on second generation.
Now when caching with same config:
- Blocks remain in their GPU/CPU positions for next inference
- Forward wrappers and RoPE patches stay intact
- Only toggles _blockswap_active flag for safety
- Reuses exact same VRAM allocation as previous run
Full cleanup only happens when cache_model=False or config changes.
Fixes OOM errors and eliminates redundant GPU transfers between cached generations.
#### Memory Management & Leak Fixes
- **Fixed BlockSwap pinned memory leak**: Removed `non_blocking` parameter that was preventing pinned memory release
- **Eliminated logging-induced memory leaks**: Modified `log_memory_state()` to avoid holding tensor references and added checkpoint history limits
- **Complete model cleanup**: Ensured proper deletion and tensor storage release when `cache_model=False`
- **Fixed lingering GPU tensors**: Resolved issues with scalar tensors and text embeddings remaining on GPU between batches with `preserve_vram`
#### Performance Optimizations
- **Direct-to-CPU model loading**: DiT/VAE weights load directly to CPU when `preserve_vram` or BlockSwap is active, avoiding VRAM spikes
- **Reactive OOM handling**: Replaced proactive memory clearing during VAE decode with reactive OOM recovery, fixing performance regression (removed the need for the keep_vae_loaded_during_decode flag)
- **Optimized cleanup functions**: Consolidated all cleanup into core functions (`manage_*`, `release_*`, `clear_*`)
#### Enhanced Debugging & Monitoring
- **Unified debug system**: Improved logging categories, icons, and hierarchical timers
- **Comprehensive memory tracking**: Per-interval VRAM peak tracking, accurate process-specific RAM monitoring
- **Debug propagation**: Debug instance now properly propagates to all submodules (sampler, VAE, etc.)
- **Improved log clarity**: Standardized message formatting, removed redundant logs, added critical operation snapshots
#### Code Architecture Improvements
- **Type hints **: Added type hints to all memory management functions
- **Centralized device management**: New `manage_model_device()` function for consistent device movement logic
- **Error handling**: Added try/except blocks to all critical memory and device operations
- **Removed private APIs**: Eliminated `torch._C._clear_cache()` usage that was triggering warning
#### Bug Fixes
- Resolved duplicate dtype detection in generation loop
- Corrected CPU offloading logic for text embeddings and sampling timesteps
- Fixed timer namespace collisions in batch processing
- Removed redundant RoPE cache clearing operations
#### Various
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove unused imports and legacy code/comments
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state() formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
This is a work-in-progress commit that consolidates a series of changes to fix memory leaks, optimize VRAM/RAM usage, improve performance, and enhance code maintainability.
* **BlockSwap Pinned Memory:** Disabled `use_non_blocking=True` for CPU-to-GPU transfers to resolve a memory leak where pinned memory was not being released.
* **Logging-Induced Leaks:** Modified `log_memory_state()` to avoid holding references to tensors during analysis and added a history limit to the checkpoint system to prevent unbounded memory growth.
* **Incomplete Model Cleanup:** Ensured models are completely deleted and their tensor storage is released when `cache_model=False`.
* **Lingering Tensors:** Fixed an issue where a scalar tensor from sampling timesteps and text embeddings remained on the GPU between batches when `preserve_vram` is active.
* **Centralized Cleanup Functions:** Introduced `clear_memory()` to replace `clear_vram_cache()` and all manual `torch.cuda.empty_cache()` calls, providing consistent VRAM/RAM cleanup logic. The function features a `full` parameter to distinguish between a fast, GPU-only cache clear (~1-5ms) for frequent operations and a full cleanup with garbage collection (~10-50ms) for critical stages.
* **Direct-to-CPU Model Loading:** Modified DiT/VAE weight loading to load directly onto the CPU when `preserve_vram` or `BlockSwap` is active, avoiding unnecessary VRAM spikes during model preparation.
* **VAE Device Management:** Created the `manage_vae_device()` helper function to centralize the logic for moving the VAE between the CPU and GPU, reducing code duplication. This also fixed a bug that incorrectly kept the VAE on the GPU when `preserve_vram` was active.
* **CPU Offloading:** Implemented logic to move text embeddings and sampling timesteps to the CPU after each batch when `preserve_vram` is active, reducing idle VRAM usage.
* **VAE Decode Performance:** Replaced proactive, frequent memory clearing during VAE decode with a reactive Out-of-Memory (OOM) handling system. This fixed a significant performance regression and eliminated the need for the `keep_vae_loaded_during_decode` flag.
* **Reduced Overhead:** Removed redundant `gc.collect()` calls from multiple locations to decrease unnecessary processing overhead.
* **Interval-Based VRAM Tracking:** Modified the logging system to reset peak VRAM statistics after each `log_memory_state()` call, enabling accurate tracking of peak memory usage for specific processing intervals (e.g., encode, inference, decode).
* **Accurate RAM Monitoring:** Added the `get_ram_usage()` function for correct process-specific RAM tracking.
* **Efficient Log Refactoring:** Refactored `log_memory_state()` into modular helper methods, optimizing tensor analysis into a single-pass `gc` iteration to improve both performance and maintainability.
* **Log Clarity:** Refined memory state and debug logging to remove redundant snapshots and add new ones for critical operations like model loading, weight loading, VAE encoding, and decoding. Standardized log message conventions.
* **Per-Batch Timers:** Implemented timer namespacing to ensure that performance timers for each batch are logged correctly without overwriting one another.
* **Error Handling:** Added `try/except` blocks to key memory and device management functions to handle edge cases and improve robustness.
* **Code Cleanup:** Removed deprecated code and outdated comments throughout the related modules.
* **Documentation:** Updated comments and function docstrings to reflect the new memory management architecture.