Major architectural change to minimize model swapping overhead by processing
all batches in three distinct phases instead of sequential per-batch processing:
- Phase 1: Encode all batches with VAE
- Phase 2: Upscale all latents with DiT
- Phase 3: Decode all latents with VAE
Core changes:
- Split monolithic generation_loop into modular functions:
- prepare_generation_context(): Shared state management
- setup_device_environment(): Device configuration
- prepare_runner(): Model loading with cache support
- encode_all_batches(): Batch VAE encoding
- upscale_all_batches(): Batch DiT upscaling
- decode_all_batches(): Batch VAE decoding
- Removed generation_step function (logic integrated into upscale phase)
- Added lazy precision initialization to avoid redundant setup
Performance improvements:
- Pre-allocated lists for memory efficiency
- Better cleanup of intermediate storage between phases
- Added unique timer names to clear_memory() to avoid naming conflicts
- Improved model state management with change detection and caching
UI/UX enhancements:
- Switched to ComfyUI's native ProgressBar with weighted phase progress
- Changed from per-batch FPS to overall average FPS (always visible)
- Improved log clarity with clear phase separators
- Added ASCII art logo to clearly identify SeedVR2 process start
- Better progress tracking with weighted percentages across three phases
Code cleanup:
- Removed deprecated timer_context from Debug class
- Removed unused time imports across multiple files
- Fixed LOCAL_RANK environment variable to handle string conversion properly
- Improved error handling with try/except/finally blocks in all phases
- Change log_memory_state() to default show_tensors=False (opt-in tensor counting), skipping expensive tensor counting during intermediate operations
- Enable detailed tensor analysis only at final cleanup
- Add early return pattern in log() method to avoid string operations when disabled
- Make reset_vram_peak() logging conditional on debug.enabled state
- Remove unused vram_info call in model_manager.py
- Fix NaMMRotaryEmbedding3d to compute only required dimensions instead of maximum (1024x128x128), enabling much higher resolution upscaling with 3B model without running OOM
- Remove slow preinitialize_rope_cache() which was causing bottleneck during model preparation - no longer needed with NaMMRotaryEmbedding3d fixed
- Clean up unnecessary memory clearing calls and improve debug logging clarity
Dtype Consistency:
- Pass compute_dtype through entire diffusion pipeline to ensure timesteps use bfloat16 instead of float32, reducing memory usage and eliminating dtype conversion overhead
- Propagate dtype from generation_loop → configure_diffusion → sampling timesteps creation
Code Deduplication:
- Removed duplicated code and made use of existing reusable functions: prepare_video_transforms(), load_text_embeddings(), calculate_optimal_batch_params()
- Improve calculate_optimal_batch_params() to prioritize temporal stability by recommending largest valid 4n+1 batch size
Memory Optimizations:
- Remove unnecessary GPU↔CPU offloading for text embeddings and timesteps during preserve_vram mode (added overhead without meaningful memory benefits)
- Simplify text embeddings cleanup to work directly with dictionary structure
- Add cuBLAS workspace clearing for better GPU memory management
Code Cleanup:
- Remove redundant self.runner = None in _internal_execute (handled by cleanup())
- Enhance user messaging for batch padding waste with clearer explanations
- Add bypass mechanism to temporarily disable BlockSwap protection during offload
- Refactor manage_model_device() to handle BlockSwap models transparently
- Restore blocks to correct devices (partial GPU config) on reload
- Simplify preserve_vram logic in infer.py by removing BlockSwap checks
This allows BlockSwap to work correctly with preserve_vram, offloading the entire model to CPU after inference and restoring only necessary blocks to GPU before next inference.
- Implement retry_on_oom() helper with single retry for all VAE operations (Upsample3D, ResnetBlock3D, InflatedCausalConv3d, GroupNorm)
- Fix regression: use BFloat16 compute/autocast for FP16 models to prevent black frames
- Improve frame padding logs to clarify model constraint (frames % 4 == 1)
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
- Merge configure_dit_model_inference and configure_vae_model_inference into single configure_model_inference function
- Simplify load_quantized_state_dict by removing unnecessary FP8 conversion logic
- Remove redundant gradient checkpointing call (training-only feature)
- Remove redundant model_weight variables
- Standardize float formatting to .2f across all debug logs for consistency
- Eliminate redundant device movement operations already handled during model creation
Updated manage_model_device() to clear InflatedCausalConv3d memory buffers
when moving VAE to CPU. These buffers were not automatically released with
.to('cpu'), causing intermediate tensors to remain in VRAM after batch processing.
Preserve BlockSwap's GPU/CPU distribution when cache_model=True to avoid unnecessary reconfiguration on subsequent runs. Also fixes regression error where tensors were not on the correct device after caching.
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state() formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
This is a work-in-progress commit that consolidates a series of changes to fix memory leaks, optimize VRAM/RAM usage, improve performance, and enhance code maintainability.
* **BlockSwap Pinned Memory:** Disabled `use_non_blocking=True` for CPU-to-GPU transfers to resolve a memory leak where pinned memory was not being released.
* **Logging-Induced Leaks:** Modified `log_memory_state()` to avoid holding references to tensors during analysis and added a history limit to the checkpoint system to prevent unbounded memory growth.
* **Incomplete Model Cleanup:** Ensured models are completely deleted and their tensor storage is released when `cache_model=False`.
* **Lingering Tensors:** Fixed an issue where a scalar tensor from sampling timesteps and text embeddings remained on the GPU between batches when `preserve_vram` is active.
* **Centralized Cleanup Functions:** Introduced `clear_memory()` to replace `clear_vram_cache()` and all manual `torch.cuda.empty_cache()` calls, providing consistent VRAM/RAM cleanup logic. The function features a `full` parameter to distinguish between a fast, GPU-only cache clear (~1-5ms) for frequent operations and a full cleanup with garbage collection (~10-50ms) for critical stages.
* **Direct-to-CPU Model Loading:** Modified DiT/VAE weight loading to load directly onto the CPU when `preserve_vram` or `BlockSwap` is active, avoiding unnecessary VRAM spikes during model preparation.
* **VAE Device Management:** Created the `manage_vae_device()` helper function to centralize the logic for moving the VAE between the CPU and GPU, reducing code duplication. This also fixed a bug that incorrectly kept the VAE on the GPU when `preserve_vram` was active.
* **CPU Offloading:** Implemented logic to move text embeddings and sampling timesteps to the CPU after each batch when `preserve_vram` is active, reducing idle VRAM usage.
* **VAE Decode Performance:** Replaced proactive, frequent memory clearing during VAE decode with a reactive Out-of-Memory (OOM) handling system. This fixed a significant performance regression and eliminated the need for the `keep_vae_loaded_during_decode` flag.
* **Reduced Overhead:** Removed redundant `gc.collect()` calls from multiple locations to decrease unnecessary processing overhead.
* **Interval-Based VRAM Tracking:** Modified the logging system to reset peak VRAM statistics after each `log_memory_state()` call, enabling accurate tracking of peak memory usage for specific processing intervals (e.g., encode, inference, decode).
* **Accurate RAM Monitoring:** Added the `get_ram_usage()` function for correct process-specific RAM tracking.
* **Efficient Log Refactoring:** Refactored `log_memory_state()` into modular helper methods, optimizing tensor analysis into a single-pass `gc` iteration to improve both performance and maintainability.
* **Log Clarity:** Refined memory state and debug logging to remove redundant snapshots and add new ones for critical operations like model loading, weight loading, VAE encoding, and decoding. Standardized log message conventions.
* **Per-Batch Timers:** Implemented timer namespacing to ensure that performance timers for each batch are logged correctly without overwriting one another.
* **Error Handling:** Added `try/except` blocks to key memory and device management functions to handle edge cases and improve robustness.
* **Code Cleanup:** Removed deprecated code and outdated comments throughout the related modules.
* **Documentation:** Updated comments and function docstrings to reflect the new memory management architecture.
- Move cache_model and enable_debug options from BlockSwap to main SeedVR2 node
- cache_model: Keep models in RAM between runs (skip reload for faster iterations)
- enable_debug: Show detailed memory/timing tracking with hierarchical display
- Replace hardcoded debug statements with unified Debug class
- Add visual icons for log categorization
- Implement parent-child timing relationships for operation breakdown
- Add VRAM/RAM usage tracking with before/after comparisons
Model caching now available for all workflows, not just BlockSwap users.
Debug system provides granular performance insights when enabled.