Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.
The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.
- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow
Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
- Add _device_str() helper to normalize MPS variants (mps:0 → MPS)
- Fix device comparison: mps:0 and mps now correctly identified as same device
- Consistent MPS logging across all memory management functions
Replaces torch.mps.is_available() with torch.backends.mps.is_available()
across all files. The torch.backends API is the official PyTorch method
for MPS detection (since PyTorch 1.12) and works reliably on:
- macOS with Apple Silicon (returns True when MPS available)
- macOS without MPS support (returns False)
- Windows/Linux (hasattr guard prevents AttributeError)
Core Fixes:
- Reset seed per batch to ensure deterministic generation across sessions and batch positions
- Fix temporal overlap logging when automatically reset to prevent incorrect frame counts
- Fix NoneType attribute error in VAE tiled encode/decode at maximum resolution (#296)
BlockSwap & Caching Architecture:
- Move BlockSwap state (_block_swap_config, _blockswap_bypass_protection) from runner to model
- Ensures state survives runner recreation during independent DiT/VAE caching scenarios
- Fix runner template caching to trigger when either DiT or VAE becomes cached (bidirectional)
- Resolves BlockSwap reload failures when only DiT was cached (#297)
Model Discovery:
- Implement case-insensitive YAML path resolution for extra_model_paths.yaml (#289-#295)
- Add debug logging for model discovery (searched paths, validation status, cache hits)
- Support any case variation (seedvr2, SEEDVR2, SeedVR2) in ComfyUI configuration
- Replace all_transformed_videos storage with lightweight batch_metadata indices
- Reconstruct transformed videos on-demand in Phase 4 only when needed
- Add missing cleanup for input_images tensor in postprocess finally block
- Fix release_tensor_memory to handle CPU/CUDA/MPS consistently
- Extract helper functions for batch preparation and 4n+1 padding
- Remove duplicate interrupt_fn key from context initialization
Add defensive checks for torch.mps.is_available() to handle PyTorch versions where the method doesn't exist on non-Mac platforms. Resolves AttributeError: module 'torch.mps' has no attribute 'is_available'
BlockSwap now uses minimal, pressure-based cache clearing instead of forced deep cleans:
- Clear only when VRAM < 5% free (was 15%, more aggressive threshold)
- Use minimal GPU cache clear (deep=False) for reduced overhead
- Applied after each block/IO swap when pressure detected
Also fixes peft dependency: Add explicit peft>=0.17.0 in requirements for diffusers>=0.33.1 compatibility in ComfyUI ecosystem where other plugins may install outdated peft versions.
- VAE: Add tensor_offload_device support for tiled encode/decode accumulation buffers to reduce VRAM usage
- BlockSwap: Re-enable adaptive memory clearing under pressure to prevent OOM at high resolutions
- Memory: Fix offload_target device type error when set to 'none'
- Interface: Rename 'pixels' input to 'image' for ComfyUI convention alignment
- Interface: Fix singular/plural frame text display
- Optimize color operations for VRAM efficiency with safe precision handling
- Limit clear_memory calls to OOM retry and final cleanup only
- Fix model cleanup regression ensuring proper GPU memory release
- Switch to deterministic VAE encoding (mode vs sample) to eliminate high-resolution noise artifacts
- Make VAE encode tiling optional (disabled by default) since deterministic encoding resolves artifacts
- Add tile debug visualization feature with adaptive scaling and color-coded boundaries
- Remove redundant dtype conversions in VAE code for better performance
- Minor code cleanup and documentation update
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
* VAE encoding: seed+1M for deterministic sampling without quality loss
* DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id
Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)
- Add indent_level param to debug.log() and manage_tensor()
- Replace all hardcoded spaces with indent_level (0/1/2)
- Fix encode_all_batches memory flow: encode before storage
- Minor whitespace cleanup
Major improvements to device handling, model offloading, and code clarity:
Device Management:
- Replace string-based devices with torch.device objects throughout
- Add explicit offload device parameters: dit_offload_device, vae_offload_device, tensor_offload_device
- Remove preserve_vram in favor of explicit offload control
- Improve get_device() to return torch.device objects consistently
- Enhance get_device_list() with smart MPS-only system handling
Context & Pipeline:
- Merge setup_device_environment and prepare_generation_context into single setup_generation_context
- Simplify LOCAL_RANK handling (set to '0' for single-GPU mode)
- Store device configuration on runner for submodule access
Tensor & Model Management:
- Add manage_tensor_device() for consistent tensor movement with logging
- Update manage_model_device() to use torch.device objects
- Add validation for BlockSwap and caching configurations
- Rename cache_in_ram to cache_model (more accurate naming)
CLI & Interface:
- Remove --preserve_vram flag (breaking change)
- Add --vae_offload_device and --tensor_offload_device flags
- Update ComfyUI node parameters with validation and better tooltips
- Improve device selection UI with offload_device options
Major refactor introducing:
- GlobalModelCache for cross-node model sharing with dynamic config updates
* Models cached by node ID, reused across different upscaler instances
* Config changes (torch.compile, BlockSwap, tiling) handled dynamically
* Runner templates cached when both models present
- Complete alpha channel processing rewrite
* New edge-guided alpha upscaling in src/core/alpha_upscaling.py
* Remove broken VAE RGBA adapter approach
* Alpha properly extracted, upscaled, and merged with RGB
- BlockSwap enhancements
* Support I/O-only swapping without transformer block offloading
* Separate memory reporting for I/O components vs transformer blocks
* Clearer logging of GPU vs CPU placement
* New is_blockswap_enabled() utility
- Memory management improvements
* Add release_tensor_collection() for batch tensor cleanup
* Proper tensor memory release throughout pipeline
* Better cleanup in postprocess phase
- Code quality and refactoring
* Extensive docstrings with type hints across all modules
* Generic _update_model_config() reduces code duplication
* Better function signatures and parameter documentation
* Improved separation of concerns in model_manager.py
* Clearer debug logging categories and messages
Major features:
- RGBA/Alpha channel support throughout upscaling pipeline
* New VAE RGBA adapter for dynamic 3→4 channel conversion
* Detect RGBA input and adapt VAE after weight loading
* Preserve alpha during encoding/decoding phases
* Split alpha for RGB-only color correction, then recombine
* ComfyUI node now accepts mask input and outputs mask
* Frame count validation between pixels and mask
Configuration management:
- Smart in-place config updates without model reload
- Compare BlockSwap and torch.compile configs for changes
- Cleanup old materialized models before creating new meta structures
- Preserve config attributes (_dit_compile_args, _vae_compile_args, _dit_block_swap_config) for change detection
- Better config change logging with descriptive before/after states
- Handle model reloads separately from config-only updates
Logging improvements:
- Consistent 2-space indentation for hierarchical log messages
- More detailed configuration change descriptions
- Better visibility of RGBA processing steps
- Improved batch optimization tips formatting
Memory and cleanup:
- Handle meta device properly in cleanup functions
- Consistent attribute cleanup across dit/vae
- Clear compile/blockswap configs only when models not cached
- Improved model movement checks for meta device
VAE torch.compile enhancements:
- Disable compilation for InflatedCausalConv3d (dynamic shapes)
- Apply to both encoder and decoder to prevent recompilation issues
- Fixed critical bug where BlockSwap configuration changes were not applied to cached models
- Apply BlockSwap immediately in _handle_blockswap_config instead of deferring to materialization phase
- Renamed offload_io_components to swap_io_components across entire codebase for consistency
- Removed unused _pending_blockswap_config attribute
Core Features:
- Add torch.compile support for DiT (20-40% speedup) and VAE (15-25% speedup)
- Ensure BlockSwap compatibility by applying BlockSwap before torch.compile
- Add SeedVR2TorchCompileSettings node for ComfyUI configuration
- Add CLI arguments: --compile_dit, --compile_vae, --compile_backend, --compile_mode, --compile_fullgraph, --compile_dynamic, --compile_dynamo_cache_size_limit, --compile_dynamo_recompile_limit
Optimizations:
- Optimize na.py and other backend files for torch.compile compatibility (replace .tolist() with tensor operations)
- Remove .item() calls from attention modules (max_seqlen_q/k now accept tensors)
- Add @torch._dynamo.disable decorators to timing/debug methods to prevent compilation warnings
- Replace torch.repeat with torch.repeat_interleave in modulation.py for better performance
Bug Fixes:
- Fix timing stack cleanup in memory_manager.py (end timer before early returns)
- Fix BlockSwap timing reporting (_get_swap_start_time and _log_swap_timing now properly excluded from compilation)
Code Quality:
- Add comprehensive type hints and docstrings across all modified modules
- Standardize BlockSwap timer names (blockswap_block_*, blockswap_io_*)
BREAKING CHANGE: Replaced BlockSwap and ExtraArgs nodes with DiT/VAE loader nodes.
Enables multi-GPU placement, independent caching, and granular memory control.
New Features:
- Add SeedVR2LoadDiTModel and SeedVR2LoadVAEModel loader nodes
- Support different devices for DiT and VAE (multi-GPU load balancing)
- Independent model caching (cache_model_dit/cache_model_vae)
- Separate encode/decode VAE tiling with independent tile configurations
- User-defined DiT and VAE model selection from registry or disk
Breaking Changes:
- Deprecated SeedVR2BlockSwap node (integrated into DiT loader)
- Deprecated SeedVR2ExtraArgs node (split into DiT/VAE loaders & Upscaler node)
- BlockSwap configuration now part of DiT loader node
Improvements:
- Follow ComfyUI conventions (pixels input, IMAGE output type)
- Update download_weight() to support user-defined DiT/VAE models
- Improve tooltips and default parameters across all nodes
- Set show_tensors=False in debug logging for performance
- Dual-device context tracking throughout pipeline
- CLI updated to match new dual-model API
- VAE loads at encoding phase, DiT loads at upscaling phase
- BlockSwap config changes for cached models deferred to DiT phase
- Add comprehensive cleanup of temporary loading attributes
- Fix critical issue where BlockSwap blocks remained on GPU when model caching enabled
- Restore BlockSwap summary display (regression fix)
- Improve model change detection to only cleanup DiT instead of full cleanup
- Standardize _model_name attribute and cleanup flow
- Add Phase 4 post-processing to timing breakdown
- Change DiT final memory clear to deep=True for better cleanup
- Split decode_all_batches into decode (Phase 3) and postprocess_all_batches (Phase 4)
- Add phase-specific cleanup functions: cleanup_dit() and cleanup_vae()
- Each phase now handles its own resource cleanup in finally blocks
- Add cleanup_text_embeddings() helper to eliminate code duplication
- Remove redundant cleanup from comfyui_node normal flow
- Update pipeline from 3-phase to 4-phase architecture (encode → upscale → decode → postprocess)
- Improve memory efficiency by releasing resources immediately when no longer needed
- Update module docstrings to reflect new architecture
- Add suppress_tensor_warnings() to constants.py to centralize tensor/numpy warning handling
- Fix type hints and List/Tuple imports
- Improve GGUFTensor.__torch_function__ to prevent recursion, add debug null checks
- Fix GGUFTensor.to() to properly preserve tensor_shape attribute
- Enhance GGUF error handling with proper exceptions instead of just warnings
- Simplify gguf_ops dequantize_weight by removing redundant fallback paths
- Add memory logging and proper CUDA sync cleanup in GGUF loading
- Translate French comments to English in euler.py
- Remove unnecessary GGUF special handling in memory_manager
- Fix _propagate_debug_to_modules to handle None debug instance
- Implement GGUF model loading with Q3_K_M through Q8_K_M quantization support
- Add GGUFTensor wrapper to preserve quantization and enable on-demand dequantization
- Maintain tensors in quantized format to reduce VRAM usage
- Add GGUF dequantization operations for inference
- Update model registry to include GGUF variants for 3B/7B models
- Fix wavelet blur radius limit to prevent OOM at high resolutions (max 1/8 of image dimension)
- Add safety clamp [-1,1] for SDR color range to prevent numerical errors
- This is a WIP commit as some additional cleaning/testing is needed
- Add type hints throughout for better code maintainability
- Changed VAE tiling default to False to restore original behavior
- Fixed BlockSwap bypass mode during cache cleanup to allow proper CPU offloading
- Added runner parameter to all manage_model_device calls for BlockSwap detection
- Removed conditional movement flags, always offload when preserve_vram=True
- Fixed device comparison to use full device strings (cuda:0) not just types
- Added debug logging for BlockSwap device skip scenarios
- Renamed model directories for clarity: dit -> dit_7b, dit_v2 -> dit_3b
- Converted all absolute imports to relative imports throughout codebase
- Removed sys.path.append() manipulations that caused namespace conflicts
- Updated YAML configs to reference renamed model directories
- Simplified model variant detection logic using new directory names
- Standardized function calls with named arguments for better clarity
- Fixed generation context initialization and interrupt handling
This resolves import conflicts with other ComfyUI nodes (e.g., Basic data handling)
that use sys.path manipulation, making the module properly isolated and compatible.
Fixes#29, #114, #136
Major architectural change to minimize model swapping overhead by processing
all batches in three distinct phases instead of sequential per-batch processing:
- Phase 1: Encode all batches with VAE
- Phase 2: Upscale all latents with DiT
- Phase 3: Decode all latents with VAE
Core changes:
- Split monolithic generation_loop into modular functions:
- prepare_generation_context(): Shared state management
- setup_device_environment(): Device configuration
- prepare_runner(): Model loading with cache support
- encode_all_batches(): Batch VAE encoding
- upscale_all_batches(): Batch DiT upscaling
- decode_all_batches(): Batch VAE decoding
- Removed generation_step function (logic integrated into upscale phase)
- Added lazy precision initialization to avoid redundant setup
Performance improvements:
- Pre-allocated lists for memory efficiency
- Better cleanup of intermediate storage between phases
- Added unique timer names to clear_memory() to avoid naming conflicts
- Improved model state management with change detection and caching
UI/UX enhancements:
- Switched to ComfyUI's native ProgressBar with weighted phase progress
- Changed from per-batch FPS to overall average FPS (always visible)
- Improved log clarity with clear phase separators
- Added ASCII art logo to clearly identify SeedVR2 process start
- Better progress tracking with weighted percentages across three phases
Code cleanup:
- Removed deprecated timer_context from Debug class
- Removed unused time imports across multiple files
- Fixed LOCAL_RANK environment variable to handle string conversion properly
- Improved error handling with try/except/finally blocks in all phases
- Change log_memory_state() to default show_tensors=False (opt-in tensor counting), skipping expensive tensor counting during intermediate operations
- Enable detailed tensor analysis only at final cleanup
- Add early return pattern in log() method to avoid string operations when disabled
- Make reset_vram_peak() logging conditional on debug.enabled state
- Remove unused vram_info call in model_manager.py
- Fix NaMMRotaryEmbedding3d to compute only required dimensions instead of maximum (1024x128x128), enabling much higher resolution upscaling with 3B model without running OOM
- Remove slow preinitialize_rope_cache() which was causing bottleneck during model preparation - no longer needed with NaMMRotaryEmbedding3d fixed
- Clean up unnecessary memory clearing calls and improve debug logging clarity
Dtype Consistency:
- Pass compute_dtype through entire diffusion pipeline to ensure timesteps use bfloat16 instead of float32, reducing memory usage and eliminating dtype conversion overhead
- Propagate dtype from generation_loop → configure_diffusion → sampling timesteps creation
Code Deduplication:
- Removed duplicated code and made use of existing reusable functions: prepare_video_transforms(), load_text_embeddings(), calculate_optimal_batch_params()
- Improve calculate_optimal_batch_params() to prioritize temporal stability by recommending largest valid 4n+1 batch size
Memory Optimizations:
- Remove unnecessary GPU↔CPU offloading for text embeddings and timesteps during preserve_vram mode (added overhead without meaningful memory benefits)
- Simplify text embeddings cleanup to work directly with dictionary structure
- Add cuBLAS workspace clearing for better GPU memory management
Code Cleanup:
- Remove redundant self.runner = None in _internal_execute (handled by cleanup())
- Enhance user messaging for batch padding waste with clearer explanations
- Add bypass mechanism to temporarily disable BlockSwap protection during offload
- Refactor manage_model_device() to handle BlockSwap models transparently
- Restore blocks to correct devices (partial GPU config) on reload
- Simplify preserve_vram logic in infer.py by removing BlockSwap checks
This allows BlockSwap to work correctly with preserve_vram, offloading the entire model to CPU after inference and restoring only necessary blocks to GPU before next inference.
- Implement retry_on_oom() helper with single retry for all VAE operations (Upsample3D, ResnetBlock3D, InflatedCausalConv3d, GroupNorm)
- Fix regression: use BFloat16 compute/autocast for FP16 models to prevent black frames
- Improve frame padding logs to clarify model constraint (frames % 4 == 1)
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
- Merge configure_dit_model_inference and configure_vae_model_inference into single configure_model_inference function
- Simplify load_quantized_state_dict by removing unnecessary FP8 conversion logic
- Remove redundant gradient checkpointing call (training-only feature)
- Remove redundant model_weight variables
- Standardize float formatting to .2f across all debug logs for consistency
- Eliminate redundant device movement operations already handled during model creation
Updated manage_model_device() to clear InflatedCausalConv3d memory buffers
when moving VAE to CPU. These buffers were not automatically released with
.to('cpu'), causing intermediate tensors to remain in VRAM after batch processing.
Preserve BlockSwap's GPU/CPU distribution when cache_model=True to avoid unnecessary reconfiguration on subsequent runs. Also fixes regression error where tensors were not on the correct device after caching.
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state() formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
This is a work-in-progress commit that consolidates a series of changes to fix memory leaks, optimize VRAM/RAM usage, improve performance, and enhance code maintainability.
* **BlockSwap Pinned Memory:** Disabled `use_non_blocking=True` for CPU-to-GPU transfers to resolve a memory leak where pinned memory was not being released.
* **Logging-Induced Leaks:** Modified `log_memory_state()` to avoid holding references to tensors during analysis and added a history limit to the checkpoint system to prevent unbounded memory growth.
* **Incomplete Model Cleanup:** Ensured models are completely deleted and their tensor storage is released when `cache_model=False`.
* **Lingering Tensors:** Fixed an issue where a scalar tensor from sampling timesteps and text embeddings remained on the GPU between batches when `preserve_vram` is active.
* **Centralized Cleanup Functions:** Introduced `clear_memory()` to replace `clear_vram_cache()` and all manual `torch.cuda.empty_cache()` calls, providing consistent VRAM/RAM cleanup logic. The function features a `full` parameter to distinguish between a fast, GPU-only cache clear (~1-5ms) for frequent operations and a full cleanup with garbage collection (~10-50ms) for critical stages.
* **Direct-to-CPU Model Loading:** Modified DiT/VAE weight loading to load directly onto the CPU when `preserve_vram` or `BlockSwap` is active, avoiding unnecessary VRAM spikes during model preparation.
* **VAE Device Management:** Created the `manage_vae_device()` helper function to centralize the logic for moving the VAE between the CPU and GPU, reducing code duplication. This also fixed a bug that incorrectly kept the VAE on the GPU when `preserve_vram` was active.
* **CPU Offloading:** Implemented logic to move text embeddings and sampling timesteps to the CPU after each batch when `preserve_vram` is active, reducing idle VRAM usage.
* **VAE Decode Performance:** Replaced proactive, frequent memory clearing during VAE decode with a reactive Out-of-Memory (OOM) handling system. This fixed a significant performance regression and eliminated the need for the `keep_vae_loaded_during_decode` flag.
* **Reduced Overhead:** Removed redundant `gc.collect()` calls from multiple locations to decrease unnecessary processing overhead.
* **Interval-Based VRAM Tracking:** Modified the logging system to reset peak VRAM statistics after each `log_memory_state()` call, enabling accurate tracking of peak memory usage for specific processing intervals (e.g., encode, inference, decode).
* **Accurate RAM Monitoring:** Added the `get_ram_usage()` function for correct process-specific RAM tracking.
* **Efficient Log Refactoring:** Refactored `log_memory_state()` into modular helper methods, optimizing tensor analysis into a single-pass `gc` iteration to improve both performance and maintainability.
* **Log Clarity:** Refined memory state and debug logging to remove redundant snapshots and add new ones for critical operations like model loading, weight loading, VAE encoding, and decoding. Standardized log message conventions.
* **Per-Batch Timers:** Implemented timer namespacing to ensure that performance timers for each batch are logged correctly without overwriting one another.
* **Error Handling:** Added `try/except` blocks to key memory and device management functions to handle edge cases and improve robustness.
* **Code Cleanup:** Removed deprecated code and outdated comments throughout the related modules.
* **Documentation:** Updated comments and function docstrings to reflect the new memory management architecture.