- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
- Standardize tooltip format with multi-line descriptions and bullet points
- Add comprehensive output tooltips for all nodes (DiT, VAE, torch.compile, upscaler)
- Enhance node descriptions with detailed capability summaries
- Simplify CLI tile size arguments to single integers (converted internally to tuples)
- Remove OneOrTwoValues argparse class for cleaner implementation
- Fix encode_tiled tooltip (was incorrectly referencing decoding)
- Clarify color correction purpose (corrects upscaling color shifts)
- Add multi-GPU offloading information to all offload_device tooltips
- Improve torch.compile parameter descriptions with use cases
- Ensure CLI and ComfyUI tooltips are consistent in terminology and structure
- Fix validation cache location to respect --model_dir parameter
- Fix output path handling for directories without extensions
- Remove spurious directory creation in get_base_cache_dir
- Enhanced CLI help with usage examples and argument grouping
- Auto-display help when script invoked without arguments
- Correct type hints (device_id: str, debug: Debug)
- Remove redundant type conversions and makedirs calls
- Reorganize imports to module top for clarity
- Improved docstrings & tooltip
- Change default batch_size from 1 to 5 to match ComfyUI integration
- Add support for seedvr2_ema_3b-Q8_0.gguf model
- Validate --cuda_device arguments early in pre-parsing phase
- Check device IDs exist and are within available GPU range
- Fail fast with clear error messages showing available devices
- Auto-detect output format per file type (mp4 for videos, png for images)
- Add visual separators between processed files for better readability
- Simplify FPS calculation to use wall-clock time for real-world throughput
- Consolidate banner/footer into shared Debug methods
- Update offload device args to support multi-GPU (cpu/cuda:N)
- Standardize terminology: 'upscaling' instead of 'video upscaling'
- Remove code duplication between CLI and ComfyUI implementations
- Consistent quote style (double quotes) throughout CLI
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)
Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable
Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability
Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
- Add max_resolution parameter (default: 0 = no limit) to both CLI and ComfyUI
- After new_resolution scales shortest edge, max_resolution ensures no edge exceeds limit
- Scales down proportionally if constraint violated
- Maintains backward compatibility with default value of 0
- Add dit_offload_device parameter for proper blockSwap configuration
- Ensure consistent dtype management throughout CLI and ComfyUI (float32 input with bfloat16 pipeline)
- Translate all French comments to English
- Add comprehensive docstrings and section headers
- Remove obsolete use_non_blocking and enable_debug parameters
- Add error handling and validation
- Implement prepend_video_frames() for artifact reduction at video start
- Add blend_overlapping_frames() with Hann window for smooth transitions
- Expose temporal_overlap (0-16) and prepend_frames (0-32) in ComfyUI node
- Unify CLI and ComfyUI to use shared prepend/overlap functions
- Add comprehensive logging for frame adjustments (prepend/overlap/padding)
CFG (Classifier-Free Guidance) does not work with SeedVR2's distilled
one-step diffusion model. The model was trained to produce final results
in a single step without iterative guidance.
- Remove cfg_scale parameter from ComfyUI node and CLI interface
- Force internal cfg_scale to 1.0 in upscale_all_batches()
This avoids artifacts introduced when users changed cfg_scale away from 1.0.
- Switch to deterministic VAE encoding (mode vs sample) to eliminate high-resolution noise artifacts
- Make VAE encode tiling optional (disabled by default) since deterministic encoding resolves artifacts
- Add tile debug visualization feature with adaptive scaling and color-coded boundaries
- Remove redundant dtype conversions in VAE code for better performance
- Minor code cleanup and documentation update
VAE Changes:
- Enable encode tiling by default (prevents noise artifacts at high resolution)
- Increase tile size to 1024px (down from 512px) for optimal quality
- Increase tile overlap to 128px for better blending
Dtype Pipeline:
- Hardcode compute_dtype to bfloat16 for consistent quality/performance/VRAM balance
- Ensure all pipeline steps are using compute_dtype when relevant
- Refactor code for improved performance and memory management
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
* VAE encoding: seed+1M for deterministic sampling without quality loss
* DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id
Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)
- Add indent_level param to debug.log() and manage_tensor()
- Replace all hardcoded spaces with indent_level (0/1/2)
- Fix encode_all_batches memory flow: encode before storage
- Minor whitespace cleanup
Major improvements to device handling, model offloading, and code clarity:
Device Management:
- Replace string-based devices with torch.device objects throughout
- Add explicit offload device parameters: dit_offload_device, vae_offload_device, tensor_offload_device
- Remove preserve_vram in favor of explicit offload control
- Improve get_device() to return torch.device objects consistently
- Enhance get_device_list() with smart MPS-only system handling
Context & Pipeline:
- Merge setup_device_environment and prepare_generation_context into single setup_generation_context
- Simplify LOCAL_RANK handling (set to '0' for single-GPU mode)
- Store device configuration on runner for submodule access
Tensor & Model Management:
- Add manage_tensor_device() for consistent tensor movement with logging
- Update manage_model_device() to use torch.device objects
- Add validation for BlockSwap and caching configurations
- Rename cache_in_ram to cache_model (more accurate naming)
CLI & Interface:
- Remove --preserve_vram flag (breaking change)
- Add --vae_offload_device and --tensor_offload_device flags
- Update ComfyUI node parameters with validation and better tooltips
- Improve device selection UI with offload_device options
- Move ensure_float32_precision() from alpha_upscaling to common/half_precision_fixes
- Update color_fix.py to use ensure_float32_precision
- Add type hints and docstrings to all half_precision_fixes functions
- Clarify batch size tips: emphasize avoiding padding waste over maximizing batch size
- Update default seed to 42 (the answer to life, the universe, and everything)
- Add LAB color transfer as new default method for superior perceptual color matching
- Add HSV hue-conditional saturation histogram matching for targeted oversaturation correction
- Add wavelet_adaptive hybrid method combining wavelet base with selective HSV correction
- Change default color correction from 'wavelet' to 'lab' for better color accuracy
- Implement full CIELAB color space conversion with D65 illuminant and histogram matching
- Add comprehensive documentation and method descriptions to color_fix.py
Major refactor introducing:
- GlobalModelCache for cross-node model sharing with dynamic config updates
* Models cached by node ID, reused across different upscaler instances
* Config changes (torch.compile, BlockSwap, tiling) handled dynamically
* Runner templates cached when both models present
- Complete alpha channel processing rewrite
* New edge-guided alpha upscaling in src/core/alpha_upscaling.py
* Remove broken VAE RGBA adapter approach
* Alpha properly extracted, upscaled, and merged with RGB
- BlockSwap enhancements
* Support I/O-only swapping without transformer block offloading
* Separate memory reporting for I/O components vs transformer blocks
* Clearer logging of GPU vs CPU placement
* New is_blockswap_enabled() utility
- Memory management improvements
* Add release_tensor_collection() for batch tensor cleanup
* Proper tensor memory release throughout pipeline
* Better cleanup in postprocess phase
- Code quality and refactoring
* Extensive docstrings with type hints across all modules
* Generic _update_model_config() reduces code duplication
* Better function signatures and parameter documentation
* Improved separation of concerns in model_manager.py
* Clearer debug logging categories and messages
- Fixed critical bug where BlockSwap configuration changes were not applied to cached models
- Apply BlockSwap immediately in _handle_blockswap_config instead of deferring to materialization phase
- Renamed offload_io_components to swap_io_components across entire codebase for consistency
- Removed unused _pending_blockswap_config attribute
Core Features:
- Add torch.compile support for DiT (20-40% speedup) and VAE (15-25% speedup)
- Ensure BlockSwap compatibility by applying BlockSwap before torch.compile
- Add SeedVR2TorchCompileSettings node for ComfyUI configuration
- Add CLI arguments: --compile_dit, --compile_vae, --compile_backend, --compile_mode, --compile_fullgraph, --compile_dynamic, --compile_dynamo_cache_size_limit, --compile_dynamo_recompile_limit
Optimizations:
- Optimize na.py and other backend files for torch.compile compatibility (replace .tolist() with tensor operations)
- Remove .item() calls from attention modules (max_seqlen_q/k now accept tensors)
- Add @torch._dynamo.disable decorators to timing/debug methods to prevent compilation warnings
- Replace torch.repeat with torch.repeat_interleave in modulation.py for better performance
Bug Fixes:
- Fix timing stack cleanup in memory_manager.py (end timer before early returns)
- Fix BlockSwap timing reporting (_get_swap_start_time and _log_swap_timing now properly excluded from compilation)
Code Quality:
- Add comprehensive type hints and docstrings across all modified modules
- Standardize BlockSwap timer names (blockswap_block_*, blockswap_io_*)
BREAKING CHANGE: Replaced BlockSwap and ExtraArgs nodes with DiT/VAE loader nodes.
Enables multi-GPU placement, independent caching, and granular memory control.
New Features:
- Add SeedVR2LoadDiTModel and SeedVR2LoadVAEModel loader nodes
- Support different devices for DiT and VAE (multi-GPU load balancing)
- Independent model caching (cache_model_dit/cache_model_vae)
- Separate encode/decode VAE tiling with independent tile configurations
- User-defined DiT and VAE model selection from registry or disk
Breaking Changes:
- Deprecated SeedVR2BlockSwap node (integrated into DiT loader)
- Deprecated SeedVR2ExtraArgs node (split into DiT/VAE loaders & Upscaler node)
- BlockSwap configuration now part of DiT loader node
Improvements:
- Follow ComfyUI conventions (pixels input, IMAGE output type)
- Update download_weight() to support user-defined DiT/VAE models
- Improve tooltips and default parameters across all nodes
- Set show_tensors=False in debug logging for performance
- Dual-device context tracking throughout pipeline
- CLI updated to match new dual-model API
- Split decode_all_batches into decode (Phase 3) and postprocess_all_batches (Phase 4)
- Add phase-specific cleanup functions: cleanup_dit() and cleanup_vae()
- Each phase now handles its own resource cleanup in finally blocks
- Add cleanup_text_embeddings() helper to eliminate code duplication
- Remove redundant cleanup from comfyui_node normal flow
- Update pipeline from 3-phase to 4-phase architecture (encode → upscale → decode → postprocess)
- Improve memory efficiency by releasing resources immediately when no longer needed
- Update module docstrings to reflect new architecture
- Store latents and transformed videos on CPU between processing phases reducing VRAM usage to enable larger batch processing
- Make transformed video storage conditional (only when color_correction != none)
- Add all_ori_lengths tracking for consistent trimming
- Use non_blocking=False for CPU-GPU transfers to avoid pinned memory issues
Fixes https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/pull/163#issuecomment-3333686784
Model Registry Integration:
- Replace hardcoded CLI model choices with dynamic get_available_models()
- Use centralized DEFAULT_MODEL constant from model_registry
- Achieve single source of truth between CLI and ComfyUI interfaces
GGUF Optimizations:
- Create tensors directly on target device to avoid CPU->GPU copy overhead
- Remove excessive memory cleanup during tensor loading for better performance
- Always use meta init for memory-efficient model creation
- Implement precision-optimized dequantization path (GGUF → FP16 → compute dtype)
GGUF Precision Handling:
- Add GGUFQuantizedLinear/Conv2d layers with get_dequantized_weight_for_compute() to preserve precision
- Track and report quantization types (Q4_K_M, Q5_K_M, etc.) in model
- Fix __torch_function__ as classmethod to resolve deprecation warning
Code Quality:
- Restructure _load_model_weights() with modular helper functions to reduce duplication
- Improve separation of concerns between standard and GGUF weight loading
- Enhance logging to always display WARNING/ERROR messages
- Add comprehensive GGUF architecture validation
- Remove debug-only log_memory_state calls
- Fix type function definitions
- Implement GGUF model loading with Q3_K_M through Q8_K_M quantization support
- Add GGUFTensor wrapper to preserve quantization and enable on-demand dequantization
- Maintain tensors in quantized format to reduce VRAM usage
- Add GGUF dequantization operations for inference
- Update model registry to include GGUF variants for 3B/7B models
- Fix wavelet blur radius limit to prevent OOM at high resolutions (max 1/8 of image dimension)
- Add safety clamp [-1,1] for SDR color range to prevent numerical errors
- This is a WIP commit as some additional cleaning/testing is needed
- Add type hints throughout for better code maintainability
- Add input_noise_scale (0-1) to reduce artifacts at high resolutions
- Applies subtle noise before VAE encoding with progressive blend (0-50%)
- Based on GitHub issue #64 community findings
- Rename cond_noise_scale to latent_noise_scale for clarity
- Better distinguishes between input (pixel) and latent (diffusion) noise
- Maintains same functionality, just clearer naming
- Update both ComfyUI and CLI interfaces with new parameters
- Add color_correction parameter (wavelet/adain/none) to both ComfyUI node and CLI interface to choose between wavelet reconstruction (frequency-based), AdaIN (statistical matching), or no color correction
- Improve memory management after upscaling batches with targeted clear_memory calls to reduce VRAM pressure during VAE decoding
- Renamed model directories for clarity: dit -> dit_7b, dit_v2 -> dit_3b
- Converted all absolute imports to relative imports throughout codebase
- Removed sys.path.append() manipulations that caused namespace conflicts
- Updated YAML configs to reference renamed model directories
- Simplified model variant detection logic using new directory names
- Standardized function calls with named arguments for better clarity
- Fixed generation context initialization and interrupt handling
This resolves import conflicts with other ComfyUI nodes (e.g., Basic data handling)
that use sys.path manipulation, making the module properly isolated and compatible.
Fixes#29, #114, #136
Major architectural change to minimize model swapping overhead by processing
all batches in three distinct phases instead of sequential per-batch processing:
- Phase 1: Encode all batches with VAE
- Phase 2: Upscale all latents with DiT
- Phase 3: Decode all latents with VAE
Core changes:
- Split monolithic generation_loop into modular functions:
- prepare_generation_context(): Shared state management
- setup_device_environment(): Device configuration
- prepare_runner(): Model loading with cache support
- encode_all_batches(): Batch VAE encoding
- upscale_all_batches(): Batch DiT upscaling
- decode_all_batches(): Batch VAE decoding
- Removed generation_step function (logic integrated into upscale phase)
- Added lazy precision initialization to avoid redundant setup
Performance improvements:
- Pre-allocated lists for memory efficiency
- Better cleanup of intermediate storage between phases
- Added unique timer names to clear_memory() to avoid naming conflicts
- Improved model state management with change detection and caching
UI/UX enhancements:
- Switched to ComfyUI's native ProgressBar with weighted phase progress
- Changed from per-batch FPS to overall average FPS (always visible)
- Improved log clarity with clear phase separators
- Added ASCII art logo to clearly identify SeedVR2 process start
- Better progress tracking with weighted percentages across three phases
Code cleanup:
- Removed deprecated timer_context from Debug class
- Removed unused time imports across multiple files
- Fixed LOCAL_RANK environment variable to handle string conversion properly
- Improved error handling with try/except/finally blocks in all phases
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state() formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
- Fixed issue that would lead to a white border around the image due to
improper tile blending
- Added smoother blending (hann ramp) for the tiles and temporal
overlaps
- use h/w tuples to specify tiling size and overlap to allow for finer
control
- added OneOrTwoValues class to handle CLI input of single values or h/w
pairs
- Fix issues with unnecessary VAE blocks at the boundarys
- Updated VAE encode/decode logic to more closely match the original flow
- Better skip_frame log info