- Add validate_blockswap_config() in blockswap.py as single validation point
- Auto-disable BlockSwap on macOS (unified memory makes it meaningless)
- Improve error messages for missing dit_offload_device
- Update CLI and ComfyUI tooltips for BlockSwap and model caching
- Update README: BlockSwap macOS note, caching descriptions, attention backends
- Remove duplicate validation from dit_model_loader.py and inference_cli.py
Partially fixes#401 (M4 Pro macOS BlockSwap offload device error)
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.
The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.
- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow
Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
- Re-implemented `precision` control (`fp16`, `bf16`, `bf32`, `auto`) in CLI, ComfyUI node, and backend logic to respect user choice.
- Fixed `UnboundLocalError` in `apply_model_specific_config` by ensuring `compute_dtype` is always initialized before use.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by properly passing the restored `precision` argument.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` for clarity and fixed fallback logic to ensure SageAttention is correctly prioritized.
- Added explicit logging of active attention backend and execution confirmation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by removing residual `precision` usage.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, CLI, and ComfyUI nodes to fix naming confusion.
- Added strict `precision` control (`fp16`, `bf16`, `bf32`, `auto`) to CLI and internal configuration logic.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Added explicit logging ("🚀 Executing SageAttention...") to confirm kernel execution.
- Fixed bug where user-selected precision was being overridden by auto-detection defaults.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, compatibility layer, and ComfyUI node definitions to fix naming confusion.
- Fixed logic bug in `FlashAttentionVarlen` that prevented SageAttention from running even when requested (was checking for `sd` prefix instead of `sa`).
- Added robust availability and version checks for SageAttention (v2 vs v3) with improved fallback logic (SA3 -> SA2 -> FlashAttention 2 -> SDPA).
- Added explicit console logging when SageAttention kernel is first executed to verify optimization is active.
- Added `--precision` argument to CLI to allow explicit control over compute dtype (fp16, bf16, bf32, auto), enabling further performance tuning.
- Updated `FP8CompatibleDiT` wrapper to exclude `FlashAttentionVarlen` modules, preventing double-wrapping and casting issues.
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup
Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
Use PyTorch shared memory instead of pickling numpy arrays through queue.
Prevents MemoryError when transferring large results between processes.
Thank you @FurkanGozukara
- BlockSwap: Show effective/total blocks (e.g., 32/32) instead of raw requested value
- CLI: Skip CUDA device validation when CUDA_VISIBLE_DEVICES already set (worker process)
- Mac: Use direct processing instead of spawning subprocess (MPS allocator fails in child process)
- Multi-GPU: Set CUDA_VISIBLE_DEVICES before spawn so child inherits it before module-level torch import
- Remove redundant env setup in worker (now inherited from parent)
- Improve output folder naming: batch creates {folder}_upscaled/ sibling with original filenames, single file adds _upscaled suffix
- Add RGBA alpha channel detection and preservation (matches ComfyUI)
- Convert all output paths to absolute for clarity in logs
- Replace dual glob loops with single iterdir scan for cross-platform consistency
- Fixes duplicate file processing in batch mode on Windows case-insensitive filesystem
- Improves directory scanning performance 2-3x by reducing filesystem operations
- Add ComfyUI registry logo
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
- Standardize tooltip format with multi-line descriptions and bullet points
- Add comprehensive output tooltips for all nodes (DiT, VAE, torch.compile, upscaler)
- Enhance node descriptions with detailed capability summaries
- Simplify CLI tile size arguments to single integers (converted internally to tuples)
- Remove OneOrTwoValues argparse class for cleaner implementation
- Fix encode_tiled tooltip (was incorrectly referencing decoding)
- Clarify color correction purpose (corrects upscaling color shifts)
- Add multi-GPU offloading information to all offload_device tooltips
- Improve torch.compile parameter descriptions with use cases
- Ensure CLI and ComfyUI tooltips are consistent in terminology and structure
- Fix validation cache location to respect --model_dir parameter
- Fix output path handling for directories without extensions
- Remove spurious directory creation in get_base_cache_dir
- Enhanced CLI help with usage examples and argument grouping
- Auto-display help when script invoked without arguments
- Correct type hints (device_id: str, debug: Debug)
- Remove redundant type conversions and makedirs calls
- Reorganize imports to module top for clarity
- Improved docstrings & tooltip
- Change default batch_size from 1 to 5 to match ComfyUI integration
- Add support for seedvr2_ema_3b-Q8_0.gguf model
- Validate --cuda_device arguments early in pre-parsing phase
- Check device IDs exist and are within available GPU range
- Fail fast with clear error messages showing available devices
- Auto-detect output format per file type (mp4 for videos, png for images)
- Add visual separators between processed files for better readability
- Simplify FPS calculation to use wall-clock time for real-world throughput
- Consolidate banner/footer into shared Debug methods
- Update offload device args to support multi-GPU (cpu/cuda:N)
- Standardize terminology: 'upscaling' instead of 'video upscaling'
- Remove code duplication between CLI and ComfyUI implementations
- Consistent quote style (double quotes) throughout CLI
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)
Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable
Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability
Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
- Add max_resolution parameter (default: 0 = no limit) to both CLI and ComfyUI
- After new_resolution scales shortest edge, max_resolution ensures no edge exceeds limit
- Scales down proportionally if constraint violated
- Maintains backward compatibility with default value of 0
- Add dit_offload_device parameter for proper blockSwap configuration
- Ensure consistent dtype management throughout CLI and ComfyUI (float32 input with bfloat16 pipeline)
- Translate all French comments to English
- Add comprehensive docstrings and section headers
- Remove obsolete use_non_blocking and enable_debug parameters
- Add error handling and validation
- Implement prepend_video_frames() for artifact reduction at video start
- Add blend_overlapping_frames() with Hann window for smooth transitions
- Expose temporal_overlap (0-16) and prepend_frames (0-32) in ComfyUI node
- Unify CLI and ComfyUI to use shared prepend/overlap functions
- Add comprehensive logging for frame adjustments (prepend/overlap/padding)
CFG (Classifier-Free Guidance) does not work with SeedVR2's distilled
one-step diffusion model. The model was trained to produce final results
in a single step without iterative guidance.
- Remove cfg_scale parameter from ComfyUI node and CLI interface
- Force internal cfg_scale to 1.0 in upscale_all_batches()
This avoids artifacts introduced when users changed cfg_scale away from 1.0.
- Switch to deterministic VAE encoding (mode vs sample) to eliminate high-resolution noise artifacts
- Make VAE encode tiling optional (disabled by default) since deterministic encoding resolves artifacts
- Add tile debug visualization feature with adaptive scaling and color-coded boundaries
- Remove redundant dtype conversions in VAE code for better performance
- Minor code cleanup and documentation update
VAE Changes:
- Enable encode tiling by default (prevents noise artifacts at high resolution)
- Increase tile size to 1024px (down from 512px) for optimal quality
- Increase tile overlap to 128px for better blending
Dtype Pipeline:
- Hardcode compute_dtype to bfloat16 for consistent quality/performance/VRAM balance
- Ensure all pipeline steps are using compute_dtype when relevant
- Refactor code for improved performance and memory management
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
* VAE encoding: seed+1M for deterministic sampling without quality loss
* DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id
Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)
- Add indent_level param to debug.log() and manage_tensor()
- Replace all hardcoded spaces with indent_level (0/1/2)
- Fix encode_all_batches memory flow: encode before storage
- Minor whitespace cleanup
Major improvements to device handling, model offloading, and code clarity:
Device Management:
- Replace string-based devices with torch.device objects throughout
- Add explicit offload device parameters: dit_offload_device, vae_offload_device, tensor_offload_device
- Remove preserve_vram in favor of explicit offload control
- Improve get_device() to return torch.device objects consistently
- Enhance get_device_list() with smart MPS-only system handling
Context & Pipeline:
- Merge setup_device_environment and prepare_generation_context into single setup_generation_context
- Simplify LOCAL_RANK handling (set to '0' for single-GPU mode)
- Store device configuration on runner for submodule access
Tensor & Model Management:
- Add manage_tensor_device() for consistent tensor movement with logging
- Update manage_model_device() to use torch.device objects
- Add validation for BlockSwap and caching configurations
- Rename cache_in_ram to cache_model (more accurate naming)
CLI & Interface:
- Remove --preserve_vram flag (breaking change)
- Add --vae_offload_device and --tensor_offload_device flags
- Update ComfyUI node parameters with validation and better tooltips
- Improve device selection UI with offload_device options
- Move ensure_float32_precision() from alpha_upscaling to common/half_precision_fixes
- Update color_fix.py to use ensure_float32_precision
- Add type hints and docstrings to all half_precision_fixes functions
- Clarify batch size tips: emphasize avoiding padding waste over maximizing batch size
- Update default seed to 42 (the answer to life, the universe, and everything)
- Add LAB color transfer as new default method for superior perceptual color matching
- Add HSV hue-conditional saturation histogram matching for targeted oversaturation correction
- Add wavelet_adaptive hybrid method combining wavelet base with selective HSV correction
- Change default color correction from 'wavelet' to 'lab' for better color accuracy
- Implement full CIELAB color space conversion with D65 illuminant and histogram matching
- Add comprehensive documentation and method descriptions to color_fix.py
Major refactor introducing:
- GlobalModelCache for cross-node model sharing with dynamic config updates
* Models cached by node ID, reused across different upscaler instances
* Config changes (torch.compile, BlockSwap, tiling) handled dynamically
* Runner templates cached when both models present
- Complete alpha channel processing rewrite
* New edge-guided alpha upscaling in src/core/alpha_upscaling.py
* Remove broken VAE RGBA adapter approach
* Alpha properly extracted, upscaled, and merged with RGB
- BlockSwap enhancements
* Support I/O-only swapping without transformer block offloading
* Separate memory reporting for I/O components vs transformer blocks
* Clearer logging of GPU vs CPU placement
* New is_blockswap_enabled() utility
- Memory management improvements
* Add release_tensor_collection() for batch tensor cleanup
* Proper tensor memory release throughout pipeline
* Better cleanup in postprocess phase
- Code quality and refactoring
* Extensive docstrings with type hints across all modules
* Generic _update_model_config() reduces code duplication
* Better function signatures and parameter documentation
* Improved separation of concerns in model_manager.py
* Clearer debug logging categories and messages