- Add --video_backend flag: 'opencv' (default) or 'ffmpeg'
- Add --10bit flag: enables x265/yuv420p10le for reduced banding
- Without --10bit, ffmpeg uses x264/yuv420p for max compatibility
- FFMPEGVideoWriter class with cv2.VideoWriter-compatible interface
- Validates ffmpeg availability before encoding
Based on PR #409 by thehhmdb
- Add validate_blockswap_config() in blockswap.py as single validation point
- Auto-disable BlockSwap on macOS (unified memory makes it meaningless)
- Improve error messages for missing dit_offload_device
- Update CLI and ComfyUI tooltips for BlockSwap and model caching
- Update README: BlockSwap macOS note, caching descriptions, attention backends
- Remove duplicate validation from dit_model_loader.py and inference_cli.py
Partially fixes#401 (M4 Pro macOS BlockSwap offload device error)
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.
The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.
- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow
Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup
Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
Core Fixes:
- Reset seed per batch to ensure deterministic generation across sessions and batch positions
- Fix temporal overlap logging when automatically reset to prevent incorrect frame counts
- Fix NoneType attribute error in VAE tiled encode/decode at maximum resolution (#296)
BlockSwap & Caching Architecture:
- Move BlockSwap state (_block_swap_config, _blockswap_bypass_protection) from runner to model
- Ensures state survives runner recreation during independent DiT/VAE caching scenarios
- Fix runner template caching to trigger when either DiT or VAE becomes cached (bidirectional)
- Resolves BlockSwap reload failures when only DiT was cached (#297)
Model Discovery:
- Implement case-insensitive YAML path resolution for extra_model_paths.yaml (#289-#295)
- Add debug logging for model discovery (searched paths, validation status, cache hits)
- Support any case variation (seedvr2, SEEDVR2, SeedVR2) in ComfyUI configuration
- Fix: OpenCV memory layout error in tile debug visualization (#283)
- Fix: macOS MPS allocator fallback for safetensors loading (#290)
- Fix: Windows log buffering with flush=True (#278)
- Fix: ComfyUI registry icon URL (raw.githubusercontent)
- Feature: Version display in node name and CLI/ComfyUI header
- Feature: GitHub Sponsors support link
- License: Migrate from MIT to Apache 2.0 to match Bytedance Seed official repo
- Improve output folder naming: batch creates {folder}_upscaled/ sibling with original filenames, single file adds _upscaled suffix
- Add RGBA alpha channel detection and preservation (matches ComfyUI)
- Convert all output paths to absolute for clarity in logs
- Replace dual glob loops with single iterdir scan for cross-platform consistency
- Fixes duplicate file processing in batch mode on Windows case-insensitive filesystem
- Improves directory scanning performance 2-3x by reducing filesystem operations
- Add ComfyUI registry logo
- Add ROCm/HIP detection to prevent NVIDIA-specific cuDNN workaround on AMD systems
- Add defensive hasattr checks for torch.cuda and cudnn.is_available()
- Add try-except fallback in _conv_forward to handle cuDNN call failures gracefully
- Resolves 'GET was unable to find an engine' errors on PyTorch 2.9+ dev builds
- Resolves 'ATen not compiled with cuDNN support' errors on ROCm platforms
- Bump version to 2.5.7
- Replace split-stack-mean with unflatten in unconcat_coalesce
- Corrects computation order to eliminate plastic/high-specular artifacts
- Maintains torch.compile compatibility (no .item() graph breaks)
- Applied to both dit_3b and dit_7b models
- Replace all_transformed_videos storage with lightweight batch_metadata indices
- Reconstruct transformed videos on-demand in Phase 4 only when needed
- Add missing cleanup for input_images tensor in postprocess finally block
- Fix release_tensor_memory to handle CPU/CUDA/MPS consistently
- Extract helper functions for batch preparation and 4n+1 padding
- Remove duplicate interrupt_fn key from context initialization
- Fix AdaIN non-contiguous tensor error by using reshape() instead of view()
- Add cuDNN availability checks to prevent ROCm 'ATen not compiled with cuDNN' error
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)
Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable
Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability
Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
**Architecture & Performance:**
- Unified debug system with categorized logging and memory tracking
- FP8 models now stay in FP8, convert to BF16 only for math operations (faster, less memory)
- Fixed memory leaks
- Removed ComfyUI dependency for standalone compatibility
- Better VRAM management between batches
**New Features:**
- VAE tiling for larger/longer video upscaling
- Multi-repo model support (numz/ and AInVFX/)
- Model auto-discovery in ComfyUI folder
- CLI: BlockSwap options, temporal overlap blending, prepend_frames for better first frames
- `enable_debug` and `cache_model` moved to main node
**Code Organization:**
- New modular structure: `constants.py`, `model_registry.py`, `debug.py`
- Removed legacy code and dead files
- Added mixed FP8 models to fix 7B artifacts