- Add _device_str() helper to normalize MPS variants (mps:0 → MPS)
- Fix device comparison: mps:0 and mps now correctly identified as same device
- Consistent MPS logging across all memory management functions
Fixes#340 - Installation error with PyTorch 2.7+cu126 and triton_windows
Replace diffusers.models.autoencoders.vae imports with local implementations
of DecoderOutput and DiagonalGaussianDistribution to avoid triggering the
bitsandbytes -> triton.ops import chain that fails on newer triton versions.
- Use in-place clamp/mul/add in postprocess normalization
- Use in-place clamp and Normalize in video transform pipeline
- Consolidate optimized channel permute functions in performance.py
- Use native permute() instead of einops rearrange in VAE encode/decode
- In-place wavelet decomposition and reconstruction operations
- Remove unused temporal_latent_blending function
- BlockSwap: Show effective/total blocks (e.g., 32/32) instead of raw requested value
- CLI: Skip CUDA device validation when CUDA_VISIBLE_DEVICES already set (worker process)
Replaces torch.mps.is_available() with torch.backends.mps.is_available()
across all files. The torch.backends API is the official PyTorch method
for MPS detection (since PyTorch 1.12) and works reliably on:
- macOS with Apple Silicon (returns True when MPS available)
- macOS without MPS support (returns False)
- Windows/Linux (hasattr guard prevents AttributeError)
Eliminates memory spike for long videos by:
- Pre-allocating final_video in Phase 3 before any batch processing
- Writing VAE-decoded RGB directly to final_video (no batch_samples accumulation)
- For RGBA: allocating 4 channels upfront, writing RGB in Phase 3, alpha in Phase 4
- Processing batches in-place through color correction and normalization
- Moving temporal overlap blending to Phase 3 decode phase
Fixes#130 (DefaultCPUAllocator: not enough memory for long video)
- Use torch.nn.attention.sdpa_kernel() (new API) with CUDNN_ATTENTION backend
- Fallback to torch.backends.cuda.sdp_kernel() when not available
Thanks to @eadwu for the original PR #317
Core Fixes:
- Reset seed per batch to ensure deterministic generation across sessions and batch positions
- Fix temporal overlap logging when automatically reset to prevent incorrect frame counts
- Fix NoneType attribute error in VAE tiled encode/decode at maximum resolution (#296)
BlockSwap & Caching Architecture:
- Move BlockSwap state (_block_swap_config, _blockswap_bypass_protection) from runner to model
- Ensures state survives runner recreation during independent DiT/VAE caching scenarios
- Fix runner template caching to trigger when either DiT or VAE becomes cached (bidirectional)
- Resolves BlockSwap reload failures when only DiT was cached (#297)
Model Discovery:
- Implement case-insensitive YAML path resolution for extra_model_paths.yaml (#289-#295)
- Add debug logging for model discovery (searched paths, validation status, cache hits)
- Support any case variation (seedvr2, SEEDVR2, SeedVR2) in ComfyUI configuration
- Fix: OpenCV memory layout error in tile debug visualization (#283)
- Fix: macOS MPS allocator fallback for safetensors loading (#290)
- Fix: Windows log buffering with flush=True (#278)
- Fix: ComfyUI registry icon URL (raw.githubusercontent)
- Feature: Version display in node name and CLI/ComfyUI header
- Feature: GitHub Sponsors support link
- License: Migrate from MIT to Apache 2.0 to match Bytedance Seed official repo
- Add ROCm/HIP detection to prevent NVIDIA-specific cuDNN workaround on AMD systems
- Add defensive hasattr checks for torch.cuda and cudnn.is_available()
- Add try-except fallback in _conv_forward to handle cuDNN call failures gracefully
- Resolves 'GET was unable to find an engine' errors on PyTorch 2.9+ dev builds
- Resolves 'ATen not compiled with cuDNN support' errors on ROCm platforms
- Bump version to 2.5.7
- Replace split-stack-mean with unflatten in unconcat_coalesce
- Corrects computation order to eliminate plastic/high-specular artifacts
- Maintains torch.compile compatibility (no .item() graph breaks)
- Applied to both dit_3b and dit_7b models
- Replace all_transformed_videos storage with lightweight batch_metadata indices
- Reconstruct transformed videos on-demand in Phase 4 only when needed
- Add missing cleanup for input_images tensor in postprocess finally block
- Fix release_tensor_memory to handle CPU/CUDA/MPS consistently
- Extract helper functions for batch preparation and 4n+1 padding
- Remove duplicate interrupt_fn key from context initialization
- Fix AdaIN non-contiguous tensor error by using reshape() instead of view()
- Add cuDNN availability checks to prevent ROCm 'ATen not compiled with cuDNN' error
Add defensive checks for torch.mps.is_available() to handle PyTorch versions where the method doesn't exist on non-Mac platforms. Resolves AttributeError: module 'torch.mps' has no attribute 'is_available'
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
- Eliminate graph breaks in na.py and attention.py for full torch.compile support
* Replace cumsum-based tensor slicing with _tensor_split to avoid .item() calls
* Use torch.tensor_split with .long().cpu() for PyTorch API requirements
* Replace reshape-based averaging with split-stack-mean pattern
- Fix phase peak VRAM tracking to capture peaks during OOM retry cycles
- Standardize phase4 naming to 'postprocessing' for consistency
- Update RoPE docstrings for accuracy
- Update example workflows with icon
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
- Standardize tooltip format with multi-line descriptions and bullet points
- Add comprehensive output tooltips for all nodes (DiT, VAE, torch.compile, upscaler)
- Enhance node descriptions with detailed capability summaries
- Simplify CLI tile size arguments to single integers (converted internally to tuples)
- Remove OneOrTwoValues argparse class for cleaner implementation
- Fix encode_tiled tooltip (was incorrectly referencing decoding)
- Clarify color correction purpose (corrects upscaling color shifts)
- Add multi-GPU offloading information to all offload_device tooltips
- Improve torch.compile parameter descriptions with use cases
- Ensure CLI and ComfyUI tooltips are consistent in terminology and structure
- Fix validation cache location to respect --model_dir parameter
- Fix output path handling for directories without extensions
- Remove spurious directory creation in get_base_cache_dir
- Enhanced CLI help with usage examples and argument grouping
- Auto-display help when script invoked without arguments
- Correct type hints (device_id: str, debug: Debug)
- Remove redundant type conversions and makedirs calls
- Reorganize imports to module top for clarity
- Improved docstrings & tooltip
- Change default batch_size from 1 to 5 to match ComfyUI integration
- Add support for seedvr2_ema_3b-Q8_0.gguf model
- Auto-detect output format per file type (mp4 for videos, png for images)
- Add visual separators between processed files for better readability
- Simplify FPS calculation to use wall-clock time for real-world throughput
- Consolidate banner/footer into shared Debug methods
- Update offload device args to support multi-GPU (cpu/cuda:N)
- Standardize terminology: 'upscaling' instead of 'video upscaling'
- Remove code duplication between CLI and ComfyUI implementations
- Consistent quote style (double quotes) throughout CLI