211 Commits
Author SHA1 Message Date
Adrien Toupet 11239eed13 Release v2.5.16: Quality regression fix, older GPU compatibility fix, system info debug 2025-12-05 15:50:08 -05:00
Adrien Toupet b4d7ab89eb Fix CUBLAS_STATUS_NOT_SUPPORTED on older GPUs (GTX 970) #314
Add automatic bfloat16 → float16 SDPA fallback for GPUs without native bf16 cuBLAS support
2025-12-05 15:39:22 -05:00
Adrien Toupet f061d97fe7 Revert bfloat16 detection - was causing quality regression / keep ensure_triton_compat() 2025-12-05 15:03:40 -05:00
Adrien Toupet 43c4f00e19 Revert bfloat16 detection - was causing quality regression / keep ensure_triton_compat() 2025-12-05 15:01:58 -05:00
Adrien Toupet f8998ebd75 Revert bfloat16 detection - was causing quality regression 2025-12-05 14:55:30 -05:00
Adrien Toupet 18b44d66e1 feat: add environment info display in debug mode to help with issue reporting 2025-12-05 11:11:18 -05:00
Adrien Toupet 65c1c1b6cd Release v2.5.15: MPS fixes, autocast device type, accurate VRAM tracking, triton 3.0 compatibility 2025-12-03 13:14:51 -05:00
Adrien Toupet b40f26167c Fix MPS compatibility: disable antialias for MPS tensors, fix bfloat16 arange (#354) 2025-12-03 13:09:59 -05:00
Adrien Toupet ed53581359 Fix triton.ops compatibility for bitsandbytes 0.45+ / triton 3.0+
Fixes #340 - Installation error with PyTorch 2.7+cu126 and triton_windows

Add compatibility shim for missing triton.ops.matmul_perf_model module.
Reverts local VAE types approach
2025-12-03 12:43:35 -05:00
Adrien Toupet 71ac9ffe54 fix: use max_memory_reserved for accurate VRAM peak tracking 2025-12-03 11:51:30 -05:00
Adrien Toupet ffba05907d Fix autocast device_type error by using .type attribute instead of str() #350 2025-12-03 11:32:39 -05:00
Adrien Toupet e2faedaaa6 Release v2.5.14 - MPS device fix, VRAM swap detection, enforce physical VRAM limit 2025-12-01 00:30:49 -05:00
Adrien Toupet 5775ff0f99 Enforce VRAM limit to physical capacity - OOM instead of silent swap 2025-12-01 00:25:26 -05:00
Adrien Toupet ff937756a3 Add VRAM swap detection - show GPU+swap breakdown in peak stats, warn when swap detected 2025-11-30 23:47:20 -05:00
Adrien Toupet 5848cef05f fix(mps): normalize device strings to prevent unnecessary tensor movements
- Add _device_str() helper to normalize MPS variants (mps:0 → MPS)
- Fix device comparison: mps:0 and mps now correctly identified as same device
- Consistent MPS logging across all memory management functions
2025-11-30 21:03:26 -05:00
Adrien Toupet 19825fa9fa Release v2.5.13: Fix triton import, OOM on long videos, macOS watermark 2025-11-30 09:01:53 -05:00
Adrien Toupet 021fc7b70f Fix triton.ops import error by using local VAE types
Fixes #340 - Installation error with PyTorch 2.7+cu126 and triton_windows

Replace diffusers.models.autoencoders.vae imports with local implementations
of DecoderOutput and DiagonalGaussianDistribution to avoid triggering the
bitsandbytes -> triton.ops import chain that fails on newer triton versions.
2025-11-30 08:27:01 -05:00
Adrien Toupet b0ac880a4f Fix OOM crash on float32 conversion for long videos. Gracefully fallback to native dtype if insufficient memory. Fixes #299 2025-11-30 08:03:07 -05:00
Adrien Toupet 16508d353d release: v2.5.12 - fix color artifacts regression from in-place transform ops 2025-11-28 17:50:40 -05:00
Adrien Toupet f791631495 release: v2.5.11 - CUDNN attention, long video memory fix, LAB artifacts fix, MPS/multi-GPU improvements 2025-11-28 16:49:08 -05:00
Adrien Toupet b15982b442 perf: in-place tensor ops to reduce memory allocation overhead
- Use in-place clamp/mul/add in postprocess normalization
- Use in-place clamp and Normalize in video transform pipeline
- Consolidate optimized channel permute functions in performance.py
- Use native permute() instead of einops rearrange in VAE encode/decode
- In-place wavelet decomposition and reconstruction operations
- Remove unused temporal_latent_blending function
2025-11-28 16:40:27 -05:00
Adrien Toupet 849f6b5997 Add peak RAM tracking alongside VRAM in debug summary 2025-11-28 15:32:07 -05:00
Adrien Toupet e34bec76d5 Auto-detect bfloat16 support to fix CUBLAS_STATUS_NOT_SUPPORTED on older GPUs (#314) 2025-11-28 14:30:41 -05:00
Adrien Toupet daa13fb4ca Fix BlockSwap logging confusion and CLI worker validation
- BlockSwap: Show effective/total blocks (e.g., 32/32) instead of raw requested value
- CLI: Skip CUDA device validation when CUDA_VISIBLE_DEVICES already set (worker process)
2025-11-28 13:40:07 -05:00
Adrien Toupet be2efd474f fix: use canonical torch.backends.mps.is_available() for reliable MPS detection
Replaces torch.mps.is_available() with torch.backends.mps.is_available()
across all files. The torch.backends API is the official PyTorch method
for MPS detection (since PyTorch 1.12) and works reliably on:
- macOS with Apple Silicon (returns True when MPS available)
- macOS without MPS support (returns False)
- Windows/Linux (hasattr guard prevents AttributeError)
2025-11-28 11:49:43 -05:00
Adrien Toupet bce62382db Fix color correction reference misalignment with temporal overlap 2025-11-28 09:43:46 -05:00
Adrien Toupet f792ca5937 fix: eliminate LAB color correction tile artifacts by preprocessing with wavelet reconstruction
Fixes #324
Fixes https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/discussions/271#discussioncomment-14997355
2025-11-28 00:10:17 -05:00
Adrien Toupet 3ab2c84091 fix: stream VAE decode directly to pre-allocated final_video tensor
Eliminates memory spike for long videos by:
- Pre-allocating final_video in Phase 3 before any batch processing
- Writing VAE-decoded RGB directly to final_video (no batch_samples accumulation)
- For RGBA: allocating 4 channels upfront, writing RGB in Phase 3, alpha in Phase 4
- Processing batches in-place through color correction and normalization
- Moving temporal overlap blending to Phase 3 decode phase

Fixes #130 (DefaultCPUAllocator: not enough memory for long video)
2025-11-27 20:15:42 -05:00
Adrien Toupet 8dd0c3061a feat: Add CUDNN attention backend support with PyTorch 2.3+ API
- Use torch.nn.attention.sdpa_kernel() (new API) with CUDNN_ATTENTION backend
- Fallback to torch.backends.cuda.sdp_kernel() when not available
Thanks to @eadwu for the original PR #317
2025-11-27 15:19:55 -05:00
Adrien Toupet 79e7f41216 Revert "Fix: MPS allocator error in model.to() transfer (#305)"
This reverts commit 9424020687.
2025-11-27 14:52:12 -05:00
Adrien Toupet 9424020687 Fix: MPS allocator error in model.to() transfer (#305)
Add fallback to individual parameter movement when bulk model.to(mps)
fails with allocator errors. Complements safetensors loading fix.
2025-11-14 12:06:13 -05:00
Adrien Toupet 65dd29a865 v2.5.10: Fix determinism, BlockSwap caching, and model path resolution
Core Fixes:
- Reset seed per batch to ensure deterministic generation across sessions and batch positions
- Fix temporal overlap logging when automatically reset to prevent incorrect frame counts
- Fix NoneType attribute error in VAE tiled encode/decode at maximum resolution (#296)

BlockSwap & Caching Architecture:
- Move BlockSwap state (_block_swap_config, _blockswap_bypass_protection) from runner to model
- Ensures state survives runner recreation during independent DiT/VAE caching scenarios
- Fix runner template caching to trigger when either DiT or VAE becomes cached (bidirectional)
- Resolves BlockSwap reload failures when only DiT was cached (#297)

Model Discovery:
- Implement case-insensitive YAML path resolution for extra_model_paths.yaml (#289-#295)
- Add debug logging for model discovery (searched paths, validation status, cache hits)
- Support any case variation (seedvr2, SEEDVR2, SeedVR2) in ComfyUI configuration
2025-11-13 12:02:37 -05:00
Adrien Toupet 40d37b7248 Release v2.5.9: Bug fixes and enhancements
- Fix: OpenCV memory layout error in tile debug visualization (#283)
- Fix: macOS MPS allocator fallback for safetensors loading (#290)
- Fix: Windows log buffering with flush=True (#278)
- Fix: ComfyUI registry icon URL (raw.githubusercontent)
- Feature: Version display in node name and CLI/ComfyUI header
- Feature: GitHub Sponsors support link
- License: Migrate from MIT to Apache 2.0 to match Bytedance Seed official repo
2025-11-12 14:27:40 -05:00
Adrien Toupet b64d526bd6 fix: enhance Conv3d workaround compatibility for PyTorch dev builds and AMD ROCm
- Add ROCm/HIP detection to prevent NVIDIA-specific cuDNN workaround on AMD systems
- Add defensive hasattr checks for torch.cuda and cudnn.is_available()
- Add try-except fallback in _conv_forward to handle cuDNN call failures gracefully
- Resolves 'GET was unable to find an engine' errors on PyTorch 2.9+ dev builds
- Resolves 'ATen not compiled with cuDNN support' errors on ROCm platforms
- Bump version to 2.5.7
2025-11-10 13:04:52 -05:00
Adrien Toupet a3f98124ca Fix: Restore natural look for 7b model (v2.5.6)
- Replace split-stack-mean with unflatten in unconcat_coalesce
- Corrects computation order to eliminate plastic/high-specular artifacts
- Maintains torch.compile compatibility (no .item() graph breaks)
- Applied to both dit_3b and dit_7b models
2025-11-09 12:26:18 -05:00
Adrien Toupet b5c40fea9c v2.5.5: Fix RAM leak for long videos via on-demand reconstruction
- Replace all_transformed_videos storage with lightweight batch_metadata indices
- Reconstruct transformed videos on-demand in Phase 4 only when needed
- Add missing cleanup for input_images tensor in postprocess finally block
- Fix release_tensor_memory to handle CPU/CUDA/MPS consistently
- Extract helper functions for batch preparation and 4n+1 padding
- Remove duplicate interrupt_fn key from context initialization
2025-11-09 02:06:16 -05:00
Adrien Toupet e86d54ab3b fix: AdaIN color correction and AMD ROCm compatibility (v2.5.4)
- Fix AdaIN non-contiguous tensor error by using reshape() instead of view()
- Add cuDNN availability checks to prevent ROCm 'ATen not compiled with cuDNN' error
2025-11-08 09:11:55 -05:00
Adrien Toupet 786fb3f688 fix: correct MPS device enumeration for Apple Silicon (v2.5.3) 2025-11-08 08:25:22 -05:00
Adrien Toupet 9715d3e37a Fix: torch.mps AttributeError on Windows
Add defensive checks for torch.mps.is_available() to handle PyTorch versions where the method doesn't exist on non-Mac platforms. Resolves AttributeError: module 'torch.mps' has no attribute 'is_available'
2025-11-07 23:37:35 -05:00
Adrien Toupet 4e9ce4710e Unify parameter names across CLI and ComfyUI interface
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
2025-11-06 22:51:20 -05:00
Adrien Toupet afd5950d37 feat: torch.compile optimization and memory tracking improvements
- Eliminate graph breaks in na.py and attention.py for full torch.compile support
  * Replace cumsum-based tensor slicing with _tensor_split to avoid .item() calls
  * Use torch.tensor_split with .long().cpu() for PyTorch API requirements
  * Replace reshape-based averaging with split-stack-mean pattern
- Fix phase peak VRAM tracking to capture peaks during OOM retry cycles
- Standardize phase4 naming to 'postprocessing' for consistency
- Update RoPE docstrings for accuracy
- Update example workflows with icon
2025-11-06 17:48:39 -05:00
Adrien Toupet 20dab62dc3 feat: add uniform_batch_size for temporal consistency + unify padding logic
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
2025-11-06 14:13:22 -05:00
Adrien Toupet f8627298f7 Add peak VRAM tracking and summary display by phase 2025-11-05 21:21:23 -05:00
Adrien Toupet 4e728ad953 docs: comprehensive README overhaul for v2.5.0 release
- Remove nightly branch warning (deploying to main)
- Add Future Releases section with community engagement links
- Expand Features into 8 organized categories (Core, Model Support, Memory Optimization, Performance, Quality Control, Workflow)
- Update Requirements: document 8GB-24GB+ VRAM tiers with optimization strategies
- Improve Installation: add ComfyUI Manager as primary method, fix python_embeded syntax, use uv for manual install
- Complete Usage rewrite: document all 4 nodes (DiT/VAE loaders, Torch Compile, Main Upscaler) with parameters, tooltips, and examples
- Add BlockSwap and VAE Tiling detailed explanations with troubleshooting guides
- Provide 3 workflow templates: Basic (24GB+), Low VRAM (8-12GB), High Performance (torch.compile)
- Revamp CLI section: separate existing ComfyUI users from standalone installation, update all arguments
- Add Multi-GPU processing explanation with overlap blending example
- Remove outdated Benchmarks section
- Update Limitations: clarify 4n+1 batch size requirement, note VAE bottleneck, add best practices
- Simplify Contributing with link to CONTRIBUTING.md
- Fix device tooltips in DiT/VAE loaders (remove /CPU reference)
2025-11-05 17:27:28 -05:00
Adrien Toupet 806bb94df0 feat: unify and improve tooltip documentation across CLI and ComfyUI nodes
- Standardize tooltip format with multi-line descriptions and bullet points
- Add comprehensive output tooltips for all nodes (DiT, VAE, torch.compile, upscaler)
- Enhance node descriptions with detailed capability summaries
- Simplify CLI tile size arguments to single integers (converted internally to tuples)
- Remove OneOrTwoValues argparse class for cleaner implementation
- Fix encode_tiled tooltip (was incorrectly referencing decoding)
- Clarify color correction purpose (corrects upscaling color shifts)
- Add multi-GPU offloading information to all offload_device tooltips
- Improve torch.compile parameter descriptions with use cases
- Ensure CLI and ComfyUI tooltips are consistent in terminology and structure
2025-11-05 15:35:22 -05:00
Adrien Toupet 9b79254c39 refactor(cli): improvements and bug fixes + 3b-Q8_0.gguf support
- Fix validation cache location to respect --model_dir parameter
- Fix output path handling for directories without extensions
- Remove spurious directory creation in get_base_cache_dir
- Enhanced CLI help with usage examples and argument grouping
- Auto-display help when script invoked without arguments
- Correct type hints (device_id: str, debug: Debug)
- Remove redundant type conversions and makedirs calls
- Reorganize imports to module top for clarity
- Improved docstrings & tooltip
- Change default batch_size from 1 to 5 to match ComfyUI integration
- Add support for seedvr2_ema_3b-Q8_0.gguf model
2025-11-05 00:28:19 -05:00
Adrien Toupet 3725c1061d refactor: centralize dimension computation and logging for CLI/ComfyUI
- Add compute_generation_info() and log_generation_start() helpers
- Move prepend_frames logic from extraction to processing pipeline
- Eliminate code duplication between CLI and ComfyUI workflows
- Add consistent dimension/parameter logging for both interfaces
- Disable argparse prefix matching for safer CLI usage
2025-11-04 17:01:01 -05:00
Adrien Toupet 32a049dfd9 feat: Add CLI model caching for multi-file processing and unify device handling
- Add --cache_dit and --cache_vae flags for efficient multi-file directory processing
- Refactor processing pipeline to eliminate duplication between worker and direct modes
- Implement platform-agnostic device management (CUDA/MPS/CPU)
- Unify parameter naming: res_w→resolution, max_res_w→max_resolution across codebase
- Add smart offload device defaults when caching enabled
- Improve validation and user feedback for cache + multi-GPU scenarios
2025-11-04 15:12:28 -05:00
Adrien Toupet ad020d3803 feat(cli): improve UX with auto-format detection, FPS tracking, and consistent messaging with ComfyUI implementation
- Auto-detect output format per file type (mp4 for videos, png for images)
- Add visual separators between processed files for better readability
- Simplify FPS calculation to use wall-clock time for real-world throughput
- Consolidate banner/footer into shared Debug methods
- Update offload device args to support multi-GPU (cpu/cuda:N)
- Standardize terminology: 'upscaling' instead of 'video upscaling'
- Remove code duplication between CLI and ComfyUI implementations
- Consistent quote style (double quotes) throughout CLI
2025-11-04 11:54:57 -05:00
Adrien Toupet ad1d775eaa fix: tile debug overlay now applies to all batches and adds overlay warning 2025-11-04 09:06:13 -05:00