81 Commits
Author SHA1 Message Date
Adrien Toupet baec4b634f fix(vae): restore MPS memory leak workaround removed in v2.5.23 cleanup 2025-12-24 09:50:19 +01:00
Adrien Toupet 43e70bf637 Release v2.5.23: Security & stability improvements
- Add security protection against malicious .pth files
- Fix FFmpeg video writer hanging issues (thanks @thehhmdb)
- Enable GGUF VAE model support via conv dequantization (thanks @naxci1)
- Fix VAE slicing division by zero edge cases (thanks @naxci1)
- Resolve LAB color transfer dtype mismatch errors
- Extend Conv3d memory workaround to PyTorch 2.9+
- Fix bitsandbytes compatibility on non-Gaudi systems
- Optimize MPS memory usage (thanks @s-cerevisiae)
2025-12-24 03:02:34 +01:00
Adrien Toupet 15cb24089a Release v2.5.22: FFmpeg 10-bit video backend, MPS bicubic fix, cross-platform histogram matching
Note: index_select(out=) optimization removed as it caused color polarization on MPS; using simple indexing instead
2025-12-13 00:29:56 -05:00
Adrien Toupet f75bcc7f37 feat(cli): add ffmpeg video backend with 10-bit support
- Add --video_backend flag: 'opencv' (default) or 'ffmpeg'
- Add --10bit flag: enables x265/yuv420p10le for reduced banding
- Without --10bit, ffmpeg uses x264/yuv420p for max compatibility
- FFMPEGVideoWriter class with cv2.VideoWriter-compatible interface
- Validates ffmpeg availability before encoding

Based on PR #409 by thehhmdb
2025-12-12 21:39:47 -05:00
Adrien Toupet 84abef8de0 Release v2.5.21: fix GGUF dequant regression on MPS, eliminate CPU sync overhead on unified memory 2025-12-12 11:22:38 -05:00
Adrien Toupet bbf649d34a Release v2.5.20: expanded attention backends (FA2/FA3/SA2/SA3), macOS MPS dtype fixes, bitsandbytes ROCm shim, flash-attn DLL fallback 2025-12-12 00:40:02 -05:00
Adrien Toupet ea0fbc689d Centralize BlockSwap validation, auto-disable on macOS, update docs
- Add validate_blockswap_config() in blockswap.py as single validation point
- Auto-disable BlockSwap on macOS (unified memory makes it meaningless)
- Improve error messages for missing dit_offload_device
- Update CLI and ComfyUI tooltips for BlockSwap and model caching
- Update README: BlockSwap macOS note, caching descriptions, attention backends
- Remove duplicate validation from dit_model_loader.py and inference_cli.py

Partially fixes #401 (M4 Pro macOS BlockSwap offload device error)
2025-12-11 22:38:51 -05:00
Adrien Toupet 2911b78288 feat: Separate Flash Attention 2/3 and SageAttention 2/3 backends
- Rename attention modes: flash_attn→flash_attn_2/3, sa2/sa3→sageattn_2/3
- Add separate detection and wrappers for FA2, FA3, SA2, SA3 in compatibility.py
- FA3: Filter unsupported params (dropout_p, window_size), return tuple[0]
- SA2/SA3: Add half-precision dtype handling (convert fp32/fp8→bf16)
- SA3: Add varlen-to-batched conversion with SA2 fallback for non-uniform seqs
- Add fallback chains: FA3→FA2→SDPA, SA3→SA2→SDPA
- Update debug.py to show granular availability: FlashAttn / SageAttn
- Update all references: README, CLI, ComfyUI nodes, docstrings
2025-12-10 15:26:16 -05:00
Adrien Toupet 118c9fcbe7 Release v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking, revert VRAM limit 2025-12-10 01:52:45 -05:00
Adrien Toupet c010deeea1 Remove ineffective allow_vram_overflow setting
- PyTorch's set_per_process_memory_fraction cannot prevent WDDM paging on Windows
- Keep overflow detection and warning when VRAM exceeds physical limit
- Simplify peak memory formatting
- Remove setting from CLI, ComfyUI node, and memory_manager
2025-12-10 00:56:27 -05:00
Adrien Toupet 7cbf025561 Fix VRAM peak tracking: separate allocated vs reserved, Windows-only overflow
- Track both peak_allocated (tensor usage) and peak_reserved (cache pool) per phase
- peak_allocated resets properly between phases via reset_peak_memory_stats()
- Overflow detection/warnings now Windows-only (WDDM paging behavior)
- Remove get_memory_architecture() - replaced with simple is_mps + platform checks
- Phase summary shows: VRAM XGB allocated, YGB reserved | RAM ZGB
- Simplify MPS path (unified memory has no overflow concept)
2025-12-09 23:51:51 -05:00
Adrien Toupet 77a00f651a Fix: OOM regression from 2.5.14 strict VRAM limit (#367)
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.

The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.

- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow

Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
2025-12-09 17:12:10 -05:00
Adrien Toupet b101deb894 docs: fix contributor links and formatting in release notes 2025-12-09 01:01:54 -05:00
Adrien Toupet 06be9c9d7a Release v2.5.18: CLI streaming mode, multi-GPU streaming with caching, shared memory fix 2025-12-09 00:59:22 -05:00
Adrien Toupet a70d82e3aa Add streaming mode for memory-efficient long video processing
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup

Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
2025-12-08 22:05:20 -05:00
Adrien Toupet 3eec5847c0 Release v2.5.17: Older GPU compatibility fix 2025-12-05 21:05:30 -05:00
Adrien Toupet 11239eed13 Release v2.5.16: Quality regression fix, older GPU compatibility fix, system info debug 2025-12-05 15:50:08 -05:00
Adrien Toupet aa968cf2c9 docs: simplify contribution workflow to main branch only 2025-12-05 13:54:39 -05:00
Adrien Toupet 65c1c1b6cd Release v2.5.15: MPS fixes, autocast device type, accurate VRAM tracking, triton 3.0 compatibility 2025-12-03 13:14:51 -05:00
Adrien Toupet e2faedaaa6 Release v2.5.14 - MPS device fix, VRAM swap detection, enforce physical VRAM limit 2025-12-01 00:30:49 -05:00
Adrien Toupet 19825fa9fa Release v2.5.13: Fix triton import, OOM on long videos, macOS watermark 2025-11-30 09:01:53 -05:00
Adrien Toupet 16508d353d release: v2.5.12 - fix color artifacts regression from in-place transform ops 2025-11-28 17:50:40 -05:00
Adrien Toupet f791631495 release: v2.5.11 - CUDNN attention, long video memory fix, LAB artifacts fix, MPS/multi-GPU improvements 2025-11-28 16:49:08 -05:00
Adrien Toupet 70c7f21eb5 docs: clarify multi-GPU frame-level parallelism behavior in README 2025-11-28 14:00:59 -05:00
Adrien Toupet 65dd29a865 v2.5.10: Fix determinism, BlockSwap caching, and model path resolution
Core Fixes:
- Reset seed per batch to ensure deterministic generation across sessions and batch positions
- Fix temporal overlap logging when automatically reset to prevent incorrect frame counts
- Fix NoneType attribute error in VAE tiled encode/decode at maximum resolution (#296)

BlockSwap & Caching Architecture:
- Move BlockSwap state (_block_swap_config, _blockswap_bypass_protection) from runner to model
- Ensures state survives runner recreation during independent DiT/VAE caching scenarios
- Fix runner template caching to trigger when either DiT or VAE becomes cached (bidirectional)
- Resolves BlockSwap reload failures when only DiT was cached (#297)

Model Discovery:
- Implement case-insensitive YAML path resolution for extra_model_paths.yaml (#289-#295)
- Add debug logging for model discovery (searched paths, validation status, cache hits)
- Support any case variation (seedvr2, SEEDVR2, SeedVR2) in ComfyUI configuration
2025-11-13 12:02:37 -05:00
Adrien Toupet 40d37b7248 Release v2.5.9: Bug fixes and enhancements
- Fix: OpenCV memory layout error in tile debug visualization (#283)
- Fix: macOS MPS allocator fallback for safetensors loading (#290)
- Fix: Windows log buffering with flush=True (#278)
- Fix: ComfyUI registry icon URL (raw.githubusercontent)
- Feature: Version display in node name and CLI/ComfyUI header
- Feature: GitHub Sponsors support link
- License: Migrate from MIT to Apache 2.0 to match Bytedance Seed official repo
2025-11-12 14:27:40 -05:00
Adrien Toupet fc64968b12 fix(cli): improve output paths and add RGBA support (v2.5.8)
- Improve output folder naming: batch creates {folder}_upscaled/ sibling with original filenames, single file adds _upscaled suffix
- Add RGBA alpha channel detection and preservation (matches ComfyUI)
- Convert all output paths to absolute for clarity in logs
2025-11-10 14:39:33 -05:00
Adrien Toupet b130a33894 fix(cli): resolve Windows duplicate file bug and improve scan perf 2-3x (v2.5.8)
- Replace dual glob loops with single iterdir scan for cross-platform consistency
- Fixes duplicate file processing in batch mode on Windows case-insensitive filesystem
- Improves directory scanning performance 2-3x by reducing filesystem operations
- Add ComfyUI registry logo
2025-11-10 13:39:42 -05:00
Adrien Toupet b64d526bd6 fix: enhance Conv3d workaround compatibility for PyTorch dev builds and AMD ROCm
- Add ROCm/HIP detection to prevent NVIDIA-specific cuDNN workaround on AMD systems
- Add defensive hasattr checks for torch.cuda and cudnn.is_available()
- Add try-except fallback in _conv_forward to handle cuDNN call failures gracefully
- Resolves 'GET was unable to find an engine' errors on PyTorch 2.9+ dev builds
- Resolves 'ATen not compiled with cuDNN support' errors on ROCm platforms
- Bump version to 2.5.7
2025-11-10 13:04:52 -05:00
Adrien Toupet a3f98124ca Fix: Restore natural look for 7b model (v2.5.6)
- Replace split-stack-mean with unflatten in unconcat_coalesce
- Corrects computation order to eliminate plastic/high-specular artifacts
- Maintains torch.compile compatibility (no .item() graph breaks)
- Applied to both dit_3b and dit_7b models
2025-11-09 12:26:18 -05:00
Adrien Toupet b5c40fea9c v2.5.5: Fix RAM leak for long videos via on-demand reconstruction
- Replace all_transformed_videos storage with lightweight batch_metadata indices
- Reconstruct transformed videos on-demand in Phase 4 only when needed
- Add missing cleanup for input_images tensor in postprocess finally block
- Fix release_tensor_memory to handle CPU/CUDA/MPS consistently
- Extract helper functions for batch preparation and 4n+1 padding
- Remove duplicate interrupt_fn key from context initialization
2025-11-09 02:06:16 -05:00
Adrien Toupet e86d54ab3b fix: AdaIN color correction and AMD ROCm compatibility (v2.5.4)
- Fix AdaIN non-contiguous tensor error by using reshape() instead of view()
- Add cuDNN availability checks to prevent ROCm 'ATen not compiled with cuDNN' error
2025-11-08 09:11:55 -05:00
Adrien Toupet 786fb3f688 fix: correct MPS device enumeration for Apple Silicon (v2.5.3) 2025-11-08 08:25:22 -05:00
Adrien Toupet 65995a4c01 docs: Adding deep dive tutorial video 2025-11-07 20:09:20 -05:00
Adrien Toupet bc16f3f32a docs: tweak release notes 2025-11-07 08:40:31 -05:00
Adrien Toupet fccd34fc5d docs: release v2.5.0 with comprehensive changelog and updated usage screenshots 2025-11-06 23:59:44 -05:00
Adrien Toupet d82234d418 docs: Add individual node screenshots and update README images 2025-11-06 23:15:56 -05:00
Adrien Toupet 4e9ce4710e Unify parameter names across CLI and ComfyUI interface
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
2025-11-06 22:51:20 -05:00
Adrien Toupet 690cc39379 docs: improve best practices OOM troubleshooting guidance 2025-11-06 14:40:59 -05:00
Adrien Toupet 20dab62dc3 feat: add uniform_batch_size for temporal consistency + unify padding logic
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
2025-11-06 14:13:22 -05:00
Adrien Toupet 3b904b3798 docs: simplify installation steps to unified .venv structure 2025-11-05 20:58:20 -05:00
Adrien Toupet 7b12801af1 fix(docs): fixed broken links in README 2025-11-05 17:34:21 -05:00
Adrien Toupet 4e728ad953 docs: comprehensive README overhaul for v2.5.0 release
- Remove nightly branch warning (deploying to main)
- Add Future Releases section with community engagement links
- Expand Features into 8 organized categories (Core, Model Support, Memory Optimization, Performance, Quality Control, Workflow)
- Update Requirements: document 8GB-24GB+ VRAM tiers with optimization strategies
- Improve Installation: add ComfyUI Manager as primary method, fix python_embeded syntax, use uv for manual install
- Complete Usage rewrite: document all 4 nodes (DiT/VAE loaders, Torch Compile, Main Upscaler) with parameters, tooltips, and examples
- Add BlockSwap and VAE Tiling detailed explanations with troubleshooting guides
- Provide 3 workflow templates: Basic (24GB+), Low VRAM (8-12GB), High Performance (torch.compile)
- Revamp CLI section: separate existing ComfyUI users from standalone installation, update all arguments
- Add Multi-GPU processing explanation with overlap blending example
- Remove outdated Benchmarks section
- Update Limitations: clarify 4n+1 batch size requirement, note VAE bottleneck, add best practices
- Simplify Contributing with link to CONTRIBUTING.md
- Fix device tooltips in DiT/VAE loaders (remove /CPU reference)
2025-11-05 17:27:28 -05:00
Adrien Toupet 32a049dfd9 feat: Add CLI model caching for multi-file processing and unify device handling
- Add --cache_dit and --cache_vae flags for efficient multi-file directory processing
- Refactor processing pipeline to eliminate duplication between worker and direct modes
- Implement platform-agnostic device management (CUDA/MPS/CPU)
- Unify parameter naming: res_w→resolution, max_res_w→max_resolution across codebase
- Add smart offload device defaults when caching enabled
- Improve validation and user feedback for cache + multi-GPU scenarios
2025-11-04 15:12:28 -05:00
Adrien Toupet 9268346388 feat: CLI Add batch processing, fix multiprocessing issues, and unify model paths
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)

Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable

Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability

Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
2025-11-04 00:12:13 -05:00
Adrien Toupet b41998c5e1 docs: Clean up README formatting 2025-09-16 15:34:48 -04:00
Adrien Toupet 2c570105b1 docs: Update project documentation and credits
- Fix project names and URLs in CONTRIBUTING.md
- Add dual branch workflow (main/nightly) guidelines
- Update LICENSE with correct copyright holders
- Expand README credits to acknowledge all contributors
- Add proper contact information for maintainers
2025-09-16 15:29:39 -04:00
Adrien Toupet 6412f03484 Merge pull request #92 & #106
**Architecture & Performance:**
- Unified debug system with categorized logging and memory tracking
- FP8 models now stay in FP8, convert to BF16 only for math operations (faster, less memory)
- Fixed memory leaks
- Removed ComfyUI dependency for standalone compatibility
- Better VRAM management between batches

**New Features:**
- VAE tiling for larger/longer video upscaling
- Multi-repo model support (numz/ and AInVFX/)
- Model auto-discovery in ComfyUI folder
- CLI: BlockSwap options, temporal overlap blending, prepend_frames for better first frames
- `enable_debug` and `cache_model` moved to main node

**Code Organization:**
- New modular structure: `constants.py`, `model_registry.py`, `debug.py`
- Removed legacy code and dead files
- Added mixed FP8 models to fix 7B artifacts
2025-08-12 14:59:37 +02:00
Adrien Toupet 0029dd6283 Added nightly disclaimer 2025-08-12 08:26:25 -04:00
Adrien Toupet 28d3d5db2a Update Readme Updates section 2025-08-08 00:53:19 +02:00