421 Commits
Author SHA1 Message Date
Adrien Toupet 5a4bf428f3 Merge pull request #438 from AInVFX/main
v2.5.23: Security hardening, GGUF VAE support, FFmpeg stability, MPS optimization
v2.5.23
2025-12-23 21:09:01 -05:00
Adrien Toupet 43e70bf637 Release v2.5.23: Security & stability improvements
- Add security protection against malicious .pth files
- Fix FFmpeg video writer hanging issues (thanks @thehhmdb)
- Enable GGUF VAE model support via conv dequantization (thanks @naxci1)
- Fix VAE slicing division by zero edge cases (thanks @naxci1)
- Resolve LAB color transfer dtype mismatch errors
- Extend Conv3d memory workaround to PyTorch 2.9+
- Fix bitsandbytes compatibility on non-Gaudi systems
- Optimize MPS memory usage (thanks @s-cerevisiae)
2025-12-24 03:02:34 +01:00
Adrien Toupet 855f8b91b3 Add sponsor call-to-action to footer 2025-12-24 02:49:42 +01:00
Adrien Toupet f561743054 Fix #434: Resolve dtype mismatch in LAB color transfer during video upscaling
Add explicit dtype alignment before matrix multiplication in _rgb_to_lab_batch
and _lab_to_rgb_batch to prevent float64 promotion from torch.pow operations.
Fixes RuntimeError: expected mat1 and mat2 to have the same dtype.
2025-12-24 02:22:46 +01:00
Adrien Toupet 396f323eae Fix #437: Catch ValueError in bitsandbytes compatibility shim
Add ValueError to exception handling to catch packaging.InvalidVersion
errors during Intel Gaudi version detection on non-Gaudi systems.
2025-12-24 02:14:33 +01:00
Adrien Toupet 6226878411 Apply critical fixes from PR #421
Fixes applied:
- GGUF conv2d/conv3d dequantization in __torch_function__
  Critical fix: Makes GGUF VAE models functional by properly handling
  InflatedCausalConv3d layers that aren't replaced by layer replacement

- Division by zero protection in slicing_latent_min_size calculations
  Defensive: Prevents crashes with edge-case temporal_downsample_factor values

Changes rejected from PR #421:
- NODE_CLASS_MAPPINGS (violates ComfyUI V3 API schema)
- Triton auto-fallback to cudagraphs (cudagraphs has compatibility issues)
- VAE 4D optimization (unproven benefit)
- force_upcast default change (untested breaking change)
2025-12-24 01:58:58 +01:00
Adrien Toupet 2f8d2ccf9a Merge PR #421 from naxci1 - Optimize VAE and GGUF 2025-12-24 01:13:19 +01:00
Adrien Toupet b0f01f2d99 Fix MPS device check precision for memory leak workaround
Improves PR #428 by checking actual VAE device (self.device.type == 'mps')
instead of system MPS availability. Only clears cache when VAE operations
are running on MPS, avoiding unnecessary overhead when VAE runs on CPU.

This fix addresses the PyTorch MPS memory leak (pytorch/pytorch#155060)
where padding operations in convolutions accumulate memory during encode
and decode. Clears MPS cache after each operation to prevent accumulation.

Fixes #363 (absurd MPS VRAM usage - 52GB → 12GB)
Fixes #410 (macOS system restarts from memory exhaustion)
Fixes #415 (inability to upscale beyond 2K resolution)
May help #417 (convolution errors under memory pressure)
2025-12-24 00:46:20 +01:00
Adrien Toupet 2214f3afde Merge PR #428: Reduce MPS memory usage 2025-12-24 00:40:59 +01:00
Adrien Toupet aeebd49f7f Refine FFMPEGVideoWriter to prevent pipe blocking
Building on @thehhmdb's fix in PR #418 which identified the stderr
pipe blocking issue. This refinement simplifies the solution by
redirecting stderr to DEVNULL and adding stdin.flush() to prevent
buffering deadlocks.

Improvements:
- Simpler implementation without threading complexity
- Zero memory overhead
- Better error messages for debugging
- Maintains fix for the original hanging issue

Co-authored-by: thehhmdb <thehhmdb@users.noreply.github.com>
Fixes numz/ComfyUI-SeedVR2_VideoUpscaler#418
2025-12-24 00:07:41 +01:00
Adrien Toupet e178b72d89 Merge commit '7bb936749f2799cab03e5d1593d1815a37c5233d' 2025-12-23 23:49:01 +01:00
Adrien Toupet 8ad4c8fa4e sec: prevent RCE vulnerability in .pth model loading
Add weights_only=True to torch.load() to restrict deserialization
to tensors only, preventing arbitrary code execution via pickle
2025-12-23 23:19:51 +01:00
Adrien Toupet 241b632cfc fix: extend Conv3d workaround to PyTorch 2.9+ (fixes 3x VAE VRAM usage in 2.11+) 2025-12-21 16:21:07 -05:00
spore 27ed3333fd perf(mps): reduce memory usage by clearing cache 2025-12-20 18:32:11 +08:00
google-labs-jules[bot] c6997fd9c2 Optimize VAE/GGUF performance and fix node registration/compile bugs
- Implement node registration in `__init__.py` to fix "Node does not exist" error.
- Implement automatic fallback from `inductor` to `cudagraphs` in `torch.compile` when Triton is missing (Windows fix).
- Optimize VAE `InflatedCausalConv3d` to use 2D convolution path for spatial-only operations, improving speed.
- Optimize VAE `ResnetBlock3D` to support flattened 4D execution path to reduce reshape overhead.
- Update `VideoAutoencoderKL` to default `force_upcast=False` for FP16 inference.
- Add `conv2d`/`conv3d` dequantization support to `GGUFTensor` for GGUF model compatibility.
2025-12-15 18:22:47 +00:00
google-labs-jules[bot] 0e849d20cd Optimize VAE Decoding Speed and robustness
- Implemented "4D execution mode" in ResnetBlock3D to flatten temporal dimension when convolutions are effectively 2D, reducing reshape overhead.
- Updated InflatedCausalConv3d to support direct 4D input processing and use 2D convolution optimization path.
- Enhanced InflatedCausalConv3d check_effective_2d to strictly verify stride/dilation/padding compatibility.
- Fixed 4D input handling in InflatedCausalConv3d forward pass.
2025-12-15 16:51:05 +00:00
google-labs-jules[bot] a65ddadc00 Optimize VAE performance (FP16/2D Conv) and GGUF support
- Implemented 2D convolution optimization in `InflatedCausalConv3d` to speed up spatial-only operations by using `F.conv2d` instead of `Conv3d`.
- Added `torch.nn.functional.conv2d` and `torch.nn.functional.conv3d` to `GGUFTensor`'s `__torch_function__` dispatch to enable automatic dequantization of weights, supporting GGUF models.
- Updated `src/models/video_vae_v3/modules/attn_video_vae.py` to default `force_upcast` to `False` for better FP16 performance.
- Fixed a bug in VAE slicing logic where `slicing_latent_min_size` could become 0, now clamping it to 1.
- Updated `Upsample3D` and `Encoder3D` in `attn_video_vae.py` to use `init_causal_conv3d` for 1x1x1 convolutions, enabling the 2D optimization path.
2025-12-15 16:12:33 +00:00
google-labs-jules[bot] 48bfbae05d Optimize VAE for performance and GGUF support
- Implemented 2D convolution optimization in `InflatedCausalConv3d` to speed up spatial-only operations by using `F.conv2d` instead of `Conv3d`.
- Added `torch.nn.functional.conv2d` and `torch.nn.functional.conv3d` to `GGUFTensor`'s `__torch_function__` dispatch to enable automatic dequantization of weights, supporting GGUF models.
- Updated `src/models/video_vae_v3/modules/attn_video_vae.py` to default `force_upcast` to `False` for better FP16 performance.
- Fixed a bug in VAE slicing logic where `slicing_latent_min_size` could become 0, now clamping it to 1.
2025-12-15 16:03:17 +00:00
google-labs-jules[bot] 04475bdc24 Optimize VAE and improve GGUF support for 50-series GPUs
- Implemented 2D convolution optimization in `InflatedCausalConv3d` to speed up spatial-only operations by reshaping effectively 2D tensors and using `F.conv2d` instead of `Conv3d`.
- Added `torch.nn.functional.conv2d` and `torch.nn.functional.conv3d` to `GGUFTensor`'s `__torch_function__` dispatch to enable automatic dequantization of weights, allowing the VAE to utilize GGUF quantization.
- Fixed a bug in `VideoAutoencoderKL` where `slicing_latent_min_size` could become 0 with small split sizes, now clamping it to a minimum of 1.
2025-12-15 15:38:14 +00:00
thehhmdb 7bb936749f To prevent ffmpeg from hanging, patched FFMPEGVideoWriter to continuously consume ffmpeg stderr in a background thread, flush stdin on write, and raise a clear error (including stderr) on BrokenPipe; release now joins the thread and logs stderr on non-zero exit. 2025-12-14 15:50:27 +00:00
HB2k 4fc3296c81 Merge branch 'numz:main' into main 2025-12-13 14:07:53 +04:00
Adrien Toupet d69b65f7e4 Merge pull request #412 from AInVFX/main
v2.5.22: CLI FFmpeg 10-bit video backend, MPS bicubic fix, cross-platform histogram matching
v2.5.22
2025-12-13 00:36:12 -05:00
Adrien Toupet 15cb24089a Release v2.5.22: FFmpeg 10-bit video backend, MPS bicubic fix, cross-platform histogram matching
Note: index_select(out=) optimization removed as it caused color polarization on MPS; using simple indexing instead
2025-12-13 00:29:56 -05:00
Adrien Toupet c52280881a refactor: replace scatter_ with argsort+index_select for better cross-platform compatibility
- Replace tensor.scatter_() with torch.argsort() + torch.index_select(out=)
- Uses fundamental PyTorch ops for improved reliability across CUDA/ROCm/MPS
- Aggressive early tensor deletion to minimize memory overhead
- Affects _histogram_matching_channel and _histogram_match_1d in color_fix.py

Related: #351
2025-12-13 00:06:15 -05:00
Adrien Toupet 4b0b7d58b6 fix(cli): validate ffmpeg availability at startup
Move ffmpeg check from FFMPEGVideoWriter to argument validation phase.
Prevents wasted GPU processing time when ffmpeg backend is selected
but ffmpeg is not installed.
2025-12-12 23:39:16 -05:00
Adrien Toupet 39d8d4bf19 Refine MPS bicubic fix to use try/except for version compatibility (#408)
- Use try/except instead of blanket MPS check for bicubic+antialias
- PyTorch 2.8.0+ MPS: native fast path (no overhead)
- PyTorch < 2.8.0 MPS: CPU fallback on NotImplementedError
2025-12-12 23:19:37 -05:00
Adrien Toupet f2f4916c05 Fix MPS bicubic+antialias error for RGBA upscaling (#408)
- Add CPU fallback for F.interpolate with antialias=True on MPS (aten::_upsample_bicubic2d_aa not implemented)
- Revert torch.mps.synchronize() calls introduced in v2.5.21 for consistent behavior with CUDA pipeline
2025-12-12 23:08:16 -05:00
Adrien Toupet f75bcc7f37 feat(cli): add ffmpeg video backend with 10-bit support
- Add --video_backend flag: 'opencv' (default) or 'ffmpeg'
- Add --10bit flag: enables x265/yuv420p10le for reduced banding
- Without --10bit, ffmpeg uses x264/yuv420p for max compatibility
- FFMPEGVideoWriter class with cv2.VideoWriter-compatible interface
- Validates ffmpeg availability before encoding

Based on PR #409 by thehhmdb
2025-12-12 21:39:47 -05:00
thehhmdb 0c2a546c12 Add option to use ffmpeg and 10-bit video to reduce blocking and banding 2025-12-12 20:43:15 -05:00
Adrien Toupet 32f9900ecd Merge pull request #407 from AInVFX/main
v2.5.21: fix GGUF dequant regression, MPS performance optimizations
v2.5.21
2025-12-12 11:28:07 -05:00
Adrien Toupet 84abef8de0 Release v2.5.21: fix GGUF dequant regression on MPS, eliminate CPU sync overhead on unified memory 2025-12-12 11:22:38 -05:00
Adrien Toupet f3136dd20c Fix GGUF dequantization shape error on MPS (#403)
Skip GGUF quantized buffers in _force_nadit_precision - these must
remain in packed format for on-the-fly dequantization during inference.
2025-12-12 11:05:31 -05:00
Adrien Toupet 93a6355517 perf(mps): eliminate sync overhead from CPU tensor offload on unified memory
- Skip CPU tensor offload on MPS (no memory benefit, causes sync stall)
- Keep input_images and final_video on MPS device
- Add explicit MPS sync at phase boundaries for accurate timing
- Preload text embeddings before Phase 1 to avoid Phase 2 stall
- Skip model→CPU movement before deletion on MPS cleanup
2025-12-12 10:52:55 -05:00
HB2k 5c07a92b33 Merge branch 'numz:main' into main 2025-12-12 13:46:50 +04:00
Adrien Toupet a1486a30fe Merge pull request #402 from AInVFX/main
v2.5.20: expanded attention backends (FA2/FA3/SA2/SA3), macOS MPS dtype fixes, bitsandbytes ROCm shim
v2.5.20
2025-12-12 00:48:00 -05:00
Adrien Toupet bbf649d34a Release v2.5.20: expanded attention backends (FA2/FA3/SA2/SA3), macOS MPS dtype fixes, bitsandbytes ROCm shim, flash-attn DLL fallback 2025-12-12 00:40:02 -05:00
Adrien Toupet b852d5fb22 Fix bitsandbytes kernel registration conflict on ROCm systems (#362)
Add ensure_bitsandbytes_safe() shim to handle broken/partial bitsandbytes
installations that cause PyTorch kernel registration conflicts when diffusers
attempts to re-import the module.

On ROCm systems without proper binaries, bitsandbytes registers kernels during
import then fails. When diffusers later imports it, the duplicate registration
causes: 'RuntimeError: already a kernel registered...int8_mm_dequant'

The shim pre-tests bitsandbytes import and stubs it only if broken, allowing
working installations to function normally for other nodes.
2025-12-12 00:10:29 -05:00
Adrien Toupet 7ba37c0557 Remove NVIDIA CUDA classifier for macOS compatibility (#395)
Package supports both CUDA and MPS - classifier was causing
ComfyUI Manager to show false 'GPU not supported' warning on macOS
2025-12-11 23:40:16 -05:00
Adrien Toupet fa2e3e79f8 Fix MPS/macOS compatibility for GGUF models (#401)
- CompatibleDiT now converts ALL model params to compute_dtype on MPS
  (previously only FP8 - GGUF models had mixed FP16/BF16 causing hangs)
- Replace MPS autocast with explicit dtype conversion in VAE encode/decode
- Skip DiT autocast on MPS (CompatibleDiT handles dtype internally)
- Guard call_rope_with_stability CUDA autocast for non-CUDA devices
- Add weights_only=True to torch.load (FutureWarning fix)
- Rename FP8CompatibleDiT → CompatibleDiT

Addresses M4 Pro macOS hang at EulerSampler 0% with GGUF models
2025-12-11 23:30:33 -05:00
Adrien Toupet ea0fbc689d Centralize BlockSwap validation, auto-disable on macOS, update docs
- Add validate_blockswap_config() in blockswap.py as single validation point
- Auto-disable BlockSwap on macOS (unified memory makes it meaningless)
- Improve error messages for missing dit_offload_device
- Update CLI and ComfyUI tooltips for BlockSwap and model caching
- Update README: BlockSwap macOS note, caching descriptions, attention backends
- Remove duplicate validation from dit_model_loader.py and inference_cli.py

Partially fixes #401 (M4 Pro macOS BlockSwap offload device error)
2025-12-11 22:38:51 -05:00
Adrien Toupet 2911b78288 feat: Separate Flash Attention 2/3 and SageAttention 2/3 backends
- Rename attention modes: flash_attn→flash_attn_2/3, sa2/sa3→sageattn_2/3
- Add separate detection and wrappers for FA2, FA3, SA2, SA3 in compatibility.py
- FA3: Filter unsupported params (dropout_p, window_size), return tuple[0]
- SA2/SA3: Add half-precision dtype handling (convert fp32/fp8→bf16)
- SA3: Add varlen-to-batched conversion with SA2 fallback for non-uniform seqs
- Add fallback chains: FA3→FA2→SDPA, SA3→SA2→SDPA
- Update debug.py to show granular availability: FlashAttn / SageAttn
- Update all references: README, CLI, ComfyUI nodes, docstrings
2025-12-10 15:26:16 -05:00
Adrien Toupet bcfbca6ae3 feat: add SageAttention (sa2/sa3) support, centralize attention wrappers
- Add sa2/sa3 attention modes for SageAttention v2/v3 kernels
- Centralize call_flash_attn_varlen and call_sage_attn_varlen in compatibility.py
- Remove duplicated attention wrapper code from dit_3b/dit_7b attention.py
- Rename validate_flash_attention_availability to validate_attention_mode
- Remove unnecessary precision control feature (auto/fp16/bf16/bf32)
- Remove unused detect_high_end_system() and log_system_capabilities()
- Update startup logging to show SageAttention availability status
- Update CLI and ComfyUI node to expose sa2/sa3 options
2025-12-10 13:58:46 -05:00
naxci1 ab3284982f Add SageAttention optimization (PR #387) 2025-12-10 11:40:40 -05:00
Adrien Toupet e842538cfa Fix graceful fallback from flash-attn #376
Add compatibility shims for corrupted/missing flash_attn and xformers DLLs.
Stubs include proper __spec__ to prevent importlib.util.find_spec() crashes.
Force-verify flash_attn_2_cuda at startup; fall back to SDPA if unavailable.
2025-12-10 09:45:20 -05:00
Adrien Toupet 2006fa3f6c Merge pull request #390 from AInVFX/main
v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking
v2.5.19
2025-12-10 01:55:47 -05:00
Adrien Toupet 118c9fcbe7 Release v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking, revert VRAM limit 2025-12-10 01:52:45 -05:00
Adrien Toupet 6106681563 Fix graceful fallback from flash-attn #376
Add compatibility shims for corrupted flash_attn/xformers DLLs.
Force-verify flash_attn_2_cuda at startup; fall back to SDPA if unavailable.
2025-12-10 01:44:49 -05:00
Adrien Toupet c010deeea1 Remove ineffective allow_vram_overflow setting
- PyTorch's set_per_process_memory_fraction cannot prevent WDDM paging on Windows
- Keep overflow detection and warning when VRAM exceeds physical limit
- Simplify peak memory formatting
- Remove setting from CLI, ComfyUI node, and memory_manager
2025-12-10 00:56:27 -05:00
Adrien Toupet 7cbf025561 Fix VRAM peak tracking: separate allocated vs reserved, Windows-only overflow
- Track both peak_allocated (tensor usage) and peak_reserved (cache pool) per phase
- peak_allocated resets properly between phases via reset_peak_memory_stats()
- Overflow detection/warnings now Windows-only (WDDM paging behavior)
- Remove get_memory_architecture() - replaced with simple is_mps + platform checks
- Phase summary shows: VRAM XGB allocated, YGB reserved | RAM ZGB
- Simplify MPS path (unified memory has no overflow concept)
2025-12-09 23:51:51 -05:00
Adrien Toupet 5c60716c47 Refactor: centralize backend detection, fix architecture-aware VRAM overflow reporting 2025-12-09 21:06:12 -05:00