386 Commits
Author SHA1 Message Date
Adrien Toupet a1486a30fe Merge pull request #402 from AInVFX/main
v2.5.20: expanded attention backends (FA2/FA3/SA2/SA3), macOS MPS dtype fixes, bitsandbytes ROCm shim
v2.5.20
2025-12-12 00:48:00 -05:00
Adrien Toupet bbf649d34a Release v2.5.20: expanded attention backends (FA2/FA3/SA2/SA3), macOS MPS dtype fixes, bitsandbytes ROCm shim, flash-attn DLL fallback 2025-12-12 00:40:02 -05:00
Adrien Toupet b852d5fb22 Fix bitsandbytes kernel registration conflict on ROCm systems (#362)
Add ensure_bitsandbytes_safe() shim to handle broken/partial bitsandbytes
installations that cause PyTorch kernel registration conflicts when diffusers
attempts to re-import the module.

On ROCm systems without proper binaries, bitsandbytes registers kernels during
import then fails. When diffusers later imports it, the duplicate registration
causes: 'RuntimeError: already a kernel registered...int8_mm_dequant'

The shim pre-tests bitsandbytes import and stubs it only if broken, allowing
working installations to function normally for other nodes.
2025-12-12 00:10:29 -05:00
Adrien Toupet 7ba37c0557 Remove NVIDIA CUDA classifier for macOS compatibility (#395)
Package supports both CUDA and MPS - classifier was causing
ComfyUI Manager to show false 'GPU not supported' warning on macOS
2025-12-11 23:40:16 -05:00
Adrien Toupet fa2e3e79f8 Fix MPS/macOS compatibility for GGUF models (#401)
- CompatibleDiT now converts ALL model params to compute_dtype on MPS
  (previously only FP8 - GGUF models had mixed FP16/BF16 causing hangs)
- Replace MPS autocast with explicit dtype conversion in VAE encode/decode
- Skip DiT autocast on MPS (CompatibleDiT handles dtype internally)
- Guard call_rope_with_stability CUDA autocast for non-CUDA devices
- Add weights_only=True to torch.load (FutureWarning fix)
- Rename FP8CompatibleDiT → CompatibleDiT

Addresses M4 Pro macOS hang at EulerSampler 0% with GGUF models
2025-12-11 23:30:33 -05:00
Adrien Toupet ea0fbc689d Centralize BlockSwap validation, auto-disable on macOS, update docs
- Add validate_blockswap_config() in blockswap.py as single validation point
- Auto-disable BlockSwap on macOS (unified memory makes it meaningless)
- Improve error messages for missing dit_offload_device
- Update CLI and ComfyUI tooltips for BlockSwap and model caching
- Update README: BlockSwap macOS note, caching descriptions, attention backends
- Remove duplicate validation from dit_model_loader.py and inference_cli.py

Partially fixes #401 (M4 Pro macOS BlockSwap offload device error)
2025-12-11 22:38:51 -05:00
Adrien Toupet 2911b78288 feat: Separate Flash Attention 2/3 and SageAttention 2/3 backends
- Rename attention modes: flash_attn→flash_attn_2/3, sa2/sa3→sageattn_2/3
- Add separate detection and wrappers for FA2, FA3, SA2, SA3 in compatibility.py
- FA3: Filter unsupported params (dropout_p, window_size), return tuple[0]
- SA2/SA3: Add half-precision dtype handling (convert fp32/fp8→bf16)
- SA3: Add varlen-to-batched conversion with SA2 fallback for non-uniform seqs
- Add fallback chains: FA3→FA2→SDPA, SA3→SA2→SDPA
- Update debug.py to show granular availability: FlashAttn / SageAttn
- Update all references: README, CLI, ComfyUI nodes, docstrings
2025-12-10 15:26:16 -05:00
Adrien Toupet bcfbca6ae3 feat: add SageAttention (sa2/sa3) support, centralize attention wrappers
- Add sa2/sa3 attention modes for SageAttention v2/v3 kernels
- Centralize call_flash_attn_varlen and call_sage_attn_varlen in compatibility.py
- Remove duplicated attention wrapper code from dit_3b/dit_7b attention.py
- Rename validate_flash_attention_availability to validate_attention_mode
- Remove unnecessary precision control feature (auto/fp16/bf16/bf32)
- Remove unused detect_high_end_system() and log_system_capabilities()
- Update startup logging to show SageAttention availability status
- Update CLI and ComfyUI node to expose sa2/sa3 options
2025-12-10 13:58:46 -05:00
naxci1 ab3284982f Add SageAttention optimization (PR #387) 2025-12-10 11:40:40 -05:00
Adrien Toupet e842538cfa Fix graceful fallback from flash-attn #376
Add compatibility shims for corrupted/missing flash_attn and xformers DLLs.
Stubs include proper __spec__ to prevent importlib.util.find_spec() crashes.
Force-verify flash_attn_2_cuda at startup; fall back to SDPA if unavailable.
2025-12-10 09:45:20 -05:00
Adrien Toupet 2006fa3f6c Merge pull request #390 from AInVFX/main
v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking
v2.5.19
2025-12-10 01:55:47 -05:00
Adrien Toupet 118c9fcbe7 Release v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking, revert VRAM limit 2025-12-10 01:52:45 -05:00
Adrien Toupet 6106681563 Fix graceful fallback from flash-attn #376
Add compatibility shims for corrupted flash_attn/xformers DLLs.
Force-verify flash_attn_2_cuda at startup; fall back to SDPA if unavailable.
2025-12-10 01:44:49 -05:00
Adrien Toupet c010deeea1 Remove ineffective allow_vram_overflow setting
- PyTorch's set_per_process_memory_fraction cannot prevent WDDM paging on Windows
- Keep overflow detection and warning when VRAM exceeds physical limit
- Simplify peak memory formatting
- Remove setting from CLI, ComfyUI node, and memory_manager
2025-12-10 00:56:27 -05:00
Adrien Toupet 7cbf025561 Fix VRAM peak tracking: separate allocated vs reserved, Windows-only overflow
- Track both peak_allocated (tensor usage) and peak_reserved (cache pool) per phase
- peak_allocated resets properly between phases via reset_peak_memory_stats()
- Overflow detection/warnings now Windows-only (WDDM paging behavior)
- Remove get_memory_architecture() - replaced with simple is_mps + platform checks
- Phase summary shows: VRAM XGB allocated, YGB reserved | RAM ZGB
- Simplify MPS path (unified memory has no overflow concept)
2025-12-09 23:51:51 -05:00
Adrien Toupet 5c60716c47 Refactor: centralize backend detection, fix architecture-aware VRAM overflow reporting 2025-12-09 21:06:12 -05:00
Adrien Toupet 77a00f651a Fix: OOM regression from 2.5.14 strict VRAM limit (#367)
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.

The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.

- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow

Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
2025-12-09 17:12:10 -05:00
Adrien Toupet 30bc924043 Update header logo design (thanks @naxci1, closes #378) 2025-12-09 14:07:46 -05:00
Adrien Toupet e65e7fa418 Remove dead flash attention wrapper from FP8CompatibleDiT
The wrapper methods (_apply_flash_attention_optimization and related)
matched NaDiT attention modules by name but required qkv or q_proj+k_proj+v_proj
attributes to optimize. NaDiT uses proj_qkv instead, so the optimization
path was never taken - always falling back to original forward.

FlashAttentionVarlen already handles flash_attn vs sdpa switching via
its attention_mode attribute, making this wrapper redundant.

Removes ~200 lines of dead code.
2025-12-09 12:29:15 -05:00
google-labs-jules[bot] 8f79511d21 Fix SageAttention, restore strict precision control, and fix crashes (v2)
- Re-implemented `precision` control (`fp16`, `bf16`, `bf32`, `auto`) in CLI, ComfyUI node, and backend logic to respect user choice.
- Fixed `UnboundLocalError` in `apply_model_specific_config` by ensuring `compute_dtype` is always initialized before use.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by properly passing the restored `precision` argument.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` for clarity and fixed fallback logic to ensure SageAttention is correctly prioritized.
- Added explicit logging of active attention backend and execution confirmation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
2025-12-09 17:09:27 +00:00
google-labs-jules[bot] 9a57539d0a Fix SageAttention naming, restore strict precision control, and fix crashes
- Re-implemented `precision` control (`fp16`, `bf16`, `bf32`, `auto`) in CLI, ComfyUI node, and backend logic to respect user choice.
- Fixed `UnboundLocalError` in `apply_model_specific_config` by ensuring `compute_dtype` is always initialized before use.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by properly passing the restored `precision` argument.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` for clarity and fixed fallback logic to ensure SageAttention is correctly prioritized.
- Added explicit logging of active attention backend and execution confirmation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
2025-12-09 15:27:29 +00:00
google-labs-jules[bot] a5e8406afa Fix SageAttention naming, enforce auto-precision, and enhance active mode logging
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` and CLI by removing all residual `precision` variable usage.
2025-12-09 14:26:57 +00:00
google-labs-jules[bot] 5e1e18e9d5 Fix SageAttention naming, enforce auto-precision, and enhance active mode logging
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` and CLI by removing all residual `precision` variable usage.
2025-12-09 12:17:51 +00:00
google-labs-jules[bot] fdd76e8c3f Fix SageAttention naming, enforce auto-precision, and enhance active mode logging
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by removing residual `precision` usage.
2025-12-09 11:58:12 +00:00
google-labs-jules[bot] 20f9132365 Fix SageAttention naming, enhance active mode logging, and enforce auto-precision
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
2025-12-09 11:43:58 +00:00
google-labs-jules[bot] afa47d4cc8 Fix SageAttention naming and add strict precision control
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, CLI, and ComfyUI nodes to fix naming confusion.
- Added strict `precision` control (`fp16`, `bf16`, `bf32`, `auto`) to CLI and internal configuration logic.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Added explicit logging ("🚀 Executing SageAttention...") to confirm kernel execution.
- Fixed bug where user-selected precision was being overridden by auto-detection defaults.
- Updated `src/interfaces/video_upscaler.py` to ensure precision setting is correctly propagated from ComfyUI node to generation context.
2025-12-09 11:12:07 +00:00
google-labs-jules[bot] 955164a5ba Fix SageAttention naming and add strict precision control
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, CLI, and ComfyUI nodes to fix naming confusion.
- Added strict `precision` control (`fp16`, `bf16`, `bf32`, `auto`) to CLI and internal configuration logic.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Added explicit logging ("🚀 Executing SageAttention...") to confirm kernel execution.
- Fixed bug where user-selected precision was being overridden by auto-detection defaults.
2025-12-09 10:50:54 +00:00
google-labs-jules[bot] 9ecacc6081 Fix SageAttention naming and logic bugs, add CLI precision option
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, compatibility layer, and ComfyUI node definitions to fix naming confusion.
- Fixed logic bug in `FlashAttentionVarlen` that prevented SageAttention from running even when requested (was checking for `sd` prefix instead of `sa`).
- Added robust availability and version checks for SageAttention (v2 vs v3) with improved fallback logic (SA3 -> SA2 -> FlashAttention 2 -> SDPA).
- Added explicit console logging when SageAttention kernel is first executed to verify optimization is active.
- Added `--precision` argument to CLI to allow explicit control over compute dtype (fp16, bf16, bf32, auto), enabling further performance tuning.
- Updated `FP8CompatibleDiT` wrapper to exclude `FlashAttentionVarlen` modules, preventing double-wrapping and casting issues.
2025-12-09 09:19:22 +00:00
Adrien Toupet a06afb5956 Merge pull request #384 from AInVFX/main
v2.5.18: CLI streaming mode, multi-GPU streaming with caching, shared memory fix
v2.5.18
2025-12-09 01:04:46 -05:00
Adrien Toupet b101deb894 docs: fix contributor links and formatting in release notes 2025-12-09 01:01:54 -05:00
Adrien Toupet 06be9c9d7a Release v2.5.18: CLI streaming mode, multi-GPU streaming with caching, shared memory fix 2025-12-09 00:59:22 -05:00
Adrien Toupet 4e96a5c366 fix: allow model caching with multi-GPU streaming (workers cache internally) 2025-12-09 00:48:18 -05:00
Adrien Toupet 4817beb148 fix: multi-GPU streaming log shows GPU count, workers log with [GPU N] prefix 2025-12-09 00:36:25 -05:00
Adrien Toupet 0b132b02ff refactor: multi-GPU workers stream video segments internally with model caching 2025-12-09 00:13:30 -05:00
Adrien Toupet f7e4fc677e Fix multi-GPU shared memory race condition with barrier sync 2025-12-08 22:29:12 -05:00
Adrien Toupet a70d82e3aa Add streaming mode for memory-efficient long video processing
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup

Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
2025-12-08 22:05:20 -05:00
Adrien Toupet bbd7e5ac02 Fix multiprocessing MemoryError for large video outputs (#372)
Use PyTorch shared memory instead of pickling numpy arrays through queue.
Prevents MemoryError when transferring large results between processes.
Thank you @FurkanGozukara
2025-12-08 14:05:41 -05:00
HB2k e105d6d457 Merge pull request #1 from naxci1/seedvr2-optimization-sageattn
SeedVR2 Optimization and SageAttention Support
2025-12-08 23:03:43 +04:00
google-labs-jules[bot] e03889bf5b feat: add SageAttention, precision controls, and 5070ti optimization 2025-12-08 19:01:40 +00:00
Adrien Toupet 58bc9e8bc9 Merge pull request #373 from AInVFX/main
v2.5.17: Proper bf16 detection for older GPUs #314
v2.5.17
2025-12-05 21:07:21 -05:00
Adrien Toupet 3eec5847c0 Release v2.5.17: Older GPU compatibility fix 2025-12-05 21:05:30 -05:00
Adrien Toupet eae3aac60d Fix CUBLAS_STATUS_NOT_SUPPORTED on older GPUs via bf16 probe (again\!) (#314) 2025-12-05 20:10:36 -05:00
Adrien Toupet 0a660065f0 Merge pull request #371 from AInVFX/main
v2.5.16: Older GPU compatibility fix, quality regression fix, debug improvements
v2.5.16
2025-12-05 15:52:34 -05:00
Adrien Toupet 11239eed13 Release v2.5.16: Quality regression fix, older GPU compatibility fix, system info debug 2025-12-05 15:50:08 -05:00
Adrien Toupet b4d7ab89eb Fix CUBLAS_STATUS_NOT_SUPPORTED on older GPUs (GTX 970) #314
Add automatic bfloat16 → float16 SDPA fallback for GPUs without native bf16 cuBLAS support
2025-12-05 15:39:22 -05:00
Adrien Toupet f061d97fe7 Revert bfloat16 detection - was causing quality regression / keep ensure_triton_compat() 2025-12-05 15:03:40 -05:00
Adrien Toupet 43c4f00e19 Revert bfloat16 detection - was causing quality regression / keep ensure_triton_compat() 2025-12-05 15:01:58 -05:00
Adrien Toupet f8998ebd75 Revert bfloat16 detection - was causing quality regression 2025-12-05 14:55:30 -05:00
Adrien Toupet aa968cf2c9 docs: simplify contribution workflow to main branch only 2025-12-05 13:54:39 -05:00
Adrien Toupet d78f6c268d Merge branch 'main' of https://github.com/ainvfx/ComfyUI-SeedVR2_VideoUpscaler 2025-12-05 11:12:32 -05:00