Add ensure_bitsandbytes_safe() shim to handle broken/partial bitsandbytes
installations that cause PyTorch kernel registration conflicts when diffusers
attempts to re-import the module.
On ROCm systems without proper binaries, bitsandbytes registers kernels during
import then fails. When diffusers later imports it, the duplicate registration
causes: 'RuntimeError: already a kernel registered...int8_mm_dequant'
The shim pre-tests bitsandbytes import and stubs it only if broken, allowing
working installations to function normally for other nodes.
- CompatibleDiT now converts ALL model params to compute_dtype on MPS
(previously only FP8 - GGUF models had mixed FP16/BF16 causing hangs)
- Replace MPS autocast with explicit dtype conversion in VAE encode/decode
- Skip DiT autocast on MPS (CompatibleDiT handles dtype internally)
- Guard call_rope_with_stability CUDA autocast for non-CUDA devices
- Add weights_only=True to torch.load (FutureWarning fix)
- Rename FP8CompatibleDiT → CompatibleDiT
Addresses M4 Pro macOS hang at EulerSampler 0% with GGUF models
- Add validate_blockswap_config() in blockswap.py as single validation point
- Auto-disable BlockSwap on macOS (unified memory makes it meaningless)
- Improve error messages for missing dit_offload_device
- Update CLI and ComfyUI tooltips for BlockSwap and model caching
- Update README: BlockSwap macOS note, caching descriptions, attention backends
- Remove duplicate validation from dit_model_loader.py and inference_cli.py
Partially fixes#401 (M4 Pro macOS BlockSwap offload device error)
Add compatibility shims for corrupted/missing flash_attn and xformers DLLs.
Stubs include proper __spec__ to prevent importlib.util.find_spec() crashes.
Force-verify flash_attn_2_cuda at startup; fall back to SDPA if unavailable.
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.
The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.
- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow
Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
The wrapper methods (_apply_flash_attention_optimization and related)
matched NaDiT attention modules by name but required qkv or q_proj+k_proj+v_proj
attributes to optimize. NaDiT uses proj_qkv instead, so the optimization
path was never taken - always falling back to original forward.
FlashAttentionVarlen already handles flash_attn vs sdpa switching via
its attention_mode attribute, making this wrapper redundant.
Removes ~200 lines of dead code.
- Re-implemented `precision` control (`fp16`, `bf16`, `bf32`, `auto`) in CLI, ComfyUI node, and backend logic to respect user choice.
- Fixed `UnboundLocalError` in `apply_model_specific_config` by ensuring `compute_dtype` is always initialized before use.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by properly passing the restored `precision` argument.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` for clarity and fixed fallback logic to ensure SageAttention is correctly prioritized.
- Added explicit logging of active attention backend and execution confirmation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Re-implemented `precision` control (`fp16`, `bf16`, `bf32`, `auto`) in CLI, ComfyUI node, and backend logic to respect user choice.
- Fixed `UnboundLocalError` in `apply_model_specific_config` by ensuring `compute_dtype` is always initialized before use.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by properly passing the restored `precision` argument.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` for clarity and fixed fallback logic to ensure SageAttention is correctly prioritized.
- Added explicit logging of active attention backend and execution confirmation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` and CLI by removing all residual `precision` variable usage.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` and CLI by removing all residual `precision` variable usage.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by removing residual `precision` usage.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, CLI, and ComfyUI nodes to fix naming confusion.
- Added strict `precision` control (`fp16`, `bf16`, `bf32`, `auto`) to CLI and internal configuration logic.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Added explicit logging ("🚀 Executing SageAttention...") to confirm kernel execution.
- Fixed bug where user-selected precision was being overridden by auto-detection defaults.
- Updated `src/interfaces/video_upscaler.py` to ensure precision setting is correctly propagated from ComfyUI node to generation context.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, CLI, and ComfyUI nodes to fix naming confusion.
- Added strict `precision` control (`fp16`, `bf16`, `bf32`, `auto`) to CLI and internal configuration logic.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Added explicit logging ("🚀 Executing SageAttention...") to confirm kernel execution.
- Fixed bug where user-selected precision was being overridden by auto-detection defaults.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, compatibility layer, and ComfyUI node definitions to fix naming confusion.
- Fixed logic bug in `FlashAttentionVarlen` that prevented SageAttention from running even when requested (was checking for `sd` prefix instead of `sa`).
- Added robust availability and version checks for SageAttention (v2 vs v3) with improved fallback logic (SA3 -> SA2 -> FlashAttention 2 -> SDPA).
- Added explicit console logging when SageAttention kernel is first executed to verify optimization is active.
- Added `--precision` argument to CLI to allow explicit control over compute dtype (fp16, bf16, bf32, auto), enabling further performance tuning.
- Updated `FP8CompatibleDiT` wrapper to exclude `FlashAttentionVarlen` modules, preventing double-wrapping and casting issues.
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup
Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
Use PyTorch shared memory instead of pickling numpy arrays through queue.
Prevents MemoryError when transferring large results between processes.
Thank you @FurkanGozukara