365 Commits
Author SHA1 Message Date
Adrien Toupet 2006fa3f6c Merge pull request #390 from AInVFX/main
v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking
v2.5.19
2025-12-10 01:55:47 -05:00
Adrien Toupet 118c9fcbe7 Release v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking, revert VRAM limit 2025-12-10 01:52:45 -05:00
Adrien Toupet 6106681563 Fix graceful fallback from flash-attn #376
Add compatibility shims for corrupted flash_attn/xformers DLLs.
Force-verify flash_attn_2_cuda at startup; fall back to SDPA if unavailable.
2025-12-10 01:44:49 -05:00
Adrien Toupet c010deeea1 Remove ineffective allow_vram_overflow setting
- PyTorch's set_per_process_memory_fraction cannot prevent WDDM paging on Windows
- Keep overflow detection and warning when VRAM exceeds physical limit
- Simplify peak memory formatting
- Remove setting from CLI, ComfyUI node, and memory_manager
2025-12-10 00:56:27 -05:00
Adrien Toupet 7cbf025561 Fix VRAM peak tracking: separate allocated vs reserved, Windows-only overflow
- Track both peak_allocated (tensor usage) and peak_reserved (cache pool) per phase
- peak_allocated resets properly between phases via reset_peak_memory_stats()
- Overflow detection/warnings now Windows-only (WDDM paging behavior)
- Remove get_memory_architecture() - replaced with simple is_mps + platform checks
- Phase summary shows: VRAM XGB allocated, YGB reserved | RAM ZGB
- Simplify MPS path (unified memory has no overflow concept)
2025-12-09 23:51:51 -05:00
Adrien Toupet 5c60716c47 Refactor: centralize backend detection, fix architecture-aware VRAM overflow reporting 2025-12-09 21:06:12 -05:00
Adrien Toupet 77a00f651a Fix: OOM regression from 2.5.14 strict VRAM limit (#367)
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.

The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.

- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow

Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
2025-12-09 17:12:10 -05:00
Adrien Toupet 30bc924043 Update header logo design (thanks @naxci1, closes #378) 2025-12-09 14:07:46 -05:00
Adrien Toupet e65e7fa418 Remove dead flash attention wrapper from FP8CompatibleDiT
The wrapper methods (_apply_flash_attention_optimization and related)
matched NaDiT attention modules by name but required qkv or q_proj+k_proj+v_proj
attributes to optimize. NaDiT uses proj_qkv instead, so the optimization
path was never taken - always falling back to original forward.

FlashAttentionVarlen already handles flash_attn vs sdpa switching via
its attention_mode attribute, making this wrapper redundant.

Removes ~200 lines of dead code.
2025-12-09 12:29:15 -05:00
Adrien Toupet a06afb5956 Merge pull request #384 from AInVFX/main
v2.5.18: CLI streaming mode, multi-GPU streaming with caching, shared memory fix
v2.5.18
2025-12-09 01:04:46 -05:00
Adrien Toupet b101deb894 docs: fix contributor links and formatting in release notes 2025-12-09 01:01:54 -05:00
Adrien Toupet 06be9c9d7a Release v2.5.18: CLI streaming mode, multi-GPU streaming with caching, shared memory fix 2025-12-09 00:59:22 -05:00
Adrien Toupet 4e96a5c366 fix: allow model caching with multi-GPU streaming (workers cache internally) 2025-12-09 00:48:18 -05:00
Adrien Toupet 4817beb148 fix: multi-GPU streaming log shows GPU count, workers log with [GPU N] prefix 2025-12-09 00:36:25 -05:00
Adrien Toupet 0b132b02ff refactor: multi-GPU workers stream video segments internally with model caching 2025-12-09 00:13:30 -05:00
Adrien Toupet f7e4fc677e Fix multi-GPU shared memory race condition with barrier sync 2025-12-08 22:29:12 -05:00
Adrien Toupet a70d82e3aa Add streaming mode for memory-efficient long video processing
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup

Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
2025-12-08 22:05:20 -05:00
Adrien Toupet bbd7e5ac02 Fix multiprocessing MemoryError for large video outputs (#372)
Use PyTorch shared memory instead of pickling numpy arrays through queue.
Prevents MemoryError when transferring large results between processes.
Thank you @FurkanGozukara
2025-12-08 14:05:41 -05:00
Adrien Toupet 58bc9e8bc9 Merge pull request #373 from AInVFX/main
v2.5.17: Proper bf16 detection for older GPUs #314
v2.5.17
2025-12-05 21:07:21 -05:00
Adrien Toupet 3eec5847c0 Release v2.5.17: Older GPU compatibility fix 2025-12-05 21:05:30 -05:00
Adrien Toupet eae3aac60d Fix CUBLAS_STATUS_NOT_SUPPORTED on older GPUs via bf16 probe (again\!) (#314) 2025-12-05 20:10:36 -05:00
Adrien Toupet 0a660065f0 Merge pull request #371 from AInVFX/main
v2.5.16: Older GPU compatibility fix, quality regression fix, debug improvements
v2.5.16
2025-12-05 15:52:34 -05:00
Adrien Toupet 11239eed13 Release v2.5.16: Quality regression fix, older GPU compatibility fix, system info debug 2025-12-05 15:50:08 -05:00
Adrien Toupet b4d7ab89eb Fix CUBLAS_STATUS_NOT_SUPPORTED on older GPUs (GTX 970) #314
Add automatic bfloat16 → float16 SDPA fallback for GPUs without native bf16 cuBLAS support
2025-12-05 15:39:22 -05:00
Adrien Toupet f061d97fe7 Revert bfloat16 detection - was causing quality regression / keep ensure_triton_compat() 2025-12-05 15:03:40 -05:00
Adrien Toupet 43c4f00e19 Revert bfloat16 detection - was causing quality regression / keep ensure_triton_compat() 2025-12-05 15:01:58 -05:00
Adrien Toupet f8998ebd75 Revert bfloat16 detection - was causing quality regression 2025-12-05 14:55:30 -05:00
Adrien Toupet aa968cf2c9 docs: simplify contribution workflow to main branch only 2025-12-05 13:54:39 -05:00
Adrien Toupet d78f6c268d Merge branch 'main' of https://github.com/ainvfx/ComfyUI-SeedVR2_VideoUpscaler 2025-12-05 11:12:32 -05:00
Adrien Toupet 18b44d66e1 feat: add environment info display in debug mode to help with issue reporting 2025-12-05 11:11:18 -05:00
Adrien Toupet f68fe920b8 Merge pull request #358 from AInVFX/main
v2.5.15: MPS compatibility fixes, autocast device type, VRAM tracking, triton 3.0+ compatibility
v2.5.15
2025-12-03 13:16:46 -05:00
Adrien Toupet 65c1c1b6cd Release v2.5.15: MPS fixes, autocast device type, accurate VRAM tracking, triton 3.0 compatibility 2025-12-03 13:14:51 -05:00
Adrien Toupet b40f26167c Fix MPS compatibility: disable antialias for MPS tensors, fix bfloat16 arange (#354) 2025-12-03 13:09:59 -05:00
Adrien Toupet ed53581359 Fix triton.ops compatibility for bitsandbytes 0.45+ / triton 3.0+
Fixes #340 - Installation error with PyTorch 2.7+cu126 and triton_windows

Add compatibility shim for missing triton.ops.matmul_perf_model module.
Reverts local VAE types approach
2025-12-03 12:43:35 -05:00
Adrien Toupet 71ac9ffe54 fix: use max_memory_reserved for accurate VRAM peak tracking 2025-12-03 11:51:30 -05:00
Adrien Toupet ffba05907d Fix autocast device_type error by using .type attribute instead of str() #350 2025-12-03 11:32:39 -05:00
Adrien Toupet d4dd5e747d Merge pull request #344 from AInVFX/main
v2.5.14: MPS device fix, VRAM swap detection, enforce physical VRAM limit
v2.5.14
2025-12-01 00:33:53 -05:00
Adrien Toupet e2faedaaa6 Release v2.5.14 - MPS device fix, VRAM swap detection, enforce physical VRAM limit 2025-12-01 00:30:49 -05:00
Adrien Toupet 5775ff0f99 Enforce VRAM limit to physical capacity - OOM instead of silent swap 2025-12-01 00:25:26 -05:00
Adrien Toupet ff937756a3 Add VRAM swap detection - show GPU+swap breakdown in peak stats, warn when swap detected 2025-11-30 23:47:20 -05:00
Adrien Toupet 5848cef05f fix(mps): normalize device strings to prevent unnecessary tensor movements
- Add _device_str() helper to normalize MPS variants (mps:0 → MPS)
- Fix device comparison: mps:0 and mps now correctly identified as same device
- Consistent MPS logging across all memory management functions
2025-11-30 21:03:26 -05:00
Adrien Toupet f5b902b8b0 Merge pull request #341 from AInVFX/main
v2.5.13: Fix triton import error, OOM on long video float32 conversion, macOS CLI watermark
v2.5.13
2025-11-30 09:04:04 -05:00
Adrien Toupet 19825fa9fa Release v2.5.13: Fix triton import, OOM on long videos, macOS watermark 2025-11-30 09:01:53 -05:00
Adrien Toupet 021fc7b70f Fix triton.ops import error by using local VAE types
Fixes #340 - Installation error with PyTorch 2.7+cu126 and triton_windows

Replace diffusers.models.autoencoders.vae imports with local implementations
of DecoderOutput and DiagonalGaussianDistribution to avoid triggering the
bitsandbytes -> triton.ops import chain that fails on newer triton versions.
2025-11-30 08:27:01 -05:00
Adrien Toupet b0ac880a4f Fix OOM crash on float32 conversion for long videos. Gracefully fallback to native dtype if insufficient memory. Fixes #299 2025-11-30 08:03:07 -05:00
Adrien Toupet 6930e0f13a Fix CLI MPS watermark error on macOS (fixes #336) 2025-11-30 07:37:51 -05:00
Adrien Toupet 112275d4e5 Merge pull request #333 from AInVFX/main
release: v2.5.12 - Regression fix
v2.5.12
2025-11-28 17:51:53 -05:00
Adrien Toupet 16508d353d release: v2.5.12 - fix color artifacts regression from in-place transform ops 2025-11-28 17:50:40 -05:00
Adrien Toupet 6f784b1324 Merge pull request #332 from AInVFX/main
v2.5.11: CUDNN attention, streaming VAE decode, LAB artifact fix, MPS/multi-GPU improvements
v2.5.11
2025-11-28 17:09:48 -05:00
Adrien Toupet f791631495 release: v2.5.11 - CUDNN attention, long video memory fix, LAB artifacts fix, MPS/multi-GPU improvements 2025-11-28 16:49:08 -05:00