Adrien Toupet
2006fa3f6c
Merge pull request #390 from AInVFX/main
...
v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking
v2.5.19
2025-12-10 01:55:47 -05:00
Adrien Toupet
118c9fcbe7
Release v2.5.19: new logo, remove dead flash-attn wrapper, graceful DLL fallback, improved VRAM tracking, revert VRAM limit
2025-12-10 01:52:45 -05:00
Adrien Toupet
6106681563
Fix graceful fallback from flash-attn #376
...
Add compatibility shims for corrupted flash_attn/xformers DLLs.
Force-verify flash_attn_2_cuda at startup; fall back to SDPA if unavailable.
2025-12-10 01:44:49 -05:00
Adrien Toupet
c010deeea1
Remove ineffective allow_vram_overflow setting
...
- PyTorch's set_per_process_memory_fraction cannot prevent WDDM paging on Windows
- Keep overflow detection and warning when VRAM exceeds physical limit
- Simplify peak memory formatting
- Remove setting from CLI, ComfyUI node, and memory_manager
2025-12-10 00:56:27 -05:00
Adrien Toupet
7cbf025561
Fix VRAM peak tracking: separate allocated vs reserved, Windows-only overflow
...
- Track both peak_allocated (tensor usage) and peak_reserved (cache pool) per phase
- peak_allocated resets properly between phases via reset_peak_memory_stats()
- Overflow detection/warnings now Windows-only (WDDM paging behavior)
- Remove get_memory_architecture() - replaced with simple is_mps + platform checks
- Phase summary shows: VRAM XGB allocated, YGB reserved | RAM ZGB
- Simplify MPS path (unified memory has no overflow concept)
2025-12-09 23:51:51 -05:00
Adrien Toupet
5c60716c47
Refactor: centralize backend detection, fix architecture-aware VRAM overflow reporting
2025-12-09 21:06:12 -05:00
Adrien Toupet
77a00f651a
Fix: OOM regression from 2.5.14 strict VRAM limit ( #367 )
...
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.
The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.
- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow
Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
2025-12-09 17:12:10 -05:00
Adrien Toupet
30bc924043
Update header logo design (thanks @naxci1, closes #378 )
2025-12-09 14:07:46 -05:00
Adrien Toupet
e65e7fa418
Remove dead flash attention wrapper from FP8CompatibleDiT
...
The wrapper methods (_apply_flash_attention_optimization and related)
matched NaDiT attention modules by name but required qkv or q_proj+k_proj+v_proj
attributes to optimize. NaDiT uses proj_qkv instead, so the optimization
path was never taken - always falling back to original forward.
FlashAttentionVarlen already handles flash_attn vs sdpa switching via
its attention_mode attribute, making this wrapper redundant.
Removes ~200 lines of dead code.
2025-12-09 12:29:15 -05:00
Adrien Toupet
a06afb5956
Merge pull request #384 from AInVFX/main
...
v2.5.18: CLI streaming mode, multi-GPU streaming with caching, shared memory fix
v2.5.18
2025-12-09 01:04:46 -05:00
Adrien Toupet
b101deb894
docs: fix contributor links and formatting in release notes
2025-12-09 01:01:54 -05:00
Adrien Toupet
06be9c9d7a
Release v2.5.18: CLI streaming mode, multi-GPU streaming with caching, shared memory fix
2025-12-09 00:59:22 -05:00
Adrien Toupet
4e96a5c366
fix: allow model caching with multi-GPU streaming (workers cache internally)
2025-12-09 00:48:18 -05:00
Adrien Toupet
4817beb148
fix: multi-GPU streaming log shows GPU count, workers log with [GPU N] prefix
2025-12-09 00:36:25 -05:00
Adrien Toupet
0b132b02ff
refactor: multi-GPU workers stream video segments internally with model caching
2025-12-09 00:13:30 -05:00
Adrien Toupet
f7e4fc677e
Fix multi-GPU shared memory race condition with barrier sync
2025-12-08 22:29:12 -05:00
Adrien Toupet
a70d82e3aa
Add streaming mode for memory-efficient long video processing
...
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup
Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
2025-12-08 22:05:20 -05:00
Adrien Toupet
bbd7e5ac02
Fix multiprocessing MemoryError for large video outputs ( #372 )
...
Use PyTorch shared memory instead of pickling numpy arrays through queue.
Prevents MemoryError when transferring large results between processes.
Thank you @FurkanGozukara
2025-12-08 14:05:41 -05:00
Adrien Toupet
58bc9e8bc9
Merge pull request #373 from AInVFX/main
...
v2.5.17: Proper bf16 detection for older GPUs #314
v2.5.17
2025-12-05 21:07:21 -05:00
Adrien Toupet
3eec5847c0
Release v2.5.17: Older GPU compatibility fix
2025-12-05 21:05:30 -05:00
Adrien Toupet
eae3aac60d
Fix CUBLAS_STATUS_NOT_SUPPORTED on older GPUs via bf16 probe (again\!) ( #314 )
2025-12-05 20:10:36 -05:00
Adrien Toupet
0a660065f0
Merge pull request #371 from AInVFX/main
...
v2.5.16: Older GPU compatibility fix, quality regression fix, debug improvements
v2.5.16
2025-12-05 15:52:34 -05:00
Adrien Toupet
11239eed13
Release v2.5.16: Quality regression fix, older GPU compatibility fix, system info debug
2025-12-05 15:50:08 -05:00
Adrien Toupet
b4d7ab89eb
Fix CUBLAS_STATUS_NOT_SUPPORTED on older GPUs (GTX 970) #314
...
Add automatic bfloat16 → float16 SDPA fallback for GPUs without native bf16 cuBLAS support
2025-12-05 15:39:22 -05:00
Adrien Toupet
f061d97fe7
Revert bfloat16 detection - was causing quality regression / keep ensure_triton_compat()
2025-12-05 15:03:40 -05:00
Adrien Toupet
43c4f00e19
Revert bfloat16 detection - was causing quality regression / keep ensure_triton_compat()
2025-12-05 15:01:58 -05:00
Adrien Toupet
f8998ebd75
Revert bfloat16 detection - was causing quality regression
2025-12-05 14:55:30 -05:00
Adrien Toupet
aa968cf2c9
docs: simplify contribution workflow to main branch only
2025-12-05 13:54:39 -05:00
Adrien Toupet
d78f6c268d
Merge branch 'main' of https://github.com/ainvfx/ComfyUI-SeedVR2_VideoUpscaler
2025-12-05 11:12:32 -05:00
Adrien Toupet
18b44d66e1
feat: add environment info display in debug mode to help with issue reporting
2025-12-05 11:11:18 -05:00
Adrien Toupet
f68fe920b8
Merge pull request #358 from AInVFX/main
...
v2.5.15: MPS compatibility fixes, autocast device type, VRAM tracking, triton 3.0+ compatibility
v2.5.15
2025-12-03 13:16:46 -05:00
Adrien Toupet
65c1c1b6cd
Release v2.5.15: MPS fixes, autocast device type, accurate VRAM tracking, triton 3.0 compatibility
2025-12-03 13:14:51 -05:00
Adrien Toupet
b40f26167c
Fix MPS compatibility: disable antialias for MPS tensors, fix bfloat16 arange ( #354 )
2025-12-03 13:09:59 -05:00
Adrien Toupet
ed53581359
Fix triton.ops compatibility for bitsandbytes 0.45+ / triton 3.0+
...
Fixes #340 - Installation error with PyTorch 2.7+cu126 and triton_windows
Add compatibility shim for missing triton.ops.matmul_perf_model module.
Reverts local VAE types approach
2025-12-03 12:43:35 -05:00
Adrien Toupet
71ac9ffe54
fix: use max_memory_reserved for accurate VRAM peak tracking
2025-12-03 11:51:30 -05:00
Adrien Toupet
ffba05907d
Fix autocast device_type error by using .type attribute instead of str() #350
2025-12-03 11:32:39 -05:00
Adrien Toupet
d4dd5e747d
Merge pull request #344 from AInVFX/main
...
v2.5.14: MPS device fix, VRAM swap detection, enforce physical VRAM limit
v2.5.14
2025-12-01 00:33:53 -05:00
Adrien Toupet
e2faedaaa6
Release v2.5.14 - MPS device fix, VRAM swap detection, enforce physical VRAM limit
2025-12-01 00:30:49 -05:00
Adrien Toupet
5775ff0f99
Enforce VRAM limit to physical capacity - OOM instead of silent swap
2025-12-01 00:25:26 -05:00
Adrien Toupet
ff937756a3
Add VRAM swap detection - show GPU+swap breakdown in peak stats, warn when swap detected
2025-11-30 23:47:20 -05:00
Adrien Toupet
5848cef05f
fix(mps): normalize device strings to prevent unnecessary tensor movements
...
- Add _device_str() helper to normalize MPS variants (mps:0 → MPS)
- Fix device comparison: mps:0 and mps now correctly identified as same device
- Consistent MPS logging across all memory management functions
2025-11-30 21:03:26 -05:00
Adrien Toupet
f5b902b8b0
Merge pull request #341 from AInVFX/main
...
v2.5.13: Fix triton import error, OOM on long video float32 conversion, macOS CLI watermark
v2.5.13
2025-11-30 09:04:04 -05:00
Adrien Toupet
19825fa9fa
Release v2.5.13: Fix triton import, OOM on long videos, macOS watermark
2025-11-30 09:01:53 -05:00
Adrien Toupet
021fc7b70f
Fix triton.ops import error by using local VAE types
...
Fixes #340 - Installation error with PyTorch 2.7+cu126 and triton_windows
Replace diffusers.models.autoencoders.vae imports with local implementations
of DecoderOutput and DiagonalGaussianDistribution to avoid triggering the
bitsandbytes -> triton.ops import chain that fails on newer triton versions.
2025-11-30 08:27:01 -05:00
Adrien Toupet
b0ac880a4f
Fix OOM crash on float32 conversion for long videos. Gracefully fallback to native dtype if insufficient memory. Fixes #299
2025-11-30 08:03:07 -05:00
Adrien Toupet
6930e0f13a
Fix CLI MPS watermark error on macOS ( fixes #336 )
2025-11-30 07:37:51 -05:00
Adrien Toupet
112275d4e5
Merge pull request #333 from AInVFX/main
...
release: v2.5.12 - Regression fix
v2.5.12
2025-11-28 17:51:53 -05:00
Adrien Toupet
16508d353d
release: v2.5.12 - fix color artifacts regression from in-place transform ops
2025-11-28 17:50:40 -05:00
Adrien Toupet
6f784b1324
Merge pull request #332 from AInVFX/main
...
v2.5.11: CUDNN attention, streaming VAE decode, LAB artifact fix, MPS/multi-GPU improvements
v2.5.11
2025-11-28 17:09:48 -05:00
Adrien Toupet
f791631495
release: v2.5.11 - CUDNN attention, long video memory fix, LAB artifacts fix, MPS/multi-GPU improvements
2025-11-28 16:49:08 -05:00