Compare commits

..
33 Commits
Author SHA1 Message Date
Adrien Toupet 4490bd1f48 Merge pull request #441 from AInVFX/main
v2.5.24: Restores the MPS memory leak workaround that was accidentally removed during code cleanup in v2.5.23
2025-12-24 09:52:32 +01:00
Adrien Toupet baec4b634f fix(vae): restore MPS memory leak workaround removed in v2.5.23 cleanup 2025-12-24 09:50:19 +01:00
Adrien Toupet 5a4bf428f3 Merge pull request #438 from AInVFX/main
v2.5.23: Security hardening, GGUF VAE support, FFmpeg stability, MPS optimization
2025-12-23 21:09:01 -05:00
Adrien Toupet 43e70bf637 Release v2.5.23: Security & stability improvements
- Add security protection against malicious .pth files
- Fix FFmpeg video writer hanging issues (thanks @thehhmdb)
- Enable GGUF VAE model support via conv dequantization (thanks @naxci1)
- Fix VAE slicing division by zero edge cases (thanks @naxci1)
- Resolve LAB color transfer dtype mismatch errors
- Extend Conv3d memory workaround to PyTorch 2.9+
- Fix bitsandbytes compatibility on non-Gaudi systems
- Optimize MPS memory usage (thanks @s-cerevisiae)
2025-12-24 03:02:34 +01:00
Adrien Toupet 855f8b91b3 Add sponsor call-to-action to footer 2025-12-24 02:49:42 +01:00
Adrien Toupet f561743054 Fix #434: Resolve dtype mismatch in LAB color transfer during video upscaling
Add explicit dtype alignment before matrix multiplication in _rgb_to_lab_batch
and _lab_to_rgb_batch to prevent float64 promotion from torch.pow operations.
Fixes RuntimeError: expected mat1 and mat2 to have the same dtype.
2025-12-24 02:22:46 +01:00
Adrien Toupet 396f323eae Fix #437: Catch ValueError in bitsandbytes compatibility shim
Add ValueError to exception handling to catch packaging.InvalidVersion
errors during Intel Gaudi version detection on non-Gaudi systems.
2025-12-24 02:14:33 +01:00
Adrien Toupet 6226878411 Apply critical fixes from PR #421
Fixes applied:
- GGUF conv2d/conv3d dequantization in __torch_function__
  Critical fix: Makes GGUF VAE models functional by properly handling
  InflatedCausalConv3d layers that aren't replaced by layer replacement

- Division by zero protection in slicing_latent_min_size calculations
  Defensive: Prevents crashes with edge-case temporal_downsample_factor values

Changes rejected from PR #421:
- NODE_CLASS_MAPPINGS (violates ComfyUI V3 API schema)
- Triton auto-fallback to cudagraphs (cudagraphs has compatibility issues)
- VAE 4D optimization (unproven benefit)
- force_upcast default change (untested breaking change)
2025-12-24 01:58:58 +01:00
Adrien Toupet 2f8d2ccf9a Merge PR #421 from naxci1 - Optimize VAE and GGUF 2025-12-24 01:13:19 +01:00
Adrien Toupet b0f01f2d99 Fix MPS device check precision for memory leak workaround
Improves PR #428 by checking actual VAE device (self.device.type == 'mps')
instead of system MPS availability. Only clears cache when VAE operations
are running on MPS, avoiding unnecessary overhead when VAE runs on CPU.

This fix addresses the PyTorch MPS memory leak (pytorch/pytorch#155060)
where padding operations in convolutions accumulate memory during encode
and decode. Clears MPS cache after each operation to prevent accumulation.

Fixes #363 (absurd MPS VRAM usage - 52GB → 12GB)
Fixes #410 (macOS system restarts from memory exhaustion)
Fixes #415 (inability to upscale beyond 2K resolution)
May help #417 (convolution errors under memory pressure)
2025-12-24 00:46:20 +01:00
Adrien Toupet 2214f3afde Merge PR #428: Reduce MPS memory usage 2025-12-24 00:40:59 +01:00
Adrien Toupet aeebd49f7f Refine FFMPEGVideoWriter to prevent pipe blocking
Building on @thehhmdb's fix in PR #418 which identified the stderr
pipe blocking issue. This refinement simplifies the solution by
redirecting stderr to DEVNULL and adding stdin.flush() to prevent
buffering deadlocks.

Improvements:
- Simpler implementation without threading complexity
- Zero memory overhead
- Better error messages for debugging
- Maintains fix for the original hanging issue

Co-authored-by: thehhmdb <thehhmdb@users.noreply.github.com>
Fixes numz/ComfyUI-SeedVR2_VideoUpscaler#418
2025-12-24 00:07:41 +01:00
Adrien Toupet e178b72d89 Merge commit '7bb936749f2799cab03e5d1593d1815a37c5233d' 2025-12-23 23:49:01 +01:00
Adrien Toupet 8ad4c8fa4e sec: prevent RCE vulnerability in .pth model loading
Add weights_only=True to torch.load() to restrict deserialization
to tensors only, preventing arbitrary code execution via pickle
2025-12-23 23:19:51 +01:00
Adrien Toupet 241b632cfc fix: extend Conv3d workaround to PyTorch 2.9+ (fixes 3x VAE VRAM usage in 2.11+) 2025-12-21 16:21:07 -05:00
spore 27ed3333fd perf(mps): reduce memory usage by clearing cache 2025-12-20 18:32:11 +08:00
google-labs-jules[bot] c6997fd9c2 Optimize VAE/GGUF performance and fix node registration/compile bugs
- Implement node registration in `__init__.py` to fix "Node does not exist" error.
- Implement automatic fallback from `inductor` to `cudagraphs` in `torch.compile` when Triton is missing (Windows fix).
- Optimize VAE `InflatedCausalConv3d` to use 2D convolution path for spatial-only operations, improving speed.
- Optimize VAE `ResnetBlock3D` to support flattened 4D execution path to reduce reshape overhead.
- Update `VideoAutoencoderKL` to default `force_upcast=False` for FP16 inference.
- Add `conv2d`/`conv3d` dequantization support to `GGUFTensor` for GGUF model compatibility.
2025-12-15 18:22:47 +00:00
google-labs-jules[bot] 0e849d20cd Optimize VAE Decoding Speed and robustness
- Implemented "4D execution mode" in ResnetBlock3D to flatten temporal dimension when convolutions are effectively 2D, reducing reshape overhead.
- Updated InflatedCausalConv3d to support direct 4D input processing and use 2D convolution optimization path.
- Enhanced InflatedCausalConv3d check_effective_2d to strictly verify stride/dilation/padding compatibility.
- Fixed 4D input handling in InflatedCausalConv3d forward pass.
2025-12-15 16:51:05 +00:00
google-labs-jules[bot] a65ddadc00 Optimize VAE performance (FP16/2D Conv) and GGUF support
- Implemented 2D convolution optimization in `InflatedCausalConv3d` to speed up spatial-only operations by using `F.conv2d` instead of `Conv3d`.
- Added `torch.nn.functional.conv2d` and `torch.nn.functional.conv3d` to `GGUFTensor`'s `__torch_function__` dispatch to enable automatic dequantization of weights, supporting GGUF models.
- Updated `src/models/video_vae_v3/modules/attn_video_vae.py` to default `force_upcast` to `False` for better FP16 performance.
- Fixed a bug in VAE slicing logic where `slicing_latent_min_size` could become 0, now clamping it to 1.
- Updated `Upsample3D` and `Encoder3D` in `attn_video_vae.py` to use `init_causal_conv3d` for 1x1x1 convolutions, enabling the 2D optimization path.
2025-12-15 16:12:33 +00:00
google-labs-jules[bot] 48bfbae05d Optimize VAE for performance and GGUF support
- Implemented 2D convolution optimization in `InflatedCausalConv3d` to speed up spatial-only operations by using `F.conv2d` instead of `Conv3d`.
- Added `torch.nn.functional.conv2d` and `torch.nn.functional.conv3d` to `GGUFTensor`'s `__torch_function__` dispatch to enable automatic dequantization of weights, supporting GGUF models.
- Updated `src/models/video_vae_v3/modules/attn_video_vae.py` to default `force_upcast` to `False` for better FP16 performance.
- Fixed a bug in VAE slicing logic where `slicing_latent_min_size` could become 0, now clamping it to 1.
2025-12-15 16:03:17 +00:00
google-labs-jules[bot] 04475bdc24 Optimize VAE and improve GGUF support for 50-series GPUs
- Implemented 2D convolution optimization in `InflatedCausalConv3d` to speed up spatial-only operations by reshaping effectively 2D tensors and using `F.conv2d` instead of `Conv3d`.
- Added `torch.nn.functional.conv2d` and `torch.nn.functional.conv3d` to `GGUFTensor`'s `__torch_function__` dispatch to enable automatic dequantization of weights, allowing the VAE to utilize GGUF quantization.
- Fixed a bug in `VideoAutoencoderKL` where `slicing_latent_min_size` could become 0 with small split sizes, now clamping it to a minimum of 1.
2025-12-15 15:38:14 +00:00
thehhmdb 7bb936749f To prevent ffmpeg from hanging, patched FFMPEGVideoWriter to continuously consume ffmpeg stderr in a background thread, flush stdin on write, and raise a clear error (including stderr) on BrokenPipe; release now joins the thread and logs stderr on non-zero exit. 2025-12-14 15:50:27 +00:00
HB2k 4fc3296c81 Merge branch 'numz:main' into main 2025-12-13 14:07:53 +04:00
Adrien Toupet d69b65f7e4 Merge pull request #412 from AInVFX/main
v2.5.22: CLI FFmpeg 10-bit video backend, MPS bicubic fix, cross-platform histogram matching
2025-12-13 00:36:12 -05:00
Adrien Toupet 15cb24089a Release v2.5.22: FFmpeg 10-bit video backend, MPS bicubic fix, cross-platform histogram matching
Note: index_select(out=) optimization removed as it caused color polarization on MPS; using simple indexing instead
2025-12-13 00:29:56 -05:00
Adrien Toupet c52280881a refactor: replace scatter_ with argsort+index_select for better cross-platform compatibility
- Replace tensor.scatter_() with torch.argsort() + torch.index_select(out=)
- Uses fundamental PyTorch ops for improved reliability across CUDA/ROCm/MPS
- Aggressive early tensor deletion to minimize memory overhead
- Affects _histogram_matching_channel and _histogram_match_1d in color_fix.py

Related: #351
2025-12-13 00:06:15 -05:00
Adrien Toupet 4b0b7d58b6 fix(cli): validate ffmpeg availability at startup
Move ffmpeg check from FFMPEGVideoWriter to argument validation phase.
Prevents wasted GPU processing time when ffmpeg backend is selected
but ffmpeg is not installed.
2025-12-12 23:39:16 -05:00
Adrien Toupet 39d8d4bf19 Refine MPS bicubic fix to use try/except for version compatibility (#408)
- Use try/except instead of blanket MPS check for bicubic+antialias
- PyTorch 2.8.0+ MPS: native fast path (no overhead)
- PyTorch < 2.8.0 MPS: CPU fallback on NotImplementedError
2025-12-12 23:19:37 -05:00
Adrien Toupet f2f4916c05 Fix MPS bicubic+antialias error for RGBA upscaling (#408)
- Add CPU fallback for F.interpolate with antialias=True on MPS (aten::_upsample_bicubic2d_aa not implemented)
- Revert torch.mps.synchronize() calls introduced in v2.5.21 for consistent behavior with CUDA pipeline
2025-12-12 23:08:16 -05:00
Adrien Toupet f75bcc7f37 feat(cli): add ffmpeg video backend with 10-bit support
- Add --video_backend flag: 'opencv' (default) or 'ffmpeg'
- Add --10bit flag: enables x265/yuv420p10le for reduced banding
- Without --10bit, ffmpeg uses x264/yuv420p for max compatibility
- FFMPEGVideoWriter class with cv2.VideoWriter-compatible interface
- Validates ffmpeg availability before encoding

Based on PR #409 by thehhmdb
2025-12-12 21:39:47 -05:00
thehhmdb 0c2a546c12 Add option to use ffmpeg and 10-bit video to reduce blocking and banding 2025-12-12 20:43:15 -05:00
HB2k 5c07a92b33 Merge branch 'numz:main' into main 2025-12-12 13:46:50 +04:00
HB2k d114e4958a Merge pull request #2 from naxci1/seedvr2-optimization-sageattn-13563569499211263226
Fix SageAttention naming/logic and add precision control
2025-12-09 13:23:49 +04:00
12 changed files with 210 additions and 52 deletions
+30 -3
View File
@@ -36,6 +36,29 @@ We're actively working on improvements and new features. To stay informed:
## 🚀 Release Notes
**2025.12.24 - Version 2.5.24**
- **🍎 Fix: MPS memory leak regression** - Restored MPS cache clearing after VAE encode/decode operations that was accidentally removed during code cleanup in v2.5.23
**2025.12.24 - Version 2.5.23**
- **🔒 Security: Prevent code execution in model loading** - Added protection against malicious .pth files by restricting deserialization to tensors only
- **🎥 Fix: FFmpeg video writer reliability** - Resolved ffmpeg process hanging issues by redirecting stderr and adding buffer flush, with improved error messages for debugging *(thanks [@thehhmdb](https://github.com/thehhmdb))*
- **⚡ Fix: GGUF VAE model support** - Enabled automatic weight dequantization for convolution operations, making GGUF-quantized VAE models fully functional *(thanks [@naxci1](https://github.com/naxci1))*
- **🛡️ Fix: VAE slicing edge cases** - Protected against division by zero crashes when using small split sizes with high temporal downsampling *(thanks [@naxci1](https://github.com/naxci1))*
- **🎨 Fix: LAB color transfer precision** - Resolved dtype mismatch errors during video upscaling by ensuring consistent float types before matrix operations
- **🔧 Fix: PyTorch 2.9+ compatibility** - Extended Conv3d memory workaround to all PyTorch 2.9+ versions, fixing 3x VRAM usage on newer PyTorch releases
- **📦 Fix: Bitsandbytes compatibility** - Added ValueError exception handling for Intel Gaudi version detection failures on non-Gaudi systems
- **🍎 MPS: Memory optimization** - Reduced memory usage during encode/decode operations on Apple Silicon *(thanks [@s-cerevisiae](https://github.com/s-cerevisiae))*
**2025.12.13 - Version 2.5.22**
- **🎬 CLI: FFmpeg video backend with 10-bit support** - New `--video_backend ffmpeg` and `--10bit` flags enable x265 encoding with 10-bit color depth, reducing banding artifacts in gradients compared to 8-bit OpenCV output *(based on PR by [@thehhmdb](https://github.com/thehhmdb) - thank you!)*
- **🍎 Fix: MPS bicubic upscaling compatibility** - Added CPU fallback for bicubic+antialias interpolation on PyTorch versions before 2.8.0, resolving RGBA alpha upscaling errors on Apple Silicon
- **⚡ Fix: Cross-platform histogram matching** - Replaced scatter_ operation with argsort+index_select for improved reliability across CUDA, ROCm, and MPS backends
- **🧹 MPS: Remove sync overhead** - Reverted unnecessary `torch.mps.synchronize()` calls introduced in v2.5.21 for consistent behavior with CUDA pipeline
**2025.12.12 - Version 2.5.21**
- **🛠️ Fix: GGUF dequantization error on MPS** - Resolved shape mismatch error introduced in 2.5.20 by skipping GGUF quantized buffers in precision conversion - these must remain in packed format for on-the-fly dequantization during inference
@@ -812,14 +835,16 @@ python inference_cli.py image.jpg
# Basic video upscaling with temporal consistency
python inference_cli.py video.mp4 --resolution 720 --batch_size 33
# Streaming mode for long videos (memory-efficient)
# Streaming mode for long videos (memory-efficient) with 10-bit video output (requires FFMPEG)
# Processes video in chunks of 330 frames to avoid loading entire video into RAM
# Use --temporal_overlap to ensure smooth transitions between chunks
python inference_cli.py long_video.mp4 \
--resolution 1080 \
--batch_size 33 \
--chunk_size 330 \
--temporal_overlap 3
--temporal_overlap 3 \
--video_backend ffmpeg \
--10bit
# Multi-GPU processing with temporal overlap
python inference_cli.py video.mp4 \
@@ -866,6 +891,8 @@ python inference_cli.py media_folder/ \
- `<input>`: Input file (.mp4, .avi, .png, .jpg, etc.) or directory
- `--output`: Output path (default: auto-generated in 'output/' directory)
- `--output_format`: Output format: 'mp4' (video) or 'png' (image sequence). Default: auto-detect from input type
- `--video_backend`: Video encoder backend: 'opencv' (default) or 'ffmpeg' (requires ffmpeg in PATH)
- `--10bit`: Save 10-bit video with x265 codec and yuv420p10le pixel format (reduces banding in gradients). Without this flag, ffmpeg uses x264 (yuv420p) for maximum compatibility. Requires --video_backend ffmpeg
- `--model_dir`: Model directory (default: ./models/SEEDVR2)
**Model Selection:**
@@ -1019,7 +1046,7 @@ For detailed contribution guidelines, see [CONTRIBUTING.md](CONTRIBUTING.md).
This ComfyUI implementation is a collaborative project by **[NumZ](https://github.com/numz)** and **[AInVFX](https://www.youtube.com/@AInVFX)** (Adrien Toupet), based on the original [SeedVR2](https://github.com/ByteDance-Seed/SeedVR) by ByteDance Seed Team.
Special thanks to our community contributors including [naxci1](https://github.com/naxci1), [benjaminherb](https://github.com/benjaminherb), [cmeka](https://github.com/cmeka), [FurkanGozukara](https://github.com/FurkanGozukara), [JohnAlcatraz](https://github.com/JohnAlcatraz), [lihaoyun6](https://github.com/lihaoyun6), [Luchuanzhao](https://github.com/Luchuanzhao), [Luke2642](https://github.com/Luke2642), [proxyid](https://github.com/proxyid), [q5sys](https://github.com/q5sys), and many others for their improvements, bug fixes, and testing.
Special thanks to our community contributors including [naxci1](https://github.com/naxci1), [thehhmdb](https://github.com/thehhmdb), [s-cerevisiae](https://github.com/s-cerevisiae), [benjaminherb](https://github.com/benjaminherb), [cmeka](https://github.com/cmeka), [FurkanGozukara](https://github.com/FurkanGozukara), [JohnAlcatraz](https://github.com/JohnAlcatraz), [lihaoyun6](https://github.com/lihaoyun6), [Luchuanzhao](https://github.com/Luchuanzhao), [Luke2642](https://github.com/Luke2642), [proxyid](https://github.com/proxyid), [q5sys](https://github.com/q5sys), and many others for their improvements, bug fixes, and testing.
## 📜 License
+103 -9
View File
@@ -108,6 +108,8 @@ else:
import torch
import cv2
import numpy as np
import subprocess
import shutil
# Project imports
from src.utils.downloads import download_weight
@@ -132,6 +134,81 @@ from src.utils.debug import Debug
from src.optimization.memory_manager import clear_memory, get_gpu_backend, is_cuda_available
debug = Debug(enabled=False) # Will be enabled via --debug CLI flag
# =============================================================================
# FFMPEG Class
# =============================================================================
class FFMPEGVideoWriter:
"""
Video writer using ffmpeg subprocess for encoding with 10-bit support.
Provides cv2.VideoWriter-compatible interface (write, isOpened, release) while
using ffmpeg for encoding. Enables 10-bit output (yuv420p10le with x265) which
reduces banding artifacts in gradients compared to 8-bit opencv output.
Args:
path: Output video file path
width: Frame width in pixels
height: Frame height in pixels
fps: Frames per second
use_10bit: If True, uses x265 codec with yuv420p10le pixel format.
If False, uses x264 with yuv420p (default: False)
Raises:
RuntimeError: If ffmpeg is not found in system PATH
Note:
Frames must be passed to write() in BGR format (same as cv2.VideoWriter).
Internally converts to RGB for ffmpeg rawvideo input.
"""
def __init__(self, path: str, width: int, height: int, fps: float, use_10bit: bool = False):
pix_fmt = 'yuv420p10le' if use_10bit else 'yuv420p'
codec = 'libx265' if use_10bit else 'libx264'
self.proc = subprocess.Popen(
['ffmpeg', '-y', '-f', 'rawvideo', '-pix_fmt', 'rgb24',
'-s', f'{width}x{height}', '-r', str(fps), '-i', '-',
'-c:v', codec, '-pix_fmt', pix_fmt, '-preset', 'medium', '-crf', '12', path],
stdin=subprocess.PIPE, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL
)
def write(self, frame_bgr: np.ndarray):
if not self.isOpened():
raise RuntimeError("FFMPEGVideoWriter: ffmpeg process is not running")
frame_rgb = cv2.cvtColor(frame_bgr, cv2.COLOR_BGR2RGB)
try:
self.proc.stdin.write(frame_rgb.astype(np.uint8).tobytes())
self.proc.stdin.flush() # Critical: prevent buffering issues
except BrokenPipeError:
raise RuntimeError(
"FFMPEGVideoWriter: ffmpeg process terminated unexpectedly. "
"Check video path, codec support, and disk space."
)
def isOpened(self) -> bool:
return self.proc is not None and self.proc.poll() is None
def release(self):
if self.proc:
try:
self.proc.stdin.close()
except Exception:
pass # Ignore errors on close
self.proc.wait()
if self.proc.returncode != 0:
debug.log(
f"ffmpeg exited with code {self.proc.returncode}. "
"Check output file for corruption.",
level="WARNING", force=True, category="file"
)
self.proc = None
# =============================================================================
# Device Management Helpers
# =============================================================================
@@ -447,7 +524,8 @@ def process_single_file(input_path: str, args: argparse.Namespace, device_list:
if is_png:
save_frames_to_image(result, output_path, base_name)
else:
video_writer = save_frames_to_video(result, output_path, fps)
video_writer = save_frames_to_video(result, output_path, fps,
video_backend=args.video_backend, use_10bit=args.use_10bit)
if video_writer is not None:
video_writer.release()
@@ -475,7 +553,8 @@ def process_single_file(input_path: str, args: argparse.Namespace, device_list:
if is_png:
save_frames_to_image(result, output_path, base_name, start_index=frames_written)
else:
video_writer = save_frames_to_video(result, output_path, fps, writer=video_writer)
video_writer = save_frames_to_video(result, output_path, fps, writer=video_writer,
video_backend=args.video_backend, use_10bit=args.use_10bit)
frames_written += result.shape[0]
del result
@@ -658,7 +737,9 @@ def save_frames_to_video(
frames_tensor: torch.Tensor,
output_path: str,
fps: float = 30.0,
writer: Optional[cv2.VideoWriter] = None
writer: Optional[cv2.VideoWriter] = None,
video_backend: str = "opencv",
use_10bit: bool = False
) -> Optional[cv2.VideoWriter]:
"""
Save frames tensor to MP4 video file.
@@ -683,10 +764,13 @@ def save_frames_to_video(
T, H, W, C = frames_np.shape
if writer is None:
debug.log(f"Saving {T} frames to video: {output_path}", category="file")
debug.log(f"Saving {T} frames to video: {output_path} (backend={video_backend})", category="file")
os.makedirs(Path(output_path).parent, exist_ok=True)
fourcc = cv2.VideoWriter_fourcc(*'mp4v')
writer = cv2.VideoWriter(output_path, fourcc, fps, (W, H))
if video_backend == "ffmpeg":
writer = FFMPEGVideoWriter(output_path, W, H, fps, use_10bit)
else:
fourcc = cv2.VideoWriter_fourcc(*'mp4v')
writer = cv2.VideoWriter(output_path, fourcc, fps, (W, H))
if not writer.isOpened():
raise ValueError(f"Cannot create video writer for: {output_path}")
@@ -1236,8 +1320,8 @@ Examples:
Basic video upscaling with temporal consistency:
python {invocation} video.mp4 --resolution 720 --batch_size 33
Streaming mode for long videos:
python {invocation} long_video.mp4 --resolution 1080 --batch_size 33 --chunk_size 330 --temporal_overlap 3
Streaming mode for long videos with 10-bit video output (requires FFMPEG):
python {invocation} long_video.mp4 --resolution 1080 --batch_size 33 --chunk_size 330 --temporal_overlap 3 --video_backend ffmpeg --10bit
Multi-GPU processing with temporal overlap:
python {invocation} video.mp4 --cuda_device 0,1 --resolution 1080 --batch_size 81 --uniform_batch_size --temporal_overlap 3 --prepend_frames 4
@@ -1250,7 +1334,6 @@ Examples:
Batch directory processing:
python {invocation} media_folder/ --output processed/ --cuda_device 0 --cache_dit --cache_vae --dit_offload_device cpu --vae_offload_device cpu --resolution 1080 --max_resolution 1920
"""
parser = argparse.ArgumentParser(
@@ -1268,6 +1351,11 @@ Examples:
help="Output path (default: auto-generated in 'output/' directory)")
io_group.add_argument("--output_format", type=str, default=None, choices=["mp4", "png", None],
help="Output format: 'mp4' (video) or 'png' (image sequence). Default: auto-detect from input type")
io_group.add_argument("--video_backend", type=str, default="opencv", choices=["opencv", "ffmpeg"],
help="Video encoder backend: 'opencv' (default) or 'ffmpeg' (requires ffmpeg in PATH)")
io_group.add_argument("--10bit", dest="use_10bit", action="store_true",
help="Save 10-bit video with x265 codec (reduces banding). Without this flag, "
"ffmpeg uses x264 for maximum compatibility. Requires --video_backend ffmpeg")
io_group.add_argument("--model_dir", type=str, default=None,
help=f"Model directory (default: ./models/{SEEDVR2_FOLDER_NAME})")
@@ -1444,6 +1532,12 @@ def main() -> None:
debug.log(f"VAE decode tile overlap ({args.vae_decode_tile_overlap}) must be smaller than tile size ({args.vae_decode_tile_size})", level="ERROR", category="vae", force=True)
sys.exit(1)
# Validate ffmpeg availability if selected
if args.video_backend == "ffmpeg" and shutil.which("ffmpeg") is None:
debug.log("--video_backend ffmpeg requires ffmpeg in PATH. Install ffmpeg or use --video_backend opencv",
level="ERROR", category="setup", force=True)
sys.exit(1)
# Inform about caching defaults
if args.cache_dit and args.dit_offload_device == "none":
offload_target = "system memory (CPU)" if get_gpu_backend() != "mps" else "unified memory"
+1 -1
View File
@@ -1,7 +1,7 @@
[project]
name = "seedvr2_videoupscaler"
description = "SeedVR2 official ComfyUI integration: ByteDance-Seed's one-step diffusion-based video/image upscaling with memory-efficient inference"
version = "2.5.21"
version = "2.5.24"
authors = [
{name = "numz"},
{name = "adrientoupet"}
+17 -7
View File
@@ -337,13 +337,23 @@ def edge_guided_alpha_upscale(
rgb_edges = detect_edges_batch(images=rgb_normalized, method='sobel', debug=debug)
# Step 1: Initial bicubic upscale provides smooth base before edge refinement
alpha_upscaled = F.interpolate(
input_alpha,
size=(H_out, W_out),
mode='bicubic',
align_corners=False,
antialias=True
).clamp(0, 1)
# MPS on PyTorch < 2.8 doesn't support bicubic+antialias - use CPU fallback
try:
alpha_upscaled = F.interpolate(
input_alpha,
size=(H_out, W_out),
mode='bicubic',
align_corners=False,
antialias=True
).clamp(0, 1)
except NotImplementedError:
alpha_upscaled = F.interpolate(
input_alpha.cpu(),
size=(H_out, W_out),
mode='bicubic',
align_corners=False,
antialias=True
).to(device).clamp(0, 1)
if is_binary_mask:
if debug:
-8
View File
@@ -533,10 +533,6 @@ def encode_all_batches(
manage_model_device(model=runner.vae, target_device=ctx['vae_offload_device'],
model_name="VAE", debug=debug, reason="VAE offload", runner=runner)
# MPS: sync to get accurate timing and free memory before Phase 2
if ctx['vae_device'].type == 'mps':
torch.mps.synchronize()
debug.end_timer("phase1_encoding", "Phase 1: VAE encoding complete", show_breakdown=True)
debug.log_memory_state("After phase 1 (VAE encoding)", show_tensors=False)
@@ -1054,10 +1050,6 @@ def decode_all_batches(
if 'all_upscaled_latents' in ctx:
release_tensor_collection(ctx['all_upscaled_latents'])
del ctx['all_upscaled_latents']
# MPS: sync to get accurate timing and free memory before Phase 4
if ctx['vae_device'].type == 'mps':
torch.mps.synchronize()
debug.end_timer("phase3_decoding", "Phase 3: VAE decoding complete", show_breakdown=True)
debug.log_memory_state("After phase 3 (VAE decoding)", show_tensors=False)
+15 -1
View File
@@ -146,7 +146,7 @@ def load_quantized_state_dict(checkpoint_path: str, device: torch.device = torch
handle_prefix="model.diffusion_model."
)
elif checkpoint_path.endswith('.pth'):
state = torch.load(checkpoint_path, map_location=device_str, mmap=True)
state = torch.load(checkpoint_path, map_location=device_str, mmap=True, weights_only=True)
else:
raise ValueError(f"Unsupported checkpoint format. Expected .safetensors or .pth, got: {checkpoint_path}")
@@ -393,6 +393,20 @@ class GGUFTensor(torch.Tensor):
if debug:
debug.log(f"Error in {func.__name__} dequantization: {e}", level="WARNING", category="dit", force=True)
raise
# Handle conv2d/conv3d operations (critical for GGUF VAE models)
# Conv3d layers (InflatedCausalConv3d) are not replaced by layer replacement
if func in {torch.nn.functional.conv2d, torch.nn.functional.conv3d}:
if len(args) >= 2 and isinstance(args[1], cls): # weight is second arg
try:
weight_tensor = args[1]
dequantized_weight = weight_tensor.dequantize(device=args[0].device, dtype=args[0].dtype)
new_args = (args[0], dequantized_weight) + args[2:]
return func(*new_args, **kwargs)
except Exception as e:
if debug:
debug.log(f"Error in conv dequantization: {e}", level="WARNING", category="dit", force=True)
raise
# For ALL other operations, delegate to parent WITHOUT dequantization
# This includes .cpu(), .to(), .device, .dtype, .shape, etc.
@@ -1093,7 +1093,7 @@ class VideoAutoencoderKL(diffusers.AutoencoderKL):
):
extra_cond_dim = kwargs.pop("extra_cond_dim") if "extra_cond_dim" in kwargs else None
self.slicing_sample_min_size = slicing_sample_min_size
self.slicing_latent_min_size = slicing_sample_min_size // (2**temporal_scale_num)
self.slicing_latent_min_size = max(1, slicing_sample_min_size // (2**temporal_scale_num))
super().__init__(
in_channels=in_channels,
@@ -1224,6 +1224,10 @@ class VideoAutoencoderKL(diffusers.AutoencoderKL):
output = causal_conv_gather_outputs(output)
# MPS memory leak workaround (pytorch/pytorch#155060)
if self.device.type == 'mps':
torch.mps.empty_cache()
# Only transfer back if needed
return output if output.device == x.device else output.to(x.device)
@@ -1240,6 +1244,10 @@ class VideoAutoencoderKL(diffusers.AutoencoderKL):
output = self.decoder(_z, memory_state=memory_state)
output = causal_conv_gather_outputs(output)
# MPS memory leak workaround (pytorch/pytorch#155060)
if self.device.type == 'mps':
torch.mps.empty_cache()
# Only transfer back if needed
return output if output.device == z.device else output.to(z.device)
@@ -1710,7 +1718,7 @@ class VideoAutoencoderKLWrapper(VideoAutoencoderKL):
if split_size is not None:
self.enable_slicing()
self.slicing_sample_min_size = split_size
self.slicing_latent_min_size = split_size // self.temporal_downsample_factor
self.slicing_latent_min_size = max(1, split_size // self.temporal_downsample_factor)
else:
self.disable_slicing()
for module in self.modules():
+3 -3
View File
@@ -733,7 +733,7 @@ class VideoAutoencoderKL(nn.Module):
if slicing_sample_min_size is None:
slicing_sample_min_size = temporal_downsample_factor
self.slicing_sample_min_size = slicing_sample_min_size
self.slicing_latent_min_size = slicing_sample_min_size // (2**temporal_scale_num)
self.slicing_latent_min_size = max(1, slicing_sample_min_size // (2**temporal_scale_num))
# pass init params to Encoder
self.encoder = Encoder3D(
@@ -886,7 +886,7 @@ class VideoAutoencoderKL(nn.Module):
if split_size is not None:
self.enable_slicing()
self.slicing_sample_min_size = split_size
self.slicing_latent_min_size = split_size // self.temporal_downsample_factor
self.slicing_latent_min_size = max(1, split_size // self.temporal_downsample_factor)
else:
self.disable_slicing()
for module in self.modules():
@@ -950,7 +950,7 @@ class VideoAutoencoderKLWrapper(VideoAutoencoderKL):
self.disable_slicing()
self.slicing_sample_min_size = split_size
if split_size is not None:
self.slicing_latent_min_size = split_size // self.temporal_downsample_factor
self.slicing_latent_min_size = max(1, split_size // self.temporal_downsample_factor)
for module in self.modules():
if isinstance(module, InflatedCausalConv3d):
module.set_memory_device(memory_device)
+6 -5
View File
@@ -98,8 +98,8 @@ def ensure_bitsandbytes_safe():
try:
import bitsandbytes
# Success - bitsandbytes works, other nodes can use it
except (ImportError, OSError, RuntimeError):
# Installation broken or not present - create stub
except (ImportError, OSError, RuntimeError, ValueError):
# Installation broken, not present, or version detection failed - create stub
stub = types.ModuleType('bitsandbytes')
stub.__spec__ = importlib.machinery.ModuleSpec('bitsandbytes', None)
stub.__file__ = None
@@ -592,11 +592,11 @@ def validate_gguf_availability(operation: str = "load GGUF model", debug=None) -
raise RuntimeError(f"GGUF library required to {operation}")
# 4. NVIDIA Conv3d Memory Bug - Workaround for PyTorch 2.9-2.10 + cuDNN >= 91002
# 4. NVIDIA Conv3d Memory Bug - Workaround for PyTorch >= 2.9 + cuDNN >= 91002
def _check_conv3d_memory_bug():
"""
Check if Conv3d memory bug workaround needed.
Bug: PyTorch 2.9-2.10 with cuDNN >= 91002 uses 3x memory for Conv3d
Bug: PyTorch 2.9+ with cuDNN >= 91002 uses 3x memory for Conv3d
with fp16/bfloat16 due to buggy dispatch layer.
"""
try:
@@ -622,7 +622,8 @@ def _check_conv3d_memory_bug():
parts = version_str.split('.')
torch_version = tuple(int(p) for p in parts[:2])
if not ((2, 9) <= torch_version <= (2, 10)):
# Bug affects PyTorch 2.9 and later versions
if torch_version < (2, 9):
return False
if not hasattr(torch.backends.cudnn, 'version'):
+21 -9
View File
@@ -381,6 +381,8 @@ def _rgb_to_lab_batch(rgb: Tensor, device: torch.device, matrix: Tensor, epsilon
rgb_flat = rgb_linear.permute(0, 2, 3, 1).reshape(-1, 3)
del rgb_linear
# Ensure dtype consistency for matrix multiplication
rgb_flat = rgb_flat.to(dtype=matrix.dtype)
xyz_flat = torch.matmul(rgb_flat, matrix.T)
del rgb_flat
@@ -452,6 +454,8 @@ def _lab_to_rgb_batch(lab: Tensor, device: torch.device, matrix_inv: Tensor, eps
xyz_flat = xyz.permute(0, 2, 3, 1).reshape(-1, 3)
del xyz
# Ensure dtype consistency for matrix multiplication
xyz_flat = xyz_flat.to(dtype=matrix_inv.dtype)
rgb_linear_flat = torch.matmul(xyz_flat, matrix_inv.T)
del xyz_flat
@@ -490,6 +494,7 @@ def _histogram_matching_channel(source: Tensor, reference: Tensor, device: torch
# Sort both arrays
source_sorted, source_indices = torch.sort(source_flat)
reference_sorted, _ = torch.sort(reference_flat)
del reference_flat
# Quantile mapping
n_source = len(source_sorted)
@@ -503,12 +508,15 @@ def _histogram_matching_channel(source: Tensor, reference: Tensor, device: torch
ref_indices = (source_quantiles * (n_reference - 1)).long()
ref_indices.clamp_(0, n_reference - 1)
matched_sorted = reference_sorted[ref_indices]
del source_quantiles, ref_indices
del source_quantiles, ref_indices, reference_sorted
# Reconstruct with matched values
matched_flat = torch.empty_like(source_flat)
matched_flat.scatter_(0, source_indices, matched_sorted)
del source_flat, reference_flat, source_sorted, source_indices, reference_sorted, matched_sorted
del source_sorted, source_flat
# Reconstruct using argsort (portable across CUDA/ROCm/MPS)
inverse_indices = torch.argsort(source_indices)
del source_indices
matched_flat = matched_sorted[inverse_indices]
del matched_sorted, inverse_indices
return matched_flat.reshape(original_shape)
@@ -748,11 +756,15 @@ def _histogram_match_1d(source: Tensor, reference: Tensor, device: torch.device)
ref_indices = (source_quantiles * (n_reference - 1)).long()
ref_indices.clamp_(0, n_reference - 1)
matched_sorted = reference_sorted[ref_indices]
del source_quantiles, ref_indices
del source_quantiles, ref_indices, reference_sorted
matched = torch.empty_like(source)
matched.scatter_(0, source_indices, matched_sorted)
del source_sorted, source_indices, reference_sorted, matched_sorted
del source_sorted
# Reconstruct using argsort (portable across CUDA/ROCm/MPS)
inverse_indices = torch.argsort(source_indices)
del source_indices
matched = matched_sorted[inverse_indices]
del matched_sorted, inverse_indices
return matched
+1 -1
View File
@@ -4,7 +4,7 @@ Only includes constants actually used in the codebase
"""
# Version information
__version__ = "2.5.21"
__version__ = "2.5.24"
import os
import warnings
+3 -3
View File
@@ -78,7 +78,7 @@ class Debug:
"device": "🖥️", # Device info
"file": "📂", # File operations
"alpha": "👻", # Alpha operations
"star": "⭐", # Star
"starlove": "⭐💝", # Star + love
"dialogue": "💬", # Dialogue
"none" : "",
}
@@ -259,9 +259,9 @@ class Debug:
"""Print the footer with links - always displayed"""
self.log("", category="none", force=True)
self.log("────────────────────────", category="none", force=True)
self.log("Questions? Updates? Watch the videos, star the repo & join us!", category="dialogue", force=True)
self.log("Questions? Updates? Watch, star & sponsor if you can!", category="dialogue", force=True)
self.log("https://www.youtube.com/@AInVFX", category="generation", force=True)
self.log("https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler", category="star", force=True)
self.log("https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler", category="starlove", force=True)
@torch._dynamo.disable # Skip tracing to avoid time.time() warnings
def start_timer(self, name: str, force: bool = False) -> None: