Centralize BlockSwap validation, auto-disable on macOS, update docs
- Add validate_blockswap_config() in blockswap.py as single validation point - Auto-disable BlockSwap on macOS (unified memory makes it meaningless) - Improve error messages for missing dit_offload_device - Update CLI and ComfyUI tooltips for BlockSwap and model caching - Update README: BlockSwap macOS note, caching descriptions, attention backends - Remove duplicate validation from dit_model_loader.py and inference_cli.py Partially fixes #401 (M4 Pro macOS BlockSwap offload device error)
This commit is contained in:
@@ -4,7 +4,7 @@
|
||||
|
||||
Official release of [SeedVR2](https://github.com/ByteDance-Seed/SeedVR) for ComfyUI that enables high-quality video and image upscaling.
|
||||
|
||||
Can run as **Multi-GPU standalone CLI** too, see [🖥️ Run as Standalone](#️-run-as-standalone-cli) section.
|
||||
Can run as **Multi-GPU standalone CLI** too, see [🖥️ Run as Standalone](#-run-as-standalone-cli) section.
|
||||
|
||||
[](https://youtu.be/MBtWYXq_r60)
|
||||
|
||||
@@ -14,8 +14,8 @@ Can run as **Multi-GPU standalone CLI** too, see [🖥️ Run as Standalone](#
|
||||
|
||||
## 📋 Quick Access
|
||||
|
||||
- [🆙 Future Releases](#-future-releases)
|
||||
- [🚀 Updates](#-updates)
|
||||
- [🆙 Future Work](#-future-work)
|
||||
- [🚀 Release Notes](#-release-notes)
|
||||
- [🎯 Features](#-features)
|
||||
- [🔧 Requirements](#-requirements)
|
||||
- [📦 Installation](#-installation)
|
||||
@@ -26,7 +26,7 @@ Can run as **Multi-GPU standalone CLI** too, see [🖥️ Run as Standalone](#
|
||||
- [🙏 Credits](#-credits)
|
||||
- [📜 License](#-license)
|
||||
|
||||
## 🆙 Future Releases
|
||||
## 🆙 Future Work
|
||||
|
||||
We're actively working on improvements and new features. To stay informed:
|
||||
|
||||
@@ -34,7 +34,7 @@ We're actively working on improvements and new features. To stay informed:
|
||||
- **💬 Join the Community**: Learn from others, share your workflows, and get help in the [Discussions](https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/discussions)
|
||||
- **🔮 Next Model Survey**: We're looking for community input on the next open-source super-powerful generic restoration model. Share your suggestions in [Issue #164](https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/issues/164)
|
||||
|
||||
## 🚀 Updates
|
||||
## 🚀 Release Notes
|
||||
|
||||
**2025.12.10 - Version 2.5.19**
|
||||
|
||||
@@ -232,7 +232,7 @@ We're actively working on improvements and new features. To stay informed:
|
||||
|
||||
**2025.07.03**
|
||||
|
||||
- 🛠️ Can run as **standalone mode** with **Multi GPU** see [🖥️ Run as Standalone](#️-run-as-standalone-cli)
|
||||
- 🛠️ Can run as **standalone mode** with **Multi GPU** see [🖥️ Run as Standalone](#run-as-standalone-cli)
|
||||
|
||||
**2025.06.30**
|
||||
|
||||
@@ -279,8 +279,8 @@ We're actively working on improvements and new features. To stay informed:
|
||||
### Performance Features
|
||||
- **torch.compile Integration**: Optional 20-40% DiT speedup and 15-25% VAE speedup with PyTorch 2.0+ compilation
|
||||
- **Multi-GPU CLI**: Distribute workload across multiple GPUs with automatic temporal overlap blending
|
||||
- **Model Caching**: Keep models loaded in memory for faster batch processing
|
||||
- **Flexible Attention Backends**: Choose between PyTorch SDPA (stable, always available) or Flash Attention 2 (faster on supported hardware)
|
||||
- **Model Caching**: Keep models loaded between generations for single-GPU directory processing or multi-GPU streaming
|
||||
- **Flexible Attention Backends**: Choose between PyTorch SDPA (stable, always available), Flash Attention 2/3, or SageAttention 2/3 for faster computation on supported hardware
|
||||
|
||||
### Quality Control
|
||||
- **Advanced Color Correction**: Five methods including LAB (recommended for highest fidelity), wavelet, wavelet adaptive, HSV, and AdaIN
|
||||
@@ -309,7 +309,7 @@ With the current optimizations (tiling, BlockSwap, GGUF quantization), SeedVR2 c
|
||||
- **Python**: 3.12+ (Python 3.12 and 3.13 tested and recommended)
|
||||
- **PyTorch**: 2.0+ for torch.compile support (optional but recommended)
|
||||
- **Triton**: Required for torch.compile with inductor backend (optional)
|
||||
- **Flash Attention 2**: Provides faster attention computation on supported hardware (optional, falls back to PyTorch SDPA)
|
||||
- **Flash Attention / SageAttention**: Flash Attention 2 (Ampere+), Flash Attention 3 (Hopper+), SageAttention 2 or SageAttention 3 (Blackwell) provide faster attention computation on supported hardware (optional, falls back to PyTorch SDPA)
|
||||
|
||||
## 📦 Installation
|
||||
|
||||
@@ -434,7 +434,11 @@ Configure the DiT (Diffusion Transformer) model for video upscaling.
|
||||
|
||||
**BlockSwap Explained:**
|
||||
|
||||
BlockSwap enables running large models on GPUs with limited VRAM by dynamically swapping transformer blocks between GPU and CPU memory during inference. Here's how it works:
|
||||
BlockSwap enables running large models on GPUs with limited VRAM by dynamically swapping transformer blocks between GPU and CPU memory during inference.
|
||||
|
||||
> **Note:** BlockSwap is not available on macOS. Apple Silicon Macs use unified memory architecture where GPU and CPU share the same memory pool, making BlockSwap meaningless. The option will be automatically disabled with a warning if requested on macOS.
|
||||
|
||||
Here's how it works:
|
||||
|
||||
- **What it does**: Keeps only the currently-needed transformer blocks on the GPU, while storing the rest on CPU or another device
|
||||
- **When to use it**: When you get OOM (Out of Memory) errors during the upscaling phase
|
||||
@@ -870,9 +874,8 @@ python inference_cli.py media_folder/ \
|
||||
**Memory Management:**
|
||||
- `--dit_offload_device`: Device to offload DiT model: 'none' (keep on GPU), 'cpu', or 'cuda:X' (default: none)
|
||||
- `--vae_offload_device`: Device to offload VAE model: 'none', 'cpu', or 'cuda:X' (default: none)
|
||||
- `--blocks_to_swap`: Number of transformer blocks to swap (0=disabled, 3B: 0-32, 7B: 0-36). Requires dit_offload_device (default: 0)
|
||||
- `--swap_io_components`: Offload I/O components for additional VRAM savings. Requires dit_offload_device
|
||||
- `--use_non_blocking`: Use non-blocking memory transfers for BlockSwap (recommended)
|
||||
- `--blocks_to_swap`: Number of transformer blocks to swap (0=disabled, 3B: 0-32, 7B: 0-36). Requires dit_offload_device (default: 0). Not available on macOS.
|
||||
- `--swap_io_components`: Offload I/O components for additional VRAM savings. Requires dit_offload_device. Not available on macOS.
|
||||
|
||||
**VAE Tiling:**
|
||||
- `--vae_encode_tiled`: Enable VAE encode tiling to reduce VRAM during encoding
|
||||
@@ -896,8 +899,8 @@ python inference_cli.py media_folder/ \
|
||||
- `--compile_dynamo_recompile_limit`: Max recompilation attempts before fallback (default: 128)
|
||||
|
||||
**Model Caching (batch processing):**
|
||||
- `--cache_dit`: Cache DiT model between files (single GPU only, speeds up directory processing)
|
||||
- `--cache_vae`: Cache VAE model between files (single GPU only, speeds up directory processing)
|
||||
- `--cache_dit`: Keep DiT model in memory between generations. Works with single-GPU directory processing or multi-GPU streaming (`--chunk_size`). Requires `--dit_offload_device`
|
||||
- `--cache_vae`: Keep VAE model in memory between generations. Works with single-GPU directory processing or multi-GPU streaming (`--chunk_size`). Requires `--vae_offload_device`
|
||||
|
||||
**Multi-GPU:**
|
||||
- `--cuda_device`: CUDA device id(s). Single id (e.g., '0') or comma-separated list '0,1' for multi-GPU
|
||||
@@ -1000,7 +1003,7 @@ For detailed contribution guidelines, see [CONTRIBUTING.md](CONTRIBUTING.md).
|
||||
|
||||
This ComfyUI implementation is a collaborative project by **[NumZ](https://github.com/numz)** and **[AInVFX](https://www.youtube.com/@AInVFX)** (Adrien Toupet), based on the original [SeedVR2](https://github.com/ByteDance-Seed/SeedVR) by ByteDance Seed Team.
|
||||
|
||||
Special thanks to our community contributors including [benjaminherb](https://github.com/benjaminherb), [cmeka](https://github.com/cmeka), [FurkanGozukara](https://github.com/FurkanGozukara), [JohnAlcatraz](https://github.com/JohnAlcatraz), [lihaoyun6](https://github.com/lihaoyun6), [Luchuanzhao](https://github.com/Luchuanzhao), [Luke2642](https://github.com/Luke2642), [naxci1](https://github.com/naxci1), [q5sys](https://github.com/q5sys), and many others for their improvements, bug fixes, and testing.
|
||||
Special thanks to our community contributors including [naxci1](https://github.com/naxci1), [benjaminherb](https://github.com/benjaminherb), [cmeka](https://github.com/cmeka), [FurkanGozukara](https://github.com/FurkanGozukara), [JohnAlcatraz](https://github.com/JohnAlcatraz), [lihaoyun6](https://github.com/lihaoyun6), [Luchuanzhao](https://github.com/Luchuanzhao), [Luke2642](https://github.com/Luke2642), [proxyid](https://github.com/proxyid), [q5sys](https://github.com/q5sys), and many others for their improvements, bug fixes, and testing.
|
||||
|
||||
## 📜 License
|
||||
|
||||
|
||||
+7
-22
@@ -1327,9 +1327,10 @@ Examples:
|
||||
blockswap_group = parser.add_argument_group('Memory optimization (BlockSwap)')
|
||||
blockswap_group.add_argument("--blocks_to_swap", type=int, default=0,
|
||||
help="Transformer blocks to swap for VRAM savings. 0-32 (3B) or 0-36 (7B). "
|
||||
"Requires --dit_offload_device. Default: 0 (disabled)")
|
||||
"Requires --dit_offload_device. Not available on macOS. Default: 0 (disabled)")
|
||||
blockswap_group.add_argument("--swap_io_components", action="store_true",
|
||||
help="Offload DiT I/O layers for extra VRAM savings. Requires --dit_offload_device")
|
||||
help="Offload DiT I/O layers for extra VRAM savings. Requires --dit_offload_device. "
|
||||
"Not available on macOS")
|
||||
|
||||
# VAE Tiling
|
||||
vae_group = parser.add_argument_group('VAE tiling (for high resolution upscale)')
|
||||
@@ -1374,9 +1375,11 @@ Examples:
|
||||
# Model Caching (for batch processing)
|
||||
cache_group = parser.add_argument_group('Model caching (batch processing)')
|
||||
cache_group.add_argument("--cache_dit", action="store_true",
|
||||
help="Cache DiT model between files (single GPU only, speeds up directory processing)")
|
||||
help="Keep DiT model in memory between generations. Works with single-GPU directory processing "
|
||||
"or multi-GPU streaming (--chunk_size). Requires --dit_offload_device")
|
||||
cache_group.add_argument("--cache_vae", action="store_true",
|
||||
help="Cache VAE model between files (single GPU only, speeds up directory processing)")
|
||||
help="Keep VAE model in memory between generations. Works with single-GPU directory processing "
|
||||
"or multi-GPU streaming (--chunk_size). Requires --vae_offload_device")
|
||||
|
||||
# Debugging
|
||||
debug_group = parser.add_argument_group('Debugging')
|
||||
@@ -1435,24 +1438,6 @@ def main() -> None:
|
||||
debug.log(f"VAE decode tile overlap ({args.vae_decode_tile_overlap}) must be smaller than tile size ({args.vae_decode_tile_size})", level="ERROR", category="vae", force=True)
|
||||
sys.exit(1)
|
||||
|
||||
# Validate BlockSwap configuration - either blocks_to_swap or swap_io_components requires dit_offload_device
|
||||
blockswap_enabled = args.blocks_to_swap > 0 or args.swap_io_components
|
||||
if blockswap_enabled and args.dit_offload_device == "none":
|
||||
config_details = []
|
||||
if args.blocks_to_swap > 0:
|
||||
config_details.append(f"blocks_to_swap={args.blocks_to_swap}")
|
||||
if args.swap_io_components:
|
||||
config_details.append("swap_io_components=True")
|
||||
|
||||
debug.log(
|
||||
f"BlockSwap enabled ({', '.join(config_details)}) but dit_offload_device='none'. "
|
||||
"BlockSwap requires dit_offload_device to be set (typically 'cpu'). "
|
||||
"Either set --dit_offload_device cpu or disable BlockSwap "
|
||||
"(--blocks_to_swap 0 and do not use --swap_io_components)",
|
||||
level="ERROR", category="blockswap", force=True
|
||||
)
|
||||
sys.exit(1)
|
||||
|
||||
# Inform about caching defaults
|
||||
if args.cache_dit and args.dit_offload_device == "none":
|
||||
offload_target = "system memory (CPU)" if get_gpu_backend() != "mps" else "unified memory"
|
||||
|
||||
@@ -74,7 +74,7 @@ from ..optimization.compatibility import (
|
||||
TRITON_AVAILABLE,
|
||||
validate_attention_mode
|
||||
)
|
||||
from ..optimization.blockswap import is_blockswap_enabled, apply_block_swap_to_dit, cleanup_blockswap
|
||||
from ..optimization.blockswap import is_blockswap_enabled, validate_blockswap_config, apply_block_swap_to_dit, cleanup_blockswap
|
||||
from ..optimization.memory_manager import cleanup_dit, cleanup_vae
|
||||
from ..utils.constants import find_model_file
|
||||
|
||||
@@ -795,6 +795,14 @@ def configure_runner(
|
||||
if debug is None:
|
||||
raise ValueError("Debug instance must be provided to configure_runner")
|
||||
|
||||
# Validate BlockSwap configuration early (before any model loading)
|
||||
block_swap_config = validate_blockswap_config(
|
||||
block_swap_config=block_swap_config,
|
||||
dit_device=ctx['dit_device'],
|
||||
dit_offload_device=ctx.get('dit_offload_device'),
|
||||
debug=debug
|
||||
)
|
||||
|
||||
# Phase 1: Initialize cache and get cached models
|
||||
cache_context = _initialize_cache_context(
|
||||
dit_cache, vae_cache, dit_id, vae_id,
|
||||
|
||||
@@ -66,7 +66,8 @@ class SeedVR2LoadDiTModel(io.ComfyNode):
|
||||
"• 3B model: 0-32 blocks\n"
|
||||
"• 7B model: 0-36 blocks\n"
|
||||
"\n"
|
||||
"Requires offload_device to be set and different from device."
|
||||
"Requires offload_device to be set and different from device.\n"
|
||||
"Not available on macOS (unified memory architecture)."
|
||||
)
|
||||
),
|
||||
io.Boolean.Input("swap_io_components",
|
||||
@@ -74,7 +75,8 @@ class SeedVR2LoadDiTModel(io.ComfyNode):
|
||||
optional=True,
|
||||
tooltip=(
|
||||
"Offload input/output embeddings and normalization layers to reduce VRAM.\n"
|
||||
"Requires offload_device to be set and different from device."
|
||||
"Requires offload_device to be set and different from device.\n"
|
||||
"Not available on macOS (unified memory architecture)."
|
||||
)
|
||||
),
|
||||
io.Combo.Input("offload_device",
|
||||
@@ -152,16 +154,8 @@ class SeedVR2LoadDiTModel(io.ComfyNode):
|
||||
NodeOutput containing configuration dictionary for SeedVR2 main node
|
||||
|
||||
Raises:
|
||||
ValueError: If BlockSwap is enabled but offload_device is invalid
|
||||
ValueError: If cache_model is enabled but offload_device is not set
|
||||
"""
|
||||
# Validate BlockSwap configuration
|
||||
if (blocks_to_swap > 0 or swap_io_components) and (offload_device == "none" or offload_device == device):
|
||||
raise ValueError(
|
||||
"BlockSwap requires offload_device to be set and different from device. "
|
||||
f"Current: device='{device}', offload_device='{offload_device}'. "
|
||||
"Please set offload_device to a different device (e.g., 'cpu' or another GPU)."
|
||||
)
|
||||
|
||||
# Validate cache_model configuration
|
||||
if cache_model and offload_device == "none":
|
||||
raise ValueError(
|
||||
|
||||
@@ -349,13 +349,12 @@ class SeedVR2VideoUpscaler(io.ComfyNode):
|
||||
|
||||
block_swap_config = None
|
||||
if blocks_to_swap > 0 or swap_io_components:
|
||||
# Convert offload device string to torch.device for BlockSwap
|
||||
block_swap_config = {
|
||||
"blocks_to_swap": blocks_to_swap,
|
||||
"swap_io_components": swap_io_components,
|
||||
}
|
||||
if dit_offload_str != "none":
|
||||
block_swap_config = {
|
||||
"blocks_to_swap": blocks_to_swap,
|
||||
"swap_io_components": swap_io_components,
|
||||
"offload_device": torch.device(dit_offload_str)
|
||||
}
|
||||
block_swap_config["offload_device"] = torch.device(dit_offload_str)
|
||||
|
||||
# Device configuration for offloading - convert "none" to None, else torch.device
|
||||
vae_offload_str = vae.get("offload_device", "none")
|
||||
|
||||
@@ -47,6 +47,78 @@ def is_blockswap_enabled(config: Optional[Dict[str, Any]]) -> bool:
|
||||
return blocks_to_swap > 0 or swap_io_components
|
||||
|
||||
|
||||
def validate_blockswap_config(
|
||||
block_swap_config: Optional[Dict[str, Any]],
|
||||
dit_device: 'torch.device',
|
||||
dit_offload_device: Optional['torch.device'],
|
||||
debug: 'Debug'
|
||||
) -> Optional[Dict[str, Any]]:
|
||||
"""
|
||||
Validate and potentially modify BlockSwap configuration.
|
||||
|
||||
Performs platform-specific validation and configuration adjustment:
|
||||
- On macOS (MPS): Auto-disables BlockSwap since unified memory makes it meaningless
|
||||
- On other platforms: Validates that offload_device is properly configured
|
||||
|
||||
This is the single authoritative validation point for BlockSwap configuration,
|
||||
called early in configure_runner() before any model loading.
|
||||
|
||||
Args:
|
||||
block_swap_config: BlockSwap configuration dictionary (may be None)
|
||||
dit_device: Target device for DiT model inference
|
||||
dit_offload_device: Device for offloading DiT blocks (may be None)
|
||||
debug: Debug instance for logging warnings/errors
|
||||
|
||||
Returns:
|
||||
Validated/modified block_swap_config (may be None or modified copy)
|
||||
|
||||
Raises:
|
||||
ValueError: If BlockSwap is enabled but offload_device is invalid (non-MPS only)
|
||||
"""
|
||||
if not is_blockswap_enabled(block_swap_config):
|
||||
return block_swap_config
|
||||
|
||||
blocks_to_swap = block_swap_config.get("blocks_to_swap", 0)
|
||||
swap_io_components = block_swap_config.get("swap_io_components", False)
|
||||
|
||||
# Check for macOS unified memory - BlockSwap is meaningless there
|
||||
if dit_device.type == "mps":
|
||||
debug.log(
|
||||
f"BlockSwap disabled: macOS uses unified memory (no separate VRAM/RAM). "
|
||||
f"Ignoring blocks_to_swap={blocks_to_swap}, swap_io_components={swap_io_components}",
|
||||
level="WARNING", category="blockswap", force=True
|
||||
)
|
||||
# Return disabled config
|
||||
return {
|
||||
**block_swap_config,
|
||||
"blocks_to_swap": 0,
|
||||
"swap_io_components": False
|
||||
}
|
||||
|
||||
# Validate offload_device is set and different from dit_device
|
||||
offload_device_valid = (
|
||||
dit_offload_device is not None and
|
||||
str(dit_offload_device) != str(dit_device)
|
||||
)
|
||||
|
||||
if not offload_device_valid:
|
||||
config_details = []
|
||||
if blocks_to_swap > 0:
|
||||
config_details.append(f"blocks_to_swap={blocks_to_swap}")
|
||||
if swap_io_components:
|
||||
config_details.append("swap_io_components=True")
|
||||
|
||||
offload_str = str(dit_offload_device) if dit_offload_device else "none"
|
||||
raise ValueError(
|
||||
f"BlockSwap enabled ({', '.join(config_details)}) but dit_offload_device is invalid. "
|
||||
f"Current: device='{dit_device}', dit_offload_device='{offload_str}'. "
|
||||
f"BlockSwap requires offload_device on the DiT Model to be set and different from device. "
|
||||
f"Set --dit_offload_device cpu or disable BlockSwap."
|
||||
)
|
||||
|
||||
return block_swap_config
|
||||
|
||||
|
||||
# Timing helpers marked to skip torch.compile tracing
|
||||
# These functions are excluded from Dynamo's graph tracing to avoid warnings
|
||||
# about non-traceable builtins like time.time(), but they still execute normally
|
||||
|
||||
Reference in New Issue
Block a user