Building on @thehhmdb's fix in PR #418 which identified the stderr
pipe blocking issue. This refinement simplifies the solution by
redirecting stderr to DEVNULL and adding stdin.flush() to prevent
buffering deadlocks.
Improvements:
- Simpler implementation without threading complexity
- Zero memory overhead
- Better error messages for debugging
- Maintains fix for the original hanging issue
Co-authored-by: thehhmdb <thehhmdb@users.noreply.github.com>
Fixesnumz/ComfyUI-SeedVR2_VideoUpscaler#418
Move ffmpeg check from FFMPEGVideoWriter to argument validation phase.
Prevents wasted GPU processing time when ffmpeg backend is selected
but ffmpeg is not installed.
- Add --video_backend flag: 'opencv' (default) or 'ffmpeg'
- Add --10bit flag: enables x265/yuv420p10le for reduced banding
- Without --10bit, ffmpeg uses x264/yuv420p for max compatibility
- FFMPEGVideoWriter class with cv2.VideoWriter-compatible interface
- Validates ffmpeg availability before encoding
Based on PR #409 by thehhmdb
- Skip CPU tensor offload on MPS (no memory benefit, causes sync stall)
- Keep input_images and final_video on MPS device
- Add explicit MPS sync at phase boundaries for accurate timing
- Preload text embeddings before Phase 1 to avoid Phase 2 stall
- Skip model→CPU movement before deletion on MPS cleanup
- Add validate_blockswap_config() in blockswap.py as single validation point
- Auto-disable BlockSwap on macOS (unified memory makes it meaningless)
- Improve error messages for missing dit_offload_device
- Update CLI and ComfyUI tooltips for BlockSwap and model caching
- Update README: BlockSwap macOS note, caching descriptions, attention backends
- Remove duplicate validation from dit_model_loader.py and inference_cli.py
Partially fixes#401 (M4 Pro macOS BlockSwap offload device error)
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.
The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.
- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow
Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
- Re-implemented `precision` control (`fp16`, `bf16`, `bf32`, `auto`) in CLI, ComfyUI node, and backend logic to respect user choice.
- Fixed `UnboundLocalError` in `apply_model_specific_config` by ensuring `compute_dtype` is always initialized before use.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by properly passing the restored `precision` argument.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` for clarity and fixed fallback logic to ensure SageAttention is correctly prioritized.
- Added explicit logging of active attention backend and execution confirmation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Fixed `NameError` crash in `SeedVR2VideoUpscaler` by removing residual `precision` usage.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across CLI, configuration, and internal logic to resolve naming confusion with Stable Diffusion.
- Removed manual `precision` argument from CLI and ComfyUI node to enforce auto-optimization and simplify usage.
- Added explicit logging of the "Active Attention Mode" at the end of the CLI process to confirm which backend was actually used.
- Added "🚀 Executing SageAttention..." console log on first kernel execution to verify optimization activation.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping, fixing a potential performance bottleneck.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, CLI, and ComfyUI nodes to fix naming confusion.
- Added strict `precision` control (`fp16`, `bf16`, `bf32`, `auto`) to CLI and internal configuration logic.
- Implemented robust fallback logic for SageAttention (SA3 -> SA2 -> Flash Attention 2 -> SDPA) with version checks.
- Updated `FP8CompatibleDiT` to exclude `FlashAttentionVarlen` modules from unnecessary wrapping.
- Added explicit logging ("🚀 Executing SageAttention...") to confirm kernel execution.
- Fixed bug where user-selected precision was being overridden by auto-detection defaults.
- Renamed `sd2`/`sd3` to `sa2`/`sa3` across configuration, compatibility layer, and ComfyUI node definitions to fix naming confusion.
- Fixed logic bug in `FlashAttentionVarlen` that prevented SageAttention from running even when requested (was checking for `sd` prefix instead of `sa`).
- Added robust availability and version checks for SageAttention (v2 vs v3) with improved fallback logic (SA3 -> SA2 -> FlashAttention 2 -> SDPA).
- Added explicit console logging when SageAttention kernel is first executed to verify optimization is active.
- Added `--precision` argument to CLI to allow explicit control over compute dtype (fp16, bf16, bf32, auto), enabling further performance tuning.
- Updated `FP8CompatibleDiT` wrapper to exclude `FlashAttentionVarlen` modules, preventing double-wrapping and casting issues.
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup
Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
Use PyTorch shared memory instead of pickling numpy arrays through queue.
Prevents MemoryError when transferring large results between processes.
Thank you @FurkanGozukara
- BlockSwap: Show effective/total blocks (e.g., 32/32) instead of raw requested value
- CLI: Skip CUDA device validation when CUDA_VISIBLE_DEVICES already set (worker process)
- Mac: Use direct processing instead of spawning subprocess (MPS allocator fails in child process)
- Multi-GPU: Set CUDA_VISIBLE_DEVICES before spawn so child inherits it before module-level torch import
- Remove redundant env setup in worker (now inherited from parent)
- Improve output folder naming: batch creates {folder}_upscaled/ sibling with original filenames, single file adds _upscaled suffix
- Add RGBA alpha channel detection and preservation (matches ComfyUI)
- Convert all output paths to absolute for clarity in logs
- Replace dual glob loops with single iterdir scan for cross-platform consistency
- Fixes duplicate file processing in batch mode on Windows case-insensitive filesystem
- Improves directory scanning performance 2-3x by reducing filesystem operations
- Add ComfyUI registry logo
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
- Standardize tooltip format with multi-line descriptions and bullet points
- Add comprehensive output tooltips for all nodes (DiT, VAE, torch.compile, upscaler)
- Enhance node descriptions with detailed capability summaries
- Simplify CLI tile size arguments to single integers (converted internally to tuples)
- Remove OneOrTwoValues argparse class for cleaner implementation
- Fix encode_tiled tooltip (was incorrectly referencing decoding)
- Clarify color correction purpose (corrects upscaling color shifts)
- Add multi-GPU offloading information to all offload_device tooltips
- Improve torch.compile parameter descriptions with use cases
- Ensure CLI and ComfyUI tooltips are consistent in terminology and structure
- Fix validation cache location to respect --model_dir parameter
- Fix output path handling for directories without extensions
- Remove spurious directory creation in get_base_cache_dir
- Enhanced CLI help with usage examples and argument grouping
- Auto-display help when script invoked without arguments
- Correct type hints (device_id: str, debug: Debug)
- Remove redundant type conversions and makedirs calls
- Reorganize imports to module top for clarity
- Improved docstrings & tooltip
- Change default batch_size from 1 to 5 to match ComfyUI integration
- Add support for seedvr2_ema_3b-Q8_0.gguf model
- Validate --cuda_device arguments early in pre-parsing phase
- Check device IDs exist and are within available GPU range
- Fail fast with clear error messages showing available devices
- Auto-detect output format per file type (mp4 for videos, png for images)
- Add visual separators between processed files for better readability
- Simplify FPS calculation to use wall-clock time for real-world throughput
- Consolidate banner/footer into shared Debug methods
- Update offload device args to support multi-GPU (cpu/cuda:N)
- Standardize terminology: 'upscaling' instead of 'video upscaling'
- Remove code duplication between CLI and ComfyUI implementations
- Consistent quote style (double quotes) throughout CLI
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)
Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable
Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability
Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
- Add max_resolution parameter (default: 0 = no limit) to both CLI and ComfyUI
- After new_resolution scales shortest edge, max_resolution ensures no edge exceeds limit
- Scales down proportionally if constraint violated
- Maintains backward compatibility with default value of 0
- Add dit_offload_device parameter for proper blockSwap configuration
- Ensure consistent dtype management throughout CLI and ComfyUI (float32 input with bfloat16 pipeline)
- Translate all French comments to English
- Add comprehensive docstrings and section headers
- Remove obsolete use_non_blocking and enable_debug parameters
- Add error handling and validation
- Implement prepend_video_frames() for artifact reduction at video start
- Add blend_overlapping_frames() with Hann window for smooth transitions
- Expose temporal_overlap (0-16) and prepend_frames (0-32) in ComfyUI node
- Unify CLI and ComfyUI to use shared prepend/overlap functions
- Add comprehensive logging for frame adjustments (prepend/overlap/padding)
CFG (Classifier-Free Guidance) does not work with SeedVR2's distilled
one-step diffusion model. The model was trained to produce final results
in a single step without iterative guidance.
- Remove cfg_scale parameter from ComfyUI node and CLI interface
- Force internal cfg_scale to 1.0 in upscale_all_batches()
This avoids artifacts introduced when users changed cfg_scale away from 1.0.
- Switch to deterministic VAE encoding (mode vs sample) to eliminate high-resolution noise artifacts
- Make VAE encode tiling optional (disabled by default) since deterministic encoding resolves artifacts
- Add tile debug visualization feature with adaptive scaling and color-coded boundaries
- Remove redundant dtype conversions in VAE code for better performance
- Minor code cleanup and documentation update
VAE Changes:
- Enable encode tiling by default (prevents noise artifacts at high resolution)
- Increase tile size to 1024px (down from 512px) for optimal quality
- Increase tile overlap to 128px for better blending
Dtype Pipeline:
- Hardcode compute_dtype to bfloat16 for consistent quality/performance/VRAM balance
- Ensure all pipeline steps are using compute_dtype when relevant
- Refactor code for improved performance and memory management
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
* VAE encoding: seed+1M for deterministic sampling without quality loss
* DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id
Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)