Commit Graph
266 Commits
Author SHA1 Message Date
Adrien Toupet 01f4e55b87 docs: add usage screenshots 2025-11-07 00:00:16 -05:00
Adrien Toupet fccd34fc5d docs: release v2.5.0 with comprehensive changelog and updated usage screenshots 2025-11-06 23:59:44 -05:00
Adrien Toupet d82234d418 docs: Add individual node screenshots and update README images 2025-11-06 23:15:56 -05:00
Adrien Toupet b0d85aae51 fix: node position in example workflow 2025-11-06 22:57:32 -05:00
Adrien Toupet 4e9ce4710e Unify parameter names across CLI and ComfyUI interface
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
2025-11-06 22:51:20 -05:00
Adrien Toupet afd5950d37 feat: torch.compile optimization and memory tracking improvements
- Eliminate graph breaks in na.py and attention.py for full torch.compile support
  * Replace cumsum-based tensor slicing with _tensor_split to avoid .item() calls
  * Use torch.tensor_split with .long().cpu() for PyTorch API requirements
  * Replace reshape-based averaging with split-stack-mean pattern
- Fix phase peak VRAM tracking to capture peaks during OOM retry cycles
- Standardize phase4 naming to 'postprocessing' for consistency
- Update RoPE docstrings for accuracy
- Update example workflows with icon
2025-11-06 17:48:39 -05:00
Adrien Toupet 690cc39379 docs: improve best practices OOM troubleshooting guidance 2025-11-06 14:40:59 -05:00
Adrien Toupet 20dab62dc3 feat: add uniform_batch_size for temporal consistency + unify padding logic
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
2025-11-06 14:13:22 -05:00
Adrien Toupet e3594b5fd1 Add example assets 2025-11-05 23:17:05 -05:00
Adrien Toupet f8627298f7 Add peak VRAM tracking and summary display by phase 2025-11-05 21:21:23 -05:00
Adrien Toupet 3b904b3798 docs: simplify installation steps to unified .venv structure 2025-11-05 20:58:20 -05:00
Adrien Toupet 7b12801af1 fix(docs): fixed broken links in README 2025-11-05 17:34:21 -05:00
Adrien Toupet 4e728ad953 docs: comprehensive README overhaul for v2.5.0 release
- Remove nightly branch warning (deploying to main)
- Add Future Releases section with community engagement links
- Expand Features into 8 organized categories (Core, Model Support, Memory Optimization, Performance, Quality Control, Workflow)
- Update Requirements: document 8GB-24GB+ VRAM tiers with optimization strategies
- Improve Installation: add ComfyUI Manager as primary method, fix python_embeded syntax, use uv for manual install
- Complete Usage rewrite: document all 4 nodes (DiT/VAE loaders, Torch Compile, Main Upscaler) with parameters, tooltips, and examples
- Add BlockSwap and VAE Tiling detailed explanations with troubleshooting guides
- Provide 3 workflow templates: Basic (24GB+), Low VRAM (8-12GB), High Performance (torch.compile)
- Revamp CLI section: separate existing ComfyUI users from standalone installation, update all arguments
- Add Multi-GPU processing explanation with overlap blending example
- Remove outdated Benchmarks section
- Update Limitations: clarify 4n+1 batch size requirement, note VAE bottleneck, add best practices
- Simplify Contributing with link to CONTRIBUTING.md
- Fix device tooltips in DiT/VAE loaders (remove /CPU reference)
2025-11-05 17:27:28 -05:00
Adrien Toupet 806bb94df0 feat: unify and improve tooltip documentation across CLI and ComfyUI nodes
- Standardize tooltip format with multi-line descriptions and bullet points
- Add comprehensive output tooltips for all nodes (DiT, VAE, torch.compile, upscaler)
- Enhance node descriptions with detailed capability summaries
- Simplify CLI tile size arguments to single integers (converted internally to tuples)
- Remove OneOrTwoValues argparse class for cleaner implementation
- Fix encode_tiled tooltip (was incorrectly referencing decoding)
- Clarify color correction purpose (corrects upscaling color shifts)
- Add multi-GPU offloading information to all offload_device tooltips
- Improve torch.compile parameter descriptions with use cases
- Ensure CLI and ComfyUI tooltips are consistent in terminology and structure
2025-11-05 15:35:22 -05:00
Adrien Toupet 326489d94c fix CLI: move Debug import after CUDA allocator config to fix batch processing errors 2025-11-05 14:27:06 -05:00
Adrien Toupet 9b79254c39 refactor(cli): improvements and bug fixes + 3b-Q8_0.gguf support
- Fix validation cache location to respect --model_dir parameter
- Fix output path handling for directories without extensions
- Remove spurious directory creation in get_base_cache_dir
- Enhanced CLI help with usage examples and argument grouping
- Auto-display help when script invoked without arguments
- Correct type hints (device_id: str, debug: Debug)
- Remove redundant type conversions and makedirs calls
- Reorganize imports to module top for clarity
- Improved docstrings & tooltip
- Change default batch_size from 1 to 5 to match ComfyUI integration
- Add support for seedvr2_ema_3b-Q8_0.gguf model
2025-11-05 00:28:19 -05:00
Adrien Toupet 3725c1061d refactor: centralize dimension computation and logging for CLI/ComfyUI
- Add compute_generation_info() and log_generation_start() helpers
- Move prepend_frames logic from extraction to processing pipeline
- Eliminate code duplication between CLI and ComfyUI workflows
- Add consistent dimension/parameter logging for both interfaces
- Disable argparse prefix matching for safer CLI usage
2025-11-04 17:01:01 -05:00
Adrien Toupet 1691e657b4 fix(cli): add CUDA device validation before torch initialization
- Validate --cuda_device arguments early in pre-parsing phase
- Check device IDs exist and are within available GPU range
- Fail fast with clear error messages showing available devices
2025-11-04 15:52:05 -05:00
Adrien Toupet 32a049dfd9 feat: Add CLI model caching for multi-file processing and unify device handling
- Add --cache_dit and --cache_vae flags for efficient multi-file directory processing
- Refactor processing pipeline to eliminate duplication between worker and direct modes
- Implement platform-agnostic device management (CUDA/MPS/CPU)
- Unify parameter naming: res_w→resolution, max_res_w→max_resolution across codebase
- Add smart offload device defaults when caching enabled
- Improve validation and user feedback for cache + multi-GPU scenarios
2025-11-04 15:12:28 -05:00
Adrien Toupet ad020d3803 feat(cli): improve UX with auto-format detection, FPS tracking, and consistent messaging with ComfyUI implementation
- Auto-detect output format per file type (mp4 for videos, png for images)
- Add visual separators between processed files for better readability
- Simplify FPS calculation to use wall-clock time for real-world throughput
- Consolidate banner/footer into shared Debug methods
- Update offload device args to support multi-GPU (cpu/cuda:N)
- Standardize terminology: 'upscaling' instead of 'video upscaling'
- Remove code duplication between CLI and ComfyUI implementations
- Consistent quote style (double quotes) throughout CLI
2025-11-04 11:54:57 -05:00
Adrien Toupet ad1d775eaa fix: tile debug overlay now applies to all batches and adds overlay warning 2025-11-04 09:06:13 -05:00
Adrien Toupet 9268346388 feat: CLI Add batch processing, fix multiprocessing issues, and unify model paths
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)

Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable

Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability

Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
2025-11-04 00:12:13 -05:00
Adrien Toupet ce8225fd48 feat: Add max_resolution parameter to limit output dimensions
- Add max_resolution parameter (default: 0 = no limit) to both CLI and ComfyUI
- After new_resolution scales shortest edge, max_resolution ensures no edge exceeds limit
- Scales down proportionally if constraint violated
- Maintains backward compatibility with default value of 0
2025-11-03 20:17:32 -05:00
Adrien Toupet a900af96fd optimize: Implement adaptive GPU cache clearing for BlockSwap efficiency
BlockSwap now uses minimal, pressure-based cache clearing instead of forced deep cleans:
- Clear only when VRAM < 5% free (was 15%, more aggressive threshold)
- Use minimal GPU cache clear (deep=False) for reduced overhead
- Applied after each block/IO swap when pressure detected

Also fixes peft dependency: Add explicit peft>=0.17.0 in requirements for diffusers>=0.33.1 compatibility in ComfyUI ecosystem where other plugins may install outdated peft versions.
2025-11-03 17:29:58 -05:00
Adrien Toupet baf0058924 Add VAE tiled tensor offload + memory fixes
- VAE: Add tensor_offload_device support for tiled encode/decode accumulation buffers to reduce VRAM usage
- BlockSwap: Re-enable adaptive memory clearing under pressure to prevent OOM at high resolutions
- Memory: Fix offload_target device type error when set to 'none'
- Interface: Rename 'pixels' input to 'image' for ComfyUI convention alignment
- Interface: Fix singular/plural frame text display
2025-11-03 13:31:41 -05:00
Adrien Toupet 77cb6ff684 refactor(cli): inference_cli to match ComfyUI integration
- Add dit_offload_device parameter for proper blockSwap configuration
- Ensure consistent dtype management throughout CLI and ComfyUI (float32 input with bfloat16 pipeline)
- Translate all French comments to English
- Add comprehensive docstrings and section headers
- Remove obsolete use_non_blocking and enable_debug parameters
- Add error handling and validation
2025-10-28 01:01:38 -04:00
Adrien Toupet 5772255aed refactor: split generation.py and model_manager.py into 4 focused modules
- generation.py → generation_phases.py (4-phase pipeline logic) + generation_utils.py (setup/helpers)
- model_manager.py → model_configuration.py (config/caching) + model_loader.py (weight loading/GGUF)
- Renamed functions and code cleanup
2025-10-28 00:04:33 -04:00
Adrien Toupet e8376ddd6d refactor: split generation.py and model_manager.py into 4 focused modules
- generation.py → generation_phases.py (4-phase pipeline logic) + generation_utils.py (setup/helpers)
- model_manager.py → model_configuration.py (config/caching) + model_loader.py (weight loading/GGUF)
- Renamed functions and code cleanup
2025-10-28 00:04:08 -04:00
Adrien Toupet f182de79fa feat: Add temporal_overlap & prepend_frames to ComfyUI with shared logic
- Implement prepend_video_frames() for artifact reduction at video start
- Add blend_overlapping_frames() with Hann window for smooth transitions
- Expose temporal_overlap (0-16) and prepend_frames (0-32) in ComfyUI node
- Unify CLI and ComfyUI to use shared prepend/overlap functions
- Add comprehensive logging for frame adjustments (prepend/overlap/padding)
2025-10-27 21:57:38 -04:00
Adrien Toupet 3a4a4900df feat: optimize SeedVR2 memory management and color ops
- Optimize color operations for VRAM efficiency with safe precision handling
- Limit clear_memory calls to OOM retry and final cleanup only
- Fix model cleanup regression ensuring proper GPU memory release
2025-10-27 14:58:20 -04:00
Adrien Toupet 003122ebcd feat: implement lossless arbitrary resolution with padding - replace DivisibleCrop with DivisiblePad to eliminate data loss, track true dimensions for post-processing trim, change default resolution to 1080p with step=2 for flexibility 2025-10-27 01:14:57 -04:00
Adrien Toupet b0f03e54e5 fix: make FSDP imports conditional for AMD ROCm compatibility
Resolves ModuleNotFoundError on PyTorch ROCm 7 nightly builds by:
- Adding conditional imports for torch.distributed.fsdp in advanced.py
- Removing unused _is_fsdp_flattened import from meta_init_utils.py
- Removing unused log_runtime decorator from decorators.py

FSDP (Fully Sharded Data Parallel) features are only needed for multi-GPU
training and are not available in some PyTorch builds including AMD ROCm 7.
This change makes FSDP optional while preserving all inference functionality.
2025-10-26 23:49:52 -04:00
Adrien Toupet 0198834299 Remove cfg_scale parameter (incompatible with distilled one-step model)
CFG (Classifier-Free Guidance) does not work with SeedVR2's distilled
one-step diffusion model. The model was trained to produce final results
in a single step without iterative guidance.

- Remove cfg_scale parameter from ComfyUI node and CLI interface
- Force internal cfg_scale to 1.0 in upscale_all_batches()

This avoids artifacts introduced when users changed cfg_scale away from 1.0.
2025-10-26 22:54:45 -04:00
Adrien Toupet 7f56cdb272 fix: revert DiT dtype conversion and optimize VRAM usage
- Remove forced bfloat16 conversion for DiT, use native model dtype
- Fixed materialization logic to use offload_device when configured
- Fix BlockSwap VRAM leak by moving IO components to offload during BlockSwap cleanup

These changes restore performance and fix memory regressions
introduced in previous commits.
2025-10-26 22:11:34 -04:00
Adrien Toupet 70087c0b93 Improve VAE encoding stability and add tile debugging
- Switch to deterministic VAE encoding (mode vs sample) to eliminate high-resolution noise artifacts
- Make VAE encode tiling optional (disabled by default) since deterministic encoding resolves artifacts
- Add tile debug visualization feature with adaptive scaling and color-coded boundaries
- Remove redundant dtype conversions in VAE code for better performance
- Minor code cleanup and documentation update
2025-10-25 21:11:41 -04:00
Adrien Toupet 01cbdf8bc3 Optimize VAE defaults and standardize dtype pipeline for quality/performance
VAE Changes:
- Enable encode tiling by default (prevents noise artifacts at high resolution)
- Increase tile size to 1024px (down from 512px) for optimal quality
- Increase tile overlap to 128px for better blending

Dtype Pipeline:
- Hardcode compute_dtype to bfloat16 for consistent quality/performance/VRAM balance
- Ensure all pipeline steps are using compute_dtype when relevant
- Refactor code for improved performance and memory management
2025-10-22 23:56:14 -04:00
Adrien Toupet c9dce827c0 feat: Add deterministic generation with seed control and CFG scale parameter
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
  * VAE encoding: seed+1M for deterministic sampling without quality loss
  * DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id

Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)
2025-10-21 13:31:54 -04:00
Adrien Toupet e735c2ad56 feat: V3 migration with GGUF fixes and attention optimizations
Major Changes:
- Migrate all nodes to ComfyUI V3 schema (stateless design, new IO types)
- Fix GGUF weight caching VRAM leak (non-persistent buffers + _apply override)
- Fix GGUF torch.compile compatibility (@torch._dynamo.disable on dequant)
- Centralize compatibility checks (Flash/Triton/GGUF/Conv3d in compatibility.py)
- Make flash_attn optional with graceful SDPA fallback
- Add attention_mode UI option (sdpa/flash_attn) to DiT loader
- Fix color correction batch padding error (trim input_video consistently)

Code Quality:
- Remove internal_execute for clarity (stateless node design)
- Add startup logging for optimization status
- Improve error messages with installation instructions
- Add get_unique_id() for V3 node-specific caching
- Standardize parameter names (dit_cache/vae_cache)
2025-10-21 00:10:37 -04:00
Adrien Toupet 22e2854531 fix: improve progress bar accuracy based on typical phase timings (encode 20%, upscale 25%, decode 50%, post 5%) 2025-10-19 23:54:22 -04:00
Adrien Toupet bb3dd1013b Fix: Add Conv3d memory workaround for PyTorch 2.9+ cuDNN bug
Implements workaround for NVIDIA Conv3d 3x memory bug in PyTorch 2.9-2.10
with cuDNN >= 91002. Uses direct cuDNN calls to bypass buggy dispatch layer.
Reference: https://github.com/comfyanonymous/ComfyUI/pull/10373
2025-10-19 23:18:51 -04:00
Adrien Toupet 9893a1d7b0 refactor: Add indent_level parameter to debug logging system
- Add indent_level param to debug.log() and manage_tensor()
- Replace all hardcoded spaces with indent_level (0/1/2)
- Fix encode_all_batches memory flow: encode before storage
- Minor whitespace cleanup
2025-10-19 10:13:17 -04:00
Adrien Toupet a8d7153bc3 Refactor: unify tensor management and enforce compute_dtype throughout pipeline
- Rename manage_tensor_device -> manage_tensor with unified device/dtype handling
- Convert VAE outputs (float16) to compute_dtype (bfloat16) immediately after encode/decode
- Align alpha channel to compute_dtype at RGBA concatenation point
- Maintain float32 precision for alpha processing numerical stability
- Optimize conversions during offload operations to minimize overhead
- speed/VRAM improvements through reduced dtype conversions and better consistency
2025-10-19 08:41:45 -04:00
Adrien Toupet 97c7b9cd12 refactor: Overhaul device and memory management architecture
Major improvements to device handling, model offloading, and code clarity:

Device Management:
- Replace string-based devices with torch.device objects throughout
- Add explicit offload device parameters: dit_offload_device, vae_offload_device, tensor_offload_device
- Remove preserve_vram in favor of explicit offload control
- Improve get_device() to return torch.device objects consistently
- Enhance get_device_list() with smart MPS-only system handling

Context & Pipeline:
- Merge setup_device_environment and prepare_generation_context into single setup_generation_context
- Simplify LOCAL_RANK handling (set to '0' for single-GPU mode)
- Store device configuration on runner for submodule access

Tensor & Model Management:
- Add manage_tensor_device() for consistent tensor movement with logging
- Update manage_model_device() to use torch.device objects
- Add validation for BlockSwap and caching configurations
- Rename cache_in_ram to cache_model (more accurate naming)

CLI & Interface:
- Remove --preserve_vram flag (breaking change)
- Add --vae_offload_device and --tensor_offload_device flags
- Update ComfyUI node parameters with validation and better tooltips
- Improve device selection UI with offload_device options
2025-10-17 17:24:53 -04:00
Adrien Toupet d37573167c feat: optimize memory usage with CPU-based streaming in decode and postprocess phases
- Phase 3: Move decoded samples to CPU immediately after decoding to prevent VRAM accumulation across batches
- Phase 4: Load samples back to GPU at processing start for fast color correction (consistent with other phases)
- Implement streaming architecture: pre-allocate final_video tensor and write directly without accumulation
- Add progressive memory release: free batch_samples and transformed_videos as they're processed
- Store total_frames in context early for accurate pre-allocation
- Ensure consistent VRAM usage regardless of batch count and eliminates memory spikes for long videos
2025-10-16 16:55:41 -04:00
Adrien Toupet a46e16d8cc refactor: 42, centralize float32 precision utilities and improve logging
- Move ensure_float32_precision() from alpha_upscaling to common/half_precision_fixes
- Update color_fix.py to use ensure_float32_precision
- Add type hints and docstrings to all half_precision_fixes functions
- Clarify batch size tips: emphasize avoiding padding waste over maximizing batch size
- Update default seed to 42 (the answer to life, the universe, and everything)
2025-10-16 14:30:47 -04:00
Adrien Toupet 97b4de0229 perf: optimize dtype pipeline for better quality and performance
- Remove lossy bfloat16→float16 conversions throughout pipeline
- Eliminate unnecessary dtype conversions per batch (euler sampler, decode, postprocess)
- Keep native bfloat16 precision until final float32 output for ComfyUI compatibility
- Minor VRAM and speed savings + improved output quality
2025-10-15 21:21:41 -04:00
Adrien Toupet 155e9e9ef0 fix: add matplotlib dependency for diffuser compatibility
- Resolves diffusers import error (matplotlib.spec not set)
- Remove redundant __version__, __author__, __description__ from __init__.py (already in pyproject.toml)
2025-10-15 15:12:58 -04:00
Adrien Toupet b83a1c30b0 Add pyproject.toml and publish.yml from main 2025-10-15 14:10:17 -04:00
Adrien Toupet f1ac7d786e Add perceptually-accurate color correction methods
- Add LAB color transfer as new default method for superior perceptual color matching
- Add HSV hue-conditional saturation histogram matching for targeted oversaturation correction
- Add wavelet_adaptive hybrid method combining wavelet base with selective HSV correction
- Change default color correction from 'wavelet' to 'lab' for better color accuracy
- Implement full CIELAB color space conversion with D65 illuminant and histogram matching
- Add comprehensive documentation and method descriptions to color_fix.py
2025-10-15 13:14:48 -04:00
Adrien Toupet e48c2da776 feat: global model cache, alpha upscaling rewrite, and quality improvements
Major refactor introducing:

- GlobalModelCache for cross-node model sharing with dynamic config updates
  * Models cached by node ID, reused across different upscaler instances
  * Config changes (torch.compile, BlockSwap, tiling) handled dynamically
  * Runner templates cached when both models present

- Complete alpha channel processing rewrite
  * New edge-guided alpha upscaling in src/core/alpha_upscaling.py
  * Remove broken VAE RGBA adapter approach
  * Alpha properly extracted, upscaled, and merged with RGB

- BlockSwap enhancements
  * Support I/O-only swapping without transformer block offloading
  * Separate memory reporting for I/O components vs transformer blocks
  * Clearer logging of GPU vs CPU placement
  * New is_blockswap_enabled() utility

- Memory management improvements
  * Add release_tensor_collection() for batch tensor cleanup
  * Proper tensor memory release throughout pipeline
  * Better cleanup in postprocess phase

- Code quality and refactoring
  * Extensive docstrings with type hints across all modules
  * Generic _update_model_config() reduces code duplication
  * Better function signatures and parameter documentation
  * Improved separation of concerns in model_manager.py
  * Clearer debug logging categories and messages
2025-10-15 00:51:56 -04:00