Commit Graph
52 Commits
Author SHA1 Message Date
Adrien Toupet 4e9ce4710e Unify parameter names across CLI and ComfyUI interface
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
2025-11-06 22:51:20 -05:00
Adrien Toupet 20dab62dc3 feat: add uniform_batch_size for temporal consistency + unify padding logic
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
2025-11-06 14:13:22 -05:00
Adrien Toupet 806bb94df0 feat: unify and improve tooltip documentation across CLI and ComfyUI nodes
- Standardize tooltip format with multi-line descriptions and bullet points
- Add comprehensive output tooltips for all nodes (DiT, VAE, torch.compile, upscaler)
- Enhance node descriptions with detailed capability summaries
- Simplify CLI tile size arguments to single integers (converted internally to tuples)
- Remove OneOrTwoValues argparse class for cleaner implementation
- Fix encode_tiled tooltip (was incorrectly referencing decoding)
- Clarify color correction purpose (corrects upscaling color shifts)
- Add multi-GPU offloading information to all offload_device tooltips
- Improve torch.compile parameter descriptions with use cases
- Ensure CLI and ComfyUI tooltips are consistent in terminology and structure
2025-11-05 15:35:22 -05:00
Adrien Toupet 326489d94c fix CLI: move Debug import after CUDA allocator config to fix batch processing errors 2025-11-05 14:27:06 -05:00
Adrien Toupet 9b79254c39 refactor(cli): improvements and bug fixes + 3b-Q8_0.gguf support
- Fix validation cache location to respect --model_dir parameter
- Fix output path handling for directories without extensions
- Remove spurious directory creation in get_base_cache_dir
- Enhanced CLI help with usage examples and argument grouping
- Auto-display help when script invoked without arguments
- Correct type hints (device_id: str, debug: Debug)
- Remove redundant type conversions and makedirs calls
- Reorganize imports to module top for clarity
- Improved docstrings & tooltip
- Change default batch_size from 1 to 5 to match ComfyUI integration
- Add support for seedvr2_ema_3b-Q8_0.gguf model
2025-11-05 00:28:19 -05:00
Adrien Toupet 3725c1061d refactor: centralize dimension computation and logging for CLI/ComfyUI
- Add compute_generation_info() and log_generation_start() helpers
- Move prepend_frames logic from extraction to processing pipeline
- Eliminate code duplication between CLI and ComfyUI workflows
- Add consistent dimension/parameter logging for both interfaces
- Disable argparse prefix matching for safer CLI usage
2025-11-04 17:01:01 -05:00
Adrien Toupet 1691e657b4 fix(cli): add CUDA device validation before torch initialization
- Validate --cuda_device arguments early in pre-parsing phase
- Check device IDs exist and are within available GPU range
- Fail fast with clear error messages showing available devices
2025-11-04 15:52:05 -05:00
Adrien Toupet 32a049dfd9 feat: Add CLI model caching for multi-file processing and unify device handling
- Add --cache_dit and --cache_vae flags for efficient multi-file directory processing
- Refactor processing pipeline to eliminate duplication between worker and direct modes
- Implement platform-agnostic device management (CUDA/MPS/CPU)
- Unify parameter naming: res_w→resolution, max_res_w→max_resolution across codebase
- Add smart offload device defaults when caching enabled
- Improve validation and user feedback for cache + multi-GPU scenarios
2025-11-04 15:12:28 -05:00
Adrien Toupet ad020d3803 feat(cli): improve UX with auto-format detection, FPS tracking, and consistent messaging with ComfyUI implementation
- Auto-detect output format per file type (mp4 for videos, png for images)
- Add visual separators between processed files for better readability
- Simplify FPS calculation to use wall-clock time for real-world throughput
- Consolidate banner/footer into shared Debug methods
- Update offload device args to support multi-GPU (cpu/cuda:N)
- Standardize terminology: 'upscaling' instead of 'video upscaling'
- Remove code duplication between CLI and ComfyUI implementations
- Consistent quote style (double quotes) throughout CLI
2025-11-04 11:54:57 -05:00
Adrien Toupet 9268346388 feat: CLI Add batch processing, fix multiprocessing issues, and unify model paths
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)

Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable

Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability

Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
2025-11-04 00:12:13 -05:00
Adrien Toupet ce8225fd48 feat: Add max_resolution parameter to limit output dimensions
- Add max_resolution parameter (default: 0 = no limit) to both CLI and ComfyUI
- After new_resolution scales shortest edge, max_resolution ensures no edge exceeds limit
- Scales down proportionally if constraint violated
- Maintains backward compatibility with default value of 0
2025-11-03 20:17:32 -05:00
Adrien Toupet 77cb6ff684 refactor(cli): inference_cli to match ComfyUI integration
- Add dit_offload_device parameter for proper blockSwap configuration
- Ensure consistent dtype management throughout CLI and ComfyUI (float32 input with bfloat16 pipeline)
- Translate all French comments to English
- Add comprehensive docstrings and section headers
- Remove obsolete use_non_blocking and enable_debug parameters
- Add error handling and validation
2025-10-28 01:01:38 -04:00
Adrien Toupet e8376ddd6d refactor: split generation.py and model_manager.py into 4 focused modules
- generation.py → generation_phases.py (4-phase pipeline logic) + generation_utils.py (setup/helpers)
- model_manager.py → model_configuration.py (config/caching) + model_loader.py (weight loading/GGUF)
- Renamed functions and code cleanup
2025-10-28 00:04:08 -04:00
Adrien Toupet f182de79fa feat: Add temporal_overlap & prepend_frames to ComfyUI with shared logic
- Implement prepend_video_frames() for artifact reduction at video start
- Add blend_overlapping_frames() with Hann window for smooth transitions
- Expose temporal_overlap (0-16) and prepend_frames (0-32) in ComfyUI node
- Unify CLI and ComfyUI to use shared prepend/overlap functions
- Add comprehensive logging for frame adjustments (prepend/overlap/padding)
2025-10-27 21:57:38 -04:00
Adrien Toupet 003122ebcd feat: implement lossless arbitrary resolution with padding - replace DivisibleCrop with DivisiblePad to eliminate data loss, track true dimensions for post-processing trim, change default resolution to 1080p with step=2 for flexibility 2025-10-27 01:14:57 -04:00
Adrien Toupet 0198834299 Remove cfg_scale parameter (incompatible with distilled one-step model)
CFG (Classifier-Free Guidance) does not work with SeedVR2's distilled
one-step diffusion model. The model was trained to produce final results
in a single step without iterative guidance.

- Remove cfg_scale parameter from ComfyUI node and CLI interface
- Force internal cfg_scale to 1.0 in upscale_all_batches()

This avoids artifacts introduced when users changed cfg_scale away from 1.0.
2025-10-26 22:54:45 -04:00
Adrien Toupet 70087c0b93 Improve VAE encoding stability and add tile debugging
- Switch to deterministic VAE encoding (mode vs sample) to eliminate high-resolution noise artifacts
- Make VAE encode tiling optional (disabled by default) since deterministic encoding resolves artifacts
- Add tile debug visualization feature with adaptive scaling and color-coded boundaries
- Remove redundant dtype conversions in VAE code for better performance
- Minor code cleanup and documentation update
2025-10-25 21:11:41 -04:00
Adrien Toupet 01cbdf8bc3 Optimize VAE defaults and standardize dtype pipeline for quality/performance
VAE Changes:
- Enable encode tiling by default (prevents noise artifacts at high resolution)
- Increase tile size to 1024px (down from 512px) for optimal quality
- Increase tile overlap to 128px for better blending

Dtype Pipeline:
- Hardcode compute_dtype to bfloat16 for consistent quality/performance/VRAM balance
- Ensure all pipeline steps are using compute_dtype when relevant
- Refactor code for improved performance and memory management
2025-10-22 23:56:14 -04:00
Adrien Toupet c9dce827c0 feat: Add deterministic generation with seed control and CFG scale parameter
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
  * VAE encoding: seed+1M for deterministic sampling without quality loss
  * DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id

Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)
2025-10-21 13:31:54 -04:00
Adrien Toupet e735c2ad56 feat: V3 migration with GGUF fixes and attention optimizations
Major Changes:
- Migrate all nodes to ComfyUI V3 schema (stateless design, new IO types)
- Fix GGUF weight caching VRAM leak (non-persistent buffers + _apply override)
- Fix GGUF torch.compile compatibility (@torch._dynamo.disable on dequant)
- Centralize compatibility checks (Flash/Triton/GGUF/Conv3d in compatibility.py)
- Make flash_attn optional with graceful SDPA fallback
- Add attention_mode UI option (sdpa/flash_attn) to DiT loader
- Fix color correction batch padding error (trim input_video consistently)

Code Quality:
- Remove internal_execute for clarity (stateless node design)
- Add startup logging for optimization status
- Improve error messages with installation instructions
- Add get_unique_id() for V3 node-specific caching
- Standardize parameter names (dit_cache/vae_cache)
2025-10-21 00:10:37 -04:00
Adrien Toupet 9893a1d7b0 refactor: Add indent_level parameter to debug logging system
- Add indent_level param to debug.log() and manage_tensor()
- Replace all hardcoded spaces with indent_level (0/1/2)
- Fix encode_all_batches memory flow: encode before storage
- Minor whitespace cleanup
2025-10-19 10:13:17 -04:00
Adrien Toupet 97c7b9cd12 refactor: Overhaul device and memory management architecture
Major improvements to device handling, model offloading, and code clarity:

Device Management:
- Replace string-based devices with torch.device objects throughout
- Add explicit offload device parameters: dit_offload_device, vae_offload_device, tensor_offload_device
- Remove preserve_vram in favor of explicit offload control
- Improve get_device() to return torch.device objects consistently
- Enhance get_device_list() with smart MPS-only system handling

Context & Pipeline:
- Merge setup_device_environment and prepare_generation_context into single setup_generation_context
- Simplify LOCAL_RANK handling (set to '0' for single-GPU mode)
- Store device configuration on runner for submodule access

Tensor & Model Management:
- Add manage_tensor_device() for consistent tensor movement with logging
- Update manage_model_device() to use torch.device objects
- Add validation for BlockSwap and caching configurations
- Rename cache_in_ram to cache_model (more accurate naming)

CLI & Interface:
- Remove --preserve_vram flag (breaking change)
- Add --vae_offload_device and --tensor_offload_device flags
- Update ComfyUI node parameters with validation and better tooltips
- Improve device selection UI with offload_device options
2025-10-17 17:24:53 -04:00
Adrien Toupet a46e16d8cc refactor: 42, centralize float32 precision utilities and improve logging
- Move ensure_float32_precision() from alpha_upscaling to common/half_precision_fixes
- Update color_fix.py to use ensure_float32_precision
- Add type hints and docstrings to all half_precision_fixes functions
- Clarify batch size tips: emphasize avoiding padding waste over maximizing batch size
- Update default seed to 42 (the answer to life, the universe, and everything)
2025-10-16 14:30:47 -04:00
Adrien Toupet f1ac7d786e Add perceptually-accurate color correction methods
- Add LAB color transfer as new default method for superior perceptual color matching
- Add HSV hue-conditional saturation histogram matching for targeted oversaturation correction
- Add wavelet_adaptive hybrid method combining wavelet base with selective HSV correction
- Change default color correction from 'wavelet' to 'lab' for better color accuracy
- Implement full CIELAB color space conversion with D65 illuminant and histogram matching
- Add comprehensive documentation and method descriptions to color_fix.py
2025-10-15 13:14:48 -04:00
Adrien Toupet e48c2da776 feat: global model cache, alpha upscaling rewrite, and quality improvements
Major refactor introducing:

- GlobalModelCache for cross-node model sharing with dynamic config updates
  * Models cached by node ID, reused across different upscaler instances
  * Config changes (torch.compile, BlockSwap, tiling) handled dynamically
  * Runner templates cached when both models present

- Complete alpha channel processing rewrite
  * New edge-guided alpha upscaling in src/core/alpha_upscaling.py
  * Remove broken VAE RGBA adapter approach
  * Alpha properly extracted, upscaled, and merged with RGB

- BlockSwap enhancements
  * Support I/O-only swapping without transformer block offloading
  * Separate memory reporting for I/O components vs transformer blocks
  * Clearer logging of GPU vs CPU placement
  * New is_blockswap_enabled() utility

- Memory management improvements
  * Add release_tensor_collection() for batch tensor cleanup
  * Proper tensor memory release throughout pipeline
  * Better cleanup in postprocess phase

- Code quality and refactoring
  * Extensive docstrings with type hints across all modules
  * Generic _update_model_config() reduces code duplication
  * Better function signatures and parameter documentation
  * Improved separation of concerns in model_manager.py
  * Clearer debug logging categories and messages
2025-10-15 00:51:56 -04:00
Adrien Toupet fb2b6c7d4c Fix BlockSwap not applying on cached models and rename offload_io_components to swap_io_components
- Fixed critical bug where BlockSwap configuration changes were not applied to cached models
- Apply BlockSwap immediately in _handle_blockswap_config instead of deferring to materialization phase
- Renamed offload_io_components to swap_io_components across entire codebase for consistency
- Removed unused _pending_blockswap_config attribute
2025-10-10 17:48:12 -04:00
Adrien Toupet 9ee244d71b feat: implement torch.compile optimization with BlockSwap compatibility
Core Features:
- Add torch.compile support for DiT (20-40% speedup) and VAE (15-25% speedup)
- Ensure BlockSwap compatibility by applying BlockSwap before torch.compile
- Add SeedVR2TorchCompileSettings node for ComfyUI configuration
- Add CLI arguments: --compile_dit, --compile_vae, --compile_backend, --compile_mode, --compile_fullgraph, --compile_dynamic, --compile_dynamo_cache_size_limit, --compile_dynamo_recompile_limit

Optimizations:
- Optimize na.py and other backend files for torch.compile compatibility (replace .tolist() with tensor operations)
- Remove .item() calls from attention modules (max_seqlen_q/k now accept tensors)
- Add @torch._dynamo.disable decorators to timing/debug methods to prevent compilation warnings
- Replace torch.repeat with torch.repeat_interleave in modulation.py for better performance

Bug Fixes:
- Fix timing stack cleanup in memory_manager.py (end timer before early returns)
- Fix BlockSwap timing reporting (_get_swap_start_time and _log_swap_timing now properly excluded from compilation)

Code Quality:
- Add comprehensive type hints and docstrings across all modified modules
- Standardize BlockSwap timer names (blockswap_block_*, blockswap_io_*)
2025-10-10 12:34:49 -04:00
Adrien Toupet 5567f232af feat!: independent DiT/VAE management with separate loader nodes
BREAKING CHANGE: Replaced BlockSwap and ExtraArgs nodes with DiT/VAE loader nodes.
Enables multi-GPU placement, independent caching, and granular memory control.

New Features:
- Add SeedVR2LoadDiTModel and SeedVR2LoadVAEModel loader nodes
- Support different devices for DiT and VAE (multi-GPU load balancing)
- Independent model caching (cache_model_dit/cache_model_vae)
- Separate encode/decode VAE tiling with independent tile configurations
- User-defined DiT and VAE model selection from registry or disk

Breaking Changes:
- Deprecated SeedVR2BlockSwap node (integrated into DiT loader)
- Deprecated SeedVR2ExtraArgs node (split into DiT/VAE loaders & Upscaler node)
- BlockSwap configuration now part of DiT loader node

Improvements:
- Follow ComfyUI conventions (pixels input, IMAGE output type)
- Update download_weight() to support user-defined DiT/VAE models
- Improve tooltips and default parameters across all nodes
- Set show_tensors=False in debug logging for performance
- Dual-device context tracking throughout pipeline
- CLI updated to match new dual-model API
2025-10-08 00:49:48 -04:00
Adrien Toupet 2fab3a1caf refactor: split decode and post-processing into separate phases for cleaner architecture
- Split decode_all_batches into decode (Phase 3) and postprocess_all_batches (Phase 4)
- Add phase-specific cleanup functions: cleanup_dit() and cleanup_vae()
- Each phase now handles its own resource cleanup in finally blocks
- Add cleanup_text_embeddings() helper to eliminate code duplication
- Remove redundant cleanup from comfyui_node normal flow
- Update pipeline from 3-phase to 4-phase architecture (encode → upscale → decode → postprocess)
- Improve memory efficiency by releasing resources immediately when no longer needed
- Update module docstrings to reflect new architecture
2025-10-03 09:53:13 -04:00
Adrien Toupet 4a3efc7883 feat: add CPU offloading for intermediate data to reduce VRAM usage
- Store latents and transformed videos on CPU between processing phases reducing VRAM usage to enable larger batch processing
- Make transformed video storage conditional (only when color_correction != none)
- Add all_ori_lengths tracking for consistent trimming
- Use non_blocking=False for CPU-GPU transfers to avoid pinned memory issues
Fixes https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/pull/163#issuecomment-3333686784
2025-09-25 11:30:32 -04:00
Adrien Toupet 9b4a7dfa3e fix: unify model registry for CLI and improve GGUF implementation
Model Registry Integration:
- Replace hardcoded CLI model choices with dynamic get_available_models()
- Use centralized DEFAULT_MODEL constant from model_registry
- Achieve single source of truth between CLI and ComfyUI interfaces

GGUF Optimizations:
- Create tensors directly on target device to avoid CPU->GPU copy overhead
- Remove excessive memory cleanup during tensor loading for better performance
- Always use meta init for memory-efficient model creation
- Implement precision-optimized dequantization path (GGUF → FP16 → compute dtype)

GGUF Precision Handling:
- Add GGUFQuantizedLinear/Conv2d layers with get_dequantized_weight_for_compute() to preserve precision
- Track and report quantization types (Q4_K_M, Q5_K_M, etc.) in model
- Fix __torch_function__ as classmethod to resolve deprecation warning

Code Quality:
- Restructure _load_model_weights() with modular helper functions to reduce duplication
- Improve separation of concerns between standard and GGUF weight loading
- Enhance logging to always display WARNING/ERROR messages
- Add comprehensive GGUF architecture validation
- Remove debug-only log_memory_state calls
- Fix type function definitions
2025-09-25 00:09:56 -04:00
Adrien Toupet 0b0c87ed4a Add GGUF quantized model support (based on PR #121 from @cmeka / @lihaoyun6)
- Implement GGUF model loading with Q3_K_M through Q8_K_M quantization support
- Add GGUFTensor wrapper to preserve quantization and enable on-demand dequantization
- Maintain tensors in quantized format to reduce VRAM usage
- Add GGUF dequantization operations for inference
- Update model registry to include GGUF variants for 3B/7B models
- Fix wavelet blur radius limit to prevent OOM at high resolutions (max 1/8 of image dimension)
- Add safety clamp [-1,1] for SDR color range to prevent numerical errors
- This is a WIP commit as some additional cleaning/testing is needed
- Add type hints throughout for better code maintainability
2025-09-24 11:49:57 -04:00
Adrien Toupet 22f540a7d6 feat: Add input noise parameter and rename latent noise for clarity
- Add input_noise_scale (0-1) to reduce artifacts at high resolutions
  - Applies subtle noise before VAE encoding with progressive blend (0-50%)
  - Based on GitHub issue #64 community findings

- Rename cond_noise_scale to latent_noise_scale for clarity
  - Better distinguishes between input (pixel) and latent (diffusion) noise
  - Maintains same functionality, just clearer naming

- Update both ComfyUI and CLI interfaces with new parameters
2025-09-22 11:47:49 -04:00
Adrien Toupet e89a8b00a5 feat: promote cond_noise_scale to UI for controlling conditioning noise level helping mitigate noise artifacts in high-res upscaling 2025-09-18 23:19:35 -04:00
Adrien Toupet ee95a3a673 feat: Add configurable color correction methods and improve memory management
- Add color_correction parameter (wavelet/adain/none) to both ComfyUI node and CLI interface to choose between wavelet reconstruction (frequency-based), AdaIN (statistical matching), or no color correction
- Improve memory management after upscaling batches with targeted clear_memory calls to reduce VRAM pressure during VAE decoding
2025-09-16 22:21:45 -04:00
Adrien Toupet 1518ecc1a5 refactor: Fix ComfyUI node conflicts via relative imports and clearer model structure
- Renamed model directories for clarity: dit -> dit_7b, dit_v2 -> dit_3b
- Converted all absolute imports to relative imports throughout codebase
- Removed sys.path.append() manipulations that caused namespace conflicts
- Updated YAML configs to reference renamed model directories
- Simplified model variant detection logic using new directory names
- Standardized function calls with named arguments for better clarity
- Fixed generation context initialization and interrupt handling

This resolves import conflicts with other ComfyUI nodes (e.g., Basic data handling)
that use sys.path manipulation, making the module properly isolated and compatible.

Fixes #29, #114, #136
2025-09-16 16:45:31 -04:00
Adrien Toupet 477f57fd5a Refactor: Three-phase batch processing pipeline for improved performance
Major architectural change to minimize model swapping overhead by processing
all batches in three distinct phases instead of sequential per-batch processing:
- Phase 1: Encode all batches with VAE
- Phase 2: Upscale all latents with DiT
- Phase 3: Decode all latents with VAE

Core changes:
- Split monolithic generation_loop into modular functions:
  - prepare_generation_context(): Shared state management
  - setup_device_environment(): Device configuration
  - prepare_runner(): Model loading with cache support
  - encode_all_batches(): Batch VAE encoding
  - upscale_all_batches(): Batch DiT upscaling
  - decode_all_batches(): Batch VAE decoding
- Removed generation_step function (logic integrated into upscale phase)
- Added lazy precision initialization to avoid redundant setup

Performance improvements:
- Pre-allocated lists for memory efficiency
- Better cleanup of intermediate storage between phases
- Added unique timer names to clear_memory() to avoid naming conflicts
- Improved model state management with change detection and caching

UI/UX enhancements:
- Switched to ComfyUI's native ProgressBar with weighted phase progress
- Changed from per-batch FPS to overall average FPS (always visible)
- Improved log clarity with clear phase separators
- Added ASCII art logo to clearly identify SeedVR2 process start
- Better progress tracking with weighted percentages across three phases

Code cleanup:
- Removed deprecated timer_context from Debug class
- Removed unused time imports across multiple files
- Fixed LOCAL_RANK environment variable to handle string conversion properly
- Improved error handling with try/except/finally blocks in all phases
2025-09-16 14:38:57 -04:00
Adrien Toupet 5595d58597 refactor: streamline generation pipeline and improve dtype handling
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
2025-08-25 17:11:29 -04:00
Adrien Toupet 2053a80f38 refactor(WIP): Improve debug logging consistency and reduce redundancy
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state()  formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
2025-08-22 15:38:00 -04:00
Benjamin Herb 22aa85abb7 fix: Updates --prepend_frames logic 2025-08-21 08:12:07 +02:00
lihaoyun6 3b0699ead9 Merge remote-tracking branch 'upstream/nightly' 2025-08-12 22:28:36 +08:00
Benjamin Herb cd2671bbe1 fix: fixed tiling issue and added smoother blending
- Fixed issue that would lead to a white border around the image due to
  improper tile blending
- Added smoother blending (hann ramp) for the tiles and temporal
  overlaps
2025-08-10 23:54:17 +02:00
Benjamin Herb 126ebfc10e feat: add separate height/width control to vae tile size and overlap
- use h/w tuples to specify tiling size and overlap to allow for finer
    control
  - added OneOrTwoValues class to handle CLI input of single values or h/w
    pairs
2025-08-10 21:27:54 +02:00
Benjamin Herb afafa67dd8 fix: VAE tiling cleanup and fixes
- Fix issues with unnecessary VAE blocks at the boundarys
- Updated VAE encode/decode logic to more closely match the original flow
- Better skip_frame log info
2025-08-09 19:23:25 +02:00
Benjamin Herb 7737efd22a feat: Add temporal overlap with blending and prepend frame options to CLI
- Add temporal overlap with crossfading (works with multi GPU)
- Add option to prepend frames to avoid artifacts at the start
2025-08-09 04:12:37 +02:00
Benjamin Herb ad62a12a3e feat: Add VAE tiling support
- Add VAE encode/decode tiling based on @Luke2642 approach
- Fix issues when tiling sequences
- Add CLI arguments for VAE tiling configuration
2025-08-09 02:30:28 +02:00
Benjamin Herb fe67b2a042 feat: Add BlockSwap support for standalone CLI and update logging
- Add BlockSwap CLI arguments and activation
- Fix imports to work without ComfyUI dependencies
- Update to new debugging framework
2025-08-09 02:27:42 +02:00
lihaoyun6 3b530dc983 Added MPS backend support (for running on macOS) 2025-08-06 18:52:33 +08:00
luchuanzhao 59850f6e40 Fix multi GPUs OOM 2025-07-11 11:29:13 +08:00
NumZ f3c704a502 Fix linux 2025-07-03 14:41:49 +02:00