Commit Graph
62 Commits
Author SHA1 Message Date
Adrien Toupet c010deeea1 Remove ineffective allow_vram_overflow setting
- PyTorch's set_per_process_memory_fraction cannot prevent WDDM paging on Windows
- Keep overflow detection and warning when VRAM exceeds physical limit
- Simplify peak memory formatting
- Remove setting from CLI, ComfyUI node, and memory_manager
2025-12-10 00:56:27 -05:00
Adrien Toupet 7cbf025561 Fix VRAM peak tracking: separate allocated vs reserved, Windows-only overflow
- Track both peak_allocated (tensor usage) and peak_reserved (cache pool) per phase
- peak_allocated resets properly between phases via reset_peak_memory_stats()
- Overflow detection/warnings now Windows-only (WDDM paging behavior)
- Remove get_memory_architecture() - replaced with simple is_mps + platform checks
- Phase summary shows: VRAM XGB allocated, YGB reserved | RAM ZGB
- Simplify MPS path (unified memory has no overflow concept)
2025-12-09 23:51:51 -05:00
Adrien Toupet 5c60716c47 Refactor: centralize backend detection, fix architecture-aware VRAM overflow reporting 2025-12-09 21:06:12 -05:00
Adrien Toupet 77a00f651a Fix: OOM regression from 2.5.14 strict VRAM limit (#367)
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.

The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.

- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow

Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
2025-12-09 17:12:10 -05:00
Adrien Toupet 71ac9ffe54 fix: use max_memory_reserved for accurate VRAM peak tracking 2025-12-03 11:51:30 -05:00
Adrien Toupet 5775ff0f99 Enforce VRAM limit to physical capacity - OOM instead of silent swap 2025-12-01 00:25:26 -05:00
Adrien Toupet 5848cef05f fix(mps): normalize device strings to prevent unnecessary tensor movements
- Add _device_str() helper to normalize MPS variants (mps:0 → MPS)
- Fix device comparison: mps:0 and mps now correctly identified as same device
- Consistent MPS logging across all memory management functions
2025-11-30 21:03:26 -05:00
Adrien Toupet be2efd474f fix: use canonical torch.backends.mps.is_available() for reliable MPS detection
Replaces torch.mps.is_available() with torch.backends.mps.is_available()
across all files. The torch.backends API is the official PyTorch method
for MPS detection (since PyTorch 1.12) and works reliably on:
- macOS with Apple Silicon (returns True when MPS available)
- macOS without MPS support (returns False)
- Windows/Linux (hasattr guard prevents AttributeError)
2025-11-28 11:49:43 -05:00
Adrien Toupet 79e7f41216 Revert "Fix: MPS allocator error in model.to() transfer (#305)"
This reverts commit 9424020687.
2025-11-27 14:52:12 -05:00
Adrien Toupet 9424020687 Fix: MPS allocator error in model.to() transfer (#305)
Add fallback to individual parameter movement when bulk model.to(mps)
fails with allocator errors. Complements safetensors loading fix.
2025-11-14 12:06:13 -05:00
Adrien Toupet 65dd29a865 v2.5.10: Fix determinism, BlockSwap caching, and model path resolution
Core Fixes:
- Reset seed per batch to ensure deterministic generation across sessions and batch positions
- Fix temporal overlap logging when automatically reset to prevent incorrect frame counts
- Fix NoneType attribute error in VAE tiled encode/decode at maximum resolution (#296)

BlockSwap & Caching Architecture:
- Move BlockSwap state (_block_swap_config, _blockswap_bypass_protection) from runner to model
- Ensures state survives runner recreation during independent DiT/VAE caching scenarios
- Fix runner template caching to trigger when either DiT or VAE becomes cached (bidirectional)
- Resolves BlockSwap reload failures when only DiT was cached (#297)

Model Discovery:
- Implement case-insensitive YAML path resolution for extra_model_paths.yaml (#289-#295)
- Add debug logging for model discovery (searched paths, validation status, cache hits)
- Support any case variation (seedvr2, SEEDVR2, SeedVR2) in ComfyUI configuration
2025-11-13 12:02:37 -05:00
Adrien Toupet b5c40fea9c v2.5.5: Fix RAM leak for long videos via on-demand reconstruction
- Replace all_transformed_videos storage with lightweight batch_metadata indices
- Reconstruct transformed videos on-demand in Phase 4 only when needed
- Add missing cleanup for input_images tensor in postprocess finally block
- Fix release_tensor_memory to handle CPU/CUDA/MPS consistently
- Extract helper functions for batch preparation and 4n+1 padding
- Remove duplicate interrupt_fn key from context initialization
2025-11-09 02:06:16 -05:00
Adrien Toupet 786fb3f688 fix: correct MPS device enumeration for Apple Silicon (v2.5.3) 2025-11-08 08:25:22 -05:00
Adrien Toupet 9715d3e37a Fix: torch.mps AttributeError on Windows
Add defensive checks for torch.mps.is_available() to handle PyTorch versions where the method doesn't exist on non-Mac platforms. Resolves AttributeError: module 'torch.mps' has no attribute 'is_available'
2025-11-07 23:37:35 -05:00
Adrien Toupet a900af96fd optimize: Implement adaptive GPU cache clearing for BlockSwap efficiency
BlockSwap now uses minimal, pressure-based cache clearing instead of forced deep cleans:
- Clear only when VRAM < 5% free (was 15%, more aggressive threshold)
- Use minimal GPU cache clear (deep=False) for reduced overhead
- Applied after each block/IO swap when pressure detected

Also fixes peft dependency: Add explicit peft>=0.17.0 in requirements for diffusers>=0.33.1 compatibility in ComfyUI ecosystem where other plugins may install outdated peft versions.
2025-11-03 17:29:58 -05:00
Adrien Toupet baf0058924 Add VAE tiled tensor offload + memory fixes
- VAE: Add tensor_offload_device support for tiled encode/decode accumulation buffers to reduce VRAM usage
- BlockSwap: Re-enable adaptive memory clearing under pressure to prevent OOM at high resolutions
- Memory: Fix offload_target device type error when set to 'none'
- Interface: Rename 'pixels' input to 'image' for ComfyUI convention alignment
- Interface: Fix singular/plural frame text display
2025-11-03 13:31:41 -05:00
Adrien Toupet 3a4a4900df feat: optimize SeedVR2 memory management and color ops
- Optimize color operations for VRAM efficiency with safe precision handling
- Limit clear_memory calls to OOM retry and final cleanup only
- Fix model cleanup regression ensuring proper GPU memory release
2025-10-27 14:58:20 -04:00
Adrien Toupet 70087c0b93 Improve VAE encoding stability and add tile debugging
- Switch to deterministic VAE encoding (mode vs sample) to eliminate high-resolution noise artifacts
- Make VAE encode tiling optional (disabled by default) since deterministic encoding resolves artifacts
- Add tile debug visualization feature with adaptive scaling and color-coded boundaries
- Remove redundant dtype conversions in VAE code for better performance
- Minor code cleanup and documentation update
2025-10-25 21:11:41 -04:00
Adrien Toupet c9dce827c0 feat: Add deterministic generation with seed control and CFG scale parameter
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
  * VAE encoding: seed+1M for deterministic sampling without quality loss
  * DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id

Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)
2025-10-21 13:31:54 -04:00
Adrien Toupet e735c2ad56 feat: V3 migration with GGUF fixes and attention optimizations
Major Changes:
- Migrate all nodes to ComfyUI V3 schema (stateless design, new IO types)
- Fix GGUF weight caching VRAM leak (non-persistent buffers + _apply override)
- Fix GGUF torch.compile compatibility (@torch._dynamo.disable on dequant)
- Centralize compatibility checks (Flash/Triton/GGUF/Conv3d in compatibility.py)
- Make flash_attn optional with graceful SDPA fallback
- Add attention_mode UI option (sdpa/flash_attn) to DiT loader
- Fix color correction batch padding error (trim input_video consistently)

Code Quality:
- Remove internal_execute for clarity (stateless node design)
- Add startup logging for optimization status
- Improve error messages with installation instructions
- Add get_unique_id() for V3 node-specific caching
- Standardize parameter names (dit_cache/vae_cache)
2025-10-21 00:10:37 -04:00
Adrien Toupet 9893a1d7b0 refactor: Add indent_level parameter to debug logging system
- Add indent_level param to debug.log() and manage_tensor()
- Replace all hardcoded spaces with indent_level (0/1/2)
- Fix encode_all_batches memory flow: encode before storage
- Minor whitespace cleanup
2025-10-19 10:13:17 -04:00
Adrien Toupet a8d7153bc3 Refactor: unify tensor management and enforce compute_dtype throughout pipeline
- Rename manage_tensor_device -> manage_tensor with unified device/dtype handling
- Convert VAE outputs (float16) to compute_dtype (bfloat16) immediately after encode/decode
- Align alpha channel to compute_dtype at RGBA concatenation point
- Maintain float32 precision for alpha processing numerical stability
- Optimize conversions during offload operations to minimize overhead
- speed/VRAM improvements through reduced dtype conversions and better consistency
2025-10-19 08:41:45 -04:00
Adrien Toupet 97c7b9cd12 refactor: Overhaul device and memory management architecture
Major improvements to device handling, model offloading, and code clarity:

Device Management:
- Replace string-based devices with torch.device objects throughout
- Add explicit offload device parameters: dit_offload_device, vae_offload_device, tensor_offload_device
- Remove preserve_vram in favor of explicit offload control
- Improve get_device() to return torch.device objects consistently
- Enhance get_device_list() with smart MPS-only system handling

Context & Pipeline:
- Merge setup_device_environment and prepare_generation_context into single setup_generation_context
- Simplify LOCAL_RANK handling (set to '0' for single-GPU mode)
- Store device configuration on runner for submodule access

Tensor & Model Management:
- Add manage_tensor_device() for consistent tensor movement with logging
- Update manage_model_device() to use torch.device objects
- Add validation for BlockSwap and caching configurations
- Rename cache_in_ram to cache_model (more accurate naming)

CLI & Interface:
- Remove --preserve_vram flag (breaking change)
- Add --vae_offload_device and --tensor_offload_device flags
- Update ComfyUI node parameters with validation and better tooltips
- Improve device selection UI with offload_device options
2025-10-17 17:24:53 -04:00
Adrien Toupet e48c2da776 feat: global model cache, alpha upscaling rewrite, and quality improvements
Major refactor introducing:

- GlobalModelCache for cross-node model sharing with dynamic config updates
  * Models cached by node ID, reused across different upscaler instances
  * Config changes (torch.compile, BlockSwap, tiling) handled dynamically
  * Runner templates cached when both models present

- Complete alpha channel processing rewrite
  * New edge-guided alpha upscaling in src/core/alpha_upscaling.py
  * Remove broken VAE RGBA adapter approach
  * Alpha properly extracted, upscaled, and merged with RGB

- BlockSwap enhancements
  * Support I/O-only swapping without transformer block offloading
  * Separate memory reporting for I/O components vs transformer blocks
  * Clearer logging of GPU vs CPU placement
  * New is_blockswap_enabled() utility

- Memory management improvements
  * Add release_tensor_collection() for batch tensor cleanup
  * Proper tensor memory release throughout pipeline
  * Better cleanup in postprocess phase

- Code quality and refactoring
  * Extensive docstrings with type hints across all modules
  * Generic _update_model_config() reduces code duplication
  * Better function signatures and parameter documentation
  * Improved separation of concerns in model_manager.py
  * Clearer debug logging categories and messages
2025-10-15 00:51:56 -04:00
Adrien Toupet 1b00665757 feat: Add RGBA support, improve config caching, and enhance logging
Major features:
- RGBA/Alpha channel support throughout upscaling pipeline
  * New VAE RGBA adapter for dynamic 3→4 channel conversion
  * Detect RGBA input and adapt VAE after weight loading
  * Preserve alpha during encoding/decoding phases
  * Split alpha for RGB-only color correction, then recombine
  * ComfyUI node now accepts mask input and outputs mask
  * Frame count validation between pixels and mask

Configuration management:
- Smart in-place config updates without model reload
- Compare BlockSwap and torch.compile configs for changes
- Cleanup old materialized models before creating new meta structures
- Preserve config attributes (_dit_compile_args, _vae_compile_args, _dit_block_swap_config) for change detection
- Better config change logging with descriptive before/after states
- Handle model reloads separately from config-only updates

Logging improvements:
- Consistent 2-space indentation for hierarchical log messages
- More detailed configuration change descriptions
- Better visibility of RGBA processing steps
- Improved batch optimization tips formatting

Memory and cleanup:
- Handle meta device properly in cleanup functions
- Consistent attribute cleanup across dit/vae
- Clear compile/blockswap configs only when models not cached
- Improved model movement checks for meta device

VAE torch.compile enhancements:
- Disable compilation for InflatedCausalConv3d (dynamic shapes)
- Apply to both encoder and decoder to prevent recompilation issues
2025-10-11 02:26:59 -04:00
Adrien Toupet fb2b6c7d4c Fix BlockSwap not applying on cached models and rename offload_io_components to swap_io_components
- Fixed critical bug where BlockSwap configuration changes were not applied to cached models
- Apply BlockSwap immediately in _handle_blockswap_config instead of deferring to materialization phase
- Renamed offload_io_components to swap_io_components across entire codebase for consistency
- Removed unused _pending_blockswap_config attribute
2025-10-10 17:48:12 -04:00
Adrien Toupet 9ee244d71b feat: implement torch.compile optimization with BlockSwap compatibility
Core Features:
- Add torch.compile support for DiT (20-40% speedup) and VAE (15-25% speedup)
- Ensure BlockSwap compatibility by applying BlockSwap before torch.compile
- Add SeedVR2TorchCompileSettings node for ComfyUI configuration
- Add CLI arguments: --compile_dit, --compile_vae, --compile_backend, --compile_mode, --compile_fullgraph, --compile_dynamic, --compile_dynamo_cache_size_limit, --compile_dynamo_recompile_limit

Optimizations:
- Optimize na.py and other backend files for torch.compile compatibility (replace .tolist() with tensor operations)
- Remove .item() calls from attention modules (max_seqlen_q/k now accept tensors)
- Add @torch._dynamo.disable decorators to timing/debug methods to prevent compilation warnings
- Replace torch.repeat with torch.repeat_interleave in modulation.py for better performance

Bug Fixes:
- Fix timing stack cleanup in memory_manager.py (end timer before early returns)
- Fix BlockSwap timing reporting (_get_swap_start_time and _log_swap_timing now properly excluded from compilation)

Code Quality:
- Add comprehensive type hints and docstrings across all modified modules
- Standardize BlockSwap timer names (blockswap_block_*, blockswap_io_*)
2025-10-10 12:34:49 -04:00
Adrien Toupet 5567f232af feat!: independent DiT/VAE management with separate loader nodes
BREAKING CHANGE: Replaced BlockSwap and ExtraArgs nodes with DiT/VAE loader nodes.
Enables multi-GPU placement, independent caching, and granular memory control.

New Features:
- Add SeedVR2LoadDiTModel and SeedVR2LoadVAEModel loader nodes
- Support different devices for DiT and VAE (multi-GPU load balancing)
- Independent model caching (cache_model_dit/cache_model_vae)
- Separate encode/decode VAE tiling with independent tile configurations
- User-defined DiT and VAE model selection from registry or disk

Breaking Changes:
- Deprecated SeedVR2BlockSwap node (integrated into DiT loader)
- Deprecated SeedVR2ExtraArgs node (split into DiT/VAE loaders & Upscaler node)
- BlockSwap configuration now part of DiT loader node

Improvements:
- Follow ComfyUI conventions (pixels input, IMAGE output type)
- Update download_weight() to support user-defined DiT/VAE models
- Improve tooltips and default parameters across all nodes
- Set show_tensors=False in debug logging for performance
- Dual-device context tracking throughout pipeline
- CLI updated to match new dual-model API
2025-10-08 00:49:48 -04:00
Adrien Toupet 3a90290c03 feat: implement deferred model loading for better memory management
- VAE loads at encoding phase, DiT loads at upscaling phase
- BlockSwap config changes for cached models deferred to DiT phase
- Add comprehensive cleanup of temporary loading attributes
2025-10-03 17:29:51 -04:00
Adrien Toupet a9b6e4201c Fix BlockSwap GPU memory leak and restore performance metrics
- Fix critical issue where BlockSwap blocks remained on GPU when model caching enabled
- Restore BlockSwap summary display (regression fix)
- Improve model change detection to only cleanup DiT instead of full cleanup
- Standardize _model_name attribute and cleanup flow
- Add Phase 4 post-processing to timing breakdown
- Change DiT final memory clear to deep=True for better cleanup
2025-10-03 16:04:48 -04:00
Adrien Toupet 2fab3a1caf refactor: split decode and post-processing into separate phases for cleaner architecture
- Split decode_all_batches into decode (Phase 3) and postprocess_all_batches (Phase 4)
- Add phase-specific cleanup functions: cleanup_dit() and cleanup_vae()
- Each phase now handles its own resource cleanup in finally blocks
- Add cleanup_text_embeddings() helper to eliminate code duplication
- Remove redundant cleanup from comfyui_node normal flow
- Update pipeline from 3-phase to 4-phase architecture (encode → upscale → decode → postprocess)
- Improve memory efficiency by releasing resources immediately when no longer needed
- Update module docstrings to reflect new architecture
2025-10-03 09:53:13 -04:00
Adrien Toupet 1fa603899f Fix GGUF implementation: centralize warning suppression, improve type hints, enhance error handling
- Add suppress_tensor_warnings() to constants.py to centralize tensor/numpy warning handling
- Fix type hints and List/Tuple imports
- Improve GGUFTensor.__torch_function__ to prevent recursion, add debug null checks
- Fix GGUFTensor.to() to properly preserve tensor_shape attribute
- Enhance GGUF error handling with proper exceptions instead of just warnings
- Simplify gguf_ops dequantize_weight by removing redundant fallback paths
- Add memory logging and proper CUDA sync cleanup in GGUF loading
- Translate French comments to English in euler.py
- Remove unnecessary GGUF special handling in memory_manager
- Fix _propagate_debug_to_modules to handle None debug instance
2025-09-24 15:47:18 -04:00
Adrien Toupet 0b0c87ed4a Add GGUF quantized model support (based on PR #121 from @cmeka / @lihaoyun6)
- Implement GGUF model loading with Q3_K_M through Q8_K_M quantization support
- Add GGUFTensor wrapper to preserve quantization and enable on-demand dequantization
- Maintain tensors in quantized format to reduce VRAM usage
- Add GGUF dequantization operations for inference
- Update model registry to include GGUF variants for 3B/7B models
- Fix wavelet blur radius limit to prevent OOM at high resolutions (max 1/8 of image dimension)
- Add safety clamp [-1,1] for SDR color range to prevent numerical errors
- This is a WIP commit as some additional cleaning/testing is needed
- Add type hints throughout for better code maintainability
2025-09-24 11:49:57 -04:00
Adrien Toupet e6da542b7b VAE tiling defaults + Fix BlockSwap + preserve_vram compatibility and device management
- Changed VAE tiling default to False to restore original behavior
- Fixed BlockSwap bypass mode during cache cleanup to allow proper CPU offloading
- Added runner parameter to all manage_model_device calls for BlockSwap detection
- Removed conditional movement flags, always offload when preserve_vram=True
- Fixed device comparison to use full device strings (cuda:0) not just types
- Added debug logging for BlockSwap device skip scenarios
2025-09-17 23:39:03 -04:00
Adrien Toupet 1518ecc1a5 refactor: Fix ComfyUI node conflicts via relative imports and clearer model structure
- Renamed model directories for clarity: dit -> dit_7b, dit_v2 -> dit_3b
- Converted all absolute imports to relative imports throughout codebase
- Removed sys.path.append() manipulations that caused namespace conflicts
- Updated YAML configs to reference renamed model directories
- Simplified model variant detection logic using new directory names
- Standardized function calls with named arguments for better clarity
- Fixed generation context initialization and interrupt handling

This resolves import conflicts with other ComfyUI nodes (e.g., Basic data handling)
that use sys.path manipulation, making the module properly isolated and compatible.

Fixes #29, #114, #136
2025-09-16 16:45:31 -04:00
Adrien Toupet 477f57fd5a Refactor: Three-phase batch processing pipeline for improved performance
Major architectural change to minimize model swapping overhead by processing
all batches in three distinct phases instead of sequential per-batch processing:
- Phase 1: Encode all batches with VAE
- Phase 2: Upscale all latents with DiT
- Phase 3: Decode all latents with VAE

Core changes:
- Split monolithic generation_loop into modular functions:
  - prepare_generation_context(): Shared state management
  - setup_device_environment(): Device configuration
  - prepare_runner(): Model loading with cache support
  - encode_all_batches(): Batch VAE encoding
  - upscale_all_batches(): Batch DiT upscaling
  - decode_all_batches(): Batch VAE decoding
- Removed generation_step function (logic integrated into upscale phase)
- Added lazy precision initialization to avoid redundant setup

Performance improvements:
- Pre-allocated lists for memory efficiency
- Better cleanup of intermediate storage between phases
- Added unique timer names to clear_memory() to avoid naming conflicts
- Improved model state management with change detection and caching

UI/UX enhancements:
- Switched to ComfyUI's native ProgressBar with weighted phase progress
- Changed from per-batch FPS to overall average FPS (always visible)
- Improved log clarity with clear phase separators
- Added ASCII art logo to clearly identify SeedVR2 process start
- Better progress tracking with weighted percentages across three phases

Code cleanup:
- Removed deprecated timer_context from Debug class
- Removed unused time imports across multiple files
- Fixed LOCAL_RANK environment variable to handle string conversion properly
- Improved error handling with try/except/finally blocks in all phases
2025-09-16 14:38:57 -04:00
Adrien Toupet 66bcf436ab perf: optimize debug logging to reduce performance overhead
- Change log_memory_state() to default show_tensors=False (opt-in tensor counting), skipping expensive tensor counting during intermediate operations
- Enable detailed tensor analysis only at final cleanup
- Add early return pattern in log() method to avoid string operations when disabled
- Make reset_vram_peak() logging conditional on debug.enabled state
- Remove unused vram_info call in model_manager.py
2025-08-28 23:50:49 -04:00
Adrien Toupet 5e2fb76414 fix: optimize RoPE frequency computation and remove initialization bottleneck
- Fix NaMMRotaryEmbedding3d to compute only required dimensions instead of maximum (1024x128x128), enabling much higher resolution upscaling with 3B model without running OOM
- Remove slow preinitialize_rope_cache() which was causing bottleneck during model preparation - no longer needed with NaMMRotaryEmbedding3d fixed
- Clean up unnecessary memory clearing calls and improve debug logging clarity
2025-08-28 14:20:41 -04:00
Adrien Toupet feabb94e5a perf: optimize memory usage and streamline generation pipeline
Dtype Consistency:
- Pass compute_dtype through entire diffusion pipeline to ensure timesteps use bfloat16 instead of float32, reducing memory usage and eliminating dtype conversion overhead
- Propagate dtype from generation_loop → configure_diffusion → sampling timesteps creation

Code Deduplication:
- Removed duplicated code and made use of existing reusable functions: prepare_video_transforms(), load_text_embeddings(), calculate_optimal_batch_params()
- Improve calculate_optimal_batch_params() to prioritize temporal stability by recommending largest valid 4n+1 batch size

Memory Optimizations:
- Remove unnecessary GPU↔CPU offloading for text embeddings and timesteps during preserve_vram mode (added overhead without meaningful memory benefits)
- Simplify text embeddings cleanup to work directly with dictionary structure
- Add cuBLAS workspace clearing for better GPU memory management

Code Cleanup:
- Remove redundant self.runner = None in _internal_execute (handled by cleanup())
- Enhance user messaging for batch padding waste with clearer explanations
2025-08-27 14:15:41 -04:00
Adrien Toupet 267bc099e6 Fix BlockSwap compatibility with preserve_vram mode
- Add bypass mechanism to temporarily disable BlockSwap protection during offload
- Refactor manage_model_device() to handle BlockSwap models transparently
- Restore blocks to correct devices (partial GPU config) on reload
- Simplify preserve_vram logic in infer.py by removing BlockSwap checks

This allows BlockSwap to work correctly with preserve_vram, offloading the entire model to CPU after inference and restoring only necessary blocks to GPU before next inference.
2025-08-26 17:31:04 -04:00
Adrien Toupet 1ee1ee0fb5 fix: add OOM retry mechanism and fix FP16 black frames regression
- Implement retry_on_oom() helper with single retry for all VAE operations (Upsample3D, ResnetBlock3D, InflatedCausalConv3d, GroupNorm)
- Fix regression: use BFloat16 compute/autocast for FP16 models to prevent black frames
- Improve frame padding logs to clarify model constraint (frames % 4 == 1)
2025-08-26 00:03:02 -04:00
Adrien Toupet 5595d58597 refactor: streamline generation pipeline and improve dtype handling
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
2025-08-25 17:11:29 -04:00
Adrien Toupet 47552c0ed5 refactor: unify model configuration and improve logging precision
- Merge configure_dit_model_inference and configure_vae_model_inference into single configure_model_inference function
- Simplify load_quantized_state_dict by removing unnecessary FP8 conversion logic
- Remove redundant gradient checkpointing call (training-only feature)
- Remove redundant model_weight variables
- Standardize float formatting to .2f across all debug logs for consistency
- Eliminate redundant device movement operations already handled during model creation
2025-08-25 13:45:13 -04:00
Adrien Toupet 8c48aa59ea Fix VRAM memory leak in VAE decoder
Updated manage_model_device() to clear InflatedCausalConv3d memory buffers
when moving VAE to CPU. These buffers were not automatically released with
.to('cpu'), causing intermediate tensors to remain in VRAM after batch processing.
2025-08-25 10:12:40 -04:00
Adrien Toupet 627f215b57 Fix BlockSwap memory configuration being lost during model caching
Preserve BlockSwap's GPU/CPU distribution when cache_model=True to avoid unnecessary reconfiguration on subsequent runs. Also fixes regression error where tensors were not on the correct device after caching.
2025-08-24 11:38:02 -04:00
Adrien Toupet 8333fb856e refactor(WIP): complete memory management overhaul with proper error handling
Memory Management:
- Add type hints to all memory functions for better IDE support and maintainability
- Replace silent exception handling with debug logging across all operations
- Removed unnecessary CPU transfers for GPU memory release
- Introduce unified manage_model_device() for consistent device management
- Remove torch._C._clear_cache() private API usage (incompatible across PyTorch versions)

Performance & Debugging:
- Add debug timers to critical operations (clear_memory, clear_runtime_caches, etc.)
- Consolidate all cleanup code into core functions: manage_*, release_*, clear_*, complete_cleanup
- Improved timer log messages for clarity

Code Quality:
- Remove unused imports
- Remove redundant RoPE cache clearing (clear_runtime_caches handles it)
- Simplify configure_runner() by eliminating duplicate code paths
- Update infer.py to use generic device management functions
- Add release_text_embeddings() helper to deduplicate embedding cleanup
2025-08-23 01:13:45 -04:00
Adrien Toupet 2053a80f38 refactor(WIP): Improve debug logging consistency and reduce redundancy
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state()  formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
2025-08-22 15:38:00 -04:00
Adrien Toupet 9b7c113681 refactor(WIP): Memory Management Overhaul
This is a work-in-progress commit that consolidates a series of changes to fix memory leaks, optimize VRAM/RAM usage, improve performance, and enhance code maintainability.

*   **BlockSwap Pinned Memory:** Disabled `use_non_blocking=True` for CPU-to-GPU transfers to resolve a memory leak where pinned memory was not being released.
*   **Logging-Induced Leaks:** Modified `log_memory_state()` to avoid holding references to tensors during analysis and added a history limit to the checkpoint system to prevent unbounded memory growth.
*   **Incomplete Model Cleanup:** Ensured models are completely deleted and their tensor storage is released when `cache_model=False`.
*   **Lingering Tensors:** Fixed an issue where a scalar tensor from sampling timesteps and text embeddings remained on the GPU between batches when `preserve_vram` is active.

*   **Centralized Cleanup Functions:** Introduced `clear_memory()` to replace `clear_vram_cache()` and all manual `torch.cuda.empty_cache()` calls, providing consistent VRAM/RAM cleanup logic. The function features a `full` parameter to distinguish between a fast, GPU-only cache clear (~1-5ms) for frequent operations and a full cleanup with garbage collection (~10-50ms) for critical stages.
*   **Direct-to-CPU Model Loading:** Modified DiT/VAE weight loading to load directly onto the CPU when `preserve_vram` or `BlockSwap` is active, avoiding unnecessary VRAM spikes during model preparation.
*   **VAE Device Management:** Created the `manage_vae_device()` helper function to centralize the logic for moving the VAE between the CPU and GPU, reducing code duplication. This also fixed a bug that incorrectly kept the VAE on the GPU when `preserve_vram` was active.
*   **CPU Offloading:** Implemented logic to move text embeddings and sampling timesteps to the CPU after each batch when `preserve_vram` is active, reducing idle VRAM usage.

*   **VAE Decode Performance:** Replaced proactive, frequent memory clearing during VAE decode with a reactive Out-of-Memory (OOM) handling system. This fixed a significant performance regression and eliminated the need for the `keep_vae_loaded_during_decode` flag.
*   **Reduced Overhead:** Removed redundant `gc.collect()` calls from multiple locations to decrease unnecessary processing overhead.

*   **Interval-Based VRAM Tracking:** Modified the logging system to reset peak VRAM statistics after each `log_memory_state()` call, enabling accurate tracking of peak memory usage for specific processing intervals (e.g., encode, inference, decode).
*   **Accurate RAM Monitoring:** Added the `get_ram_usage()` function for correct process-specific RAM tracking.
*   **Efficient Log Refactoring:** Refactored `log_memory_state()` into modular helper methods, optimizing tensor analysis into a single-pass `gc` iteration to improve both performance and maintainability.
*   **Log Clarity:** Refined memory state and debug logging to remove redundant snapshots and add new ones for critical operations like model loading, weight loading, VAE encoding, and decoding. Standardized log message conventions.
*   **Per-Batch Timers:** Implemented timer namespacing to ensure that performance timers for each batch are logged correctly without overwriting one another.

*   **Error Handling:** Added `try/except` blocks to key memory and device management functions to handle edge cases and improve robustness.
*   **Code Cleanup:** Removed deprecated code and outdated comments throughout the related modules.
*   **Documentation:** Updated comments and function docstrings to reflect the new memory management architecture.
2025-08-22 09:20:41 -04:00
lihaoyun6 1cfde3d1c7 Added multi-gpu support for ComfyUI nodes 2025-08-15 21:06:16 +08:00
lihaoyun6 176562f829 Restored compatibility with FP8 safetensors for MPS backend; Changed the way to detect MPS device 2025-08-13 00:53:05 +08:00