Commit Graph
67 Commits
Author SHA1 Message Date
Adrien Toupet 7cbf025561 Fix VRAM peak tracking: separate allocated vs reserved, Windows-only overflow
- Track both peak_allocated (tensor usage) and peak_reserved (cache pool) per phase
- peak_allocated resets properly between phases via reset_peak_memory_stats()
- Overflow detection/warnings now Windows-only (WDDM paging behavior)
- Remove get_memory_architecture() - replaced with simple is_mps + platform checks
- Phase summary shows: VRAM XGB allocated, YGB reserved | RAM ZGB
- Simplify MPS path (unified memory has no overflow concept)
2025-12-09 23:51:51 -05:00
Adrien Toupet 5c60716c47 Refactor: centralize backend detection, fix architecture-aware VRAM overflow reporting 2025-12-09 21:06:12 -05:00
Adrien Toupet 77a00f651a Fix: OOM regression from 2.5.14 strict VRAM limit (#367)
Add allow_vram_overflow option (default: False) to make strict VRAM limit configurable.

The 2.5.14 change 'Enforce physical VRAM limit' prevented PyTorch from
overflowing to system RAM, causing OOM on workflows that previously
worked.

- Add allow_vram_overflow parameter to DiT Model Loader node
- Add --allow_vram_overflow CLI flag
- Show warning when enabled, track mid-session changes
- Suppress swap detection warning when user explicitly allows overflow

Note: Enabling overflow is a last resort - performance degrades severely
when physical VRAM is exceeded. Optimizing settings (BlockSwap, VAE tiling,
batch size, resolution, model size...) is always recommended.
2025-12-09 17:12:10 -05:00
Adrien Toupet 4e96a5c366 fix: allow model caching with multi-GPU streaming (workers cache internally) 2025-12-09 00:48:18 -05:00
Adrien Toupet 4817beb148 fix: multi-GPU streaming log shows GPU count, workers log with [GPU N] prefix 2025-12-09 00:36:25 -05:00
Adrien Toupet 0b132b02ff refactor: multi-GPU workers stream video segments internally with model caching 2025-12-09 00:13:30 -05:00
Adrien Toupet f7e4fc677e Fix multi-GPU shared memory race condition with barrier sync 2025-12-08 22:29:12 -05:00
Adrien Toupet a70d82e3aa Add streaming mode for memory-efficient long video processing
- New --chunk_size flag enables streaming mode, processing video in bounded chunks
- Supports both MP4 output (single file) and PNG sequence output while streaming
- Preserves --load_cap for total frame limiting (backward compatible)
- Model caching now works between chunks when --cache_dit/--cache_vae enabled
- Instant frame seeking with cv2.CAP_PROP_POS_FRAMES (fixes slow skip on long videos)
- Early exit for empty/exhausted videos
- Minor: function renames (save_frames_to_png → save_frames_to_image), log message cleanup

Inspired by PR #353 - thank you @disk02 for the initial chunked_mode implementation
2025-12-08 22:05:20 -05:00
Adrien Toupet bbd7e5ac02 Fix multiprocessing MemoryError for large video outputs (#372)
Use PyTorch shared memory instead of pickling numpy arrays through queue.
Prevents MemoryError when transferring large results between processes.
Thank you @FurkanGozukara
2025-12-08 14:05:41 -05:00
Adrien Toupet 18b44d66e1 feat: add environment info display in debug mode to help with issue reporting 2025-12-05 11:11:18 -05:00
Adrien Toupet 6930e0f13a Fix CLI MPS watermark error on macOS (fixes #336) 2025-11-30 07:37:51 -05:00
Adrien Toupet daa13fb4ca Fix BlockSwap logging confusion and CLI worker validation
- BlockSwap: Show effective/total blocks (e.g., 32/32) instead of raw requested value
- CLI: Skip CUDA device validation when CUDA_VISIBLE_DEVICES already set (worker process)
2025-11-28 13:40:07 -05:00
Adrien Toupet 7581e0014f Fix CLI: MPS subprocess allocator error (#290) and multi-GPU distribution (#309)
- Mac: Use direct processing instead of spawning subprocess (MPS allocator fails in child process)
- Multi-GPU: Set CUDA_VISIBLE_DEVICES before spawn so child inherits it before module-level torch import
- Remove redundant env setup in worker (now inherited from parent)
2025-11-28 12:11:42 -05:00
Adrien Toupet fc64968b12 fix(cli): improve output paths and add RGBA support (v2.5.8)
- Improve output folder naming: batch creates {folder}_upscaled/ sibling with original filenames, single file adds _upscaled suffix
- Add RGBA alpha channel detection and preservation (matches ComfyUI)
- Convert all output paths to absolute for clarity in logs
2025-11-10 14:39:33 -05:00
Adrien Toupet b130a33894 fix(cli): resolve Windows duplicate file bug and improve scan perf 2-3x (v2.5.8)
- Replace dual glob loops with single iterdir scan for cross-platform consistency
- Fixes duplicate file processing in batch mode on Windows case-insensitive filesystem
- Improves directory scanning performance 2-3x by reducing filesystem operations
- Add ComfyUI registry logo
2025-11-10 13:39:42 -05:00
Adrien Toupet 4e9ce4710e Unify parameter names across CLI and ComfyUI interface
- Rename new_resolution to resolution for consistency
- Change --input to positional input argument
- Rename --model to --dit_model for clarity
- Simplify VAE tiling flags: --vae_encode_tiled and --vae_decode_tiled
- Update all documentation and example workflows
- Maintain consistent naming convention across entire codebase
2025-11-06 22:51:20 -05:00
Adrien Toupet 20dab62dc3 feat: add uniform_batch_size for temporal consistency + unify padding logic
- Add uniform_batch_size parameter to eliminate temporal artifacts in final batch
- Unify temporal padding: single pad_video_temporal() replaces cut_videos() and prepend_video_frames()
- Improve logging: separate messages for uniform vs 4n+1 padding
- Enhance CLI: Improved dynamic examples and use actual invocation path
- README.md: standardize folder references, use seedvr2_videoupscaler folder name consistently, improve parameter documentation
2025-11-06 14:13:22 -05:00
Adrien Toupet 806bb94df0 feat: unify and improve tooltip documentation across CLI and ComfyUI nodes
- Standardize tooltip format with multi-line descriptions and bullet points
- Add comprehensive output tooltips for all nodes (DiT, VAE, torch.compile, upscaler)
- Enhance node descriptions with detailed capability summaries
- Simplify CLI tile size arguments to single integers (converted internally to tuples)
- Remove OneOrTwoValues argparse class for cleaner implementation
- Fix encode_tiled tooltip (was incorrectly referencing decoding)
- Clarify color correction purpose (corrects upscaling color shifts)
- Add multi-GPU offloading information to all offload_device tooltips
- Improve torch.compile parameter descriptions with use cases
- Ensure CLI and ComfyUI tooltips are consistent in terminology and structure
2025-11-05 15:35:22 -05:00
Adrien Toupet 326489d94c fix CLI: move Debug import after CUDA allocator config to fix batch processing errors 2025-11-05 14:27:06 -05:00
Adrien Toupet 9b79254c39 refactor(cli): improvements and bug fixes + 3b-Q8_0.gguf support
- Fix validation cache location to respect --model_dir parameter
- Fix output path handling for directories without extensions
- Remove spurious directory creation in get_base_cache_dir
- Enhanced CLI help with usage examples and argument grouping
- Auto-display help when script invoked without arguments
- Correct type hints (device_id: str, debug: Debug)
- Remove redundant type conversions and makedirs calls
- Reorganize imports to module top for clarity
- Improved docstrings & tooltip
- Change default batch_size from 1 to 5 to match ComfyUI integration
- Add support for seedvr2_ema_3b-Q8_0.gguf model
2025-11-05 00:28:19 -05:00
Adrien Toupet 3725c1061d refactor: centralize dimension computation and logging for CLI/ComfyUI
- Add compute_generation_info() and log_generation_start() helpers
- Move prepend_frames logic from extraction to processing pipeline
- Eliminate code duplication between CLI and ComfyUI workflows
- Add consistent dimension/parameter logging for both interfaces
- Disable argparse prefix matching for safer CLI usage
2025-11-04 17:01:01 -05:00
Adrien Toupet 1691e657b4 fix(cli): add CUDA device validation before torch initialization
- Validate --cuda_device arguments early in pre-parsing phase
- Check device IDs exist and are within available GPU range
- Fail fast with clear error messages showing available devices
2025-11-04 15:52:05 -05:00
Adrien Toupet 32a049dfd9 feat: Add CLI model caching for multi-file processing and unify device handling
- Add --cache_dit and --cache_vae flags for efficient multi-file directory processing
- Refactor processing pipeline to eliminate duplication between worker and direct modes
- Implement platform-agnostic device management (CUDA/MPS/CPU)
- Unify parameter naming: res_w→resolution, max_res_w→max_resolution across codebase
- Add smart offload device defaults when caching enabled
- Improve validation and user feedback for cache + multi-GPU scenarios
2025-11-04 15:12:28 -05:00
Adrien Toupet ad020d3803 feat(cli): improve UX with auto-format detection, FPS tracking, and consistent messaging with ComfyUI implementation
- Auto-detect output format per file type (mp4 for videos, png for images)
- Add visual separators between processed files for better readability
- Simplify FPS calculation to use wall-clock time for real-world throughput
- Consolidate banner/footer into shared Debug methods
- Update offload device args to support multi-GPU (cpu/cuda:N)
- Standardize terminology: 'upscaling' instead of 'video upscaling'
- Remove code duplication between CLI and ComfyUI implementations
- Consistent quote style (double quotes) throughout CLI
2025-11-04 11:54:57 -05:00
Adrien Toupet 9268346388 feat: CLI Add batch processing, fix multiprocessing issues, and unify model paths
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)

Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable

Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability

Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
2025-11-04 00:12:13 -05:00
Adrien Toupet ce8225fd48 feat: Add max_resolution parameter to limit output dimensions
- Add max_resolution parameter (default: 0 = no limit) to both CLI and ComfyUI
- After new_resolution scales shortest edge, max_resolution ensures no edge exceeds limit
- Scales down proportionally if constraint violated
- Maintains backward compatibility with default value of 0
2025-11-03 20:17:32 -05:00
Adrien Toupet 77cb6ff684 refactor(cli): inference_cli to match ComfyUI integration
- Add dit_offload_device parameter for proper blockSwap configuration
- Ensure consistent dtype management throughout CLI and ComfyUI (float32 input with bfloat16 pipeline)
- Translate all French comments to English
- Add comprehensive docstrings and section headers
- Remove obsolete use_non_blocking and enable_debug parameters
- Add error handling and validation
2025-10-28 01:01:38 -04:00
Adrien Toupet e8376ddd6d refactor: split generation.py and model_manager.py into 4 focused modules
- generation.py → generation_phases.py (4-phase pipeline logic) + generation_utils.py (setup/helpers)
- model_manager.py → model_configuration.py (config/caching) + model_loader.py (weight loading/GGUF)
- Renamed functions and code cleanup
2025-10-28 00:04:08 -04:00
Adrien Toupet f182de79fa feat: Add temporal_overlap & prepend_frames to ComfyUI with shared logic
- Implement prepend_video_frames() for artifact reduction at video start
- Add blend_overlapping_frames() with Hann window for smooth transitions
- Expose temporal_overlap (0-16) and prepend_frames (0-32) in ComfyUI node
- Unify CLI and ComfyUI to use shared prepend/overlap functions
- Add comprehensive logging for frame adjustments (prepend/overlap/padding)
2025-10-27 21:57:38 -04:00
Adrien Toupet 003122ebcd feat: implement lossless arbitrary resolution with padding - replace DivisibleCrop with DivisiblePad to eliminate data loss, track true dimensions for post-processing trim, change default resolution to 1080p with step=2 for flexibility 2025-10-27 01:14:57 -04:00
Adrien Toupet 0198834299 Remove cfg_scale parameter (incompatible with distilled one-step model)
CFG (Classifier-Free Guidance) does not work with SeedVR2's distilled
one-step diffusion model. The model was trained to produce final results
in a single step without iterative guidance.

- Remove cfg_scale parameter from ComfyUI node and CLI interface
- Force internal cfg_scale to 1.0 in upscale_all_batches()

This avoids artifacts introduced when users changed cfg_scale away from 1.0.
2025-10-26 22:54:45 -04:00
Adrien Toupet 70087c0b93 Improve VAE encoding stability and add tile debugging
- Switch to deterministic VAE encoding (mode vs sample) to eliminate high-resolution noise artifacts
- Make VAE encode tiling optional (disabled by default) since deterministic encoding resolves artifacts
- Add tile debug visualization feature with adaptive scaling and color-coded boundaries
- Remove redundant dtype conversions in VAE code for better performance
- Minor code cleanup and documentation update
2025-10-25 21:11:41 -04:00
Adrien Toupet 01cbdf8bc3 Optimize VAE defaults and standardize dtype pipeline for quality/performance
VAE Changes:
- Enable encode tiling by default (prevents noise artifacts at high resolution)
- Increase tile size to 1024px (down from 512px) for optimal quality
- Increase tile overlap to 128px for better blending

Dtype Pipeline:
- Hardcode compute_dtype to bfloat16 for consistent quality/performance/VRAM balance
- Ensure all pipeline steps are using compute_dtype when relevant
- Refactor code for improved performance and memory management
2025-10-22 23:56:14 -04:00
Adrien Toupet c9dce827c0 feat: Add deterministic generation with seed control and CFG scale parameter
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
  * VAE encoding: seed+1M for deterministic sampling without quality loss
  * DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id

Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)
2025-10-21 13:31:54 -04:00
Adrien Toupet e735c2ad56 feat: V3 migration with GGUF fixes and attention optimizations
Major Changes:
- Migrate all nodes to ComfyUI V3 schema (stateless design, new IO types)
- Fix GGUF weight caching VRAM leak (non-persistent buffers + _apply override)
- Fix GGUF torch.compile compatibility (@torch._dynamo.disable on dequant)
- Centralize compatibility checks (Flash/Triton/GGUF/Conv3d in compatibility.py)
- Make flash_attn optional with graceful SDPA fallback
- Add attention_mode UI option (sdpa/flash_attn) to DiT loader
- Fix color correction batch padding error (trim input_video consistently)

Code Quality:
- Remove internal_execute for clarity (stateless node design)
- Add startup logging for optimization status
- Improve error messages with installation instructions
- Add get_unique_id() for V3 node-specific caching
- Standardize parameter names (dit_cache/vae_cache)
2025-10-21 00:10:37 -04:00
Adrien Toupet 9893a1d7b0 refactor: Add indent_level parameter to debug logging system
- Add indent_level param to debug.log() and manage_tensor()
- Replace all hardcoded spaces with indent_level (0/1/2)
- Fix encode_all_batches memory flow: encode before storage
- Minor whitespace cleanup
2025-10-19 10:13:17 -04:00
Adrien Toupet 97c7b9cd12 refactor: Overhaul device and memory management architecture
Major improvements to device handling, model offloading, and code clarity:

Device Management:
- Replace string-based devices with torch.device objects throughout
- Add explicit offload device parameters: dit_offload_device, vae_offload_device, tensor_offload_device
- Remove preserve_vram in favor of explicit offload control
- Improve get_device() to return torch.device objects consistently
- Enhance get_device_list() with smart MPS-only system handling

Context & Pipeline:
- Merge setup_device_environment and prepare_generation_context into single setup_generation_context
- Simplify LOCAL_RANK handling (set to '0' for single-GPU mode)
- Store device configuration on runner for submodule access

Tensor & Model Management:
- Add manage_tensor_device() for consistent tensor movement with logging
- Update manage_model_device() to use torch.device objects
- Add validation for BlockSwap and caching configurations
- Rename cache_in_ram to cache_model (more accurate naming)

CLI & Interface:
- Remove --preserve_vram flag (breaking change)
- Add --vae_offload_device and --tensor_offload_device flags
- Update ComfyUI node parameters with validation and better tooltips
- Improve device selection UI with offload_device options
2025-10-17 17:24:53 -04:00
Adrien Toupet a46e16d8cc refactor: 42, centralize float32 precision utilities and improve logging
- Move ensure_float32_precision() from alpha_upscaling to common/half_precision_fixes
- Update color_fix.py to use ensure_float32_precision
- Add type hints and docstrings to all half_precision_fixes functions
- Clarify batch size tips: emphasize avoiding padding waste over maximizing batch size
- Update default seed to 42 (the answer to life, the universe, and everything)
2025-10-16 14:30:47 -04:00
Adrien Toupet f1ac7d786e Add perceptually-accurate color correction methods
- Add LAB color transfer as new default method for superior perceptual color matching
- Add HSV hue-conditional saturation histogram matching for targeted oversaturation correction
- Add wavelet_adaptive hybrid method combining wavelet base with selective HSV correction
- Change default color correction from 'wavelet' to 'lab' for better color accuracy
- Implement full CIELAB color space conversion with D65 illuminant and histogram matching
- Add comprehensive documentation and method descriptions to color_fix.py
2025-10-15 13:14:48 -04:00
Adrien Toupet e48c2da776 feat: global model cache, alpha upscaling rewrite, and quality improvements
Major refactor introducing:

- GlobalModelCache for cross-node model sharing with dynamic config updates
  * Models cached by node ID, reused across different upscaler instances
  * Config changes (torch.compile, BlockSwap, tiling) handled dynamically
  * Runner templates cached when both models present

- Complete alpha channel processing rewrite
  * New edge-guided alpha upscaling in src/core/alpha_upscaling.py
  * Remove broken VAE RGBA adapter approach
  * Alpha properly extracted, upscaled, and merged with RGB

- BlockSwap enhancements
  * Support I/O-only swapping without transformer block offloading
  * Separate memory reporting for I/O components vs transformer blocks
  * Clearer logging of GPU vs CPU placement
  * New is_blockswap_enabled() utility

- Memory management improvements
  * Add release_tensor_collection() for batch tensor cleanup
  * Proper tensor memory release throughout pipeline
  * Better cleanup in postprocess phase

- Code quality and refactoring
  * Extensive docstrings with type hints across all modules
  * Generic _update_model_config() reduces code duplication
  * Better function signatures and parameter documentation
  * Improved separation of concerns in model_manager.py
  * Clearer debug logging categories and messages
2025-10-15 00:51:56 -04:00
Adrien Toupet fb2b6c7d4c Fix BlockSwap not applying on cached models and rename offload_io_components to swap_io_components
- Fixed critical bug where BlockSwap configuration changes were not applied to cached models
- Apply BlockSwap immediately in _handle_blockswap_config instead of deferring to materialization phase
- Renamed offload_io_components to swap_io_components across entire codebase for consistency
- Removed unused _pending_blockswap_config attribute
2025-10-10 17:48:12 -04:00
Adrien Toupet 9ee244d71b feat: implement torch.compile optimization with BlockSwap compatibility
Core Features:
- Add torch.compile support for DiT (20-40% speedup) and VAE (15-25% speedup)
- Ensure BlockSwap compatibility by applying BlockSwap before torch.compile
- Add SeedVR2TorchCompileSettings node for ComfyUI configuration
- Add CLI arguments: --compile_dit, --compile_vae, --compile_backend, --compile_mode, --compile_fullgraph, --compile_dynamic, --compile_dynamo_cache_size_limit, --compile_dynamo_recompile_limit

Optimizations:
- Optimize na.py and other backend files for torch.compile compatibility (replace .tolist() with tensor operations)
- Remove .item() calls from attention modules (max_seqlen_q/k now accept tensors)
- Add @torch._dynamo.disable decorators to timing/debug methods to prevent compilation warnings
- Replace torch.repeat with torch.repeat_interleave in modulation.py for better performance

Bug Fixes:
- Fix timing stack cleanup in memory_manager.py (end timer before early returns)
- Fix BlockSwap timing reporting (_get_swap_start_time and _log_swap_timing now properly excluded from compilation)

Code Quality:
- Add comprehensive type hints and docstrings across all modified modules
- Standardize BlockSwap timer names (blockswap_block_*, blockswap_io_*)
2025-10-10 12:34:49 -04:00
Adrien Toupet 5567f232af feat!: independent DiT/VAE management with separate loader nodes
BREAKING CHANGE: Replaced BlockSwap and ExtraArgs nodes with DiT/VAE loader nodes.
Enables multi-GPU placement, independent caching, and granular memory control.

New Features:
- Add SeedVR2LoadDiTModel and SeedVR2LoadVAEModel loader nodes
- Support different devices for DiT and VAE (multi-GPU load balancing)
- Independent model caching (cache_model_dit/cache_model_vae)
- Separate encode/decode VAE tiling with independent tile configurations
- User-defined DiT and VAE model selection from registry or disk

Breaking Changes:
- Deprecated SeedVR2BlockSwap node (integrated into DiT loader)
- Deprecated SeedVR2ExtraArgs node (split into DiT/VAE loaders & Upscaler node)
- BlockSwap configuration now part of DiT loader node

Improvements:
- Follow ComfyUI conventions (pixels input, IMAGE output type)
- Update download_weight() to support user-defined DiT/VAE models
- Improve tooltips and default parameters across all nodes
- Set show_tensors=False in debug logging for performance
- Dual-device context tracking throughout pipeline
- CLI updated to match new dual-model API
2025-10-08 00:49:48 -04:00
Adrien Toupet 2fab3a1caf refactor: split decode and post-processing into separate phases for cleaner architecture
- Split decode_all_batches into decode (Phase 3) and postprocess_all_batches (Phase 4)
- Add phase-specific cleanup functions: cleanup_dit() and cleanup_vae()
- Each phase now handles its own resource cleanup in finally blocks
- Add cleanup_text_embeddings() helper to eliminate code duplication
- Remove redundant cleanup from comfyui_node normal flow
- Update pipeline from 3-phase to 4-phase architecture (encode → upscale → decode → postprocess)
- Improve memory efficiency by releasing resources immediately when no longer needed
- Update module docstrings to reflect new architecture
2025-10-03 09:53:13 -04:00
Adrien Toupet 4a3efc7883 feat: add CPU offloading for intermediate data to reduce VRAM usage
- Store latents and transformed videos on CPU between processing phases reducing VRAM usage to enable larger batch processing
- Make transformed video storage conditional (only when color_correction != none)
- Add all_ori_lengths tracking for consistent trimming
- Use non_blocking=False for CPU-GPU transfers to avoid pinned memory issues
Fixes https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler/pull/163#issuecomment-3333686784
2025-09-25 11:30:32 -04:00
Adrien Toupet 9b4a7dfa3e fix: unify model registry for CLI and improve GGUF implementation
Model Registry Integration:
- Replace hardcoded CLI model choices with dynamic get_available_models()
- Use centralized DEFAULT_MODEL constant from model_registry
- Achieve single source of truth between CLI and ComfyUI interfaces

GGUF Optimizations:
- Create tensors directly on target device to avoid CPU->GPU copy overhead
- Remove excessive memory cleanup during tensor loading for better performance
- Always use meta init for memory-efficient model creation
- Implement precision-optimized dequantization path (GGUF → FP16 → compute dtype)

GGUF Precision Handling:
- Add GGUFQuantizedLinear/Conv2d layers with get_dequantized_weight_for_compute() to preserve precision
- Track and report quantization types (Q4_K_M, Q5_K_M, etc.) in model
- Fix __torch_function__ as classmethod to resolve deprecation warning

Code Quality:
- Restructure _load_model_weights() with modular helper functions to reduce duplication
- Improve separation of concerns between standard and GGUF weight loading
- Enhance logging to always display WARNING/ERROR messages
- Add comprehensive GGUF architecture validation
- Remove debug-only log_memory_state calls
- Fix type function definitions
2025-09-25 00:09:56 -04:00
Adrien Toupet 0b0c87ed4a Add GGUF quantized model support (based on PR #121 from @cmeka / @lihaoyun6)
- Implement GGUF model loading with Q3_K_M through Q8_K_M quantization support
- Add GGUFTensor wrapper to preserve quantization and enable on-demand dequantization
- Maintain tensors in quantized format to reduce VRAM usage
- Add GGUF dequantization operations for inference
- Update model registry to include GGUF variants for 3B/7B models
- Fix wavelet blur radius limit to prevent OOM at high resolutions (max 1/8 of image dimension)
- Add safety clamp [-1,1] for SDR color range to prevent numerical errors
- This is a WIP commit as some additional cleaning/testing is needed
- Add type hints throughout for better code maintainability
2025-09-24 11:49:57 -04:00
Adrien Toupet 22f540a7d6 feat: Add input noise parameter and rename latent noise for clarity
- Add input_noise_scale (0-1) to reduce artifacts at high resolutions
  - Applies subtle noise before VAE encoding with progressive blend (0-50%)
  - Based on GitHub issue #64 community findings

- Rename cond_noise_scale to latent_noise_scale for clarity
  - Better distinguishes between input (pixel) and latent (diffusion) noise
  - Maintains same functionality, just clearer naming

- Update both ComfyUI and CLI interfaces with new parameters
2025-09-22 11:47:49 -04:00
Adrien Toupet e89a8b00a5 feat: promote cond_noise_scale to UI for controlling conditioning noise level helping mitigate noise artifacts in high-res upscaling 2025-09-18 23:19:35 -04:00
Adrien Toupet ee95a3a673 feat: Add configurable color correction methods and improve memory management
- Add color_correction parameter (wavelet/adain/none) to both ComfyUI node and CLI interface to choose between wavelet reconstruction (frequency-based), AdaIN (statistical matching), or no color correction
- Improve memory management after upscaling batches with targeted clear_memory calls to reduce VRAM pressure during VAE decoding
2025-09-16 22:21:45 -04:00