Commit Graph
34 Commits
Author SHA1 Message Date
Adrien Toupet 9715d3e37a Fix: torch.mps AttributeError on Windows
Add defensive checks for torch.mps.is_available() to handle PyTorch versions where the method doesn't exist on non-Mac platforms. Resolves AttributeError: module 'torch.mps' has no attribute 'is_available'
2025-11-07 23:37:35 -05:00
Adrien Toupet 9268346388 feat: CLI Add batch processing, fix multiprocessing issues, and unify model paths
Major Features:
- Renamed --video_path to --input supporting video files, images, and directories
- Added batch processing for directories (iterates all media files)
- Added single image upscaling with extract_frames_from_image()
- Auto-detect output format: images→PNG, videos→MP4 (overridable)
- Smart output path generation (single PNG vs frame sequences)

Critical Bug Fixes:
- Fixed 'str' object has no attribute 'type' by normalizing devices to torch.device
- Fixed 'Got unsupported ScalarType BFloat16' by converting ML dtypes to float32
- Fixed prepare_runner() signature mismatch (returned 2 values, claimed 3)
- Fixed KeyError 'cache_context' by storing cache_context in ctx
- Fixed duplicate optimization logging (3x imports) using environment variable

Performance Improvements:
- Removed mp.Manager() overhead
- Using direct mp.Queue(maxsize=0) for better throughput
- Improved multiprocessing reliability

Consistency & Quality:
- Unified model directory between CLI & ComfyUI to models/SEEDVR2 using constants
- Default CLI output folder to use ./output/
2025-11-04 00:12:13 -05:00
Adrien Toupet 01cbdf8bc3 Optimize VAE defaults and standardize dtype pipeline for quality/performance
VAE Changes:
- Enable encode tiling by default (prevents noise artifacts at high resolution)
- Increase tile size to 1024px (down from 512px) for optimal quality
- Increase tile overlap to 128px for better blending

Dtype Pipeline:
- Hardcode compute_dtype to bfloat16 for consistent quality/performance/VRAM balance
- Ensure all pipeline steps are using compute_dtype when relevant
- Refactor code for improved performance and memory management
2025-10-22 23:56:14 -04:00
Adrien Toupet c9dce827c0 feat: Add deterministic generation with seed control and CFG scale parameter
Core Changes:
- Implement deterministic generation with phase-specific seeding strategy
  * VAE encoding: seed+1M for deterministic sampling without quality loss
  * DiT upscaling: base seed for reproducible noise generation
- Add cfg_scale parameter for user control of upscaling strength (WIP)
- Fix ComfyUI V3 unique_id extraction using get_executing_context().node_id

Improvements:
- Standardize Optional['Debug'] type hints across codebase
- Make debug parameter required where it's essential (generate, infer)
- Remove legacy get_unique_id() stack inspection approach
- Add seed and cfg_scale logging for transparency
- Fix FP8CompatibleDiT parameter order consistency
- Refine input/latent noise scale steps (0.01 → 0.001 for finer control)
2025-10-21 13:31:54 -04:00
Adrien Toupet e735c2ad56 feat: V3 migration with GGUF fixes and attention optimizations
Major Changes:
- Migrate all nodes to ComfyUI V3 schema (stateless design, new IO types)
- Fix GGUF weight caching VRAM leak (non-persistent buffers + _apply override)
- Fix GGUF torch.compile compatibility (@torch._dynamo.disable on dequant)
- Centralize compatibility checks (Flash/Triton/GGUF/Conv3d in compatibility.py)
- Make flash_attn optional with graceful SDPA fallback
- Add attention_mode UI option (sdpa/flash_attn) to DiT loader
- Fix color correction batch padding error (trim input_video consistently)

Code Quality:
- Remove internal_execute for clarity (stateless node design)
- Add startup logging for optimization status
- Improve error messages with installation instructions
- Add get_unique_id() for V3 node-specific caching
- Standardize parameter names (dit_cache/vae_cache)
2025-10-21 00:10:37 -04:00
Adrien Toupet 1518ecc1a5 refactor: Fix ComfyUI node conflicts via relative imports and clearer model structure
- Renamed model directories for clarity: dit -> dit_7b, dit_v2 -> dit_3b
- Converted all absolute imports to relative imports throughout codebase
- Removed sys.path.append() manipulations that caused namespace conflicts
- Updated YAML configs to reference renamed model directories
- Simplified model variant detection logic using new directory names
- Standardized function calls with named arguments for better clarity
- Fixed generation context initialization and interrupt handling

This resolves import conflicts with other ComfyUI nodes (e.g., Basic data handling)
that use sys.path manipulation, making the module properly isolated and compatible.

Fixes #29, #114, #136
2025-09-16 16:45:31 -04:00
Adrien Toupet 5e2fb76414 fix: optimize RoPE frequency computation and remove initialization bottleneck
- Fix NaMMRotaryEmbedding3d to compute only required dimensions instead of maximum (1024x128x128), enabling much higher resolution upscaling with 3B model without running OOM
- Remove slow preinitialize_rope_cache() which was causing bottleneck during model preparation - no longer needed with NaMMRotaryEmbedding3d fixed
- Clean up unnecessary memory clearing calls and improve debug logging clarity
2025-08-28 14:20:41 -04:00
Adrien Toupet 5595d58597 refactor: streamline generation pipeline and improve dtype handling
- Merge generation setup log steps: combine "Model Configuration" and "Input Preparation" into unified "Generation Setup" step
- Simplify dtype handling: remove redundant vae_dtype variable, VAE now uses configured dtype consistently
- Remove platform-specific VAE dtype logic (MPS special case)
- Unify VAE encode/decode: remove autocast wrapper and target_dtype parameter, both now use configured dtype
- Add consistent docstrings for vae_encode and vae_decode methods
- Improve precision logging: show both DiT and VAE dtypes, rename model_dtype to dit_dtype
- Optimize text embeddings movement: only move when preserve_vram is active (BlockSwap handles model layers separately)
- Load text embeddings directly to CPU when preserve_vram is enabled
- Update logging consistency: uppercase device names, unified category names, clearer messages
- Clarify FP8 RoPE conversion log message: specify "from FP8 to BFloat16"
2025-08-25 17:11:29 -04:00
Adrien Toupet 8333fb856e refactor(WIP): complete memory management overhaul with proper error handling
Memory Management:
- Add type hints to all memory functions for better IDE support and maintainability
- Replace silent exception handling with debug logging across all operations
- Removed unnecessary CPU transfers for GPU memory release
- Introduce unified manage_model_device() for consistent device management
- Remove torch._C._clear_cache() private API usage (incompatible across PyTorch versions)

Performance & Debugging:
- Add debug timers to critical operations (clear_memory, clear_runtime_caches, etc.)
- Consolidate all cleanup code into core functions: manage_*, release_*, clear_*, complete_cleanup
- Improved timer log messages for clarity

Code Quality:
- Remove unused imports
- Remove redundant RoPE cache clearing (clear_runtime_caches handles it)
- Simplify configure_runner() by eliminating duplicate code paths
- Update infer.py to use generic device management functions
- Add release_text_embeddings() helper to deduplicate embedding cleanup
2025-08-23 01:13:45 -04:00
Adrien Toupet 2053a80f38 refactor(WIP): Improve debug logging consistency and reduce redundancy
- Remove "Force move weights to device" in FP8CompatibleDiT forward() to avoid clash with blockswap - This was used for preserve_vram but will refactor preserve_vram in a separate commit
- Remove duplicate dtype detection in generation_step (now passed from generation_loop)
- Add device checks before CPU moves to avoid redundant operations
- Improve debug.log and debug.log_memory_state()  formatting, content, and categories for better visibility
- Remove unused imports and excessive clear_memory() calls
- Clean up non-essential logs from always display
2025-08-22 15:38:00 -04:00
lihaoyun6 176562f829 Restored compatibility with FP8 safetensors for MPS backend; Changed the way to detect MPS device 2025-08-13 00:53:05 +08:00
lihaoyun6 8ae90c3af2 Keep weights and arguments on the same device when 'preserve_vram' is ON 2025-08-12 23:12:12 +08:00
lihaoyun6 3b0699ead9 Merge remote-tracking branch 'upstream/nightly' 2025-08-12 22:28:36 +08:00
Adrien Toupet d075dde799 feat: Extract model caching to main node + add hierarchical debug system
- Move cache_model and enable_debug options from BlockSwap to main SeedVR2 node

- cache_model: Keep models in RAM between runs (skip reload for faster iterations)

- enable_debug: Show detailed memory/timing tracking with hierarchical display

- Replace hardcoded debug statements with unified Debug class

- Add visual icons for log categorization

- Implement parent-child timing relationships for operation breakdown

- Add VRAM/RAM usage tracking with before/after comparisons

Model caching now available for all workflows, not just BlockSwap users.

Debug system provides granular performance insights when enabled.
2025-08-08 00:49:46 +02:00
Adrien Toupet 4118595968 Revert to 5c65c44 and fix integration issues
- Resolved merge conflicts

- Added missing constants.py and model_registry.py files

- Added graceful error handling for failed model downloads
2025-08-08 00:49:46 +02:00
lihaoyun6 3b530dc983 Added MPS backend support (for running on macOS) 2025-08-06 18:52:33 +08:00
Adrien Toupet 7b63eeea4c Revert "Models update - FP8, GGUF, Artifacts" 2025-07-24 12:53:46 -04:00
Adrien Toupet bc8f0aa6d0 Optimized RoPE stability fix and Patch RoPE for blockswap 2025-07-24 09:32:18 -04:00
Adrien Toupet b551e32ae1 Fixed scoping in _stabilize_rope_computations 2025-07-24 01:05:49 -04:00
Adrien Toupet 99938bf175 Apply RoPE stability fix universally to prevent artifacts in all precision modes 2025-07-23 22:01:55 -04:00
Adrien Toupet 09a9bf081d Fix artifacts when using mixed precision model without blockswap 2025-07-23 21:09:27 -04:00
Adrien Toupet bc292f26bd Fix artifacts when using mixed precision model without blockswap 2025-07-23 20:56:02 -04:00
Adrien Toupet 63defbd150 Tentative fix for artifacts without blockswap 2025-07-23 17:36:08 -04:00
Adrien Toupet d38a3956b9 Tentative fix for artifacts without blockswap 2025-07-23 17:21:45 -04:00
Adrien Toupet d87fca3ee2 Tentative fix for artifacts without blockswap 2025-07-23 17:10:59 -04:00
Adrien Toupet 5b38f1d9af Tentative fix for artifacts without blockswap 2025-07-23 17:01:36 -04:00
Adrien Toupet 2f02db68b4 Tentative fix for artifacts without blockswap 2025-07-23 16:52:51 -04:00
Adrien Toupet 468d8a29ba Tentative fix for artifacts without blockswap 2025-07-23 16:41:21 -04:00
Adrien Toupet 6feb33ac23 Tentative fix for artifacts without blockswap 2025-07-23 15:46:49 -04:00
Adrien Toupet 5fa2383b6a Added logging for artifact debugging 2025-07-23 12:10:42 -04:00
Adrien Toupet e30fce3a35 Optimize FP8 handling to preserve native FP8 weights for memory efficiency while converting only tensors during computation 2025-07-22 23:09:34 -04:00
Adrien Toupet a0bf32cb97 Reverted bf16 conversion logic - don't force it 2025-07-09 09:29:56 -04:00
Adrien Toupet 50dfb771e6 Add BlockSwap support 2025-07-04 21:28:47 -04:00
NumZ 2786c05faa Speed Up 30 to 50%, fix Memory Leak, refacto) 2025-06-30 12:46:41 +02:00