This commit introduces a new block swapping mechanism specifically for WanVideo models to enable running them on GPUs with limited VRAM.
A new `WanVideoBlockSwapManager` is implemented which uses a pre-allocation or "shell" strategy. Instead of moving entire blocks between CPU and GPU, this approach:
1. Pre-allocates a single "shell" block on the GPU, sized to match the largest block in the model.
2. Offloads designated model blocks to the CPU.
3. Patches the `forward` method of these offloaded blocks.
4. During inference, the patched method copies the weights (`state_dict`) from the CPU block into the GPU shell just before execution.
This method avoids the overhead of allocating and deallocating GPU memory for each block, reducing memory fragmentation and potentially improving performance and or corruption copying potentially modified blocks back to the swap space.
This commit introduces block swapping functionality for Qwen models, enabling them to run on systems with limited VRAM by offloading layers to a swap device (e.g., CPU RAM).
Key changes:
- A new `QwenBlockSwapManager` class is implemented to handle the patching of Qwen transformer blocks.
- The `apply_block_swap` function is extended to detect Qwen models and apply the swapping logic to their `transformer_blocks`.
- A model signature for Qwen is added to `model_sig.py` to correctly identify the swappable modules.
- A new diagnostic function, `log_unsupported_model_analysis`, is added to log the structure of unsupported models, aiding future development.