Release v2.5.21: fix GGUF dequant regression on MPS, eliminate CPU sync overhead on unified memory

This commit is contained in:
Adrien Toupet
2025-12-12 11:22:38 -05:00
parent f3136dd20c
commit 84abef8de0
3 changed files with 9 additions and 2 deletions
+7
View File
@@ -36,6 +36,13 @@ We're actively working on improvements and new features. To stay informed:
## 🚀 Release Notes
**2025.12.12 - Version 2.5.21**
- **🛠️ Fix: GGUF dequantization error on MPS** - Resolved shape mismatch error introduced in 2.5.20 by skipping GGUF quantized buffers in precision conversion - these must remain in packed format for on-the-fly dequantization during inference
- **🍎 MPS: Eliminate CPU sync overhead** - Skip unnecessary CPU tensor offload on Apple Silicon unified memory architecture, preventing sync stalls that caused slowdowns. Input images and output video now stay on MPS device throughout the pipeline
- **⚡ MPS: Preload text embeddings** - Load text embeddings before Phase 1 encoding to avoid sync stall at Phase 2 start, improving timing accuracy and throughput
- **🧹 MPS: Optimized model cleanup** - Skip redundant CPU movement before model deletion on unified memory
**2025.12.12 - Version 2.5.20**
- **⚡ Expanded attention backends** - Full support for Flash Attention 2 (Ampere+), Flash Attention 3 (Hopper+), SageAttention 2, and SageAttention 3 (Blackwell/RTX 50xx), with automatic fallback chains to PyTorch SDPA when unavailable *(based on PR by [@naxci1](https://github.com/naxci1) - thank you!)*
+1 -1
View File
@@ -1,7 +1,7 @@
[project]
name = "seedvr2_videoupscaler"
description = "SeedVR2 official ComfyUI integration: ByteDance-Seed's one-step diffusion-based video/image upscaling with memory-efficient inference"
version = "2.5.20"
version = "2.5.21"
authors = [
{name = "numz"},
{name = "adrientoupet"}
+1 -1
View File
@@ -4,7 +4,7 @@ Only includes constants actually used in the codebase
"""
# Version information
__version__ = "2.5.20"
__version__ = "2.5.21"
import os
import warnings