fp8 matmul for scaled models

Fp8 matmul (fp8_fast) doesn't seem feasible with unmerged LoRAs as you'd need to first upcast, then apply LoRA, then downcast back to fp8 and that is too slow. Direct adding in fp8 is also not possible since that's just not something fp8 dtypes support.
This commit is contained in:
kijai
2025-08-09 10:17:11 +03:00
parent 1757847e5f
commit 48fa904ad8
4 changed files with 65 additions and 55 deletions
+1 -1
View File
@@ -149,7 +149,7 @@ class WanVideoDiffusionForcingSampler:
gguf = model["gguf"]
transformer_options = patcher.model_options.get("transformer_options", None)
if len(patcher.patches) != 0 and transformer_options.get("linear_with_lora", False) is True:
if len(patcher.patches) != 0 and transformer_options.get("linear_patched", False) is True:
log.info(f"Using {len(patcher.patches)} LoRA weight patches for WanVideo model")
if not gguf:
convert_linear_with_lora_and_scale(transformer, patches=patcher.patches)