Compare commits

..
Author SHA1 Message Date
SolitaryThinker 445aaac585 update test 2025-07-07 02:36:29 -07:00
SolitaryThinker 2276ad7d51 update test 2025-07-07 00:50:04 -07:00
SolitaryThinker 9f0eacf35f fix scheduler 2025-07-06 19:02:15 -07:00
SolitaryThinker a9fe0b48c4 fix scheduler 2025-07-06 18:54:07 -07:00
SolitaryThinker 661cac1a4a move to cpu all at once 2025-07-06 18:26:56 -07:00
SolitaryThinker dd94fe6139 refactor 2025-07-06 18:05:05 -07:00
SolitaryThinker 075bc69d5e rewrite validation comm 2025-07-06 17:16:47 -07:00
37 changed files with 277 additions and 129 deletions
+1 -1
View File
@@ -8,7 +8,7 @@ It features a clean, consistent API that works across popular video models, maki
With FastVideo's optimizations, you can achieve more than 3x inference improvement compared to other systems.
<p align="center">
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start.html"><b> Quick Start</b></a> | 🤗 <a href="https://huggingface.co/FastVideo/FastHunyuan" target="_blank"><b>FastHunyuan</b></a> | 🤗 <a href="https://huggingface.co/FastVideo/FastMochi-diffusers" target="_blank"><b>FastMochi</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-38u6p1jqe-yDI1QJOCEnbtkLoaI5bjZQ" target="_blank"> <b>Slack</b> </a> |
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start.html"><b> Quick Start</b></a> | 🤗 <a href="https://huggingface.co/FastVideo/FastHunyuan" target="_blank"><b>FastHunyuan</b></a> | 🤗 <a href="https://huggingface.co/FastVideo/FastMochi-diffusers" target="_blank"><b>FastMochi</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-2zf6ru791-sRwI9lPIUJQq1mIeB_yjJg" target="_blank"> <b>Slack</b> </a> |
</p>
<div align="center">
@@ -8,14 +8,12 @@ You can easily use the FastVideo Docker image as a custom container on [RunPod](
Choose a GPU that supports CUDA 12.4
Pick 1 or 2 L40S GPU(s)
![RunPod CUDA selection](../../_static/images/runpod_cuda.png)
When creating your pod template, use this image:
```
ghcr.io/hao-ai-lab/fastvideo/fastvideo-dev:py3.12-latest
ghcr.io/hao-ai-lab/fastvideo/fastvideo-dev:latest
```
Paste Container Start Command to support SSH ([RunPod Docs](https://docs.runpod.io/pods/configuration/use-ssh)):
+1 -1
View File
@@ -117,4 +117,4 @@ If you're planning to contribute to FastVideo please see the following page:
If you encounter any issues during installation, please open an issue on our [GitHub repository](https://github.com/hao-ai-lab/FastVideo).
You can also join our [Slack community](https://join.slack.com/t/fastvideo/shared_invite/zt-38u6p1jqe-yDI1QJOCEnbtkLoaI5bjZQ) for additional support.
You can also join our [Slack community](https://join.slack.com/t/fastvideo/shared_invite/zt-2zf6ru791-sRwI9lPIUJQq1mIeB_yjJg) for additional support.
+1 -1
View File
@@ -12,7 +12,7 @@ This guide explains how to implement a custom diffusion pipeline in FastVideo, l
4. **Register Your Pipeline** - Make it discoverable by the framework
5. **Configure Your Pipeline** - (Coming soon)
Need help? Join our [Slack community](https://join.slack.com/t/fastvideo/shared_invite/zt-38u6p1jqe-yDI1QJOCEnbtkLoaI5bjZQ).
Need help? Join our [Slack community](https://join.slack.com/t/fastvideo/shared_invite/zt-2zf6ru791-sRwI9lPIUJQq1mIeB_yjJg).
## Step 1: Pipeline Modules
+3 -3
View File
@@ -27,7 +27,7 @@ fastvideo generate --help
### Hardware Configuration
- `--num-gpus {NUM_GPUS}`: Number of GPUs to use
- `--tp-size {TP_SIZE}`: Tensor parallelism size (only for the encoder, should not be larger than 1 if text encoder offload is enabled, as layerwise offload + prefetch is faster)
- `--tp-size {TP_SIZE}`: Tensor parallelism size (Typically should match the number of GPUs)
- `--sp-size {SP_SIZE}`: Sequence parallelism size (Typically should match the number of GPUs)
#### Video Configuration
@@ -68,7 +68,7 @@ Example configuration file (config.json):
"output_path": "outputs/",
"num_gpus": 2,
"sp_size": 2,
"tp_size": 1,
"tp_size": 2,
"num_frames": 45,
"height": 720,
"width": 1280,
@@ -102,7 +102,7 @@ prompt: "A beautiful woman in a red dress walking down a street"
output_path: "outputs/"
num_gpus: 2
sp_size: 2
tp_size: 1
tp_size: 2
num_frames: 45
height: 720
width: 1280
@@ -121,4 +121,4 @@ If the generated video doesn't match your prompt:
- Learn about using [Optimizations](#inference-optimizations)
- See [Examples](../examples/examples_inference_index.md) for more usage scenarios
- Join our [Community Discord](https://discord.gg/JA7cksDz86).
- Join our [Community Slack](https://join.slack.com/t/fastvideo/shared_invite/zt-38u6p1jqe-yDI1QJOCEnbtkLoaI5bjZQ).
- Join our [Community Slack](https://join.slack.com/t/fastvideo/shared_invite/zt-2zf6ru791-sRwI9lPIUJQq1mIeB_yjJg).
@@ -24,7 +24,6 @@ training_args=(
--num_height 480
--num_width 832
--num_frames 77
--enable_gradient_checkpointing_type "full"
)
# Parallel arguments
@@ -1,11 +1,13 @@
#!/bin/bash
#SBATCH --job-name=i2v
#SBATCH --partition=main
#SBATCH --qos=hao
#SBATCH --nodes=4
#SBATCH --ntasks=4
#SBATCH --ntasks-per-node=1
#SBATCH --gres=gpu:8
#SBATCH --cpus-per-task=128
#SBATCH --nodelist=fs-mbz-gpu-[100-850]
#SBATCH --mem=1440G
#SBATCH --output=i2v_output/i2v_%j.out
#SBATCH --error=i2v_output/i2v_%j.err
@@ -58,7 +60,6 @@ training_args=(
--num_height 480
--num_width 832
--num_frames 77
--enable_gradient_checkpointing_type "full"
)
# Parallel arguments
@@ -24,14 +24,13 @@ training_args=(
--num_height 480
--num_width 832
--num_frames 77
--enable_gradient_checkpointing_type "full"
)
# Parallel arguments
parallel_args=(
--num_gpus $NUM_GPUS
--sp_size 8
--tp_size 1
--tp_size 8
--hsdp_replicate_dim 1
--hsdp_shard_dim 8
)
@@ -1,11 +1,13 @@
#!/bin/bash
#SBATCH --job-name=i2v
#SBATCH --partition=main
#SBATCH --qos=hao
#SBATCH --nodes=4
#SBATCH --ntasks=4
#SBATCH --ntasks-per-node=1
#SBATCH --gres=gpu:8
#SBATCH --cpus-per-task=128
#SBATCH --nodelist=fs-mbz-gpu-[100-850]
#SBATCH --mem=1440G
#SBATCH --output=i2v_output/i2v_%j.out
#SBATCH --error=i2v_output/i2v_%j.err
@@ -58,14 +60,13 @@ training_args=(
--num_height 480
--num_width 832
--num_frames 77
--enable_gradient_checkpointing_type "full"
)
# Parallel arguments
parallel_args=(
--num_gpus $NUM_GPUS
--sp_size $NUM_GPUS
--tp_size 1
--tp_size $NUM_GPUS
--hsdp_replicate_dim $SLURM_JOB_NUM_NODES
--hsdp_shard_dim $NUM_GPUS
)
@@ -24,14 +24,13 @@ training_args=(
--num_height 480
--num_width 832
--num_frames 77
--enable_gradient_checkpointing_type "full"
)
# Parallel arguments
parallel_args=(
--num_gpus $NUM_GPUS
--sp_size $NUM_GPUS
--tp_size 1
--tp_size $NUM_GPUS
--hsdp_replicate_dim 1
--hsdp_shard_dim $NUM_GPUS
)
@@ -1,11 +1,13 @@
#!/bin/bash
#SBATCH --job-name=t2v
#SBATCH --partition=main
#SBATCH --qos=hao
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --ntasks-per-node=1
#SBATCH --gres=gpu:8
#SBATCH --cpus-per-task=128
#SBATCH --nodelist=fs-mbz-gpu-[100-850]
#SBATCH --mem=1440G
#SBATCH --output=t2v_output/t2v_%j.out
#SBATCH --error=t2v_output/t2v_%j.err
@@ -55,14 +57,13 @@ training_args=(
--num_height 480
--num_width 832
--num_frames 77
--enable_gradient_checkpointing_type "full"
)
# Parallel arguments
parallel_args=(
--num_gpus $NUM_GPUS
--sp_size 4
--tp_size 1
--tp_size 4
--hsdp_replicate_dim 2
--hsdp_shard_dim 4
)
@@ -23,6 +23,7 @@ If you only need to use the distributed environment without model parallelism,
you can skip the model parallel initialization and destruction steps.
"""
import contextlib
import gc
import os
import pickle
import weakref
@@ -322,6 +323,23 @@ class GroupCoordinator:
return input_
return self.device_communicator.gather(input_, dst, dim)
def gather_object(self, obj: Any, dst: int = 0) -> list[Any] | None:
"""Gather the input object.
NOTE: `dst` is the global rank of the destination rank.
"""
world_size = self.world_size
if self.world_size == 1:
return [obj]
gather_list = None
if dst == self.rank:
gather_list = [None] * world_size
torch.distributed.gather_object(obj,
gather_list,
dst,
group=self.cpu_group)
return gather_list
def all_to_all_4D(self,
input_: torch.Tensor,
scatter_dim: int = 2,
@@ -1015,6 +1033,15 @@ def cleanup_dist_env_and_memory(shutdown_ray: bool = False):
if shutdown_ray:
import ray # Lazy import Ray
ray.shutdown()
gc.collect()
from fastvideo.v1.platforms import current_platform
if not current_platform.is_cpu():
torch.cuda.empty_cache()
try:
torch._C._host_emptyCache()
except AttributeError:
logger.warning(
"torch._C._host_emptyCache() only available in Pytorch >=2.5")
def in_the_same_node_as(pg: ProcessGroup | StatelessProcessGroup,
@@ -6,6 +6,7 @@ This module provides a consolidated interface for generating videos using
diffusion models.
"""
import gc
import math
import os
import time
@@ -276,3 +277,5 @@ class VideoGenerator:
"""
self.executor.shutdown()
del self.executor
gc.collect()
torch.cuda.empty_cache()
+6 -1
View File
@@ -292,7 +292,7 @@ class FastVideoArgs:
assert self.sp_size != -1, "sp_size must be set for training"
if self.tp_size == -1:
self.tp_size = 1
self.tp_size = self.num_gpus
if self.sp_size == -1:
self.sp_size = self.num_gpus
if self.hsdp_shard_dim == -1:
@@ -305,6 +305,11 @@ class FastVideoArgs:
if self.num_gpus < max(self.tp_size, self.sp_size):
self.num_gpus = max(self.tp_size, self.sp_size)
if self.tp_size != self.sp_size:
raise ValueError(
f"tp_size ({self.tp_size}) must be equal to sp_size ({self.sp_size})"
)
if self.enable_torch_compile and self.num_gpus > 1:
logger.warning(
"Currently torch compile does not work with multi-gpu. Setting enable_torch_compile to False"
@@ -327,7 +327,8 @@ class ImageProcessorLoader(ComponentLoader):
"""Load the image processor based on the model path, and inference args."""
logger.info("Loading image processor from %s", model_path)
image_processor = AutoImageProcessor.from_pretrained(model_path, )
image_processor = AutoImageProcessor.from_pretrained(model_path,
use_fast=True)
logger.info("Loaded image processor: %s",
image_processor.__class__.__name__)
return image_processor
+2 -1
View File
@@ -239,6 +239,7 @@ class ParallelTiledVAE(ABC):
results = torch.cat(local_results, dim=0).contiguous()
del local_results
torch.cuda.empty_cache()
# first gather size to pad the results
local_size = torch.tensor([results.size(0)],
device=results.device,
@@ -252,7 +253,7 @@ class ParallelTiledVAE(ABC):
padded_results = torch.zeros(max_size, device=results.device)
padded_results[:results.size(0)] = results
del results
torch.cuda.empty_cache()
# Gather all results
gathered_dim_metadata = [None] * world_size
gathered_results = torch.zeros_like(padded_results).repeat(
+1 -2
View File
@@ -108,8 +108,7 @@ class DecodingStage(PipelineStage):
# Normalize image to [0, 1] range
image = (image / 2 + 0.5).clamp(0, 1)
# Convert to CPU float32 for compatibility
image = image.cpu().float()
image = image.float()
# Update batch with decoded image
batch.output = image
@@ -136,6 +136,7 @@ class EncodingStage(PipelineStage):
self.maybe_free_model_hooks()
self.vae.to("cpu")
torch.cuda.empty_cache()
return batch
@@ -5,6 +5,8 @@ Image encoding stages for I2V diffusion pipelines.
This module contains implementations of image encoding stages for diffusion pipelines.
"""
import torch
from fastvideo.v1.distributed import get_local_torch_device
from fastvideo.v1.fastvideo_args import FastVideoArgs
from fastvideo.v1.forward_context import set_forward_context
@@ -66,6 +68,7 @@ class ImageEncodingStage(PipelineStage):
if fastvideo_args.use_cpu_offload:
self.image_encoder.to('cpu')
torch.cuda.empty_cache()
return batch
@@ -105,7 +105,7 @@ def run_training():
"--num_latent_t", "8",
"--num_gpus", NUM_GPUS_PER_NODE_TRAINING,
"--sp_size", NUM_GPUS_PER_NODE_TRAINING,
"--tp_size", 1,
"--tp_size", NUM_GPUS_PER_NODE_TRAINING,
"--hsdp_replicate_dim", "1",
"--hsdp_shard_dim", NUM_GPUS_PER_NODE_TRAINING,
"--num_gpus", NUM_GPUS_PER_NODE_TRAINING,
+3 -3
View File
@@ -24,7 +24,7 @@ FastHunyuan-diffusers: {
"flow_shift": 17,
"seed": 1024,
"sp_size": 2,
"tp_size": 1,
"tp_size": 2,
"vae_sp": true,
"fps": 24
}
@@ -41,7 +41,7 @@ Wan2.1-T2V-1.3B-Diffusers: {
"flow_shift": 7.0,
"seed": 1024,
"sp_size": 2,
"tp_size": 1,
"tp_size": 2,
"vae_sp": True,
"fps": 24,
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards",
@@ -60,7 +60,7 @@ Wan2.1-I2V-14B-480P-Diffusers: {
"flow_shift": 7.0,
"seed": 1024,
"sp_size": 2,
"tp_size": 1,
"tp_size": 2,
"vae_sp": True,
"fps": 24,
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards",
@@ -33,7 +33,7 @@ HUNYUAN_PARAMS = {
"flow_shift": 17,
"seed": 1024,
"sp_size": 2,
"tp_size": 1,
"tp_size": 2,
"vae_sp": True,
"fps": 24,
}
@@ -50,7 +50,7 @@ WAN_T2V_PARAMS = {
"flow_shift": 7.0,
"seed": 1024,
"sp_size": 2,
"tp_size": 1,
"tp_size": 2,
"vae_sp": True,
"fps": 24,
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards",
@@ -69,7 +69,7 @@ WAN_I2V_PARAMS = {
"flow_shift": 7.0,
"seed": 1024,
"sp_size": 2,
"tp_size": 1,
"tp_size": 2,
"vae_sp": True,
"fps": 24,
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards",
@@ -238,7 +238,7 @@ def test_i2v_inference_similarity(prompt, ATTENTION_BACKEND, model_id):
logger.error("Failed to write SSIM results to file")
min_acceptable_ssim = 0.97
assert mean_ssim >= min_acceptable_ssim, f"SSIM value {mean_ssim} is below threshold {min_acceptable_ssim} for {model_id} with backend {ATTENTION_BACKEND}"
assert mean_ssim >= min_acceptable_ssim, f"SSIM value {mean_ssim} is below threshold {min_acceptable_ssim}"
@pytest.mark.parametrize("prompt", TEST_PROMPTS)
@pytest.mark.parametrize("ATTENTION_BACKEND", ["FLASH_ATTN", "TORCH_SDPA"])
@@ -337,5 +337,5 @@ def test_inference_similarity(prompt, ATTENTION_BACKEND, model_id):
if not success:
logger.error("Failed to write SSIM results to file")
min_acceptable_ssim = 0.93
assert mean_ssim >= min_acceptable_ssim, f"SSIM value {mean_ssim} is below threshold {min_acceptable_ssim} for {model_id} with backend {ATTENTION_BACKEND}"
min_acceptable_ssim = 0.95
assert mean_ssim >= min_acceptable_ssim, f"SSIM value {mean_ssim} is below threshold {min_acceptable_ssim}"
@@ -1 +1 @@
{"step_time":0.6983645600266755,"_wandb":{"runtime":107},"grad_norm":0.50390625,"avg_step_time":1.002151239803061,"_step":5,"validation_videos_50_steps":{"captions":false,"_type":"videos","count":8,"videos":[{"size":159131,"path":"media/videos/validation_videos_50_steps_0_dc447599dbe48350e9c9.mp4","_type":"video-file","sha256":"dc447599dbe48350e9c920f4971e1786bde580dd20d52b4aa147ae8d3dc564d6"},{"_type":"video-file","sha256":"4e283876ddfbf5a2cb6f8aca07a39f832b5f806fbefde8854bc73ed904ff20ee","size":160315,"path":"media/videos/validation_videos_50_steps_0_4e283876ddfbf5a2cb6f.mp4"},{"size":135225,"path":"media/videos/validation_videos_50_steps_0_78185c41e1935306e93c.mp4","_type":"video-file","sha256":"78185c41e1935306e93c2d416ee40b31d038abffb029cb5bfb11c2a634eb2fcf"},{"_type":"video-file","sha256":"27e9819d002d3f63c8918bbdc5bf2857b5effe0caf5d5b8374b9c590fc6432eb","size":197873,"path":"media/videos/validation_videos_50_steps_0_27e9819d002d3f63c891.mp4"},{"_type":"video-file","sha256":"46fe548e86144ca60a9396fcd15e8788d3a05c066c6819ccdd6a9041feaaec8f","size":170601,"path":"media/videos/validation_videos_50_steps_0_46fe548e86144ca60a93.mp4"},{"sha256":"91ec338774bec870b9c5c81be4330a8f0f2535124cb372e881c3703e6b65ed77","size":164462,"path":"media/videos/validation_videos_50_steps_0_91ec338774bec870b9c5.mp4","_type":"video-file"},{"_type":"video-file","sha256":"ee4e811080a619215fd7541b39203dfb69c2db4cdadbfdd7a234a2cff684f6f3","size":139435,"path":"media/videos/validation_videos_50_steps_0_ee4e811080a619215fd7.mp4"},{"sha256":"22e31e048ba5e5b9d6587d7306904d6edc6c618b8f77c3ceb05ba2d602309274","size":147072,"path":"media/videos/validation_videos_50_steps_0_22e31e048ba5e5b9d658.mp4","_type":"video-file"}]},"_timestamp":1.75118195270901e+09,"vsa_sparsity":0.05,"learning_rate":1e-05,"train_loss":0.19960195198655128,"_runtime":107.325113071}
{"step_time":0.6983645600266755,"_wandb":{"runtime":107},"grad_norm":1.39390625,"avg_step_time":1.002151239803061,"_step":5,"validation_videos_50_steps":{"captions":false,"_type":"videos","count":8,"videos":[{"size":159131,"path":"media/videos/validation_videos_50_steps_0_dc447599dbe48350e9c9.mp4","_type":"video-file","sha256":"dc447599dbe48350e9c920f4971e1786bde580dd20d52b4aa147ae8d3dc564d6"},{"_type":"video-file","sha256":"4e283876ddfbf5a2cb6f8aca07a39f832b5f806fbefde8854bc73ed904ff20ee","size":160315,"path":"media/videos/validation_videos_50_steps_0_4e283876ddfbf5a2cb6f.mp4"},{"size":135225,"path":"media/videos/validation_videos_50_steps_0_78185c41e1935306e93c.mp4","_type":"video-file","sha256":"78185c41e1935306e93c2d416ee40b31d038abffb029cb5bfb11c2a634eb2fcf"},{"_type":"video-file","sha256":"27e9819d002d3f63c8918bbdc5bf2857b5effe0caf5d5b8374b9c590fc6432eb","size":197873,"path":"media/videos/validation_videos_50_steps_0_27e9819d002d3f63c891.mp4"},{"_type":"video-file","sha256":"46fe548e86144ca60a9396fcd15e8788d3a05c066c6819ccdd6a9041feaaec8f","size":170601,"path":"media/videos/validation_videos_50_steps_0_46fe548e86144ca60a93.mp4"},{"sha256":"91ec338774bec870b9c5c81be4330a8f0f2535124cb372e881c3703e6b65ed77","size":164462,"path":"media/videos/validation_videos_50_steps_0_91ec338774bec870b9c5.mp4","_type":"video-file"},{"_type":"video-file","sha256":"ee4e811080a619215fd7541b39203dfb69c2db4cdadbfdd7a234a2cff684f6f3","size":139435,"path":"media/videos/validation_videos_50_steps_0_ee4e811080a619215fd7.mp4"},{"sha256":"22e31e048ba5e5b9d6587d7306904d6edc6c618b8f77c3ceb05ba2d602309274","size":147072,"path":"media/videos/validation_videos_50_steps_0_22e31e048ba5e5b9d658.mp4","_type":"video-file"}]},"_timestamp":1.75118195270901e+09,"vsa_sparsity":0.05,"learning_rate":1e-05,"train_loss":0.16960195198655128,"_runtime":107.325113071}
@@ -111,7 +111,7 @@ def test_distributed_training():
'avg_step_time': 1.0,
'grad_norm': 0.1,
'step_time': 1.0,
'train_loss': 0.001
'train_loss': 0.01
}
failures = []
@@ -1 +1 @@
{"step_time":5.501357046999999,"grad_norm":0.384765625,"train_loss":0.07890288904309273,"avg_step_time":5.831571423200001}
{"step_time":3.501357046999999,"grad_norm":0.384765625,"train_loss":0.05890288904309273,"avg_step_time":3.831571423200001}
@@ -43,7 +43,7 @@ def run_worker():
"--num_latent_t", "4",
"--num_gpus", "4",
"--sp_size", "4",
"--tp_size", "1",
"--tp_size", "4",
"--hsdp_replicate_dim", "1",
"--hsdp_shard_dim", "4",
"--train_sp_batch_size", "1",
@@ -121,10 +121,10 @@ def test_distributed_training():
wandb_summary = json.load(open(summary_file))
fields_and_thresholds = {
'avg_step_time': 6.0,
'avg_step_time': 3.0,
'grad_norm': 0.3,
'step_time': 6.0,
'train_loss': 0.0025
'step_time': 3.0,
'train_loss': 0.01
}
failures = []
+160 -79
View File
@@ -1,4 +1,5 @@
# SPDX-License-Identifier: Apache-2.0
import gc
import math
import os
import time
@@ -10,7 +11,6 @@ from typing import Any
import imageio
import numpy as np
import torch
import torchvision
from diffusers import FlowMatchEulerDiscreteScheduler
from diffusers.optimization import get_scheduler
from einops import rearrange
@@ -429,7 +429,8 @@ class TrainingPipeline(ComposedPipelineBase, ABC):
self.seed)
logger.info("Initialized random seeds with seed: %s", self.seed)
self.noise_scheduler = FlowMatchEulerDiscreteScheduler()
self.noise_scheduler = FlowMatchEulerDiscreteScheduler(
shift=self.training_args.pipeline_config.flow_shift, )
if self.training_args.resume_from_checkpoint:
self._resume_from_checkpoint()
@@ -594,6 +595,47 @@ class TrainingPipeline(ComposedPipelineBase, ABC):
logger.info("Starting validation")
# Setup validation
sampling_param, validation_dataloader, validation_steps = self._setup_validation(
training_args)
transformer.eval()
world_group = get_world_group()
# Process each validation step
for num_inference_steps in validation_steps:
logger.info("rank: %s: num_inference_steps: %s",
self.global_rank,
num_inference_steps,
local_main_process_only=False)
# Run inference for this step
local_videos, local_captions, final_video_shape = self._run_validation_step(
sampling_param, training_args, validation_dataloader,
num_inference_steps)
# Gather results from all ranks
all_videos_gathered = world_group.gather(local_videos, dst=0, dim=0)
# all_videos_gathered: [num_validation_videos * world_size, num_frames, height, width, 3]
all_captions_gathered = world_group.gather_object(local_captions,
dst=0)
# Log results (only on rank 0)
if self.global_rank == 0:
self._log_gathered_results(all_videos_gathered,
all_captions_gathered,
num_inference_steps, global_step,
training_args, sampling_param)
world_group.barrier()
# Re-enable gradients for training
training_args.inference_mode = False
transformer.train()
gc.collect()
torch.cuda.empty_cache()
def _setup_validation(
self, training_args) -> tuple[SamplingParam, DataLoader, list[int]]:
"""Setup validation parameters and data."""
# Create sampling parameters if not provided
sampling_param = SamplingParam.from_pretrained(training_args.model_path)
@@ -612,96 +654,135 @@ class TrainingPipeline(ComposedPipelineBase, ABC):
batch_size=None,
num_workers=0)
transformer.eval()
validation_steps = training_args.validation_sampling_steps.split(",")
validation_steps = [int(step) for step in validation_steps]
validation_steps = [step for step in validation_steps if step > 0]
# Log validation results for this step
world_group = get_world_group()
num_sp_groups = world_group.world_size // self.sp_group.world_size
# Process each validation prompt for each validation step
for num_inference_steps in validation_steps:
logger.info("rank: %s: num_inference_steps: %s",
return sampling_param, validation_dataloader, validation_steps
def _run_validation_step(
self, sampling_param, training_args, validation_dataloader,
num_inference_steps
) -> tuple[torch.Tensor, list[str], tuple[int, int, int, int, int]]:
"""Run validation inference for one step."""
step_video_tensors: list[torch.Tensor] = []
step_captions: list[str] = []
batch = None
for validation_batch in validation_dataloader:
batch = self._prepare_validation_batch(sampling_param,
training_args,
validation_batch,
num_inference_steps)
logger.info("rank: %s: rank_in_sp_group: %s, batch.prompt: %s",
self.global_rank,
num_inference_steps,
self.rank_in_sp_group,
batch.prompt,
local_main_process_only=False)
step_videos: list[np.ndarray] = []
step_captions: list[str] = []
for validation_batch in validation_dataloader:
batch = self._prepare_validation_batch(sampling_param,
training_args,
validation_batch,
num_inference_steps)
logger.info("rank: %s: rank_in_sp_group: %s, batch.prompt: %s",
self.global_rank,
self.rank_in_sp_group,
batch.prompt,
local_main_process_only=False)
assert batch.prompt is not None and isinstance(batch.prompt, str)
step_captions.append(batch.prompt)
assert batch.prompt is not None and isinstance(
batch.prompt, str)
step_captions.append(batch.prompt)
# Run validation inference
output_batch = self.validation_pipeline.forward(
batch, training_args)
samples = output_batch.output
logger.info("Samples device: %s", samples.device)
# Run validation inference
output_batch = self.validation_pipeline.forward(
batch, training_args)
samples = output_batch.output
if self.rank_in_sp_group != 0:
continue
if self.rank_in_sp_group != 0:
continue
# Process outputs
assert samples.shape[
0] == 1, "validation samples should have batch size 1"
video = rearrange(samples, "b c t h w -> b t h w c")
video = video * 255
step_video_tensors.append(video)
# Process outputs
video = rearrange(samples, "b c t h w -> t b c h w")
frames = []
for x in video:
x = torchvision.utils.make_grid(x, nrow=6)
x = x.transpose(0, 1).transpose(1, 2).squeeze(-1)
frames.append((x * 255).numpy().astype(np.uint8))
step_videos.append(frames)
# ValidationDataset will always pad the dataset so that the number
# of videos is a multiple of the number of sp groups. Each sp group
# will have the same number of videos
num_validation_videos = len(step_captions)
assert batch is not None
assert batch.height is not None
assert batch.width is not None
final_video_shape = (num_validation_videos, batch.num_frames,
batch.height, batch.width, 3)
logger.info("Final video shape: %s",
final_video_shape,
local_main_process_only=False)
# Only sp_group leaders (rank_in_sp_group == 0) need to send their
# results to global rank 0
if self.rank_in_sp_group == 0:
if self.global_rank == 0:
# Global rank 0 collects results from all sp_group leaders
all_videos = step_videos # Start with own results
all_captions = step_captions
# Collect validation results from all SP group leaders using
# all_gather_object.
# Prepare data for gathering - only SP group leaders have valid
# data, other ranks have duplicate data and so we send empty data.
if self.rank_in_sp_group == 0:
# SP group leaders contribute their data
local_videos = torch.cat(step_video_tensors, dim=0)
local_captions = step_captions
else:
# Other ranks contribute empty data
local_videos = torch.zeros(final_video_shape, device=self.device)
local_captions = []
# Receive from other sp_group leaders
for sp_group_idx in range(1, num_sp_groups):
src_rank = sp_group_idx * self.sp_world_size # Global rank of other sp_group leaders
recv_videos = world_group.recv_object(src=src_rank)
recv_captions = world_group.recv_object(src=src_rank)
all_videos.extend(recv_videos)
all_captions.extend(recv_captions)
return local_videos, local_captions, final_video_shape
video_filenames = []
for i, (video, caption) in enumerate(
zip(all_videos, all_captions, strict=True)):
os.makedirs(training_args.output_dir, exist_ok=True)
filename = os.path.join(
training_args.output_dir,
f"validation_step_{global_step}_inference_steps_{num_inference_steps}_video_{i}.mp4"
)
imageio.mimsave(filename, video, fps=sampling_param.fps)
video_filenames.append(filename)
def _log_gathered_results(self, all_videos_gathered, all_captions_gathered,
num_inference_steps, global_step, training_args,
sampling_param) -> None:
"""Process and log gathered validation results."""
assert all_videos_gathered is not None
assert all_captions_gathered is not None
assert len(all_captions_gathered) == get_world_group().world_size
logs = {
f"validation_videos_{num_inference_steps}_steps": [
wandb.Video(filename, caption=caption)
for filename, caption in zip(
video_filenames, all_captions, strict=True)
]
}
wandb.log(logs, step=global_step)
else:
# Other sp_group leaders send their results to global rank 0
world_group.send_object(step_videos, dst=0)
world_group.send_object(step_captions, dst=0)
all_videos_chunked_by_rank = all_videos_gathered.chunk(
get_world_group().world_size, dim=0)
num_validation_videos = all_videos_chunked_by_rank[0].shape[0]
assert num_validation_videos > 0, "mismatch in num_validation_videos and how many videos were gathered"
# Re-enable gradients for training
training_args.inference_mode = False
transformer.train()
# Flatten the gathered data (filter out empty contributions)
all_sp_rank_0_videos = []
all_sp_rank_0_captions = []
for idx in range(0,
get_world_group().world_size,
self.sp_group.world_size):
all_sp_rank_0_videos.append(all_videos_chunked_by_rank[idx])
all_sp_rank_0_captions.extend(all_captions_gathered[idx])
all_videos_tensor = torch.cat(all_sp_rank_0_videos, dim=0)
# all_videos_tensor: [num_validation_videos * num_sp_groups, num_frames, height, width, 3]
assert len(all_videos_tensor.shape) == 5
all_videos_tensor = all_videos_tensor.cpu()
all_videos_processed = []
for video in all_videos_tensor:
assert len(video.shape) == 4
frames = []
for frame in video:
frames.append(frame.numpy().astype(np.uint8))
all_videos_processed.append(frames)
all_captions = all_sp_rank_0_captions
assert len(all_videos_processed) == len(all_captions), (
f"mismatch in number of videos and captions: "
f"{len(all_videos_processed)} != {len(all_captions)}")
# Save videos and log to wandb
video_filenames = []
for i, (video, caption) in enumerate(
zip(all_videos_processed, all_captions, strict=True)):
os.makedirs(training_args.output_dir, exist_ok=True)
filename = os.path.join(
training_args.output_dir,
f"validation_step_{global_step}_inference_steps_{num_inference_steps}_video_{i}.mp4"
)
imageio.mimsave(filename, video, fps=sampling_param.fps)
video_filenames.append(filename)
logs = {
f"validation_videos_{num_inference_steps}_steps": [
wandb.Video(filename, caption=caption) for filename, caption in
zip(video_filenames, all_captions, strict=True)
]
}
wandb.log(logs, step=global_step)
+8
View File
@@ -1,6 +1,7 @@
# SPDX-License-Identifier: Apache-2.0
import contextlib
import faulthandler
import gc
import multiprocessing as mp
import os
import signal
@@ -68,6 +69,8 @@ class Worker:
torch.cuda.set_device(self.device)
# _check_if_gpu_supports_dtype(self.model_config.dtype)
gc.collect()
torch.cuda.empty_cache()
self.init_gpu_memory = torch.cuda.mem_get_info()[0]
os.environ["MASTER_ADDR"] = "localhost"
@@ -99,6 +102,9 @@ class Worker:
if hasattr(self, 'pipeline') and self.pipeline is not None:
# Clean up pipeline resources if needed
pass
# Release CUDA resources
if torch.cuda.is_available():
torch.cuda.empty_cache()
# Destroy the distributed environment
cleanup_dist_env_and_memory(shutdown_ray=False)
@@ -127,6 +133,8 @@ class Worker:
# Handle regular RPC calls
if method_name == 'execute_forward':
gc.collect()
torch.cuda.empty_cache()
forward_batch = recv_rpc['kwargs']['forward_batch']
fastvideo_args = recv_rpc['kwargs']['fastvideo_args']
output_batch = self.execute_forward(forward_batch,
+1 -1
View File
@@ -20,7 +20,7 @@ torchrun --nnodes 1 --nproc_per_node $NUM_GPUS\
--train_batch_size=4 \
--num_latent_t 20 \
--sp_size 4 \
--tp_size 1 \
--tp_size 4 \
--hsdp_replicate_dim 1 \
--hsdp_shard_dim 4 \
--num_gpus $NUM_GPUS \
@@ -3,10 +3,13 @@
num_gpus=4
export MODEL_BASE=FastVideo/FastHunyuan-Diffusers
export FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN
# Note that the tp_size and sp_size should be the same and equal to the number
# of GPUs. They are used for different parallel groups. sp_size is used for
# dit model and tp_size is used for encoder models.
fastvideo generate \
--model-path $MODEL_BASE \
--sp-size $num_gpus \
--tp-size 1 \
--tp-size $num_gpus \
--num-gpus $num_gpus \
--height 720 \
--width 1280 \
+4 -1
View File
@@ -4,10 +4,13 @@ num_gpus=4
export MODEL_BASE=hunyuanvideo-community/HunyuanVideo
export FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN
# export MODEL_BASE=hunyuanvideo-community/HunyuanVideo
# Note that the tp_size and sp_size should be the same and equal to the number
# of GPUs. They are used for different parallel groups. sp_size is used for
# dit model and tp_size is used for encoder models.
fastvideo generate \
--model-path $MODEL_BASE \
--sp-size $num_gpus \
--tp-size 1 \
--tp-size $num_gpus \
--num-gpus $num_gpus \
--height 720 \
--width 1280 \
@@ -5,10 +5,13 @@ export FASTVIDEO_ATTENTION_CONFIG=assets/mask_strategy_hunyuan.json
export FASTVIDEO_ATTENTION_BACKEND=SLIDING_TILE_ATTN
export MODEL_BASE=hunyuanvideo-community/HunyuanVideo
# export MODEL_BASE=hunyuanvideo-community/HunyuanVideo
# Note that the tp_size and sp_size should be the same and equal to the number
# of GPUs. They are used for different parallel groups. sp_size is used for
# dit model and tp_size is used for encoder models.
fastvideo generate \
--model-path $MODEL_BASE \
--sp-size ${num_gpus} \
--tp-size 1 \
--tp-size ${num_gpus} \
--height 768 \
--width 1280 \
--num-frames 117 \
+4 -1
View File
@@ -4,10 +4,13 @@ num_gpus=2
export FASTVIDEO_ATTENTION_BACKEND=
export MODEL_BASE=Wan-AI/Wan2.1-T2V-1.3B-Diffusers
# export MODEL_BASE=hunyuanvideo-community/HunyuanVideo
# Note that the tp_size and sp_size should be the same and equal to the number
# of GPUs. They are used for different parallel groups. sp_size is used for
# dit model and tp_size is used for encoder models.
fastvideo generate \
--model-path $MODEL_BASE \
--sp-size $num_gpus \
--tp-size 1 \
--tp-size $num_gpus \
--num-gpus $num_gpus \
--height 480 \
--width 832 \
+4 -1
View File
@@ -4,10 +4,13 @@ num_gpus=2
export FASTVIDEO_ATTENTION_CONFIG=assets/mask_strategy_wan.json
export FASTVIDEO_ATTENTION_BACKEND=SLIDING_TILE_ATTN
export MODEL_BASE=Wan-AI/Wan2.1-T2V-14B-Diffusers
# Note that the tp_size and sp_size should be the same and equal to the number
# of GPUs. They are used for different parallel groups. sp_size is used for
# dit model and tp_size is used for encoder models.
fastvideo generate \
--model-path $MODEL_BASE \
--sp-size $num_gpus \
--tp-size 1 \
--tp-size $num_gpus \
--num-gpus $num_gpus \
--height 768 \
--width 1280 \
+5 -2
View File
@@ -4,11 +4,14 @@ num_gpus=1
export FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN
# change model path to local dir if you want to inference using your checkpoint
export MODEL_BASE=Wan-AI/Wan2.1-T2V-1.3B-Diffusers
# export MODEL_BASE=hunyuanvideo-community/HunyuanVideo
# export MODEL_BASE=hunyuanvideo-community/HunyuanVideo
# Note that the tp_size and sp_size should be the same and equal to the number
# of GPUs. They are used for different parallel groups. sp_size is used for
# dit model and tp_size is used for encoder models.
fastvideo generate \
--model-path $MODEL_BASE \
--sp-size $num_gpus \
--tp-size 1 \
--tp-size $num_gpus \
--num-gpus $num_gpus \
--height 448 \
--width 832 \
@@ -4,6 +4,9 @@ num_gpus=2
export FASTVIDEO_ATTENTION_BACKEND=
export MODEL_BASE=Wan-AI/Wan2.1-I2V-14B-480P-Diffusers
# export MODEL_BASE=hunyuanvideo-community/HunyuanVideo
# Note that the tp_size and sp_size should be the same and equal to the number
# of GPUs. They are used for different parallel groups. sp_size is used for
# dit model and tp_size is used for encoder models.
fastvideo generate \
--model-path $MODEL_BASE \
--sp-size $num_gpus \