Compare commits
37
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
eb16ffb7b8 | ||
|
|
acd65e42e4 | ||
|
|
892d13d14d | ||
|
|
3a93f954df | ||
|
|
3e5ff5b583 | ||
|
|
84a6c32c96 | ||
|
|
ec9dd96d2e | ||
|
|
ed74b3c65c | ||
|
|
fce56124d7 | ||
|
|
4109928d27 | ||
|
|
d4ca37df9e | ||
|
|
9099e88e9a | ||
|
|
b679c8e515 | ||
|
|
da6003bc50 | ||
|
|
e811130464 | ||
|
|
d8a45e71c1 | ||
|
|
7b50887e38 | ||
|
|
89add12d3c | ||
|
|
d467c7cd35 | ||
|
|
88b2583c2c | ||
|
|
a730e43d5f | ||
|
|
edf116fa46 | ||
|
|
de3cefb5e5 | ||
|
|
e087e85e09 | ||
|
|
e1b998b6ef | ||
|
|
fb49c93dbc | ||
|
|
172f4802b4 | ||
|
|
24e57fafc9 | ||
|
|
6debd46482 | ||
|
|
f7dc36f7ea | ||
|
|
a0fb954f56 | ||
|
|
053106922c | ||
|
|
a57122c519 | ||
|
|
b393570e45 | ||
|
|
285635e8c0 | ||
|
|
58cfd71b5e | ||
|
|
3bf892b6ab |
@@ -0,0 +1,29 @@
|
||||
name: 🐞 Bug report
|
||||
description: Create a report to help us reproduce and fix the bug
|
||||
title: "[Bug] "
|
||||
labels: ['Bug']
|
||||
|
||||
body:
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Environment
|
||||
description: |
|
||||
Please share your environment with us. You can run the command **python fastvideo/utils/env_utils.py** and copy-paste its output below.
|
||||
placeholder: FastVideo version, platform, python version, cuda version...
|
||||
validations:
|
||||
required: true
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Describe the bug
|
||||
description: A clear and concise description of what the bug is.
|
||||
validations:
|
||||
required: true
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Reproduction
|
||||
description: |
|
||||
What command or script did you run? Which **model** are you using?
|
||||
placeholder: |
|
||||
A placeholder for the command.
|
||||
validations:
|
||||
required: true
|
||||
@@ -0,0 +1,17 @@
|
||||
name: 🚀 Feature request
|
||||
description: Suggest an idea for this project
|
||||
title: "[Feature] "
|
||||
|
||||
body:
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Motivation
|
||||
description: |
|
||||
A clear and concise description of the motivation of the feature.
|
||||
validations:
|
||||
required: true
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Related resources
|
||||
description: |
|
||||
If there is an official code release or third-party implementations, please also provide the information here, which would be very helpful.
|
||||
@@ -1,19 +1,155 @@
|
||||
prepare the src data
|
||||
```
|
||||
python scripts/download_hf.py --repo_id=FastVideo/Distill-30K-Data-Src --local_dir=data/Distill-30K-Data-Src --repo_type=dataset
|
||||
cd data/Distill-30K-Data-Src
|
||||
cat archive_name.tar.gz.* > Distill-30K-Data-Src.tar.gz
|
||||
rm archive_name.tar.gz.*
|
||||
tar --use-compress-program="pigz --processes 64" -xvf Distill-30K-Data-Src.tar.gz
|
||||
mv ephemeral/hao.zhang/codefolder/FastVideo-OSP/data/Distill-30K-Src/* .
|
||||
rm -r ephemeral
|
||||
rm Distill-30K-Data-Src.tar.gz
|
||||
cd ../..
|
||||
```
|
||||
Now the src data is stored in folder data/Distill-30K-Data-Src, adjust the path in `data/Distill-30K-Src/merge.txt` correspondingly
|
||||
<div align="center">
|
||||
<img src=assets/logo.jpg width="30%"/>
|
||||
</div>
|
||||
|
||||
Adjust the `SHARD_NUM` (total shards num) and `SHARD_IDX` (idx of the shard to preprocess) correspoindingly in `./scripts/preprocess_hunyuan_data.sh` for each GPU node to preprocess different shards
|
||||
FastVideo is a lightweight framework for accelerating large video diffusion models.
|
||||
|
||||
|
||||
https://github.com/user-attachments/assets/5fbc4596-56d6-43aa-98e0-da472cf8e26c
|
||||
|
||||
|
||||
|
||||
|
||||
<p align="center">
|
||||
🤗 <a href="https://huggingface.co/FastVideo/FastMochi-diffusers" target="_blank">FastMochi</a> | 🤗 <a href="https://huggingface.co/FastVideo/FastHunyuan" target="_blank">FastHunyuan</a> | 🎮 <a href="https://discord.gg/REBzDQTWWt" target="_blank"> Discord </a> | 🕹️ <a href="https://replicate.com/lucataco/fast-hunyuan-video" target="_blank"> Replicate </a>
|
||||
</p>
|
||||
|
||||
|
||||
FastVideo currently offers: (with more to come)
|
||||
|
||||
- FastHunyuan and FastMochi: consistency distilled video diffusion models for 8x inference speedup.
|
||||
- First open distillation recipes for video DiT, based on [PCM](https://github.com/G-U-N/Phased-Consistency-Model).
|
||||
- Support distilling/finetuning/inferencing state-of-the-art open video DiTs: 1. Mochi 2. Hunyuan.
|
||||
- Scalable training with FSDP, sequence parallelism, and selective activation checkpointing, with near linear scaling to 64 GPUs.
|
||||
- Memory efficient finetuning with LoRA, precomputed latent, and precomputed text embeddings.
|
||||
|
||||
Dev in progress and highly experimental.
|
||||
|
||||
## 🎥 More Demos
|
||||
|
||||
Fast-Hunyuan comparison with original Hunyuan, achieving an 8X diffusion speed boost with the FastVideo framework.
|
||||
|
||||
https://github.com/user-attachments/assets/064ac1d2-11ed-4a0c-955b-4d412a96ef30
|
||||
|
||||
Comparison between OpenAI Sora, original Hunyuan and FastHunyuan
|
||||
|
||||
https://github.com/user-attachments/assets/d323b712-3f68-42b2-952b-94f6a49c4836
|
||||
|
||||
Comparison between original FastHunyuan, LLM-INT8 quantized FastHunyuan and NF4 quantized FastHunyuan
|
||||
|
||||
https://github.com/user-attachments/assets/cf89efb5-5f68-4949-a085-f41c1ef26c94
|
||||
|
||||
## Change Log
|
||||
- ```2024/12/25```: Enable single 4090 inference for `FastHunyuan`, please rerun the installation steps to update the environment.
|
||||
- ```2024/12/17```: `FastVideo` v1.0 is released.
|
||||
|
||||
|
||||
## 🔧 Installation
|
||||
The code is tested on Python 3.10.0, CUDA 12.1 and H100.
|
||||
```
|
||||
bash ./scripts/preprocess_hunyuan_data.sh
|
||||
./env_setup.sh fastvideo
|
||||
```
|
||||
The preprocessed data will be stored in `data/HD-Hunyuan-30K-Distill-Data_Shard*`
|
||||
|
||||
## 🚀 Inference
|
||||
|
||||
### Inference FastHunyuan on single RTX4090
|
||||
We now support NF4 and LLM-INT8 quantized inference using BitsAndBytes for FastHunyuan. With NF4 quantization, inference can be performed on a single RTX 4090 GPU, requiring just 20GB of VRAM.
|
||||
```bash
|
||||
# Download the model weight
|
||||
python scripts/huggingface/download_hf.py --repo_id=FastVideo/FastHunyuan-diffusers --local_dir=data/FastHunyuan-diffusers --repo_type=model
|
||||
# CLI inference
|
||||
bash scripts/inference/inference_diffusers_hunyuan.sh
|
||||
```
|
||||
For more information about the VRAM requirements for BitsAndBytes quantization, please refer to the table below (timing measured on an H100 GPU):
|
||||
|
||||
|
||||
| Configuration | Memory to Init Transformer | Peak Memory After Init Pipeline (Denoise) | Diffusion Time | End-to-End Time |
|
||||
|--------------------------------|----------------------------|--------------------------------------------|----------------|-----------------|
|
||||
| BF16 + Pipeline CPU Offload | 23.883G | 33.744G | 81s | 121.5s |
|
||||
| INT8 + Pipeline CPU Offload | 13.911G | 27.979G | 88s | 116.7s |
|
||||
| NF4 + Pipeline CPU Offload | 9.453G | 19.26G | 78s | 114.5s |
|
||||
|
||||
|
||||
|
||||
For improved quality in generated videos, we recommend using a GPU with 80GB of memory to run the BF16 model with the original Hunyuan pipeline. To execute the inference, use the following section:
|
||||
|
||||
### FastHunyuan
|
||||
```bash
|
||||
# Download the model weight
|
||||
python scripts/huggingface/download_hf.py --repo_id=FastVideo/FastHunyuan --local_dir=data/FastHunyuan --repo_type=model
|
||||
# CLI inference
|
||||
bash scripts/inference/inference_hunyuan.sh
|
||||
```
|
||||
You can also inference FastHunyuan in the [official Hunyuan github](https://github.com/Tencent/HunyuanVideo).
|
||||
|
||||
### FastMochi
|
||||
|
||||
```bash
|
||||
# Download the model weight
|
||||
python scripts/huggingface/download_hf.py --repo_id=FastVideo/FastMochi-diffusers --local_dir=data/FastMochi-diffusers --repo_type=model
|
||||
# CLI inference
|
||||
bash scripts/inference/inference_mochi_sp.sh
|
||||
```
|
||||
|
||||
|
||||
## 🎯 Distill
|
||||
Our distillation recipe is based on [Phased Consistency Model](https://github.com/G-U-N/Phased-Consistency-Model). We did not find significant improvement using multi-phase distillation, so we keep the one phase setup similar to the original latent consistency model's recipe.
|
||||
We use the [MixKit](https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0/tree/main/all_mixkit) dataset for distillation. To avoid running the text encoder and VAE during training, we preprocess all data to generate text embeddings and VAE latents.
|
||||
Preprocessing instructions can be found [data_preprocess.md](docs/data_preprocess.md). For convenience, we also provide preprocessed data that can be downloaded directly using the following command:
|
||||
```bash
|
||||
python scripts/huggingface/download_hf.py --repo_id=FastVideo/HD-Mixkit-Finetune-Hunyuan --local_dir=data/HD-Mixkit-Finetune-Hunyuan --repo_type=dataset
|
||||
```
|
||||
Next, download the original model weights with:
|
||||
```bash
|
||||
python scripts/huggingface/download_hf.py --repo_id=genmo/mochi-1-preview --local_dir=data/mochi --repo_type=model # original mochi
|
||||
python scripts/huggingface/download_hf.py --repo_id=FastVideo/hunyuan --local_dir=data/hunyuan --repo_type=model # original hunyuan
|
||||
```
|
||||
To launch the distillation process, use the following commands:
|
||||
```
|
||||
bash scripts/distill/distill_mochi.sh # for mochi
|
||||
bash scripts/distill/distill_hunyuan.sh # for hunyuan
|
||||
```
|
||||
We also provide an optional script for distillation with adversarial loss, located at `fastvideo/distill_adv.py`. Although we tried adversarial loss, we did not observe significant improvements.
|
||||
## Finetune
|
||||
### ⚡ Full Finetune
|
||||
Ensure your data is prepared and preprocessed in the format specified in [data_preprocess.md](docs/data_preprocess.md). For convenience, we also provide a mochi preprocessed Black Myth Wukong data that can be downloaded directly:
|
||||
```bash
|
||||
python scripts/huggingface/download_hf.py --repo_id=FastVideo/Mochi-Black-Myth --local_dir=data/Mochi-Black-Myth --repo_type=dataset
|
||||
```
|
||||
Download the original model weights as specificed in [Distill Section](#-distill):
|
||||
|
||||
Then you can run the finetune with:
|
||||
```
|
||||
bash scripts/finetune/finetune_mochi.sh # for mochi
|
||||
```
|
||||
**Note that for finetuning, we did not tune the hyperparameters in the provided script**
|
||||
### ⚡ Lora Finetune
|
||||
Currently, we only provide Lora Finetune for Mochi model, the command for Lora Finetune is
|
||||
```
|
||||
bash scripts/finetune/finetune_mochi_lora.sh
|
||||
```
|
||||
### Minimum Hardware Requirement
|
||||
- 40 GB GPU memory each for 2 GPUs with lora
|
||||
- 30 GB GPU memory each for 2 GPUs with CPU offload and lora.
|
||||
### Finetune with Both Image and Video
|
||||
Our codebase support finetuning with both image and video.
|
||||
```bash
|
||||
bash scripts/finetune/finetune_hunyuan.sh
|
||||
bash scripts/finetune/finetune_mochi_lora_mix.sh
|
||||
```
|
||||
For Image-Video Mixture Fine-tuning, make sure to enable the --group_frame option in your script.
|
||||
|
||||
## 📑 Development Plan
|
||||
|
||||
- More distillation methods
|
||||
- [ ] Add Distribution Matching Distillation
|
||||
- More models support
|
||||
- [ ] Add CogvideoX model
|
||||
- Code update
|
||||
- [ ] fp8 support
|
||||
- [ ] faster load model and save model support
|
||||
|
||||
## Acknowledgement
|
||||
We learned and reused code from the following projects: [PCM](https://github.com/G-U-N/Phased-Consistency-Model), [diffusers](https://github.com/huggingface/diffusers), [OpenSoraPlan](https://github.com/PKU-YuanGroup/Open-Sora-Plan), and [xDiT](https://github.com/xdit-project/xDiT).
|
||||
|
||||
We thank MBZUAI and Anyscale for their support throughout this project.
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
|
Before Width: | Height: | Size: 26 MiB After Width: | Height: | Size: 22 MiB |
Binary file not shown.
|
After Width: | Height: | Size: 149 KiB |
Binary file not shown.
|
Before Width: | Height: | Size: 380 KiB |
+8
-9
@@ -1,9 +1,8 @@
|
||||
A hand enters the frame, pulling a sheet of plastic wrap over three balls of dough placed on a wooden surface. The plastic wrap is stretched to cover the dough more securely. The hand adjusts the wrap, ensuring that it is tight and smooth over the dough. The scene focuses on the hand's movements as it secures the edges of the plastic wrap. No new objects appear, and the camera remains stationary, focusing on the action of covering the dough.
|
||||
A vintage train snakes through the mountains, its plume of white steam rising dramatically against the jagged peaks. The cars glint in the late afternoon sun, their deep crimson and gold accents lending a touch of elegance. The tracks carve a precarious path along the cliffside, revealing glimpses of a roaring river far below. Inside, passengers peer out the large windows, their faces lit with awe as the landscape unfolds.
|
||||
A crowded rooftop bar buzzes with energy, the city skyline twinkling like a field of stars in the background. Strings of fairy lights hang above, casting a warm, golden glow over the scene. Groups of people gather around high tables, their laughter blending with the soft rhythm of live jazz. The aroma of freshly mixed cocktails and charred appetizers wafts through the air, mingling with the cool night breeze.
|
||||
En "The Matrix", Neo, interpretado por Keanu Reeves, personifica la lucha contra un sistema opresor a través de su icónica imagen, que incluye unos anteojos oscuros. Estos lentes no son solo un accesorio de moda; representan una barrera entre la realidad y la percepción. Al usar estos anteojos, Neo se sumerge en un mundo donde la verdad se oculta detrás de ilusiones y engaños. La oscuridad de los lentes simboliza la ignorancia y el control que las máquinas tienen sobre la humanidad, mientras que su propia búsqueda de la verdad lo lleva a descubrir sus auténticos poderes. La escena en que se los pone se convierte en un momento crucial, marcando su transformación de un simple programador a "El Elegido". Esta imagen se ha convertido en un ícono cultural, encapsulando el mensaje de que, al enfrentar la oscuridad, podemos encontrar la luz que nos guía hacia la libertad. Así, los anteojos de Neo se convierten en un símbolo de resistencia y autoconocimiento en un mundo manipulado.
|
||||
Medium close up. Low-angle shot. A woman in a 1950s retro dress sits in a diner bathed in neon light, surrounded by classic decor and lively chatter. The camera starts with a medium shot of her sitting at the counter, then slowly zooms in as she blows a shiny pink bubblegum bubble. The bubble swells dramatically before popping with a soft, playful burst. The scene is vibrant and nostalgic, evoking the fun and carefree spirit of the 1950s.
|
||||
Will Smith eats noodles.
|
||||
A short clip of the blonde woman taking a sip from her whiskey glass, her eyes locking with the camera as she smirks playfully. The background shows a group of people laughing and enjoying the party, with vibrant neon signs illuminating the space. The shot is taken in a way that conveys the feeling of a tipsy, carefree night out. The camera then zooms in on her face as she winks, creating a cheeky, flirtatious vibe.
|
||||
A superintelligent humanoid robot waking up. The robot has a sleek metallic body with futuristic design features. Its glowing red eyes are the focal point, emanating a sharp, intense light as it powers on. The scene is set in a dimly lit, high-tech laboratory filled with glowing control panels, robotic arms, and holographic screens. The setting emphasizes advanced technology and an atmosphere of mystery. The ambiance is eerie and dramatic, highlighting the moment of awakening and the robot's immense intelligence. Photorealistic style with a cinematic, dark sci-fi aesthetic. Aspect ratio: 16:9 --v 6.1
|
||||
A chimpanzee lead vocalist singing into a microphone on stage. The camera zooms in to show him singing. There is a spotlight on him.
|
||||
Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting.
|
||||
A lone hiker stands atop a towering cliff, silhouetted against the vast horizon. The rugged landscape stretches endlessly beneath, its earthy tones blending into the soft blues of the sky. The scene captures the spirit of exploration and human resilience. High angle, dynamic framing, with soft natural lighting emphasizing the grandeur of nature.
|
||||
A hand with delicate fingers picks up a bright yellow lemon from a wooden bowl filled with lemons and sprigs of mint against a peach-colored background. The hand gently tosses the lemon up and catches it, showcasing its smooth texture. A beige string bag sits beside the bowl, adding a rustic touch to the scene. Additional lemons, one halved, are scattered around the base of the bowl. The even lighting enhances the vibrant colors and creates a fresh, inviting atmosphere.
|
||||
A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with interest. The playful yet serene atmosphere is complemented by soft natural light filtering through the petals. Mid-shot, warm and cheerful tones.
|
||||
A superintelligent humanoid robot waking up. The robot has a sleek metallic body with futuristic design features. Its glowing red eyes are the focal point, emanating a sharp, intense light as it powers on. The scene is set in a dimly lit, high-tech laboratory filled with glowing control panels, robotic arms, and holographic screens. The setting emphasizes advanced technology and an atmosphere of mystery. The ambiance is eerie and dramatic, highlighting the moment of awakening and the robots immense intelligence. Photorealistic style with a cinematic, dark sci-fi aesthetic. Aspect ratio: 16:9 --v 6.1
|
||||
fox in the forest close-up quickly turned its head to the left
|
||||
Man walking his dog in the woods on a hot sunny day
|
||||
A majestic lion strides across the golden savanna, its powerful frame glistening under the warm afternoon sun. The tall grass ripples gently in the breeze, enhancing the lion's commanding presence. The tone is vibrant, embodying the raw energy of the wild. Low angle, steady tracking shot, cinematic.
|
||||
@@ -0,0 +1,24 @@
|
||||
# Configuration for Cog ⚙️
|
||||
# Reference: https://cog.run/yaml
|
||||
|
||||
build:
|
||||
gpu: true
|
||||
cuda: "12.1"
|
||||
python_version: "3.10"
|
||||
python_packages:
|
||||
- "torch==2.4.0"
|
||||
- "torchvision"
|
||||
- "ninja==1.11.1.3"
|
||||
- "transformers==4.46.1"
|
||||
- "git+https://github.com/huggingface/diffusers.git@bf64b32652a63a1865a0528a73a13652b201698b"
|
||||
- "accelerate==1.0.1"
|
||||
- "safetensors==0.4.5"
|
||||
- "peft==0.13.2"
|
||||
- "packaging==24.2"
|
||||
- "git+https://github.com/hao-ai-lab/FastVideo"
|
||||
|
||||
run:
|
||||
- FLASH_ATTENTION_SKIP_CUDA_BUILD=TRUE pip install flash-attn --no-build-isolation
|
||||
- curl -o /usr/local/bin/pget -L "https://github.com/replicate/pget/releases/latest/download/pget_$(uname -s)_$(uname -m)" && chmod +x /usr/local/bin/pget
|
||||
|
||||
predict: "predict.py:Predictor"
|
||||
@@ -1,7 +1,7 @@
|
||||
import gradio as gr
|
||||
import torch
|
||||
from fastvideo.model.pipeline_mochi import MochiPipeline
|
||||
from fastvideo.model.modeling_mochi import MochiTransformer3DModel
|
||||
from fastvideo.models.mochi_hf.pipeline_mochi import MochiPipeline
|
||||
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
|
||||
from diffusers import FlowMatchEulerDiscreteScheduler
|
||||
from diffusers.utils import export_to_video
|
||||
from fastvideo.distill.solver import PCMFMScheduler
|
||||
@@ -13,20 +13,20 @@ import argparse
|
||||
def init_args():
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--prompts", nargs="+", default=[])
|
||||
parser.add_argument("--num_frames", type=int, default=163)
|
||||
parser.add_argument("--num_frames", type=int, default=25)
|
||||
parser.add_argument("--height", type=int, default=480)
|
||||
parser.add_argument("--width", type=int, default=848)
|
||||
parser.add_argument("--num_inference_steps", type=int, default=64)
|
||||
parser.add_argument("--num_inference_steps", type=int, default=8)
|
||||
parser.add_argument("--guidance_scale", type=float, default=4.5)
|
||||
parser.add_argument("--model_path", type=str, default="data/mochi")
|
||||
parser.add_argument("--seed", type=int, default=42)
|
||||
parser.add_argument("--seed", type=int, default=12345)
|
||||
parser.add_argument("--transformer_path", type=str, default=None)
|
||||
parser.add_argument("--scheduler_type", type=str, default="euler")
|
||||
parser.add_argument("--scheduler_type", type=str, default="pcm_linear_quadratic")
|
||||
parser.add_argument("--lora_checkpoint_dir", type=str, default=None)
|
||||
parser.add_argument("--shift", type=float, default=8.0)
|
||||
parser.add_argument("--num_euler_timesteps", type=int, default=100)
|
||||
parser.add_argument("--linear_threshold", type=float, default=0.025)
|
||||
parser.add_argument("--linear_range", type=float, default=0.5)
|
||||
parser.add_argument("--num_euler_timesteps", type=int, default=50)
|
||||
parser.add_argument("--linear_threshold", type=float, default=0.1)
|
||||
parser.add_argument("--linear_range", type=float, default=0.75)
|
||||
parser.add_argument("--cpu_offload", action="store_true")
|
||||
return parser.parse_args()
|
||||
|
||||
@@ -36,11 +36,12 @@ def load_model(args):
|
||||
if args.scheduler_type == "euler":
|
||||
scheduler = FlowMatchEulerDiscreteScheduler()
|
||||
else:
|
||||
linear_quadratic = True if "linear_quadratic" in args.scheduler_type else False
|
||||
scheduler = PCMFMScheduler(
|
||||
1000,
|
||||
args.shift,
|
||||
args.num_euler_timesteps,
|
||||
False,
|
||||
linear_quadratic,
|
||||
args.linear_threshold,
|
||||
args.linear_range,
|
||||
)
|
||||
@@ -56,9 +57,9 @@ def load_model(args):
|
||||
args.model_path, transformer=transformer, scheduler=scheduler
|
||||
)
|
||||
pipe.enable_vae_tiling()
|
||||
pipe.to(device)
|
||||
if args.cpu_offload:
|
||||
pipe.enable_model_cpu_offload()
|
||||
# pipe.to(device)
|
||||
# if args.cpu_offload:
|
||||
pipe.enable_sequential_cpu_offload()
|
||||
return pipe
|
||||
|
||||
|
||||
@@ -77,8 +78,6 @@ def generate_video(
|
||||
if randomize_seed:
|
||||
seed = torch.randint(0, 1000000, (1,)).item()
|
||||
|
||||
pipe = load_model(args)
|
||||
print("load model successfully")
|
||||
generator = torch.Generator(device="cuda").manual_seed(seed)
|
||||
|
||||
if not use_negative_prompt:
|
||||
@@ -108,9 +107,10 @@ examples = [
|
||||
]
|
||||
|
||||
args = init_args()
|
||||
|
||||
pipe = load_model(args)
|
||||
print("load model successfully")
|
||||
with gr.Blocks() as demo:
|
||||
gr.Markdown("# Mochi Video Generation Demo")
|
||||
gr.Markdown("# Fastvideo Mochi Video Generation Demo")
|
||||
|
||||
with gr.Group():
|
||||
with gr.Row():
|
||||
@@ -141,19 +141,19 @@ with gr.Blocks() as demo:
|
||||
with gr.Row():
|
||||
num_frames = gr.Slider(
|
||||
label="Number of Frames",
|
||||
minimum=8,
|
||||
maximum=256,
|
||||
minimum=21,
|
||||
maximum=163,
|
||||
value=args.num_frames,
|
||||
)
|
||||
guidance_scale = gr.Slider(
|
||||
label="Guidance Scale",
|
||||
minimum=1,
|
||||
maximum=20,
|
||||
maximum=12,
|
||||
value=args.guidance_scale,
|
||||
)
|
||||
num_inference_steps = gr.Slider(
|
||||
label="Inference Steps",
|
||||
minimum=10,
|
||||
minimum=4,
|
||||
maximum=100,
|
||||
value=args.num_inference_steps,
|
||||
)
|
||||
@@ -201,4 +201,4 @@ with gr.Blocks() as demo:
|
||||
)
|
||||
|
||||
if __name__ == "__main__":
|
||||
demo.queue(max_size=20).launch(server_name="0.0.0.0", server_port=7860)
|
||||
demo.queue(max_size=20).launch(server_name="0.0.0.0", server_port=7860)
|
||||
@@ -0,0 +1,68 @@
|
||||
|
||||
|
||||
|
||||
## 🧱 Data Preprocess
|
||||
|
||||
To save GPU memory, we precompute text embeddings and VAE latents to eliminate the need to load the text encoder and VAE during training.
|
||||
|
||||
|
||||
We provide a sample dataset to help you get started. Download the source media using the following command:
|
||||
```bash
|
||||
python scripts/huggingface/download_hf.py --repo_id=FastVideo/Image-Vid-Finetune-Src --local_dir=data/Image-Vid-Finetune-Src --repo_type=dataset
|
||||
```
|
||||
To preprocess the dataset for fine-tuning or distillation, run:
|
||||
```
|
||||
bash scripts/preprocess/preprocess_mochi_data.sh # for mochi
|
||||
bash scripts/preprocess/preprocess_hunyuan_data.sh # for hunyuan
|
||||
```
|
||||
|
||||
The preprocessed dataset will be stored in `Image-Vid-Finetune-Mochi` or `Image-Vid-Finetune-HunYuan` correspondingly.
|
||||
|
||||
### Process your own dataset
|
||||
|
||||
If you wish to create your own dataset for finetuning or distillation, please structure you video dataset in the following format:
|
||||
|
||||
path_to_dataset_folder/
|
||||
├── media/
|
||||
│ ├── 0.jpg
|
||||
│ ├── 1.mp4
|
||||
│ ├── 2.jpg
|
||||
├── video2caption.json
|
||||
└── merge.txt
|
||||
|
||||
Format the JSON file as a list, where each item represents a media source:
|
||||
|
||||
For image media,
|
||||
```
|
||||
{
|
||||
"path": "0.jpg",
|
||||
"cap": ["captions"]
|
||||
}
|
||||
```
|
||||
For video media,
|
||||
```
|
||||
{
|
||||
"path": "1.mp4",
|
||||
"resolution": {
|
||||
"width": 848,
|
||||
"height": 480
|
||||
},
|
||||
"fps": 30.0,
|
||||
"duration": 6.033333333333333,
|
||||
"cap": [
|
||||
"caption"
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Use a txt file (merge.txt) to contain the source folder for media and the JSON file for meta information:
|
||||
|
||||
```
|
||||
path_to_media_source_foder,path_to_json_file
|
||||
```
|
||||
|
||||
Adjust the `DATA_MERGE_PATH` and `OUTPUT_DIR` in `scripts/preprocess/preprocess_****_data.sh` accordingly and run:
|
||||
```
|
||||
bash scripts/preprocess/preprocess_****_data.sh
|
||||
```
|
||||
The preprocessed data will be put into the `OUTPUT_DIR` and the `videos2caption.json` can be used in finetune and distill scripts.
|
||||
Executable
+10
@@ -0,0 +1,10 @@
|
||||
#!/bin/bash
|
||||
|
||||
# install torch
|
||||
pip install torch==2.5.0 torchvision --index-url https://download.pytorch.org/whl/cu121
|
||||
|
||||
# install FA2 and diffusers
|
||||
pip install packaging ninja && pip install flash-attn==2.7.0.post2 --no-build-isolation
|
||||
|
||||
# install fastvideo
|
||||
pip install -e .
|
||||
@@ -18,9 +18,7 @@ from tqdm import tqdm
|
||||
|
||||
class T5dataset(Dataset):
|
||||
def __init__(
|
||||
self,
|
||||
json_path,
|
||||
vae_debug,
|
||||
self, json_path, vae_debug,
|
||||
):
|
||||
self.json_path = json_path
|
||||
self.vae_debug = vae_debug
|
||||
@@ -67,9 +65,7 @@ def main(args):
|
||||
os.makedirs(os.path.join(args.output_dir, "prompt_embed"), exist_ok=True)
|
||||
os.makedirs(os.path.join(args.output_dir, "prompt_attention_mask"), exist_ok=True)
|
||||
|
||||
latents_json_path = os.path.join(
|
||||
args.output_dir, "videos2caption_temp.json"
|
||||
)
|
||||
latents_json_path = os.path.join(args.output_dir, "videos2caption_temp.json")
|
||||
train_dataset = T5dataset(latents_json_path, args.vae_debug)
|
||||
text_encoder = load_text_encoder(args.model_type, args.model_path, device=device)
|
||||
vae, autocast_type, fps = load_vae(args.model_type, args.model_path)
|
||||
|
||||
@@ -79,8 +79,6 @@ if __name__ == "__main__":
|
||||
parser.add_argument("--model_type", type=str, default="mochi")
|
||||
parser.add_argument("--data_merge_path", type=str, required=True)
|
||||
parser.add_argument("--num_frames", type=int, default=163)
|
||||
parser.add_argument("--shard_num", type=int)
|
||||
parser.add_argument("--shard_idx", type=int)
|
||||
parser.add_argument(
|
||||
"--dataloader_num_workers",
|
||||
type=int,
|
||||
|
||||
@@ -7,10 +7,7 @@ import random
|
||||
|
||||
class LatentDataset(Dataset):
|
||||
def __init__(
|
||||
self,
|
||||
json_path,
|
||||
num_latent_t,
|
||||
cfg_rate,
|
||||
self, json_path, num_latent_t, cfg_rate,
|
||||
):
|
||||
# data_merge_path: video_dir, latent_dir, prompt_embed_dir, json_path
|
||||
self.json_path = json_path
|
||||
@@ -46,7 +43,7 @@ class LatentDataset(Dataset):
|
||||
map_location="cpu",
|
||||
weights_only=True,
|
||||
)
|
||||
latent = latent.squeeze(0)[:, :self.num_latent_t]
|
||||
latent = latent.squeeze(0)[:, -self.num_latent_t :]
|
||||
if random.random() < self.cfg_rate:
|
||||
prompt_embed = self.uncond_prompt_embed
|
||||
prompt_attention_mask = self.uncond_prompt_mask
|
||||
@@ -110,7 +107,7 @@ def latent_collate_function(batch):
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
dataset = LatentDataset("data/Mochi-Synthetic-Data/merge.txt", num_latent_t=28)
|
||||
dataset = LatentDataset("data/HD-Mixkit-Finetune-Hunyuan/videos2caption.json", num_latent_t=8, cfg_rate=6)
|
||||
dataloader = torch.utils.data.DataLoader(
|
||||
dataset, batch_size=2, shuffle=False, collate_fn=latent_collate_function
|
||||
)
|
||||
|
||||
@@ -92,8 +92,6 @@ class T2V_dataset(Dataset):
|
||||
self.v_decoder = DecordInit()
|
||||
self.video_length_tolerance_range = args.video_length_tolerance_range
|
||||
self.support_Chinese = True
|
||||
self.shard_num = args.shard_num
|
||||
self.shard_idx = args.shard_idx
|
||||
if not ("mt5" in args.text_encoder_name):
|
||||
self.support_Chinese = False
|
||||
|
||||
@@ -342,11 +340,7 @@ class T2V_dataset(Dataset):
|
||||
for i in range(len(sub_list)):
|
||||
sub_list[i]["path"] = opj(folder, sub_list[i]["path"])
|
||||
cap_lists += sub_list
|
||||
total = len(cap_lists)
|
||||
shard_size = (total + self.shard_num - 1) // self.shard_num # Ceiling division
|
||||
start_idx = self.shard_idx * shard_size
|
||||
end_idx = min(start_idx + shard_size, total)
|
||||
return cap_lists[start_idx:end_idx]
|
||||
return cap_lists
|
||||
|
||||
def get_cap_list(self):
|
||||
cap_lists = self.read_jsons(self.data)
|
||||
|
||||
@@ -288,10 +288,7 @@ class LongSideResizeVideo:
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
size,
|
||||
skip_low_resolution=False,
|
||||
interpolation_mode="bilinear",
|
||||
self, size, skip_low_resolution=False, interpolation_mode="bilinear",
|
||||
):
|
||||
self.size = size
|
||||
self.skip_low_resolution = skip_low_resolution
|
||||
@@ -330,10 +327,7 @@ class CenterCropResizeVideo:
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
size,
|
||||
top_crop=False,
|
||||
interpolation_mode="bilinear",
|
||||
self, size, top_crop=False, interpolation_mode="bilinear",
|
||||
):
|
||||
if len(size) != 2:
|
||||
raise ValueError(
|
||||
@@ -374,9 +368,7 @@ class UCFCenterCropVideo:
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
size,
|
||||
interpolation_mode="bilinear",
|
||||
self, size, interpolation_mode="bilinear",
|
||||
):
|
||||
if isinstance(size, tuple):
|
||||
if len(size) != 2:
|
||||
@@ -413,9 +405,7 @@ class KineticsRandomCropResizeVideo:
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
size,
|
||||
interpolation_mode="bilinear",
|
||||
self, size, interpolation_mode="bilinear",
|
||||
):
|
||||
if isinstance(size, tuple):
|
||||
if len(size) != 2:
|
||||
@@ -436,9 +426,7 @@ class KineticsRandomCropResizeVideo:
|
||||
|
||||
class CenterCropVideo:
|
||||
def __init__(
|
||||
self,
|
||||
size,
|
||||
interpolation_mode="bilinear",
|
||||
self, size, interpolation_mode="bilinear",
|
||||
):
|
||||
if isinstance(size, tuple):
|
||||
if len(size) != 2:
|
||||
|
||||
+39
-57
@@ -27,10 +27,8 @@ import wandb
|
||||
from accelerate.utils import set_seed
|
||||
from tqdm.auto import tqdm
|
||||
from fastvideo.utils.fsdp_util import get_dit_fsdp_kwargs, apply_fsdp_checkpointing
|
||||
from diffusers import (
|
||||
FlowMatchEulerDiscreteScheduler,
|
||||
)
|
||||
from fastvideo.utils.load import load_transformer
|
||||
from diffusers import FlowMatchEulerDiscreteScheduler
|
||||
from fastvideo.utils.load import load_transformer
|
||||
from fastvideo.distill.solver import EulerSolver, extract_into_tensor
|
||||
from copy import deepcopy
|
||||
from diffusers.optimization import get_scheduler
|
||||
@@ -39,9 +37,7 @@ from fastvideo.dataset.latent_datasets import LatentDataset, latent_collate_func
|
||||
import torch.distributed as dist
|
||||
from safetensors.torch import save_file
|
||||
from peft import LoraConfig
|
||||
from torch.distributed.fsdp import (
|
||||
FullyShardedDataParallel as FSDP,
|
||||
)
|
||||
from torch.distributed.fsdp import FullyShardedDataParallel as FSDP
|
||||
from fastvideo.utils.checkpoint import (
|
||||
save_checkpoint,
|
||||
save_lora_checkpoint,
|
||||
@@ -131,6 +127,7 @@ def distill_one_step(
|
||||
ema_decay,
|
||||
pred_decay_weight,
|
||||
pred_decay_type,
|
||||
hunyuan_teacher_disable_cfg,
|
||||
):
|
||||
total_loss = 0.0
|
||||
optimizer.zero_grad()
|
||||
@@ -140,13 +137,6 @@ def distill_one_step(
|
||||
"absolute mean": 0.0,
|
||||
"absolute max": 0.0,
|
||||
}
|
||||
distill_cfg = distill_cfg.split(",")
|
||||
distill_cfg = [float(cfg) for cfg in distill_cfg]
|
||||
# randomly pick
|
||||
distill_cfg = distill_cfg[torch.randint(0, len(distill_cfg), (1,)).item()]
|
||||
distill_cfg = torch.tensor([distill_cfg], device=transformer.device, dtype=torch.bfloat16) * 1000
|
||||
if sp_size > 1:
|
||||
broadcast(distill_cfg)
|
||||
for _ in range(gradient_accumulation_steps):
|
||||
(
|
||||
latents,
|
||||
@@ -176,16 +166,18 @@ def distill_one_step(
|
||||
noisy_model_input = sigmas * noise + (1.0 - sigmas) * model_input
|
||||
# Predict the noise residual
|
||||
with torch.autocast("cuda", dtype=torch.bfloat16):
|
||||
kwargs = {
|
||||
teacher_kwargs = {
|
||||
"hidden_states": noisy_model_input,
|
||||
"encoder_hidden_states": encoder_hidden_states,
|
||||
"timestep": timesteps,
|
||||
"encoder_attention_mask": encoder_attention_mask, # B, L
|
||||
"return_dict": False,
|
||||
}
|
||||
if model_type == "hunyuan":
|
||||
kwargs["guidance"] = distill_cfg
|
||||
model_pred = transformer(**kwargs)[0]
|
||||
if hunyuan_teacher_disable_cfg:
|
||||
teacher_kwargs["guidance"] = torch.tensor(
|
||||
[1000.0], device=noisy_model_input.device, dtype=torch.bfloat16
|
||||
)
|
||||
model_pred = transformer(**teacher_kwargs)[0]
|
||||
|
||||
# if accelerator.is_main_process:
|
||||
model_pred, end_index = solver.euler_style_multiphase_pred(
|
||||
@@ -194,18 +186,14 @@ def distill_one_step(
|
||||
with torch.no_grad():
|
||||
w = distill_cfg
|
||||
with torch.autocast("cuda", dtype=torch.bfloat16):
|
||||
|
||||
kwargs = {
|
||||
"hidden_states": noisy_model_input,
|
||||
"encoder_hidden_states": encoder_hidden_states,
|
||||
"timestep": timesteps,
|
||||
"encoder_attention_mask": encoder_attention_mask, # B, L
|
||||
"return_dict": False,
|
||||
}
|
||||
if model_type == "hunyuan":
|
||||
kwargs["guidance"] = distill_cfg
|
||||
cond_teacher_output = teacher_transformer(**kwargs)[0].float()
|
||||
if not_apply_cfg_solver or model_type == "hunyuan":
|
||||
cond_teacher_output = teacher_transformer(
|
||||
noisy_model_input,
|
||||
encoder_hidden_states,
|
||||
timesteps,
|
||||
encoder_attention_mask, # B, L
|
||||
return_dict=False,
|
||||
)[0].float()
|
||||
if not_apply_cfg_solver:
|
||||
uncond_teacher_output = cond_teacher_output
|
||||
else:
|
||||
# Get teacher model prediction on noisy_latents and unconditional embedding
|
||||
@@ -225,19 +213,22 @@ def distill_one_step(
|
||||
# 20.4.12. Get target LCM prediction on x_prev, w, c, t_n
|
||||
with torch.no_grad():
|
||||
with torch.autocast("cuda", dtype=torch.bfloat16):
|
||||
kwargs = {
|
||||
"hidden_states": x_prev,
|
||||
"encoder_hidden_states": encoder_hidden_states,
|
||||
"timestep": timesteps_prev,
|
||||
"encoder_attention_mask": encoder_attention_mask, # B, L
|
||||
"return_dict": False,
|
||||
}
|
||||
if model_type == "hunyuan":
|
||||
kwargs["guidance"] = distill_cfg
|
||||
if ema_transformer is not None:
|
||||
target_pred = ema_transformer(**kwargs)[0]
|
||||
target_pred = ema_transformer(
|
||||
x_prev.float(),
|
||||
encoder_hidden_states,
|
||||
timesteps_prev,
|
||||
encoder_attention_mask, # B, L
|
||||
return_dict=False,
|
||||
)[0]
|
||||
else:
|
||||
target_pred = transformer(**kwargs)[0]
|
||||
target_pred = transformer(
|
||||
x_prev.float(),
|
||||
encoder_hidden_states,
|
||||
timesteps_prev,
|
||||
encoder_attention_mask, # B, L
|
||||
return_dict=False,
|
||||
)[0]
|
||||
|
||||
target, end_index = solver.euler_style_multiphase_pred(
|
||||
x_prev, target_pred, index, multiphase, True
|
||||
@@ -247,7 +238,7 @@ def distill_one_step(
|
||||
# loss = loss.mean()
|
||||
loss = (
|
||||
torch.mean(
|
||||
torch.sqrt((model_pred.float() - target.float()) ** 2 + huber_c**2)
|
||||
torch.sqrt((model_pred.float() - target.float()) ** 2 + huber_c ** 2)
|
||||
- huber_c
|
||||
)
|
||||
/ gradient_accumulation_steps
|
||||
@@ -373,19 +364,10 @@ def main(args):
|
||||
transformer._no_split_modules = no_split_modules
|
||||
fsdp_kwargs["auto_wrap_policy"] = fsdp_kwargs["auto_wrap_policy"](transformer)
|
||||
|
||||
transformer = FSDP(
|
||||
transformer,
|
||||
**fsdp_kwargs,
|
||||
)
|
||||
teacher_transformer = FSDP(
|
||||
teacher_transformer,
|
||||
**fsdp_kwargs,
|
||||
)
|
||||
transformer = FSDP(transformer, **fsdp_kwargs,)
|
||||
teacher_transformer = FSDP(teacher_transformer, **fsdp_kwargs,)
|
||||
if args.use_ema:
|
||||
ema_transformer = FSDP(
|
||||
ema_transformer,
|
||||
**fsdp_kwargs,
|
||||
)
|
||||
ema_transformer = FSDP(ema_transformer, **fsdp_kwargs,)
|
||||
main_print(f"--> model loaded")
|
||||
|
||||
if args.gradient_checkpointing:
|
||||
@@ -584,6 +566,7 @@ def main(args):
|
||||
args.ema_decay,
|
||||
args.pred_decay_weight,
|
||||
args.pred_decay_type,
|
||||
args.hunyuan_teacher_disable_cfg,
|
||||
)
|
||||
|
||||
step_time = time.time() - start_time
|
||||
@@ -676,8 +659,6 @@ if __name__ == "__main__":
|
||||
|
||||
# dataset & dataloader
|
||||
parser.add_argument("--data_json_path", type=str, required=True)
|
||||
parser.add_argument("--num_height", type=int, default=480)
|
||||
parser.add_argument("--num_width", type=int, default=848)
|
||||
parser.add_argument("--num_frames", type=int, default=163)
|
||||
parser.add_argument(
|
||||
"--dataloader_num_workers",
|
||||
@@ -886,7 +867,7 @@ if __name__ == "__main__":
|
||||
help="Whether to apply the cfg_solver.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--distill_cfg", type=str
|
||||
"--distill_cfg", type=float, default=3.0, help="Distillation coefficient."
|
||||
)
|
||||
# ["euler_linear_quadratic", "pcm", "pcm_linear_qudratic"]
|
||||
parser.add_argument(
|
||||
@@ -911,6 +892,7 @@ if __name__ == "__main__":
|
||||
parser.add_argument("--multi_phased_distill_schedule", type=str, default=None)
|
||||
parser.add_argument("--pred_decay_weight", type=float, default=0.0)
|
||||
parser.add_argument("--pred_decay_type", default="l1")
|
||||
parser.add_argument("--hunyuan_teacher_disable_cfg", action="store_true")
|
||||
parser.add_argument(
|
||||
"--master_weight_type",
|
||||
type=str,
|
||||
|
||||
@@ -24,7 +24,7 @@ logger = logging.get_logger(__name__) # pylint: disable=invalid-name
|
||||
|
||||
|
||||
class DiscriminatorHead(nn.Module):
|
||||
def __init__(self, input_channel, output_channel=1):
|
||||
def __init__(self, input_channel, output_channel=1, args=None):
|
||||
super().__init__()
|
||||
inner_channel = 1024
|
||||
self.conv1 = nn.Sequential(
|
||||
@@ -43,17 +43,21 @@ class DiscriminatorHead(nn.Module):
|
||||
)
|
||||
|
||||
self.conv_out = nn.Conv2d(inner_channel, output_channel, 1, 1, 0)
|
||||
|
||||
vae_spatial_scale_factor = 8
|
||||
|
||||
self.patch_height = args.num_height // vae_spatial_scale_factor // 2
|
||||
self.patch_width = args.num_width // vae_spatial_scale_factor // 2
|
||||
print("## DiscriminatorHead: patch_height: ", self.patch_height)
|
||||
print("## DiscriminatorHead: patch_width: ", self.patch_width)
|
||||
|
||||
def forward(self, x):
|
||||
b, twh, c = x.shape
|
||||
# from fastvideo.utils.logging_ import ForkedPdb
|
||||
# ForkedPdb().set_trace()
|
||||
after_patching_height = 49
|
||||
after_patching_wight = 80
|
||||
t = twh // (after_patching_height * after_patching_wight)
|
||||
x = x.view(-1, after_patching_height * after_patching_wight, c)
|
||||
|
||||
t = twh // (self.patch_height * self.patch_width)
|
||||
x = x.view(-1, self.patch_height * self.patch_width, c)
|
||||
x = x.permute(0, 2, 1)
|
||||
x = x.view(b * t, c, after_patching_height, after_patching_wight)
|
||||
x = x.view(b * t, c, self.patch_height, self.patch_width)
|
||||
x = self.conv1(x)
|
||||
x = self.conv2(x) + x
|
||||
x = self.conv_out(x)
|
||||
@@ -67,6 +71,7 @@ class Discriminator(nn.Module):
|
||||
num_h_per_head=1,
|
||||
adapter_channel_dims=[3072],
|
||||
total_layers = 48,
|
||||
args=None,
|
||||
):
|
||||
super().__init__()
|
||||
adapter_channel_dims = adapter_channel_dims * (total_layers // stride)
|
||||
@@ -77,7 +82,10 @@ class Discriminator(nn.Module):
|
||||
[
|
||||
nn.ModuleList(
|
||||
[
|
||||
DiscriminatorHead(adapter_channel)
|
||||
DiscriminatorHead(
|
||||
adapter_channel,
|
||||
args=args
|
||||
)
|
||||
for _ in range(self.num_h_per_head)
|
||||
]
|
||||
)
|
||||
@@ -105,3 +113,25 @@ class Discriminator(nn.Module):
|
||||
out = h(features[i])
|
||||
outputs.append(out)
|
||||
return outputs
|
||||
|
||||
|
||||
class DMDiscriminator(nn.Module):
|
||||
|
||||
def __init__(self):
|
||||
super().__init__()
|
||||
|
||||
self.cls_pred_branch = nn.Sequential(
|
||||
nn.Conv2d(kernel_size=4, in_channels=1280, out_channels=1280, stride=2, padding=1), # 8x8 -> 4x4
|
||||
nn.GroupNorm(num_groups=32, num_channels=1280),
|
||||
nn.SiLU(),
|
||||
nn.Conv2d(kernel_size=4, in_channels=1280, out_channels=1280, stride=4, padding=0), # 4x4 -> 1x1
|
||||
nn.GroupNorm(num_groups=32, num_channels=1280),
|
||||
nn.SiLU(),
|
||||
nn.Conv2d(kernel_size=1, in_channels=1280, out_channels=1, stride=1, padding=0), # 1x1 -> 1x1
|
||||
)
|
||||
|
||||
self.cls_pred_branch.requires_grad_(True)
|
||||
|
||||
def forward(self, features):
|
||||
print("## features shape: ", features.shape)
|
||||
return self.cls_pred_branch(features)
|
||||
|
||||
@@ -275,12 +275,7 @@ class EulerSolver:
|
||||
return x_prev
|
||||
|
||||
def euler_style_multiphase_pred(
|
||||
self,
|
||||
sample,
|
||||
model_pred,
|
||||
timestep_index,
|
||||
multiphase,
|
||||
is_target=False,
|
||||
self, sample, model_pred, timestep_index, multiphase, is_target=False,
|
||||
):
|
||||
inference_indices = np.linspace(
|
||||
0, len(self.euler_timesteps), num=multiphase, endpoint=False
|
||||
|
||||
@@ -983,4 +983,4 @@ if __name__ == "__main__":
|
||||
help="Weight type to use - fp32 or bf16.",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
main(args)
|
||||
main(args)
|
||||
File diff suppressed because it is too large
Load Diff
@@ -48,10 +48,7 @@ PROMPT_TEMPLATE_ENCODE_VIDEO = (
|
||||
NEGATIVE_PROMPT = "Aerial view, aerial view, overexposed, low quality, deformation, a poor composition, bad hands, bad teeth, bad eyes, bad limbs, distortion"
|
||||
|
||||
PROMPT_TEMPLATE = {
|
||||
"dit-llm-encode": {
|
||||
"template": PROMPT_TEMPLATE_ENCODE,
|
||||
"crop_start": 36,
|
||||
},
|
||||
"dit-llm-encode": {"template": PROMPT_TEMPLATE_ENCODE, "crop_start": 36,},
|
||||
"dit-llm-encode-video": {
|
||||
"template": PROMPT_TEMPLATE_ENCODE_VIDEO,
|
||||
"crop_start": 95,
|
||||
|
||||
@@ -876,8 +876,7 @@ class HunyuanVideoPipeline(DiffusionPipeline):
|
||||
|
||||
# 6. Prepare extra step kwargs. TODO: Logic should ideally just be moved out of the pipeline
|
||||
extra_step_kwargs = self.prepare_extra_func_kwargs(
|
||||
self.scheduler.step,
|
||||
{"generator": generator, "eta": eta},
|
||||
self.scheduler.step, {"generator": generator, "eta": eta},
|
||||
)
|
||||
|
||||
target_dtype = PRECISION_TO_TYPE[self.args.precision]
|
||||
|
||||
@@ -195,10 +195,7 @@ def add_denoise_schedule_args(parser: argparse.ArgumentParser):
|
||||
help="If reverse, learning/sampling from t=1 -> t=0.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--flow-solver",
|
||||
type=str,
|
||||
default="euler",
|
||||
help="Solver for flow matching.",
|
||||
"--flow-solver", type=str, default="euler", help="Solver for flow matching.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--use-linear-quadratic-schedule",
|
||||
@@ -360,16 +357,10 @@ def add_parallel_args(parser: argparse.ArgumentParser):
|
||||
|
||||
# ======================== Model loads ========================
|
||||
group.add_argument(
|
||||
"--ulysses-degree",
|
||||
type=int,
|
||||
default=1,
|
||||
help="Ulysses degree.",
|
||||
"--ulysses-degree", type=int, default=1, help="Ulysses degree.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--ring-degree",
|
||||
type=int,
|
||||
default=1,
|
||||
help="Ulysses degree.",
|
||||
"--ring-degree", type=int, default=1, help="Ulysses degree.",
|
||||
)
|
||||
|
||||
return parser
|
||||
|
||||
@@ -20,7 +20,7 @@ from fastvideo.models.hunyuan.text_encoder import TextEncoder
|
||||
from fastvideo.models.hunyuan.utils.data_utils import align_to
|
||||
from fastvideo.models.hunyuan.diffusion.schedulers import FlowMatchDiscreteScheduler
|
||||
from fastvideo.models.hunyuan.diffusion.pipelines import HunyuanVideoPipeline
|
||||
|
||||
from safetensors.torch import load_file as safetensors_load_file
|
||||
from fastvideo.utils.parallel_states import (
|
||||
initialize_sequence_parallel_state,
|
||||
nccl_info,
|
||||
@@ -56,7 +56,9 @@ class Inference(object):
|
||||
self.device = (
|
||||
device
|
||||
if device is not None
|
||||
else "cuda" if torch.cuda.is_available() else "cpu"
|
||||
else "cuda"
|
||||
if torch.cuda.is_available()
|
||||
else "cpu"
|
||||
)
|
||||
self.logger = logger
|
||||
self.parallel_args = parallel_args
|
||||
@@ -238,7 +240,16 @@ class Inference(object):
|
||||
if not model_path.exists():
|
||||
raise ValueError(f"model_path not exists: {model_path}")
|
||||
logger.info(f"Loading torch model {model_path}...")
|
||||
state_dict = torch.load(model_path, map_location=lambda storage, loc: storage)
|
||||
if model_path.suffix == ".safetensors":
|
||||
# Use safetensors library for .safetensors files
|
||||
state_dict = safetensors_load_file(model_path)
|
||||
elif model_path.suffix == ".pt":
|
||||
# Use torch for .pt files
|
||||
state_dict = torch.load(
|
||||
model_path, map_location=lambda storage, loc: storage
|
||||
)
|
||||
else:
|
||||
raise ValueError(f"Unsupported file format: {model_path}")
|
||||
|
||||
if bare_model == "unknown" and ("ema" in state_dict or "module" in state_dict):
|
||||
bare_model = False
|
||||
@@ -414,7 +425,8 @@ class HunyuanVideoSampler(Inference):
|
||||
raise ValueError(
|
||||
f"Seed must be an integer, a list of integers, or None, got {seed}."
|
||||
)
|
||||
generator = [torch.Generator(self.device).manual_seed(seed) for seed in seeds]
|
||||
# Peiyuan: using GPU seed will cause A100 and H100 to generate different results...
|
||||
generator = [torch.Generator("cpu").manual_seed(seed) for seed in seeds]
|
||||
out_dict["seeds"] = seeds
|
||||
|
||||
# ========================================================================
|
||||
|
||||
@@ -12,12 +12,7 @@ from fastvideo.models.flash_attn_no_pad import flash_attn_no_pad
|
||||
|
||||
|
||||
def attention(
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
drop_rate=0,
|
||||
attn_mask=None,
|
||||
causal=False,
|
||||
q, k, v, drop_rate=0, attn_mask=None, causal=False,
|
||||
):
|
||||
|
||||
qkv = torch.stack([q, k, v], dim=2)
|
||||
|
||||
@@ -18,9 +18,7 @@ from .modulate_layers import ModulateDiT, modulate, apply_gate
|
||||
from .token_refiner import SingleTokenRefiner
|
||||
from fastvideo.models.hunyuan.modules.posemb_layers import get_nd_rotary_pos_embed
|
||||
|
||||
from fastvideo.utils.parallel_states import (
|
||||
nccl_info,
|
||||
)
|
||||
from fastvideo.utils.parallel_states import nccl_info
|
||||
|
||||
|
||||
class MMDoubleStreamBlock(nn.Module):
|
||||
@@ -272,7 +270,7 @@ class MMSingleStreamBlock(nn.Module):
|
||||
head_dim = hidden_size // heads_num
|
||||
mlp_hidden_dim = int(hidden_size * mlp_width_ratio)
|
||||
self.mlp_hidden_dim = mlp_hidden_dim
|
||||
self.scale = qk_scale or head_dim**-0.5
|
||||
self.scale = qk_scale or head_dim ** -0.5
|
||||
|
||||
# qkv and mlp_in
|
||||
self.linear1 = nn.Linear(
|
||||
@@ -602,7 +600,7 @@ class HYVideoDiffusionTransformer(ModelMixin, ConfigMixin):
|
||||
timestep: torch.LongTensor,
|
||||
encoder_attention_mask: torch.Tensor,
|
||||
output_features=False,
|
||||
output_features_stride = 8,
|
||||
output_features_stride=8,
|
||||
attention_kwargs: Optional[Dict[str, Any]] = None,
|
||||
return_dict: bool = False,
|
||||
guidance=None,
|
||||
|
||||
@@ -135,10 +135,7 @@ class IndividualTokenRefiner(nn.Module):
|
||||
)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
x: torch.Tensor,
|
||||
c: torch.LongTensor,
|
||||
mask: Optional[torch.Tensor] = None,
|
||||
self, x: torch.Tensor, c: torch.LongTensor, mask: Optional[torch.Tensor] = None,
|
||||
):
|
||||
mask = mask.clone().bool()
|
||||
# avoid attention weight become NaN
|
||||
|
||||
@@ -10,9 +10,7 @@ def preprocess_text_encoder_tokenizer(args):
|
||||
|
||||
processor = AutoProcessor.from_pretrained(args.input_dir)
|
||||
model = LlavaForConditionalGeneration.from_pretrained(
|
||||
args.input_dir,
|
||||
torch_dtype=torch.float16,
|
||||
low_cpu_mem_usage=True,
|
||||
args.input_dir, torch_dtype=torch.float16, low_cpu_mem_usage=True,
|
||||
).to(0)
|
||||
|
||||
model.language_model.save_pretrained(f"{args.output_dir}")
|
||||
|
||||
@@ -229,54 +229,54 @@ def convert_diffusers_vae_to_mochi(state_dict):
|
||||
# Convert down_blocks
|
||||
down_block_layers = [3, 4, 6]
|
||||
for block in range(3):
|
||||
encoder_state_dict[f"layers.{block+4}.layers.0.weight"] = (
|
||||
original_state_dict.pop(f"{prefix}down_blocks.{block}.conv_in.conv.weight")
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.0.weight"
|
||||
] = original_state_dict.pop(f"{prefix}down_blocks.{block}.conv_in.conv.weight")
|
||||
encoder_state_dict[f"layers.{block+4}.layers.0.bias"] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.conv_in.conv.bias"
|
||||
)
|
||||
|
||||
for i in range(down_block_layers[block]):
|
||||
# Convert resnets
|
||||
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.0.weight"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.norm1.norm_layer.weight"
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.{i+1}.stack.0.weight"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.norm1.norm_layer.weight"
|
||||
)
|
||||
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.0.bias"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.norm1.norm_layer.bias"
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.{i+1}.stack.0.bias"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.norm1.norm_layer.bias"
|
||||
)
|
||||
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.2.weight"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.conv1.conv.weight"
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.{i+1}.stack.2.weight"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.conv1.conv.weight"
|
||||
)
|
||||
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.2.bias"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.conv1.conv.bias"
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.{i+1}.stack.2.bias"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.conv1.conv.bias"
|
||||
)
|
||||
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.3.weight"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.norm2.norm_layer.weight"
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.{i+1}.stack.3.weight"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.norm2.norm_layer.weight"
|
||||
)
|
||||
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.3.bias"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.norm2.norm_layer.bias"
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.{i+1}.stack.3.bias"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.norm2.norm_layer.bias"
|
||||
)
|
||||
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.5.weight"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.conv2.conv.weight"
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.{i+1}.stack.5.weight"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.conv2.conv.weight"
|
||||
)
|
||||
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.5.bias"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.conv2.conv.bias"
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{block+4}.layers.{i+1}.stack.5.bias"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}down_blocks.{block}.resnets.{i}.conv2.conv.bias"
|
||||
)
|
||||
|
||||
# Convert attentions
|
||||
@@ -348,18 +348,18 @@ def convert_diffusers_vae_to_mochi(state_dict):
|
||||
qkv_weight = torch.cat([q, k, v], dim=0)
|
||||
encoder_state_dict[f"layers.{i+7}.attn_block.attn.qkv.weight"] = qkv_weight
|
||||
|
||||
encoder_state_dict[f"layers.{i+7}.attn_block.attn.out.weight"] = (
|
||||
original_state_dict.pop(f"{prefix}block_out.attentions.{i}.to_out.0.weight")
|
||||
)
|
||||
encoder_state_dict[f"layers.{i+7}.attn_block.attn.out.bias"] = (
|
||||
original_state_dict.pop(f"{prefix}block_out.attentions.{i}.to_out.0.bias")
|
||||
)
|
||||
encoder_state_dict[f"layers.{i+7}.attn_block.norm.weight"] = (
|
||||
original_state_dict.pop(f"{prefix}block_out.norms.{i}.norm_layer.weight")
|
||||
)
|
||||
encoder_state_dict[f"layers.{i+7}.attn_block.norm.bias"] = (
|
||||
original_state_dict.pop(f"{prefix}block_out.norms.{i}.norm_layer.bias")
|
||||
)
|
||||
encoder_state_dict[
|
||||
f"layers.{i+7}.attn_block.attn.out.weight"
|
||||
] = original_state_dict.pop(f"{prefix}block_out.attentions.{i}.to_out.0.weight")
|
||||
encoder_state_dict[
|
||||
f"layers.{i+7}.attn_block.attn.out.bias"
|
||||
] = original_state_dict.pop(f"{prefix}block_out.attentions.{i}.to_out.0.bias")
|
||||
encoder_state_dict[
|
||||
f"layers.{i+7}.attn_block.norm.weight"
|
||||
] = original_state_dict.pop(f"{prefix}block_out.norms.{i}.norm_layer.weight")
|
||||
encoder_state_dict[
|
||||
f"layers.{i+7}.attn_block.norm.bias"
|
||||
] = original_state_dict.pop(f"{prefix}block_out.norms.{i}.norm_layer.bias")
|
||||
|
||||
# Convert output layers
|
||||
encoder_state_dict["output_norm.weight"] = original_state_dict.pop(
|
||||
@@ -413,45 +413,45 @@ def convert_diffusers_vae_to_mochi(state_dict):
|
||||
up_block_layers = [6, 4, 3]
|
||||
for block in range(3):
|
||||
for i in range(up_block_layers[block]):
|
||||
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.0.weight"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.norm1.norm_layer.weight"
|
||||
)
|
||||
decoder_state_dict[
|
||||
f"blocks.{block+1}.blocks.{i}.stack.0.weight"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.norm1.norm_layer.weight"
|
||||
)
|
||||
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.0.bias"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.norm1.norm_layer.bias"
|
||||
)
|
||||
decoder_state_dict[
|
||||
f"blocks.{block+1}.blocks.{i}.stack.0.bias"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.norm1.norm_layer.bias"
|
||||
)
|
||||
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.2.weight"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.conv1.conv.weight"
|
||||
)
|
||||
decoder_state_dict[
|
||||
f"blocks.{block+1}.blocks.{i}.stack.2.weight"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.conv1.conv.weight"
|
||||
)
|
||||
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.2.bias"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.conv1.conv.bias"
|
||||
)
|
||||
decoder_state_dict[
|
||||
f"blocks.{block+1}.blocks.{i}.stack.2.bias"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.conv1.conv.bias"
|
||||
)
|
||||
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.3.weight"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.norm2.norm_layer.weight"
|
||||
)
|
||||
decoder_state_dict[
|
||||
f"blocks.{block+1}.blocks.{i}.stack.3.weight"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.norm2.norm_layer.weight"
|
||||
)
|
||||
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.3.bias"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.norm2.norm_layer.bias"
|
||||
)
|
||||
decoder_state_dict[
|
||||
f"blocks.{block+1}.blocks.{i}.stack.3.bias"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.norm2.norm_layer.bias"
|
||||
)
|
||||
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.5.weight"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.conv2.conv.weight"
|
||||
)
|
||||
decoder_state_dict[
|
||||
f"blocks.{block+1}.blocks.{i}.stack.5.weight"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.conv2.conv.weight"
|
||||
)
|
||||
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.5.bias"] = (
|
||||
original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.conv2.conv.bias"
|
||||
)
|
||||
decoder_state_dict[
|
||||
f"blocks.{block+1}.blocks.{i}.stack.5.bias"
|
||||
] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.resnets.{i}.conv2.conv.bias"
|
||||
)
|
||||
decoder_state_dict[f"blocks.{block+1}.proj.weight"] = original_state_dict.pop(
|
||||
f"{prefix}up_blocks.{block}.proj.weight"
|
||||
|
||||
@@ -622,7 +622,7 @@ class MochiTransformer3DModel(ModelMixin, ConfigMixin, PeftAdapterMixin):
|
||||
timestep: torch.LongTensor,
|
||||
encoder_attention_mask: torch.Tensor,
|
||||
output_features=False,
|
||||
output_features_stride = 8,
|
||||
output_features_stride=8,
|
||||
attention_kwargs: Optional[Dict[str, Any]] = None,
|
||||
return_dict: bool = False,
|
||||
) -> torch.Tensor:
|
||||
|
||||
@@ -67,11 +67,7 @@ class MochiRMSNorm(nn.Module):
|
||||
|
||||
class MochiLayerNormContinuous(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
embedding_dim: int,
|
||||
conditioning_embedding_dim: int,
|
||||
eps=1e-5,
|
||||
bias=True,
|
||||
self, embedding_dim: int, conditioning_embedding_dim: int, eps=1e-5, bias=True,
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
@@ -81,9 +77,7 @@ class MochiLayerNormContinuous(nn.Module):
|
||||
self.norm = MochiModulatedRMSNorm(eps=eps)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
x: torch.Tensor,
|
||||
conditioning_embedding: torch.Tensor,
|
||||
self, x: torch.Tensor, conditioning_embedding: torch.Tensor,
|
||||
) -> torch.Tensor:
|
||||
input_dtype = x.dtype
|
||||
|
||||
|
||||
@@ -36,7 +36,7 @@ from diffusers.pipelines.mochi.pipeline_output import MochiPipelineOutput
|
||||
from einops import rearrange
|
||||
from fastvideo.utils.parallel_states import get_sequence_parallel_state, nccl_info
|
||||
from fastvideo.utils.communications import all_gather
|
||||
|
||||
from diffusers.loaders import Mochi1LoraLoaderMixin
|
||||
if is_torch_xla_available():
|
||||
import torch_xla.core.xla_model as xm
|
||||
|
||||
@@ -85,13 +85,13 @@ def linear_quadratic_schedule(num_steps, threshold_noise, linear_steps=None):
|
||||
]
|
||||
threshold_noise_step_diff = linear_steps - threshold_noise * num_steps
|
||||
quadratic_steps = num_steps - linear_steps
|
||||
quadratic_coef = threshold_noise_step_diff / (linear_steps * quadratic_steps**2)
|
||||
quadratic_coef = threshold_noise_step_diff / (linear_steps * quadratic_steps ** 2)
|
||||
linear_coef = threshold_noise / linear_steps - 2 * threshold_noise_step_diff / (
|
||||
quadratic_steps**2
|
||||
quadratic_steps ** 2
|
||||
)
|
||||
const = quadratic_coef * (linear_steps**2)
|
||||
const = quadratic_coef * (linear_steps ** 2)
|
||||
quadratic_sigma_schedule = [
|
||||
quadratic_coef * (i**2) + linear_coef * i + const
|
||||
quadratic_coef * (i ** 2) + linear_coef * i + const
|
||||
for i in range(linear_steps, num_steps)
|
||||
]
|
||||
sigma_schedule = linear_sigma_schedule + quadratic_sigma_schedule
|
||||
@@ -165,7 +165,7 @@ def retrieve_timesteps(
|
||||
return timesteps, num_inference_steps
|
||||
|
||||
|
||||
class MochiPipeline(DiffusionPipeline):
|
||||
class MochiPipeline(DiffusionPipeline, Mochi1LoraLoaderMixin):
|
||||
r"""
|
||||
The mochi pipeline for text-to-video generation.
|
||||
|
||||
@@ -502,7 +502,8 @@ class MochiPipeline(DiffusionPipeline):
|
||||
f" size of {batch_size}. Make sure the batch size matches the length of the generators."
|
||||
)
|
||||
|
||||
latents = randn_tensor(shape, generator=generator, device=device, dtype=dtype)
|
||||
latents = randn_tensor(shape, generator=generator, device=device, dtype=torch.float32)
|
||||
latents = latents.to(dtype)
|
||||
return latents
|
||||
|
||||
@property
|
||||
@@ -533,8 +534,8 @@ class MochiPipeline(DiffusionPipeline):
|
||||
negative_prompt: Optional[Union[str, List[str]]] = None,
|
||||
height: Optional[int] = None,
|
||||
width: Optional[int] = None,
|
||||
num_frames: int = 16,
|
||||
num_inference_steps: int = 28,
|
||||
num_frames: int = 19,
|
||||
num_inference_steps: int = 64,
|
||||
timesteps: List[int] = None,
|
||||
guidance_scale: float = 4.5,
|
||||
num_videos_per_prompt: Optional[int] = 1,
|
||||
@@ -711,17 +712,11 @@ class MochiPipeline(DiffusionPipeline):
|
||||
# check if of type FlowMatchEulerDiscreteScheduler
|
||||
if isinstance(self.scheduler, FlowMatchEulerDiscreteScheduler):
|
||||
timesteps, num_inference_steps = retrieve_timesteps(
|
||||
self.scheduler,
|
||||
num_inference_steps,
|
||||
device,
|
||||
timesteps,
|
||||
sigmas,
|
||||
self.scheduler, num_inference_steps, device, timesteps, sigmas,
|
||||
)
|
||||
else:
|
||||
timesteps, num_inference_steps = retrieve_timesteps(
|
||||
self.scheduler,
|
||||
num_inference_steps,
|
||||
device,
|
||||
self.scheduler, num_inference_steps, device,
|
||||
)
|
||||
num_warmup_steps = max(
|
||||
len(timesteps) - num_inference_steps * self.scheduler.order, 0
|
||||
@@ -729,6 +724,7 @@ class MochiPipeline(DiffusionPipeline):
|
||||
self._num_timesteps = len(timesteps)
|
||||
|
||||
# 6. Denoising loop
|
||||
self._progress_bar_config = {"disable": nccl_info.rank_within_group != 0}
|
||||
with self.progress_bar(total=num_inference_steps) as progress_bar:
|
||||
for i, t in enumerate(timesteps):
|
||||
if self.interrupt:
|
||||
|
||||
+92
-60
@@ -1,74 +1,105 @@
|
||||
import os
|
||||
import imageio
|
||||
import time
|
||||
from einops import rearrange
|
||||
|
||||
import torch
|
||||
import torchvision
|
||||
from diffusers import HunyuanVideoPipeline, HunyuanVideoTransformer3DModel, BitsAndBytesConfig
|
||||
import imageio as iio
|
||||
import math
|
||||
import numpy as np
|
||||
from pathlib import Path
|
||||
from loguru import logger
|
||||
from datetime import datetime
|
||||
import io
|
||||
import time
|
||||
import argparse
|
||||
from diffusers.utils import export_to_video
|
||||
import os
|
||||
|
||||
from fastvideo.models.hunyuan.utils.file_utils import save_videos_grid
|
||||
from fastvideo.models.hunyuan.inference import HunyuanVideoSampler
|
||||
def export_to_video_bytes(fps, frames):
|
||||
request = iio.core.Request("<bytes>", mode="w", extension=".mp4")
|
||||
pyavobject = iio.plugins.pyav.PyAVPlugin(request)
|
||||
if isinstance(frames, np.ndarray):
|
||||
frames = (np.array(frames) * 255).astype('uint8')
|
||||
else:
|
||||
frames = np.array(frames)
|
||||
new_bytes = pyavobject.write(frames, codec="libx264", fps=fps)
|
||||
out_bytes = io.BytesIO(new_bytes)
|
||||
return out_bytes
|
||||
|
||||
def export_to_video(frames, path, fps):
|
||||
video_bytes = export_to_video_bytes(fps, frames)
|
||||
video_bytes.seek(0)
|
||||
with open(path, "wb") as f:
|
||||
f.write(video_bytes.getbuffer())
|
||||
|
||||
def main(args):
|
||||
print(args)
|
||||
models_root_path = Path(args.model_path)
|
||||
if not models_root_path.exists():
|
||||
raise ValueError(f"`models_root` not exists: {models_root_path}")
|
||||
torch.manual_seed(args.seed)
|
||||
device = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
prompt_template = {
|
||||
"template": (
|
||||
"<|start_header_cid|>system<|end_header_id|>\n\nDescribe the video by detailing the following aspects: "
|
||||
"1. The main content and theme of the video."
|
||||
"2. The color, shape, size, texture, quantity, text, and spatial relationships of the contents, including objects, people, and anything else."
|
||||
"3. Actions, events, behaviors temporal relationships, physical movement changes of the contents."
|
||||
"4. Background environment, light, style, atmosphere, and qualities."
|
||||
"5. Camera angles, movements, and transitions used in the video."
|
||||
"6. Thematic and aesthetic concepts associated with the scene, i.e. realistic, futuristic, fairy tale, etc<|eot_id|>"
|
||||
"<|start_header_id|>user<|end_header_id|>\n\n{}<|eot_id|>"
|
||||
),
|
||||
"crop_start": 95,
|
||||
}
|
||||
|
||||
|
||||
model_id = args.model_path
|
||||
|
||||
# Create save folder to save the samples
|
||||
save_path = args.output_path
|
||||
os.makedirs(os.path.dirname(save_path), exist_ok=True)
|
||||
|
||||
# Load models
|
||||
hunyuan_video_sampler = HunyuanVideoSampler.from_pretrained(
|
||||
models_root_path, args=args
|
||||
)
|
||||
|
||||
# Get the updated args
|
||||
args = hunyuan_video_sampler.args
|
||||
|
||||
# Start sampling
|
||||
samples = []
|
||||
for prompt in args.prompts:
|
||||
outputs = hunyuan_video_sampler.predict(
|
||||
prompt=prompt,
|
||||
height=args.height,
|
||||
width=args.width,
|
||||
video_length=args.num_frames,
|
||||
seed=args.seed,
|
||||
negative_prompt=args.neg_prompt,
|
||||
infer_steps=args.num_inference_steps,
|
||||
guidance_scale=args.guidance_scale,
|
||||
num_videos_per_prompt=args.num_videos,
|
||||
flow_shift=args.flow_shift,
|
||||
batch_size=args.batch_size,
|
||||
embedded_guidance_scale=args.embedded_cfg_scale,
|
||||
if args.quantization == "nf4":
|
||||
quantization_config = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_quant_type="nf4", llm_int8_skip_modules=["proj_out", "norm_out"])
|
||||
transformer = HunyuanVideoTransformer3DModel.from_pretrained(
|
||||
model_id, subfolder="transformer/" ,torch_dtype=torch.bfloat16, quantization_config=quantization_config
|
||||
)
|
||||
samples.append(outputs["samples"][0])
|
||||
|
||||
for prompt, video in zip(args.prompts, samples):
|
||||
videos = rearrange(video.unsqueeze(0), "b c t h w -> t b c h w")
|
||||
outputs = []
|
||||
for x in videos:
|
||||
x = torchvision.utils.make_grid(x, nrow=6)
|
||||
x = x.transpose(0, 1).transpose(1, 2).squeeze(-1)
|
||||
outputs.append((x * 255).numpy().astype(np.uint8))
|
||||
os.makedirs(os.path.dirname(args.output_path), exist_ok=True)
|
||||
imageio.mimsave(args.output_path + f"{prompt[:100]}.mp4", outputs, fps=args.fps)
|
||||
|
||||
if args.quantization == "int8":
|
||||
quantization_config = BitsAndBytesConfig(load_in_8bit=True, llm_int8_skip_modules=["proj_out", "norm_out"])
|
||||
transformer = HunyuanVideoTransformer3DModel.from_pretrained(
|
||||
model_id, subfolder="transformer/" ,torch_dtype=torch.bfloat16, quantization_config=quantization_config
|
||||
)
|
||||
elif not args.quantization:
|
||||
transformer = HunyuanVideoTransformer3DModel.from_pretrained(
|
||||
model_id, subfolder="transformer/" ,torch_dtype=torch.bfloat16
|
||||
).to(device)
|
||||
|
||||
print("Max vram for read transofrmer:", round(torch.cuda.max_memory_allocated(device="cuda") / 1024 ** 3, 3), "GiB")
|
||||
torch.cuda.reset_max_memory_allocated(device)
|
||||
|
||||
if not args.cpu_offload:
|
||||
pipe = HunyuanVideoPipeline.from_pretrained(model_id, torch_dtype=torch.bfloat16).to(device)
|
||||
pipe.transformer = transformer
|
||||
else:
|
||||
pipe = HunyuanVideoPipeline.from_pretrained(model_id, transformer=transformer, torch_dtype=torch.bfloat16)
|
||||
torch.cuda.reset_max_memory_allocated(device)
|
||||
pipe.scheduler._shift = args.flow_shift
|
||||
pipe.vae.enable_tiling()
|
||||
if args.cpu_offload:
|
||||
pipe.enable_model_cpu_offload()
|
||||
print("Max vram for init pipeline:", round(torch.cuda.max_memory_allocated(device="cuda") / 1024 ** 3, 3), "GiB")
|
||||
with open(args.prompt) as f:
|
||||
prompts = f.readlines()
|
||||
|
||||
generator = torch.Generator("cpu").manual_seed(args.seed)
|
||||
os.makedirs(os.path.dirname(args.output_path), exist_ok=True)
|
||||
torch.cuda.reset_max_memory_allocated(device)
|
||||
for prompt in prompts:
|
||||
start_time = time.perf_counter()
|
||||
output = pipe(
|
||||
prompt=prompt,
|
||||
height = args.height,
|
||||
width = args.width,
|
||||
num_frames = args.num_frames,
|
||||
prompt_template=prompt_template,
|
||||
num_inference_steps = args.num_inference_steps,
|
||||
generator=generator,
|
||||
).frames[0]
|
||||
export_to_video(output, os.path.join(args.output_path, f"{prompt[:100]}.mp4"), fps=args.fps)
|
||||
print("Time:", round(time.perf_counter() - start_time, 2), "seconds")
|
||||
print("Max vram for denoise:", round(torch.cuda.max_memory_allocated(device="cuda") / 1024 ** 3, 3), "GiB")
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser()
|
||||
|
||||
# Basic parameters
|
||||
parser.add_argument("--prompts", nargs="+", default=[])
|
||||
parser.add_argument("--prompt", type=str, help="prompt file for inference")
|
||||
parser.add_argument("--num_frames", type=int, default=16)
|
||||
parser.add_argument("--height", type=int, default=256)
|
||||
parser.add_argument("--width", type=int, default=256)
|
||||
@@ -76,7 +107,8 @@ if __name__ == "__main__":
|
||||
parser.add_argument("--model_path", type=str, default="data/hunyuan")
|
||||
parser.add_argument("--output_path", type=str, default="./outputs/video")
|
||||
parser.add_argument("--fps", type=int, default=24)
|
||||
|
||||
parser.add_argument("--quantization", type=str, default=None)
|
||||
parser.add_argument("--cpu_offload", action="store_true")
|
||||
# Additional parameters
|
||||
parser.add_argument(
|
||||
"--denoise-type",
|
||||
@@ -164,7 +196,7 @@ if __name__ == "__main__":
|
||||
parser.add_argument("--model", type=str, default="HYVideo-T/2-cfgdistill")
|
||||
parser.add_argument("--latent-channels", type=int, default=16)
|
||||
parser.add_argument(
|
||||
"--precision", type=str, default="bf16", choices=["fp32", "fp16", "bf16"]
|
||||
"--precision", type=str, default="bf16", choices=["fp32", "fp16", "bf16", "fp8"]
|
||||
)
|
||||
parser.add_argument(
|
||||
"--rope-theta", type=int, default=256, help="Theta used in RoPE."
|
||||
@@ -205,4 +237,4 @@ if __name__ == "__main__":
|
||||
parser.add_argument("--text-len-2", type=int, default=77)
|
||||
|
||||
args = parser.parse_args()
|
||||
main(args)
|
||||
main(args)
|
||||
@@ -58,7 +58,11 @@ def main(args):
|
||||
|
||||
# Start sampling
|
||||
samples = []
|
||||
for prompt in args.prompts:
|
||||
|
||||
with open(args.prompt) as f:
|
||||
prompts = f.readlines()
|
||||
|
||||
for prompt in prompts:
|
||||
outputs = hunyuan_video_sampler.predict(
|
||||
prompt=prompt,
|
||||
height=args.height,
|
||||
@@ -73,24 +77,23 @@ def main(args):
|
||||
batch_size=args.batch_size,
|
||||
embedded_guidance_scale=args.embedded_cfg_scale,
|
||||
)
|
||||
samples.append(outputs["samples"][0])
|
||||
|
||||
for prompt, video in zip(args.prompts, samples):
|
||||
videos = rearrange(video.unsqueeze(0), "b c t h w -> t b c h w")
|
||||
videos = rearrange(outputs["samples"], "b c t h w -> t b c h w")
|
||||
outputs = []
|
||||
for x in videos:
|
||||
x = torchvision.utils.make_grid(x, nrow=6)
|
||||
x = x.transpose(0, 1).transpose(1, 2).squeeze(-1)
|
||||
outputs.append((x * 255).numpy().astype(np.uint8))
|
||||
os.makedirs(os.path.dirname(args.output_path), exist_ok=True)
|
||||
imageio.mimsave(args.output_path + f"{prompt[:100]}.mp4", outputs, fps=args.fps)
|
||||
imageio.mimsave(
|
||||
os.path.join(args.output_path, f"{prompt[:100]}.mp4"), outputs, fps=args.fps
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser()
|
||||
|
||||
# Basic parameters
|
||||
parser.add_argument("--prompts", nargs="+", default=[])
|
||||
parser.add_argument("--prompt", type=str, help="prompt file for inference")
|
||||
parser.add_argument("--num_frames", type=int, default=16)
|
||||
parser.add_argument("--height", type=int, default=256)
|
||||
parser.add_argument("--width", type=int, default=256)
|
||||
|
||||
@@ -39,7 +39,7 @@ def main(args):
|
||||
initialize_distributed()
|
||||
print(nccl_info.sp_size)
|
||||
device = torch.cuda.current_device()
|
||||
generator = torch.Generator(device).manual_seed(args.seed)
|
||||
# Peiyuan: GPU seed will cause A100 and H100 to produce different results .....
|
||||
weight_dtype = torch.bfloat16
|
||||
if args.scheduler_type == "euler":
|
||||
scheduler = FlowMatchEulerDiscreteScheduler()
|
||||
@@ -107,9 +107,9 @@ def main(args):
|
||||
encoder_attention_mask = None
|
||||
|
||||
if prompts is not None:
|
||||
videos = []
|
||||
with torch.autocast("cuda", dtype=torch.bfloat16):
|
||||
for prompt in prompts:
|
||||
generator = torch.Generator("cpu").manual_seed(args.seed)
|
||||
video = pipe(
|
||||
prompt=[prompt],
|
||||
height=args.height,
|
||||
@@ -119,9 +119,17 @@ def main(args):
|
||||
guidance_scale=args.guidance_scale,
|
||||
generator=generator,
|
||||
).frames
|
||||
videos.append(video[0])
|
||||
if nccl_info.global_rank <= 0:
|
||||
os.makedirs(args.output_path, exist_ok=True)
|
||||
suffix = prompt.split(".")[0]
|
||||
export_to_video(
|
||||
video[0],
|
||||
os.path.join(args.output_path, f"{suffix}.mp4"),
|
||||
fps=30,
|
||||
)
|
||||
else:
|
||||
with torch.autocast("cuda", dtype=torch.bfloat16):
|
||||
generator = torch.Generator("cpu").manual_seed(args.seed)
|
||||
videos = pipe(
|
||||
prompt_embeds=prompt_embeds,
|
||||
prompt_attention_mask=encoder_attention_mask,
|
||||
@@ -133,16 +141,7 @@ def main(args):
|
||||
generator=generator,
|
||||
).frames
|
||||
|
||||
if nccl_info.global_rank <= 0:
|
||||
if prompts is not None:
|
||||
# mkdir
|
||||
os.makedirs(args.output_path, exist_ok=True)
|
||||
for video, prompt in zip(videos, prompts):
|
||||
suffix = prompt.split(".")[0]
|
||||
export_to_video(
|
||||
video, os.path.join(args.output_path, f"{suffix}.mp4"), fps=30
|
||||
)
|
||||
else:
|
||||
if nccl_info.global_rank <= 0:
|
||||
export_to_video(videos[0], args.output_path + ".mp4", fps=30)
|
||||
|
||||
|
||||
|
||||
+26
-22
@@ -30,9 +30,8 @@ from accelerate.utils import set_seed
|
||||
from tqdm.auto import tqdm
|
||||
from fastvideo.utils.fsdp_util import get_dit_fsdp_kwargs, apply_fsdp_checkpointing
|
||||
from diffusers.utils import convert_unet_state_dict_to_peft
|
||||
from diffusers import (
|
||||
FlowMatchEulerDiscreteScheduler,
|
||||
)
|
||||
from diffusers import FlowMatchEulerDiscreteScheduler
|
||||
from fastvideo.utils.load import load_transformer
|
||||
from diffusers.optimization import get_scheduler
|
||||
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
|
||||
from diffusers.utils import check_min_version
|
||||
@@ -40,9 +39,7 @@ from fastvideo.dataset.latent_datasets import LatentDataset, latent_collate_func
|
||||
import torch.distributed as dist
|
||||
from safetensors.torch import save_file, load_file
|
||||
from peft import LoraConfig, get_peft_model_state_dict, set_peft_model_state_dict
|
||||
from torch.distributed.fsdp import (
|
||||
FullyShardedDataParallel as FSDP,
|
||||
)
|
||||
from torch.distributed.fsdp import FullyShardedDataParallel as FSDP
|
||||
from fastvideo.utils.checkpoint import (
|
||||
save_checkpoint,
|
||||
save_lora_checkpoint,
|
||||
@@ -102,8 +99,9 @@ def get_sigmas(noise_scheduler, device, timesteps, n_dim=4, dtype=torch.float32)
|
||||
return sigma
|
||||
|
||||
|
||||
def train_one_step_mochi(
|
||||
def train_one_step(
|
||||
transformer,
|
||||
model_type,
|
||||
optimizer,
|
||||
lr_scheduler,
|
||||
loader,
|
||||
@@ -127,7 +125,7 @@ def train_one_step_mochi(
|
||||
latents_attention_mask,
|
||||
encoder_attention_mask,
|
||||
) = next(loader)
|
||||
latents = normalize_mochi_dit_input(latents)
|
||||
latents = normalize_dit_input(model_type, latents)
|
||||
batch_size = latents.shape[0]
|
||||
noise = torch.randn_like(latents)
|
||||
u = compute_density_for_timestep_sampling(
|
||||
@@ -218,15 +216,15 @@ def main(args):
|
||||
|
||||
main_print(f"--> loading model from {args.pretrained_model_name_or_path}")
|
||||
# keep the master weight to float32
|
||||
transformer = MochiTransformer3DModel.from_pretrained(
|
||||
transformer = load_transformer(
|
||||
args.model_type,
|
||||
args.dit_model_name_or_path,
|
||||
args.pretrained_model_name_or_path,
|
||||
subfolder="transformer",
|
||||
torch_dtype=(
|
||||
torch.float32 if args.master_weight_type == "fp32" else torch.bfloat16
|
||||
),
|
||||
torch.float32 if args.master_weight_type == "fp32" else torch.bfloat16,
|
||||
)
|
||||
|
||||
if args.use_lora:
|
||||
assert args.model_type == "mochi", "LoRA is only supported for Mochi model."
|
||||
transformer.requires_grad_(False)
|
||||
transformer_lora_config = LoraConfig(
|
||||
r=args.lora_rank,
|
||||
@@ -264,7 +262,8 @@ def main(args):
|
||||
main_print(
|
||||
f"--> Initializing FSDP with sharding strategy: {args.fsdp_sharding_startegy}"
|
||||
)
|
||||
fsdp_kwargs = get_dit_fsdp_kwargs(
|
||||
fsdp_kwargs, no_split_modules = get_dit_fsdp_kwargs(
|
||||
transformer,
|
||||
args.fsdp_sharding_startegy,
|
||||
args.use_lora,
|
||||
args.use_cpu_offload,
|
||||
@@ -275,17 +274,18 @@ def main(args):
|
||||
transformer.config.lora_rank = args.lora_rank
|
||||
transformer.config.lora_alpha = args.lora_alpha
|
||||
transformer.config.lora_target_modules = ["to_k", "to_q", "to_v", "to_out.0"]
|
||||
transformer._no_split_modules = ["MochiTransformerBlock"]
|
||||
transformer._no_split_modules = [
|
||||
no_split_module.__name__ for no_split_module in no_split_modules
|
||||
]
|
||||
fsdp_kwargs["auto_wrap_policy"] = fsdp_kwargs["auto_wrap_policy"](transformer)
|
||||
|
||||
transformer = FSDP(
|
||||
transformer,
|
||||
**fsdp_kwargs,
|
||||
)
|
||||
transformer = FSDP(transformer, **fsdp_kwargs,)
|
||||
main_print(f"--> model loaded")
|
||||
|
||||
if args.gradient_checkpointing:
|
||||
apply_fsdp_checkpointing(transformer, args.selective_checkpointing)
|
||||
apply_fsdp_checkpointing(
|
||||
transformer, no_split_modules, args.selective_checkpointing
|
||||
)
|
||||
|
||||
# Set model as trainable.
|
||||
transformer.train()
|
||||
@@ -411,8 +411,9 @@ def main(args):
|
||||
next(loader)
|
||||
for step in range(init_steps + 1, args.max_train_steps + 1):
|
||||
start_time = time.time()
|
||||
loss, grad_norm = train_one_step_mochi(
|
||||
loss, grad_norm = train_one_step(
|
||||
transformer,
|
||||
args.model_type,
|
||||
optimizer,
|
||||
lr_scheduler,
|
||||
loader,
|
||||
@@ -479,7 +480,9 @@ def main(args):
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser()
|
||||
|
||||
parser.add_argument(
|
||||
"--model_type", type=str, default="mochi", help="The type of model to train."
|
||||
)
|
||||
# dataset & dataloader
|
||||
parser.add_argument("--data_json_path", type=str, required=True)
|
||||
parser.add_argument("--num_frames", type=int, default=163)
|
||||
@@ -503,6 +506,7 @@ if __name__ == "__main__":
|
||||
|
||||
# text encoder & vae & diffusion model
|
||||
parser.add_argument("--pretrained_model_name_or_path", type=str)
|
||||
parser.add_argument("--dit_model_name_or_path", type=str, default=None)
|
||||
parser.add_argument("--cache_dir", type=str, default="./cache_dir")
|
||||
|
||||
# diffusion setting
|
||||
|
||||
@@ -28,10 +28,7 @@ def save_checkpoint(model, optimizer, rank, output_dir, step, discriminator=Fals
|
||||
FullOptimStateDictConfig(offload_to_cpu=True, rank0_only=True),
|
||||
):
|
||||
cpu_state = model.state_dict()
|
||||
optim_state = FSDP.optim_state_dict(
|
||||
model,
|
||||
optimizer,
|
||||
)
|
||||
optim_state = FSDP.optim_state_dict(model, optimizer,)
|
||||
|
||||
# todo move to get_state_dict
|
||||
save_dir = os.path.join(output_dir, f"checkpoint-{step}")
|
||||
@@ -52,16 +49,10 @@ def save_checkpoint(model, optimizer, rank, output_dir, step, discriminator=Fals
|
||||
save_file(cpu_state, weight_path)
|
||||
optimizer_path = os.path.join(save_dir, "discriminator_optimizer.pt")
|
||||
torch.save(optim_state, optimizer_path)
|
||||
|
||||
|
||||
|
||||
def save_checkpoint_generator_discriminator(
|
||||
model,
|
||||
optimizer,
|
||||
discriminator,
|
||||
discriminator_optimizer,
|
||||
rank,
|
||||
output_dir,
|
||||
step,
|
||||
model, optimizer, discriminator, discriminator_optimizer, rank, output_dir, step,
|
||||
):
|
||||
with FSDP.state_dict_type(
|
||||
model,
|
||||
@@ -192,6 +183,36 @@ def resume_training_generator_discriminator(
|
||||
)
|
||||
return model, optimizer, discriminator, discriminator_optimizer, step
|
||||
|
||||
def resume_training_generator_fake_transformer(
|
||||
model,
|
||||
optimizer,
|
||||
fake_transformer,
|
||||
guidance_optimizer,
|
||||
discriminator,
|
||||
discriminator_optimizer,
|
||||
checkpoint_dir,
|
||||
rank,
|
||||
):
|
||||
step = int(checkpoint_dir.split("-")[-1])
|
||||
model_weight_dir = os.path.join(checkpoint_dir, "model_weights_state")
|
||||
model_optimizer_dir = os.path.join(checkpoint_dir, "model_optimizer_state")
|
||||
fake_model_weight_dir = os.path.join(checkpoint_dir, "fake_model_weights_state")
|
||||
fake_model_optimizer_dir = os.path.join(checkpoint_dir, "fake_model_optimizer_state")
|
||||
model, optimizer = load_sharded_model(
|
||||
model, optimizer, model_weight_dir, model_optimizer_dir
|
||||
)
|
||||
fake_transformer, guidance_optimizer = load_sharded_model(
|
||||
fake_transformer, guidance_optimizer, fake_model_weight_dir, fake_model_optimizer_dir
|
||||
)
|
||||
discriminator_ckpt_file = os.path.join(
|
||||
checkpoint_dir, "discriminator_fsdp_state", "discriminator_state.pt"
|
||||
)
|
||||
discriminator, discriminator_optimizer = load_full_state_model(
|
||||
discriminator, discriminator_optimizer, discriminator_ckpt_file, rank
|
||||
)
|
||||
return model, optimizer, fake_transformer, guidance_optimizer, discriminator, discriminator_optimizer, step
|
||||
|
||||
|
||||
|
||||
def resume_training(model, optimizer, checkpoint_dir, discriminator=False):
|
||||
weight_path = os.path.join(checkpoint_dir, "diffusion_pytorch_model.safetensors")
|
||||
@@ -230,10 +251,7 @@ def save_lora_checkpoint(transformer, optimizer, rank, output_dir, step):
|
||||
FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
|
||||
):
|
||||
full_state_dict = transformer.state_dict()
|
||||
lora_optim_state = FSDP.optim_state_dict(
|
||||
transformer,
|
||||
optimizer,
|
||||
)
|
||||
lora_optim_state = FSDP.optim_state_dict(transformer, optimizer,)
|
||||
|
||||
if rank <= 0:
|
||||
save_dir = os.path.join(output_dir, f"lora-checkpoint-{step}")
|
||||
|
||||
@@ -132,9 +132,7 @@ class SeqAllToAll4D(torch.autograd.Function):
|
||||
|
||||
|
||||
def all_to_all_4D(
|
||||
input_: torch.Tensor,
|
||||
scatter_dim: int = 2,
|
||||
gather_dim: int = 1,
|
||||
input_: torch.Tensor, scatter_dim: int = 2, gather_dim: int = 1,
|
||||
):
|
||||
return SeqAllToAll4D.apply(nccl_info.group, input_, scatter_dim, gather_dim)
|
||||
|
||||
@@ -193,9 +191,7 @@ class _AllToAll(torch.autograd.Function):
|
||||
|
||||
|
||||
def all_to_all(
|
||||
input_: torch.Tensor,
|
||||
scatter_dim: int = 2,
|
||||
gather_dim: int = 1,
|
||||
input_: torch.Tensor, scatter_dim: int = 2, gather_dim: int = 1,
|
||||
):
|
||||
return _AllToAll.apply(input_, nccl_info.group, scatter_dim, gather_dim)
|
||||
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
import platform
|
||||
|
||||
import accelerate
|
||||
import peft
|
||||
import torch
|
||||
import transformers
|
||||
from transformers.utils import is_torch_cuda_available, is_torch_npu_available
|
||||
|
||||
|
||||
VERSION = "1.2.0"
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
info = {
|
||||
"FastVideo version": VERSION,
|
||||
"Platform": platform.platform(),
|
||||
"Python version": platform.python_version(),
|
||||
"PyTorch version": torch.__version__,
|
||||
"Transformers version": transformers.__version__,
|
||||
"Accelerate version": accelerate.__version__,
|
||||
"PEFT version": peft.__version__,
|
||||
}
|
||||
|
||||
if is_torch_cuda_available():
|
||||
info["PyTorch version"] += " (GPU)"
|
||||
info["GPU type"] = torch.cuda.get_device_name()
|
||||
|
||||
if is_torch_npu_available():
|
||||
info["PyTorch version"] += " (NPU)"
|
||||
info["NPU type"] = torch.npu.get_device_name()
|
||||
info["CANN version"] = torch.version.cann
|
||||
|
||||
try:
|
||||
import bitsandbytes
|
||||
|
||||
info["Bitsandbytes version"] = bitsandbytes.__version__
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
print("\n" + "\n".join([f"- {key}: {value}" for key, value in info.items()]) + "\n")
|
||||
@@ -29,8 +29,7 @@ import functools
|
||||
|
||||
|
||||
non_reentrant_wrapper = partial(
|
||||
checkpoint_wrapper,
|
||||
checkpoint_impl=CheckpointImpl.NO_REENTRANT,
|
||||
checkpoint_wrapper, checkpoint_impl=CheckpointImpl.NO_REENTRANT,
|
||||
)
|
||||
|
||||
check_fn = lambda submodule: isinstance(submodule, MochiTransformerBlock)
|
||||
@@ -91,8 +90,7 @@ def get_dit_fsdp_kwargs(
|
||||
auto_wrap_policy = fsdp_auto_wrap_policy
|
||||
else:
|
||||
auto_wrap_policy = functools.partial(
|
||||
transformer_auto_wrap_policy,
|
||||
transformer_layer_cls=no_split_modules,
|
||||
transformer_auto_wrap_policy, transformer_layer_cls=no_split_modules,
|
||||
)
|
||||
|
||||
# we use float32 for fsdp but autocast during training
|
||||
|
||||
+6
-10
@@ -49,10 +49,7 @@ PROMPT_TEMPLATE_ENCODE_VIDEO = (
|
||||
NEGATIVE_PROMPT = "Aerial view, aerial view, overexposed, low quality, deformation, a poor composition, bad hands, bad teeth, bad eyes, bad limbs, distortion"
|
||||
|
||||
PROMPT_TEMPLATE = {
|
||||
"dit-llm-encode": {
|
||||
"template": PROMPT_TEMPLATE_ENCODE,
|
||||
"crop_start": 36,
|
||||
},
|
||||
"dit-llm-encode": {"template": PROMPT_TEMPLATE_ENCODE, "crop_start": 36,},
|
||||
"dit-llm-encode-video": {
|
||||
"template": PROMPT_TEMPLATE_ENCODE_VIDEO,
|
||||
"crop_start": 95,
|
||||
@@ -180,13 +177,13 @@ class MochiTextEncoderWrapper(nn.Module):
|
||||
os.path.join(pretrained_model_name_or_path, "text_encoder")
|
||||
).to(device)
|
||||
self.tokenizer = AutoTokenizer.from_pretrained(
|
||||
os.path.join(pretrained_model_name_or_path, "text_encoder")
|
||||
os.path.join(pretrained_model_name_or_path, "tokenizer")
|
||||
)
|
||||
self.max_sequence_length = 256
|
||||
|
||||
def encode_prompt(self, prompt):
|
||||
device = self.text_encoder.device
|
||||
dtype = self.dtype
|
||||
dtype = self.text_encoder.dtype
|
||||
|
||||
prompt = [prompt] if isinstance(prompt, str) else prompt
|
||||
batch_size = len(prompt)
|
||||
@@ -274,12 +271,11 @@ def load_transformer(
|
||||
)
|
||||
elif model_type == "hunyuan":
|
||||
transformer = HYVideoDiffusionTransformer(
|
||||
in_channels=16,
|
||||
out_channels=16,
|
||||
**hunyuan_config,
|
||||
dtype=master_weight_type,
|
||||
in_channels=16, out_channels=16, **hunyuan_config, dtype=master_weight_type,
|
||||
)
|
||||
transformer = load_hunyuan_state_dict(transformer, dit_model_name_or_path)
|
||||
if master_weight_type == torch.bfloat16:
|
||||
transformer = transformer.bfloat16()
|
||||
else:
|
||||
raise ValueError(f"Unsupported model type: {model_type}")
|
||||
return transformer
|
||||
|
||||
@@ -1,20 +1,23 @@
|
||||
from accelerate.logging import get_logger
|
||||
import torch
|
||||
|
||||
logger = get_logger(__name__)
|
||||
|
||||
|
||||
def get_optimizer(args, params_to_optimize, use_deepspeed: bool = False):
|
||||
def get_optimizer(
|
||||
params_to_optimize,
|
||||
args,
|
||||
lr=1e-5,
|
||||
betas=(0.9, 0.999),
|
||||
weight_decay=1e-3,
|
||||
eps=1e-8,
|
||||
):
|
||||
# Optimizer creation
|
||||
supported_optimizers = ["adam", "adamw", "prodigy"]
|
||||
supported_optimizers = ["adam", "adamw"]
|
||||
if args.optimizer not in supported_optimizers:
|
||||
logger.warning(
|
||||
print(
|
||||
f"Unsupported choice of optimizer: {args.optimizer}. Supported optimizers include {supported_optimizers}. Defaulting to AdamW"
|
||||
)
|
||||
args.optimizer = "adamw"
|
||||
|
||||
|
||||
if args.use_8bit_adam and not (args.optimizer.lower() not in ["adam", "adamw"]):
|
||||
logger.warning(
|
||||
print(
|
||||
f"use_8bit_adam is ignored when optimizer is not set to 'Adam' or 'AdamW'. Optimizer was "
|
||||
f"set to {args.optimizer.lower()}"
|
||||
)
|
||||
@@ -34,44 +37,20 @@ def get_optimizer(args, params_to_optimize, use_deepspeed: bool = False):
|
||||
|
||||
optimizer = optimizer_class(
|
||||
params_to_optimize,
|
||||
betas=(args.adam_beta1, args.adam_beta2),
|
||||
eps=args.adam_epsilon,
|
||||
weight_decay=args.adam_weight_decay,
|
||||
lr=lr,
|
||||
betas=betas,
|
||||
eps=eps,
|
||||
weight_decay=weight_decay,
|
||||
)
|
||||
elif args.optimizer.lower() == "adam":
|
||||
optimizer_class = bnb.optim.Adam8bit if args.use_8bit_adam else torch.optim.Adam
|
||||
|
||||
optimizer = optimizer_class(
|
||||
params_to_optimize,
|
||||
betas=(args.adam_beta1, args.adam_beta2),
|
||||
eps=args.adam_epsilon,
|
||||
weight_decay=args.adam_weight_decay,
|
||||
)
|
||||
elif args.optimizer.lower() == "prodigy":
|
||||
try:
|
||||
import prodigyopt
|
||||
except ImportError:
|
||||
raise ImportError(
|
||||
"To use Prodigy, please install the prodigyopt library: `pip install prodigyopt`"
|
||||
)
|
||||
|
||||
optimizer_class = prodigyopt.Prodigy
|
||||
|
||||
if args.learning_rate <= 0.1:
|
||||
logger.warning(
|
||||
"Learning rate is too low. When using prodigy, it's generally better to set learning rate around 1.0"
|
||||
)
|
||||
|
||||
optimizer = optimizer_class(
|
||||
params_to_optimize,
|
||||
lr=args.learning_rate,
|
||||
betas=(args.adam_beta1, args.adam_beta2),
|
||||
beta3=args.prodigy_beta3,
|
||||
weight_decay=args.adam_weight_decay,
|
||||
eps=args.adam_epsilon,
|
||||
decouple=args.prodigy_decouple,
|
||||
use_bias_correction=args.prodigy_use_bias_correction,
|
||||
safeguard_warmup=args.prodigy_safeguard_warmup,
|
||||
lr=lr,
|
||||
betas=betas,
|
||||
eps=eps,
|
||||
weight_decay=weight_decay,
|
||||
)
|
||||
|
||||
return optimizer
|
||||
|
||||
@@ -46,6 +46,7 @@ def prepare_latents(
|
||||
return latents
|
||||
|
||||
|
||||
|
||||
def sample_validation_video(
|
||||
transformer,
|
||||
vae,
|
||||
@@ -57,7 +58,6 @@ def sample_validation_video(
|
||||
num_inference_steps: int = 28,
|
||||
timesteps: List[int] = None,
|
||||
guidance_scale: float = 4.5,
|
||||
embed_guidance_scale: float = None,
|
||||
num_videos_per_prompt: Optional[int] = 1,
|
||||
generator: Optional[Union[torch.Generator, List[torch.Generator]]] = None,
|
||||
prompt_embeds: Optional[torch.Tensor] = None,
|
||||
@@ -108,17 +108,11 @@ def sample_validation_video(
|
||||
sigmas = np.array(sigmas)
|
||||
if scheduler_type == "euler":
|
||||
timesteps, num_inference_steps = retrieve_timesteps(
|
||||
scheduler,
|
||||
num_inference_steps,
|
||||
device,
|
||||
timesteps,
|
||||
sigmas,
|
||||
scheduler, num_inference_steps, device, timesteps, sigmas,
|
||||
)
|
||||
else:
|
||||
timesteps, num_inference_steps = retrieve_timesteps(
|
||||
scheduler,
|
||||
num_inference_steps,
|
||||
device,
|
||||
scheduler, num_inference_steps, device,
|
||||
)
|
||||
num_warmup_steps = max(len(timesteps) - num_inference_steps * scheduler.order, 0)
|
||||
|
||||
@@ -139,16 +133,14 @@ def sample_validation_video(
|
||||
# broadcast to batch dimension in a way that's compatible with ONNX/Core ML
|
||||
timestep = t.expand(latent_model_input.shape[0])
|
||||
with torch.autocast("cuda", dtype=torch.bfloat16):
|
||||
kwargs = {
|
||||
"hidden_states": latent_model_input,
|
||||
"encoder_hidden_states": prompt_embeds,
|
||||
"timestep": timestep,
|
||||
"encoder_attention_mask": prompt_attention_mask,
|
||||
"return_dict": False,
|
||||
}
|
||||
if embed_guidance_scale > 0.0:
|
||||
kwargs["guidance"] = torch.tensor([embed_guidance_scale], device=device, dtype=torch.bfloat16) * 1000
|
||||
noise_pred = transformer(**kwargs)[0]
|
||||
noise_pred = transformer(
|
||||
hidden_states=latent_model_input,
|
||||
encoder_hidden_states=prompt_embeds,
|
||||
timestep=timestep,
|
||||
encoder_attention_mask=prompt_attention_mask,
|
||||
return_dict=False,
|
||||
)[0]
|
||||
|
||||
# Mochi CFG + Sampling runs in FP32
|
||||
noise_pred = noise_pred.to(torch.float32)
|
||||
if do_classifier_free_guidance:
|
||||
@@ -309,11 +301,10 @@ def log_validation(
|
||||
scheduler_type=scheduler_type,
|
||||
num_frames=args.num_frames,
|
||||
# Peiyuan TODO: remove hardcode
|
||||
height=args.num_height,
|
||||
width=args.num_width,
|
||||
height=480,
|
||||
width=848,
|
||||
num_inference_steps=validation_sampling_step,
|
||||
guidance_scale=0.0 if args.model_type == "hunyuan" else validation_guidance_scale,
|
||||
embed_guidance_scale= validation_guidance_scale if args.model_type == "hunyuan" else 0.0,
|
||||
guidance_scale=validation_guidance_scale,
|
||||
generator=generator,
|
||||
prompt_embeds=prompt_embeds,
|
||||
prompt_attention_mask=prompt_attention_mask,
|
||||
|
||||
@@ -1,12 +0,0 @@
|
||||
python3 fastvideo/sample/sample_t2v_hunyuan_no_sp.py \
|
||||
--height 500 \
|
||||
--width 700 \
|
||||
--num_frames 29 \
|
||||
--num_inference_steps 50 \
|
||||
--guidance_scale 1 \
|
||||
--embedded_cfg_scale 6 \
|
||||
--flow-reverse \
|
||||
--prompts "A cat walks on the grass, realistic style." \
|
||||
--prompts "A dog runs in the park, realistic style." \
|
||||
--seed 42 \
|
||||
--output_path outputs_video/hunyuan/
|
||||
@@ -1,39 +0,0 @@
|
||||
num_gpus=4
|
||||
|
||||
torchrun --nnodes=1 --nproc_per_node=$num_gpus --master_port 29503 \
|
||||
fastvideo/sample/sample_t2v_hunyuan.py \
|
||||
--height 512 \
|
||||
--width 512 \
|
||||
--num_frames 29 \
|
||||
--num_inference_steps 4 \
|
||||
--guidance_scale 1 \
|
||||
--embedded_cfg_scale 6 \
|
||||
--flow_shift 17 \
|
||||
--flow-reverse \
|
||||
--prompts "A man on stage claps his hands together while facing the audience. The audience, visible in the foreground, holds up mobile devices to record the event, capturing the moment from various angles. The background features a large banner with text identifying the man on stage. Throughout the sequence, the man's expression remains engaged and directed towards the audience. The camera angle remains constant, focusing on capturing the interaction between the man on stage and the audience."\
|
||||
--seed 12345 \
|
||||
--output_path outputs_video/hunyuan/
|
||||
|
||||
|
||||
tensor(-0.1065, device='cuda:0', dtype=torch.float16)
|
||||
tensor(-0.0034, device='cuda:2', dtype=torch.float16)
|
||||
|
||||
tensor(-0.0230, device='cuda:0', dtype=torch.float16)
|
||||
>>> weight[0, :768].mean()
|
||||
tensor(-0.0367, device='cuda:0', dtype=torch.float16)
|
||||
|
||||
num_gpus=1
|
||||
|
||||
torchrun --nnodes=1 --nproc_per_node=$num_gpus --master_port 29503 \
|
||||
fastvideo/sample/sample_t2v_hunyuan.py \
|
||||
--height 480 \
|
||||
--width 848 \
|
||||
--num_frames 93 \
|
||||
--num_inference_steps 50 \
|
||||
--guidance_scale 1 \
|
||||
--embedded_cfg_scale 6 \
|
||||
--flow_shift 17 \
|
||||
--flow-reverse \
|
||||
--prompts "A man on stage claps his hands together while facing the audience. The audience, visible in the foreground, holds up mobile devices to record the event, capturing the moment from various angles. The background features a large banner with text identifying the man on stage. Throughout the sequence, the man's expression remains engaged and directed towards the audience. The camera angle remains constant, focusing on capturing the interaction between the man on stage and the audience."\
|
||||
--seed 12345 \
|
||||
--output_path outputs_video/hunyuan/
|
||||
@@ -1,20 +0,0 @@
|
||||
|
||||
|
||||
|
||||
num_gpus=4
|
||||
|
||||
torchrun --nnodes=1 --nproc_per_node=$num_gpus --master_port 29503 \
|
||||
fastvideo/sample/sample_t2v_mochi.py \
|
||||
--model_path data/mochi \
|
||||
--prompt_embed_path "data/synthetic_debug2/prompt_embed/2.pt" \
|
||||
--encoder_attention_mask_path "data/synthetic_debug2/prompt_attention_mask/1.pt" \
|
||||
--num_frames 163 \
|
||||
--height 480 \
|
||||
--width 848 \
|
||||
--num_inference_steps 32 \
|
||||
--guidance_scale 4.5 \
|
||||
--output_path outputs_video/debug \
|
||||
--shift 8 \
|
||||
--seed 12345 \
|
||||
--scheduler_type "pcm_linear_quadratic"
|
||||
|
||||
@@ -1,11 +0,0 @@
|
||||
python fastvideo/sample/sample_t2v_mochi_no_sp.py \
|
||||
--model_path data/mochi \
|
||||
--prompts "A hand enters the frame, pulling a sheet of plastic wrap over three balls of dough placed on a wooden surface. The plastic wrap is stretched to cover the dough more securely. The hand adjusts the wrap, ensuring that it is tight and smooth over the dough. The scene focuses on the hand's movements as it secures the edges of the plastic wrap. No new objects appear, and the camera remains stationary, focusing on the action of covering the dough." \
|
||||
--num_frames 163 \
|
||||
--height 480 \
|
||||
--width 848 \
|
||||
--num_inference_steps 64 \
|
||||
--guidance_scale 0.0 \
|
||||
--seed 12346 \
|
||||
--transformer_path data/outputs/debug/checkpoint-100/transformer \
|
||||
--output_path outputs_video/single_no_guidance
|
||||
+130
@@ -0,0 +1,130 @@
|
||||
# Prediction interface for Cog ⚙️
|
||||
# https://cog.run/python
|
||||
|
||||
from cog import BasePredictor, Input, Path
|
||||
import os
|
||||
import time
|
||||
import torch
|
||||
import imageio
|
||||
import argparse
|
||||
import subprocess
|
||||
import torchvision
|
||||
import numpy as np
|
||||
from einops import rearrange
|
||||
|
||||
MODEL_CACHE = 'FastHunyuan'
|
||||
os.environ['MODEL_BASE'] = './'+MODEL_CACHE
|
||||
from fastvideo.models.hunyuan.inference import HunyuanVideoSampler
|
||||
MODEL_URL = "https://weights.replicate.delivery/default/FastVideo/FastHunyuan/model.tar"
|
||||
|
||||
def download_weights(url, dest):
|
||||
start = time.time()
|
||||
print("downloading url: ", url)
|
||||
print("downloading to: ", dest)
|
||||
subprocess.check_call(["pget", "-xf", url, dest], close_fds=False)
|
||||
print("downloading took: ", time.time() - start)
|
||||
|
||||
class Predictor(BasePredictor):
|
||||
def setup(self):
|
||||
"""Load the model into memory"""
|
||||
print("Model Base: " + os.environ['MODEL_BASE'])
|
||||
# Download weights
|
||||
if not os.path.exists(MODEL_CACHE):
|
||||
download_weights(MODEL_URL, MODEL_CACHE)
|
||||
|
||||
self.device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||
args = argparse.Namespace(
|
||||
num_frames=125,
|
||||
height=720,
|
||||
width=1280,
|
||||
num_inference_steps=6,
|
||||
fps=24,
|
||||
denoise_type='flow',
|
||||
seed=1024,
|
||||
neg_prompt=None,
|
||||
guidance_scale=1.0,
|
||||
embedded_cfg_scale=6.0,
|
||||
flow_shift=17,
|
||||
batch_size=1,
|
||||
num_videos=1,
|
||||
load_key='module',
|
||||
use_cpu_offload=False,
|
||||
dit_weight='FastHunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt',
|
||||
reproduce=True,
|
||||
disable_autocast=False,
|
||||
flow_reverse=True,
|
||||
flow_solver='euler',
|
||||
use_linear_quadratic_schedule=False,
|
||||
linear_schedule_end=25,
|
||||
model='HYVideo-T/2-cfgdistill',
|
||||
latent_channels=16,
|
||||
precision='bf16',
|
||||
rope_theta=256,
|
||||
vae='884-16c-hy',
|
||||
vae_precision='fp16',
|
||||
vae_tiling=True,
|
||||
text_encoder='llm',
|
||||
text_encoder_precision='fp16',
|
||||
text_states_dim=4096,
|
||||
text_len=256,
|
||||
tokenizer='llm',
|
||||
prompt_template='dit-llm-encode',
|
||||
prompt_template_video='dit-llm-encode-video',
|
||||
hidden_state_skip_layer=2,
|
||||
apply_final_norm=False,
|
||||
text_encoder_2='clipL',
|
||||
text_encoder_precision_2='fp16',
|
||||
text_states_dim_2=768,
|
||||
tokenizer_2='clipL',
|
||||
text_len_2=77,
|
||||
model_path=MODEL_CACHE,
|
||||
)
|
||||
self.model = HunyuanVideoSampler.from_pretrained(MODEL_CACHE, args=args)
|
||||
|
||||
def predict(
|
||||
self,
|
||||
prompt: str = Input(description="Text prompt for video generation", default="A cat walks on the grass, realistic style."),
|
||||
negative_prompt: str = Input(description="Text prompt to specify what you don't want in the video.", default=""),
|
||||
width: int = Input(description="Width of output video", default=1280, ge=256),
|
||||
height: int = Input(description="Height of output video", default=720, ge=256),
|
||||
num_frames: int = Input(description="Number of frames to generate", default=125, ge=16),
|
||||
num_inference_steps: int = Input(description="Number of denoising steps", default=6, ge=1, le=50),
|
||||
guidance_scale: float = Input(description="Classifier free guidance scale", default=1.0, ge=0.1, le=10.0),
|
||||
embedded_cfg_scale: float = Input(description="Embedded classifier free guidance scale", default=6.0, ge=0.1, le=10.0),
|
||||
flow_shift: int = Input(description="Flow shift parameter", default=17, ge=1, le=20),
|
||||
fps: int = Input(description="Frames per second of output video", default=24, ge=1, le=60),
|
||||
seed: int = Input(description="0 for Random seed. Set for reproducible generation", default=0),
|
||||
) -> Path:
|
||||
"""Run video generation"""
|
||||
if seed <=0:
|
||||
seed = int.from_bytes(os.urandom(2), "big")
|
||||
print(f"Using seed: {seed}")
|
||||
|
||||
outputs = self.model.predict(
|
||||
prompt=prompt,
|
||||
height=height,
|
||||
width=width,
|
||||
video_length=num_frames,
|
||||
seed=seed,
|
||||
negative_prompt=negative_prompt,
|
||||
infer_steps=num_inference_steps,
|
||||
guidance_scale=guidance_scale,
|
||||
embedded_guidance_scale=embedded_cfg_scale,
|
||||
flow_shift=flow_shift,
|
||||
flow_reverse=True,
|
||||
batch_size=1,
|
||||
num_videos_per_prompt=1,
|
||||
)
|
||||
|
||||
# Process output video
|
||||
videos = rearrange(outputs["samples"], "b c t h w -> t b c h w")
|
||||
frames = []
|
||||
for x in videos:
|
||||
x = torchvision.utils.make_grid(x, nrow=6)
|
||||
x = x.transpose(0, 1).transpose(1, 2).squeeze(-1)
|
||||
frames.append((x * 255).numpy().astype(np.uint8))
|
||||
|
||||
# Save video
|
||||
output_path = Path("/tmp/output.mp4")
|
||||
imageio.mimsave(str(output_path), frames, fps=fps)
|
||||
return Path(output_path)
|
||||
+1
-1
@@ -21,7 +21,7 @@ dependencies = [
|
||||
"timm==1.0.11", "torchdiffeq==0.2.4", "torchmetrics==1.5.1", "tqdm==4.66.5", "urllib3==2.2.0", "uvicorn==0.32.0",
|
||||
"scikit-video==1.1.11", "imageio-ffmpeg==0.5.1", "sentencepiece==0.2.0", "beautifulsoup4==4.12.3", "ftfy==6.3.0",
|
||||
"moviepy==1.0.3", "wandb==0.18.5", "tensorboard==2.18.0", "pydantic==2.9.2", "gradio==5.3.0", "huggingface_hub==0.26.1", "protobuf==5.28.3",
|
||||
"watch", "gpustat", "peft==0.13.2", "liger_kernel==0.4.1", "einops==0.8.0", "wheel==0.44.0", "loguru"]
|
||||
"watch", "gpustat", "peft==0.13.2", "liger_kernel==0.4.1", "einops==0.8.0", "wheel==0.44.0", "loguru", "diffusers==0.32.0", "bitsandbytes"]
|
||||
|
||||
|
||||
[tool.setuptools.packages.find]
|
||||
|
||||
@@ -1,25 +1,25 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
|
||||
DATA_DIR=./data
|
||||
|
||||
torchrun --nnodes 1 --nproc_per_node 8\
|
||||
fastvideo/distill_adv.py\
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-30K-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 4\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=640\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
@@ -30,7 +30,7 @@ torchrun --nnodes 1 --nproc_per_node 8\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32_adv"\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
@@ -1,31 +1,21 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
DATA_DIR=./data
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
torchrun --nnodes 1 --nproc_per_node 4\
|
||||
fastvideo/distill_adv.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-HunYuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-HunYuan/validation"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 32\
|
||||
--num_latent_t 8\
|
||||
--sp_size 4 \
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
@@ -50,4 +40,5 @@ torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
--not_apply_cfg_solver \
|
||||
--master_weight_type "bf16"
|
||||
@@ -1,21 +1,11 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
export WANDB_API_KEY=7abd9763ff10869f9e88526d919c2639766139a8
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
DATA_DIR=./data
|
||||
|
||||
torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
torchrun --nnodes 1 --nproc_per_node 4\
|
||||
fastvideo/distill_dmd.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
@@ -25,30 +15,34 @@ torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 32 \
|
||||
--sp_size 2 \
|
||||
--num_latent_t 8\
|
||||
--sp_size 4\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--max_train_steps=30000\
|
||||
--learning_rate=1e-5\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "4,6,8" \
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--num_inference_steps 4 \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_16_HD_random_cfg"\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_height 720 \
|
||||
--num_width 1280 \
|
||||
--num_frames 125 \
|
||||
--num_frames 125 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "3.0,4.0,5.0,6.0" \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver \
|
||||
--distill_cfg "1.0,2.0,3.0,4.0,5.0,6.0"
|
||||
--master_weight_type "bf16" \
|
||||
--optimizer "AdamW" \
|
||||
--use_8bit_adam \
|
||||
--generator_update_steps 3
|
||||
@@ -1,53 +1,48 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
export WANDB_API_KEY=7abd9763ff10869f9e88526d919c2639766139a8
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
DATA_DIR=./data
|
||||
|
||||
torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
torchrun --nnodes 1 --nproc_per_node 4\
|
||||
fastvideo/distill_dmd.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-HunYuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-HunYuan/validation"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 32 \
|
||||
--sp_size 2 \
|
||||
--num_latent_t 8\
|
||||
--sp_size 4\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--max_train_steps=30000\
|
||||
--learning_rate=1e-5\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--num_inference_steps 8 \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_16_HD"\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_height 720 \
|
||||
--num_width 1280 \
|
||||
--num_frames 125 \
|
||||
--num_frames 125 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
--not_apply_cfg_solver \
|
||||
--master_weight_type "bf16" \
|
||||
--optimizer "AdamW" \
|
||||
--use_8bit_adam \
|
||||
--generator_update_steps 3
|
||||
@@ -1,41 +1,32 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
export WANDB_API_KEY=7abd9763ff10869f9e88526d919c2639766139a8
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
DATA_DIR=./data
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
torchrun --nnodes 1 --nproc_per_node 4\
|
||||
fastvideo/distill_dmd.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--num_latent_t 8\
|
||||
--sp_size 4\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=480\
|
||||
--learning_rate=1e-6\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=30000\
|
||||
--learning_rate=1e-5\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--checkpointing_steps=1\
|
||||
--validation_steps 1\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--num_inference_steps 4 \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
@@ -43,9 +34,16 @@ torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--num_height 784 \
|
||||
--num_width 1280 \
|
||||
--num_frames 29 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
--not_apply_cfg_solver \
|
||||
--master_weight_type "bf16" \
|
||||
--optimizer "AdamW" \
|
||||
--use_8bit_adam \
|
||||
--generator_update_steps 5 \
|
||||
--run_pod
|
||||
@@ -0,0 +1,38 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_MODE=online
|
||||
|
||||
torchrun --nnodes 1 --nproc_per_node 4 \
|
||||
fastvideo/distill.py \
|
||||
--seed 42 \
|
||||
--pretrained_model_name_or_path data/mochi \
|
||||
--model_type "mochi" \
|
||||
--cache_dir data/.cache \
|
||||
--data_json_path data/Merge-30k-Data/video2caption.json \
|
||||
--validation_prompt_dir data/Image-Vid-Finetune-Mochi/validation \
|
||||
--gradient_checkpointing \
|
||||
--train_batch_size=1 \
|
||||
--num_latent_t 28 \
|
||||
--sp_size 4 \
|
||||
--train_sp_batch_size 2 \
|
||||
--dataloader_num_workers 4 \
|
||||
--gradient_accumulation_steps=1 \
|
||||
--max_train_steps=4000 \
|
||||
--learning_rate=1e-6 \
|
||||
--mixed_precision=bf16 \
|
||||
--checkpointing_steps=64 \
|
||||
--validation_steps 1 \
|
||||
--validation_sampling_steps 8 \
|
||||
--checkpoints_total_limit 3 \
|
||||
--allow_tf32 \
|
||||
--ema_start_step 0 \
|
||||
--cfg 0.0 \
|
||||
--log_validation \
|
||||
--output_dir="data/outputs/lq_euler_50_thres0.1_lrg_0.75_phase1_lr1e-6_repro" \
|
||||
--tracker_project_name PCM \
|
||||
--num_frames 163 \
|
||||
--scheduler_type pcm_linear_quadratic \
|
||||
--validation_guidance_scale 0.5,1.5,2.5 \
|
||||
--num_euler_timesteps 50 \
|
||||
--linear_quadratic_threshold 0.1 \
|
||||
--linear_range 0.75 \
|
||||
--multi_phased_distill_schedule 4000-1
|
||||
@@ -0,0 +1,39 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_MODE=online
|
||||
|
||||
torchrun --nnodes 1 --nproc_per_node 4 \
|
||||
fastvideo/distill_dmd.py \
|
||||
--seed 42 \
|
||||
--pretrained_model_name_or_path data/mochi \
|
||||
--model_type "mochi" \
|
||||
--cache_dir data/.cache \
|
||||
--data_json_path data/Mochi-425-Data/videos2caption.json \
|
||||
--validation_prompt_dir data/Mochi-425-Data/validation \
|
||||
--gradient_checkpointing \
|
||||
--train_batch_size=1 \
|
||||
--num_latent_t 16 \
|
||||
--sp_size 4 \
|
||||
--train_sp_batch_size 1 \
|
||||
--generator_update_steps=1 \
|
||||
--dataloader_num_workers 4 \
|
||||
--gradient_accumulation_steps=1 \
|
||||
--max_train_steps=4000 \
|
||||
--learning_rate=1e-6 \
|
||||
--mixed_precision=bf16 \
|
||||
--checkpointing_steps=64 \
|
||||
--validation_steps=64 \
|
||||
--validation_sampling_steps 6 \
|
||||
--checkpoints_total_limit 3 \
|
||||
--allow_tf32 \
|
||||
--ema_start_step 0 \
|
||||
--cfg 0.0 \
|
||||
--log_validation \
|
||||
--output_dir="data/outputs/lq_euler_50_thres0.1_lrg_0.75_phase1_lr1e-6_repro" \
|
||||
--tracker_project_name PCM \
|
||||
--num_frames 93 \
|
||||
--scheduler_type pcm_linear_quadratic \
|
||||
--validation_guidance_scale 0.5,1.5,2.5 \
|
||||
--num_euler_timesteps 50 \
|
||||
--linear_quadratic_threshold 0.1 \
|
||||
--linear_range 0.75 \
|
||||
--multi_phased_distill_schedule 4000-1
|
||||
@@ -1,5 +0,0 @@
|
||||
## How to Distill Hunyuan
|
||||
|
||||
python scripts/download_hf.py --repo_id FastVideo/hunyuan --local_dir data/hunyuan --repo_type model
|
||||
python scripts/download_hf.py --repo_id FastVideo/Hunyuan-Distill-Data --local_dir data/Hunyuan-Distill-Data --repo_type=dataset
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_MODE=online
|
||||
|
||||
torchrun --nnodes 1 --nproc_per_node 8 \
|
||||
fastvideo/train.py \
|
||||
--seed 42 \
|
||||
--pretrained_model_name_or_path data/hunyuan \
|
||||
--dit_model_name_or_path data/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir data/.cache \
|
||||
--data_json_path data/Image-Vid-Finetune-HunYuan/videos2caption.json \
|
||||
--validation_prompt_dir data/Image-Vid-Finetune-HunYuan/validation \
|
||||
--gradient_checkpointing \
|
||||
--train_batch_size=1 \
|
||||
--num_latent_t 24 \
|
||||
--sp_size 4 \
|
||||
--train_sp_batch_size 1 \
|
||||
--dataloader_num_workers 4 \
|
||||
--gradient_accumulation_steps=1 \
|
||||
--max_train_steps=2000 \
|
||||
--learning_rate=5e-6 \
|
||||
--mixed_precision=bf16 \
|
||||
--checkpointing_steps=200 \
|
||||
--validation_steps 100 \
|
||||
--validation_sampling_steps 64 \
|
||||
--checkpoints_total_limit 3 \
|
||||
--allow_tf32 \
|
||||
--ema_start_step 0 \
|
||||
--cfg 0.0 \
|
||||
--ema_decay 0.999 \
|
||||
--log_validation \
|
||||
--output_dir=data/outputs/HSH-Taylor-Finetune-Hunyuan \
|
||||
--tracker_project_name HSH-Taylor-Finetune-Hunyuan \
|
||||
--num_frames 93 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--group_frame
|
||||
@@ -0,0 +1,32 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_MODE=online
|
||||
|
||||
torchrun --nnodes 1 --nproc_per_node 4 \
|
||||
fastvideo/train.py \
|
||||
--seed 42 \
|
||||
--pretrained_model_name_or_path data/mochi \
|
||||
--cache_dir data/.cache \
|
||||
--data_json_path data/Mochi-Black-Myth/videos2caption.json \
|
||||
--validation_prompt_dir data/Mochi-Black-Myth/validation \
|
||||
--gradient_checkpointing \
|
||||
--train_batch_size=1 \
|
||||
--num_latent_t 16 \
|
||||
--sp_size 1 \
|
||||
--train_sp_batch_size 1 \
|
||||
--dataloader_num_workers 4 \
|
||||
--gradient_accumulation_steps=1 \
|
||||
--max_train_steps=2000 \
|
||||
--learning_rate=5e-6 \
|
||||
--mixed_precision=bf16 \
|
||||
--checkpointing_steps=200 \
|
||||
--validation_steps 100 \
|
||||
--validation_sampling_steps 64 \
|
||||
--checkpoints_total_limit 3 \
|
||||
--allow_tf32 \
|
||||
--ema_start_step 0 \
|
||||
--cfg 0.0 \
|
||||
--ema_decay 0.999 \
|
||||
--log_validation \
|
||||
--output_dir=data/outputs/Black-Myth-Finetune \
|
||||
--tracker_project_name Black-Myth-Finetune \
|
||||
--num_frames 93
|
||||
@@ -0,0 +1,37 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_MODE=online
|
||||
|
||||
torchrun --nnodes 1 --nproc_per_node 2 \
|
||||
fastvideo/train.py \
|
||||
--seed 42 \
|
||||
--pretrained_model_name_or_path data/mochi \
|
||||
--cache_dir data/.cache \
|
||||
--data_json_path data/Mochi-Black-Myth/videos2caption.json \
|
||||
--validation_prompt_dir data/Mochi-Black-Myth/validation \
|
||||
--gradient_checkpointing \
|
||||
--train_batch_size 1 \
|
||||
--num_latent_t 14 \
|
||||
--sp_size 2 \
|
||||
--train_sp_batch_size 1 \
|
||||
--dataloader_num_workers 1 \
|
||||
--gradient_accumulation_steps 2 \
|
||||
--max_train_steps 2000 \
|
||||
--learning_rate 5e-6 \
|
||||
--mixed_precision bf16 \
|
||||
--checkpointing_steps 200 \
|
||||
--validation_steps 100 \
|
||||
--validation_sampling_steps 64 \
|
||||
--checkpoints_total_limit 3 \
|
||||
--allow_tf32 \
|
||||
--ema_start_step 0 \
|
||||
--cfg 0.0 \
|
||||
--ema_decay 0.999 \
|
||||
--log_validation \
|
||||
--output_dir=data/outputs/Black-Myth-Lora-FT \
|
||||
--tracker_project_name Black-Myth-Lora-Finetune \
|
||||
--num_frames 91 \
|
||||
--lora_rank 128 \
|
||||
--lora_alpha 256 \
|
||||
--master_weight_type "bf16" \
|
||||
--use_lora \
|
||||
--use_cpu_offload
|
||||
@@ -0,0 +1,37 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_MODE=online
|
||||
|
||||
CUDA_VISIBLE_DEVICES=5 torchrun --nnodes 1 --nproc_per_node 1 \
|
||||
fastvideo/train.py \
|
||||
--seed 42 \
|
||||
--pretrained_model_name_or_path data/mochi \
|
||||
--cache_dir data/.cache \
|
||||
--data_json_path data/Image-Vid-Finetune-Mochi/videos2caption.json \
|
||||
--validation_prompt_dir data/Image-Vid-Finetune-Mochi/validation \
|
||||
--gradient_checkpointing \
|
||||
--train_batch_size=1 \
|
||||
--num_latent_t 14 \
|
||||
--sp_size 1 \
|
||||
--train_sp_batch_size 1 \
|
||||
--dataloader_num_workers 1 \
|
||||
--gradient_accumulation_steps=1 \
|
||||
--max_train_steps=2000 \
|
||||
--learning_rate=5e-6 \
|
||||
--mixed_precision=bf16 \
|
||||
--checkpointing_steps=200 \
|
||||
--validation_steps 100 \
|
||||
--validation_sampling_steps 64 \
|
||||
--checkpoints_total_limit 3 \
|
||||
--allow_tf32 \
|
||||
--ema_start_step 0 \
|
||||
--cfg 0.0 \
|
||||
--ema_decay 0.999 \
|
||||
--log_validation \
|
||||
--output_dir=data/outputs/HSH-Taylor-Finetune-Lora \
|
||||
--tracker_project_name HSH-Taylor-Finetune-Lora \
|
||||
--num_frames 91 \
|
||||
--group_frame \
|
||||
--lora_rank 128 \
|
||||
--lora_alpha 256 \
|
||||
--master_weight_type "bf16" \
|
||||
--use_lora
|
||||
@@ -1,36 +0,0 @@
|
||||
torchrun --nnodes 4 --nproc_per_node 4 \
|
||||
--node_rank=3 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=172.23.30.15:29500 \
|
||||
fastvideo/train.py \
|
||||
--seed 42 \
|
||||
--pretrained_model_name_or_path data/mochi \
|
||||
--cache_dir "data/.cache" \
|
||||
--data_json_path "data/BLACK-MYTH-Finetune-Dataset/videos2caption.json" \
|
||||
--validation_prompt_dir "data/BLACK-MYTH-Finetune-Dataset/validation_prompt_embed_mask" \
|
||||
--uncond_prompt_dir "data/BLACK-MYTH-Finetune-Dataset/uncond_prompt_embed_mask" \
|
||||
--gradient_checkpointing \
|
||||
--train_batch_size=1 \
|
||||
--num_latent_t 16 \
|
||||
--sp_size 4 \
|
||||
--train_sp_batch_size 4 \
|
||||
--dataloader_num_workers 4 \
|
||||
--gradient_accumulation_steps=1 \
|
||||
--max_train_steps=4000 \
|
||||
--learning_rate=2e-5 \
|
||||
--mixed_precision="bf16" \
|
||||
--checkpointing_steps=500 \
|
||||
--validation_steps=100 \
|
||||
--validation_sampling_steps 64 \
|
||||
--checkpoints_total_limit 3 \
|
||||
--allow_tf32 \
|
||||
--ema_start_step 0 \
|
||||
--cfg 0.0 \
|
||||
--ema_decay 0.999 \
|
||||
--log_validation \
|
||||
--output_dir="data/outputs/black_myth_correct_video_mask" \
|
||||
--weighting_scheme "uniform" \
|
||||
--num_frames 91 \
|
||||
--selective_checkpointing 1.0
|
||||
|
||||
@@ -1,13 +0,0 @@
|
||||
num_gpus=1
|
||||
|
||||
torchrun --nproc_per_node=$num_gpus fastvideo/sample/generate_synthetic.py \
|
||||
--model_path data/mochi \
|
||||
--num_frames 1 \
|
||||
--height 480 \
|
||||
--width 848 \
|
||||
--num_inference_steps 4 \
|
||||
--guidance_scale 4.5 \
|
||||
--prompt_path "data/prompt.txt" \
|
||||
--dataset_output_dir data/synthetic_debug2
|
||||
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
from huggingface_hub import HfApi
|
||||
|
||||
api = HfApi()
|
||||
|
||||
api.upload_folder(
|
||||
folder_path="data/Black-Myth-Taylor-Src",
|
||||
repo_id="FastVideo/Image-Vid-Finetune-Src",
|
||||
repo_type="dataset",
|
||||
)
|
||||
@@ -1,92 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 2\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=480\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 2\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=480\
|
||||
--learning_rate=3e-7\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_lr_3e-7"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,51 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift25_bs_32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 25 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,51 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift33_bs_32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 33 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,51 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=3e-7\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32_lr3e-7"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,51 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,51 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32_euler20"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 20 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,51 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-30K-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32_moredata"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,53 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-HunYuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-HunYuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 32 \
|
||||
--sp_size 2 \
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift25_bs_16_HD"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_height 720 \
|
||||
--num_width 1280 \
|
||||
--num_frames 125 \
|
||||
--shift 25 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,53 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-HunYuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-HunYuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 32 \
|
||||
--sp_size 2 \
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift33_bs_16_HD"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_height 720 \
|
||||
--num_width 1280 \
|
||||
--num_frames 125 \
|
||||
--shift 33 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,92 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 2\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=480\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift27"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 27 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 2\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=480\
|
||||
--learning_rate=3e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_lr_3e-6"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,58 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 32 \
|
||||
--sp_size 2 \
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "4,6,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_lq_0.1_0.67_bs_16_HD"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_height 720 \
|
||||
--num_width 1280 \
|
||||
--num_frames 125 \
|
||||
--scheduler_type pcm_linear_quadratic \
|
||||
--linear_quadratic_threshold 0.1 \
|
||||
--linear_range 0.67 \
|
||||
--validation_guidance_scale "6.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver \
|
||||
--distill_cfg "6.0" \
|
||||
|
||||
|
||||
@@ -1,58 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 32 \
|
||||
--sp_size 2 \
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "4,6,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_lq_0.05_0.67_bs_16_HD"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_height 720 \
|
||||
--num_width 1280 \
|
||||
--num_frames 125 \
|
||||
--scheduler_type pcm_linear_quadratic \
|
||||
--linear_quadratic_threshold 0.05 \
|
||||
--linear_range 0.67 \
|
||||
--validation_guidance_scale "6.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver \
|
||||
--distill_cfg "6.0"
|
||||
|
||||
|
||||
@@ -1,56 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 4 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/HD-Mixkit-Finetune-Hunyuan/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 32 \
|
||||
--sp_size 2 \
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=2000\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "4,6,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_lq_0.1_0.67_bs_16_HD_random_cfg"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_height 720 \
|
||||
--num_width 1280 \
|
||||
--num_frames 125 \
|
||||
--scheduler_type pcm_linear_quadratic \
|
||||
--linear_quadratic_threshold 0.1 \
|
||||
--linear_range 0.67 \
|
||||
--validation_guidance_scale "3.0,4.0,5.0,6.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver \
|
||||
--distill_cfg "1.0,2.0,3.0,4.0,5.0,6.0"
|
||||
@@ -1,95 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 2\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=480\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_teacher_no_cfg"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver \
|
||||
--hunyuan_teacher_disable_cfg
|
||||
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 2\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=480\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_ema"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver \
|
||||
--use_ema
|
||||
|
||||
@@ -1,92 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=1\
|
||||
--max_train_steps=320\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_16_nosp"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 2\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=320\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_16_sp"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,53 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_student3_batchsize32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver \
|
||||
--hunyuan_student_cfg_embed 3
|
||||
|
||||
@@ -1,53 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_student4_batchsize32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver \
|
||||
--hunyuan_student_cfg_embed 4
|
||||
|
||||
@@ -1,51 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=3e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase1_shift17_bs_32_lr_3e-6"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-1" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -1,51 +0,0 @@
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export WANDB_DIR="$HOME"
|
||||
export WANDB_MODE=online
|
||||
export WANDB_API_KEY=4f6de3765d6464f43e0506ec7d785641af645e73
|
||||
export LD_LIBRARY_PATH=/opt/amazon/efa/lib:/opt/aws-ofi-nccl/lib:$LD_LIBRARY_PATH
|
||||
export FI_PROVIDER=efa
|
||||
export FI_EFA_USE_DEVICE_RDMA=1
|
||||
export NCCL_PROTO=simple
|
||||
|
||||
DATA_DIR=/data
|
||||
IP=10.4.139.86
|
||||
|
||||
torchrun --nnodes 2 --nproc_per_node 8\
|
||||
--node_rank=0 \
|
||||
--rdzv_id=456 \
|
||||
--rdzv_backend=c10d \
|
||||
--rdzv_endpoint=$IP:29500 \
|
||||
fastvideo/distill.py\
|
||||
--seed 42\
|
||||
--pretrained_model_name_or_path $DATA_DIR/hunyuan\
|
||||
--dit_model_name_or_path $DATA_DIR/hunyuan/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt\
|
||||
--model_type "hunyuan" \
|
||||
--cache_dir "$DATA_DIR/.cache"\
|
||||
--data_json_path "$DATA_DIR/Hunyuan-Distill-Data/videos2caption.json"\
|
||||
--validation_prompt_dir "$DATA_DIR/Hunyuan-Distill-Data/validation"\
|
||||
--gradient_checkpointing\
|
||||
--train_batch_size=1\
|
||||
--num_latent_t 24\
|
||||
--sp_size 1\
|
||||
--train_sp_batch_size 1\
|
||||
--dataloader_num_workers 4\
|
||||
--gradient_accumulation_steps=2\
|
||||
--max_train_steps=640\
|
||||
--learning_rate=1e-6\
|
||||
--mixed_precision="bf16"\
|
||||
--checkpointing_steps=64\
|
||||
--validation_steps 64\
|
||||
--validation_sampling_steps "2,4,8" \
|
||||
--checkpoints_total_limit 3\
|
||||
--allow_tf32\
|
||||
--ema_start_step 0\
|
||||
--cfg 0.0\
|
||||
--log_validation\
|
||||
--output_dir="$DATA_DIR/outputs/hy_phase2_shift17_bs_32"\
|
||||
--tracker_project_name Hunyuan_Distill \
|
||||
--num_frames 93 \
|
||||
--shift 17 \
|
||||
--validation_guidance_scale "1.0" \
|
||||
--num_euler_timesteps 50 \
|
||||
--multi_phased_distill_schedule "4000-2" \
|
||||
--not_apply_cfg_solver
|
||||
@@ -0,0 +1,20 @@
|
||||
#!/bin/bash
|
||||
|
||||
num_gpus=1
|
||||
export MODEL_BASE="data/FastHunyuan-diffusers"
|
||||
torchrun --nnodes=1 --nproc_per_node=$num_gpus --master_port 12345 \
|
||||
fastvideo/sample/sample_t2v_diffusers_hunyuan.py \
|
||||
--height 720 \
|
||||
--width 1280 \
|
||||
--num_frames 45 \
|
||||
--num_inference_steps 6 \
|
||||
--guidance_scale 1 \
|
||||
--embedded_cfg_scale 6 \
|
||||
--flow_shift 17 \
|
||||
--flow-reverse \
|
||||
--prompt ./assets/prompt.txt \
|
||||
--seed 1024 \
|
||||
--output_path outputs_video/hunyuan_quant/nf4/ \
|
||||
--model_path $MODEL_BASE \
|
||||
--quantization "nf4" \
|
||||
--cpu_offload
|
||||
@@ -0,0 +1,19 @@
|
||||
#!/bin/bash
|
||||
|
||||
num_gpus=4
|
||||
export MODEL_BASE=data/FastHunyuan
|
||||
torchrun --nnodes=1 --nproc_per_node=$num_gpus --master_port 29503 \
|
||||
fastvideo/sample/sample_t2v_hunyuan.py \
|
||||
--height 720 \
|
||||
--width 1280 \
|
||||
--num_frames 125 \
|
||||
--num_inference_steps 6 \
|
||||
--guidance_scale 1 \
|
||||
--embedded_cfg_scale 6 \
|
||||
--flow_shift 17 \
|
||||
--flow-reverse \
|
||||
--prompt ./assets/prompt.txt \
|
||||
--seed 1024 \
|
||||
--output_path outputs_video/hunyuan/cfg6/ \
|
||||
--model_path $MODEL_BASE \
|
||||
--dit-weight ${MODEL_BASE}/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt
|
||||
@@ -0,0 +1,19 @@
|
||||
#!/bin/bash
|
||||
|
||||
num_gpus=4
|
||||
|
||||
torchrun --nnodes=1 --nproc_per_node=$num_gpus --master_port 29503 \
|
||||
fastvideo/sample/sample_t2v_mochi.py \
|
||||
--model_path data/FastMochi-diffusers \
|
||||
--prompt_path "assets/prompt.txt" \
|
||||
--num_frames 163 \
|
||||
--height 480 \
|
||||
--width 848 \
|
||||
--num_inference_steps 8 \
|
||||
--guidance_scale 1.5 \
|
||||
--output_path outputs_video/mochi_sp/ \
|
||||
--seed 1024 \
|
||||
--scheduler_type "pcm_linear_quadratic" \
|
||||
--linear_threshold 0.1 \
|
||||
--linear_range 0.75
|
||||
|
||||
@@ -1,11 +1,9 @@
|
||||
# export WANDB_MODE="offline"
|
||||
GPU_NUM=8
|
||||
SHARD_NUM=8
|
||||
SHARD_IDX=0
|
||||
GPU_NUM=1 # 2,4,8
|
||||
MODEL_PATH="data/hunyuan"
|
||||
MODEL_TYPE="hunyuan"
|
||||
DATA_MERGE_PATH="data/Distill-30K-Src/merge.txt"
|
||||
OUTPUT_DIR="data/HD-Hunyuan-30K-Distill-Data_Shard${SHARD_IDX}"
|
||||
DATA_MERGE_PATH="data/Image-Vid-Finetune-Src/merge.txt"
|
||||
OUTPUT_DIR="data/Image-Vid-Finetune-HunYuan"
|
||||
VALIDATION_PATH="assets/prompt.txt"
|
||||
|
||||
torchrun --nproc_per_node=$GPU_NUM \
|
||||
@@ -13,26 +11,20 @@ torchrun --nproc_per_node=$GPU_NUM \
|
||||
--model_path $MODEL_PATH \
|
||||
--data_merge_path $DATA_MERGE_PATH \
|
||||
--train_batch_size=1 \
|
||||
--max_height=720 \
|
||||
--max_width=1280 \
|
||||
--num_frames=129 \
|
||||
--max_height=480 \
|
||||
--max_width=848 \
|
||||
--num_frames=93 \
|
||||
--dataloader_num_workers 1 \
|
||||
--output_dir=$OUTPUT_DIR \
|
||||
--model_type $MODEL_TYPE \
|
||||
--train_fps 24 \
|
||||
--shard_num=$SHARD_NUM \
|
||||
--shard_idx=$SHARD_IDX
|
||||
--train_fps 24
|
||||
|
||||
|
||||
|
||||
torchrun --nproc_per_node=$GPU_NUM \
|
||||
torchrun --nproc_per_node=$GPU_NUM \
|
||||
fastvideo/data_preprocess/preprocess_text_embeddings.py \
|
||||
--model_type $MODEL_TYPE \
|
||||
--model_path $MODEL_PATH \
|
||||
--output_dir=$OUTPUT_DIR
|
||||
|
||||
|
||||
|
||||
torchrun --nproc_per_node=1 \
|
||||
fastvideo/data_preprocess/preprocess_validation_text_embeddings.py \
|
||||
--model_type $MODEL_TYPE \
|
||||
@@ -0,0 +1,33 @@
|
||||
# export WANDB_MODE="offline"
|
||||
GPU_NUM=1 # 2,4,8
|
||||
MODEL_PATH="data/FastMochi-diffusers"
|
||||
MODEL_TYPE="mochi"
|
||||
DATA_MERGE_PATH="data/Image-Vid-Finetune-Src/merge.txt"
|
||||
OUTPUT_DIR="data/Image-Vid-Finetune-Mochi"
|
||||
VALIDATION_PATH="assets/prompt.txt"
|
||||
|
||||
torchrun --nproc_per_node=$GPU_NUM \
|
||||
fastvideo/data_preprocess/preprocess_vae_latents.py \
|
||||
--model_path $MODEL_PATH \
|
||||
--data_merge_path $DATA_MERGE_PATH \
|
||||
--train_batch_size=1 \
|
||||
--max_height=480 \
|
||||
--max_width=848 \
|
||||
--num_frames=93 \
|
||||
--dataloader_num_workers 1 \
|
||||
--output_dir=$OUTPUT_DIR \
|
||||
--model_type $MODEL_TYPE \
|
||||
--train_fps 24
|
||||
|
||||
torchrun --nproc_per_node=$GPU_NUM \
|
||||
fastvideo/data_preprocess/preprocess_text_embeddings.py \
|
||||
--model_type $MODEL_TYPE \
|
||||
--model_path $MODEL_PATH \
|
||||
--output_dir=$OUTPUT_DIR
|
||||
|
||||
torchrun --nproc_per_node=1 \
|
||||
fastvideo/data_preprocess/preprocess_validation_text_embeddings.py \
|
||||
--model_type $MODEL_TYPE \
|
||||
--model_path $MODEL_PATH \
|
||||
--output_dir=$OUTPUT_DIR \
|
||||
--validation_prompt_txt $VALIDATION_PATH
|
||||
@@ -1,9 +0,0 @@
|
||||
from huggingface_hub import HfApi
|
||||
|
||||
api = HfApi()
|
||||
|
||||
api.upload_folder(
|
||||
folder_path="/ephemeral/hao.zhang/codefolder/FastVideo-OSP/data/Hunyuan-Mixkit-Data",
|
||||
repo_id="FastVideo/Hunyuan-Distill-Data",
|
||||
repo_type="dataset",
|
||||
)
|
||||
@@ -1,38 +0,0 @@
|
||||
import torch
|
||||
from fastvideo.models.hunyuan.diffusion.pipelines.pipeline_hunyuan_video import (
|
||||
HunyuanVideoPipeline,
|
||||
)
|
||||
from fastvideo.models.hunyuan.modules.models import HYVideoDiffusionTransformer
|
||||
from fastvideo.models.hunyuan.vae.autoencoder_kl_causal_3d import AutoencoderKLCausal3D
|
||||
|
||||
|
||||
transformer = HYVideoDiffusionTransformer.from_pretrained(
|
||||
"data/hyvideo-diffusers", torch_dtype=torch.bfloat16, subfolder="transformer"
|
||||
)
|
||||
vae = AutoencoderKLCausal3D.from_pretrained(
|
||||
"data/hyvideo-diffusers", torch_dtype=torch.float16, subfolder="vae"
|
||||
)
|
||||
|
||||
pipe = HunyuanVideoPipeline.from_pretrained(
|
||||
"data/hyvideo-diffusers", transformer=transformer, vae=vae
|
||||
)
|
||||
pipe = pipe.to("cuda")
|
||||
pipe.vae.enable_tiling()
|
||||
|
||||
prompt = "Close-up, A little girl wearing a red hoodie in winter strikes a match. The sky is dark, there is a layer of snow on the ground, and it is still snowing lightly. The flame of the match flickers, illuminating the girl's face intermittently."
|
||||
|
||||
result = pipe(
|
||||
prompt,
|
||||
height=512,
|
||||
width=512,
|
||||
video_length=29,
|
||||
)
|
||||
|
||||
import PIL.Image
|
||||
from diffusers.utils import export_to_video
|
||||
|
||||
output = result.videos[0].permute(1, 2, 3, 0).detach().cpu().numpy()
|
||||
output = (output * 255).clip(0, 255).astype("uint8")
|
||||
output = [PIL.Image.fromarray(x) for x in output]
|
||||
|
||||
export_to_video(output, "output.mp4", fps=24)
|
||||
Reference in New Issue
Block a user