* Update Flow * Update Flow * Update Flow * add image recaptioning * Fix bug in t2v * update train_reward_lora.py * update reward training * Update V5.1 and mix multi text_encoders to one pipeline * Update V5.1 training Code * Update ComfyUI * Update Comment * Delete files * update reward training * Update Readme * fix extract frames in compute_semantic_consistency * Update Readme && Remove to in prediction * Update Demo * Update Readme * Update ui * support vae gradient checkpointing in reward training * Update Training Readme --------- Co-authored-by: hkunzhe <huangkunzhe.hkz@alibaba-inc.com>
9.8 KiB
Executable File
9.8 KiB
Executable File
Training Code
The default training commands for the different versions are as follows:
We can choose whether to use deep speed in EasyAnimate, which can save a lot of video memory.
The metadata_control.json is a little different from normal json in EasyAnimate, you need to add a control_file_path, and DWPose is suggested as tool to generate control file.
[
{
"file_path": "train/00000001.mp4",
"control_file_path": "control/00000001.mp4",
"text": "A group of young men in suits and sunglasses are walking down a city street.",
"type": "video"
},
{
"file_path": "train/00000002.jpg",
"control_file_path": "control/00000002.jpg",
"text": "A group of young men in suits and sunglasses are walking down a city street.",
"type": "image"
},
.....
]
Some parameters in the sh file can be confusing, and they are explained in this document:
enable_bucketis used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution.random_frame_cropis used for random cropping on video frames to simulate videos with different frame counts.random_hw_adaptis used to enable automatic height and width scaling for images and videos. Whenrandom_hw_adaptis enabled, the training images will have their height and width set toimage_sample_sizeas the maximum andmin(video_sample_size, 512)as the minimum. For training videos, the height and width will be set toimage_sample_sizeas the maximum andmin(video_sample_size, 512)as the minimum.- For example, when
random_hw_adaptis enabled, withvideo_sample_n_frames=49,video_sample_size=1024, andimage_sample_size=1024, the resolution of image inputs for training is512x512to1024x1024, and the resolution of video inputs for training is512x512x49to1024x1024x49. - For example, when
random_hw_adaptis enabled, withvideo_sample_n_frames=49,video_sample_size=1024, andimage_sample_size=256, the resolution of image inputs for training is256x256to1024x1024, and the resolution of video inputs for training is256x256x49.
- For example, when
training_with_video_token_lengthspecifies training the model according to token length. For training images and videos, the height and width will be set toimage_sample_sizeas the maximum andvideo_sample_sizeas the minimum.- For example, when
training_with_video_token_lengthis enabled, withvideo_sample_n_frames=49,token_sample_size=1024,video_sample_size=1024, andimage_sample_size=256, the resolution of image inputs for training is256x256to1024x1024, and the resolution of video inputs for training is256x256x49to1024x1024x49. - For example, when
training_with_video_token_lengthis enabled, withvideo_sample_n_frames=49,token_sample_size=512,video_sample_size=1024, andimage_sample_size=256, the resolution of image inputs for training is256x256to1024x1024, and the resolution of video inputs for training is256x256x49to1024x1024x9. - The token length for a video with dimensions 512x512 and 49 frames is 13,312. We need to set the
token_sample_size = 512.- At 512x512 resolution, the number of video frames is 49 (~= 512 * 512 * 49 / 512 / 512).
- At 768x768 resolution, the number of video frames is 21 (~= 512 * 512 * 49 / 768 / 768).
- At 1024x1024 resolution, the number of video frames is 9 (~= 512 * 512 * 49 / 1024 / 1024).
- These resolutions combined with their corresponding lengths allow the model to generate videos of different sizes.
- For example, when
loss_type: The loss type for training. Currently, flow is used in v5.1, ddpm is used in v5 and v4, sigma is used in v3, v2 and v1.
EasyAnimateV5.1 without deepspeed:
export MODEL_NAME="models/Diffusion_Transformer/EasyAnimateV5.1-12b-zh-Control"
export DATASET_NAME="datasets/internal_datasets/"
export DATASET_META_NAME="datasets/internal_datasets/metadata.json"
export NCCL_IB_DISABLE=1
export NCCL_P2P_DISABLE=1
NCCL_DEBUG=INFO
# When train model with multi machines, use "--config_file accelerate.yaml" instead of "--mixed_precision='bf16'".
accelerate launch --mixed_precision="bf16" scripts/train_control.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATASET_NAME \
--train_data_meta=$DATASET_META_NAME \
--config_path "config/easyanimate_video_v5.1_magvit_qwen.yaml" \
--image_sample_size=1024 \
--video_sample_size=256 \
--token_sample_size=512 \
--video_sample_stride=3 \
--video_sample_n_frames=49 \
--train_batch_size=1 \
--video_repeat=1 \
--gradient_accumulation_steps=1 \
--dataloader_num_workers=8 \
--num_train_epochs=100 \
--checkpointing_steps=100 \
--learning_rate=2e-05 \
--lr_scheduler="constant_with_warmup" \
--lr_warmup_steps=100 \
--seed=42 \
--output_dir="output_dir" \
--gradient_checkpointing \
--mixed_precision="bf16" \
--adam_weight_decay=5e-3 \
--adam_epsilon=1e-10 \
--vae_mini_batch=1 \
--max_grad_norm=0.05 \
--random_hw_adapt \
--training_with_video_token_length \
--enable_bucket \
--uniform_sampling \
--loss_type="flow" \
--train_mode="control_ref" \
--control_ref_image="first_frame" \
--trainable_modules "."
EasyAnimateV5.1 with deepspeed:
export MODEL_NAME="models/Diffusion_Transformer/EasyAnimateV5.1-12b-zh-Control"
export DATASET_NAME="datasets/internal_datasets/"
export DATASET_META_NAME="datasets/internal_datasets/metadata.json"
export NCCL_IB_DISABLE=1
export NCCL_P2P_DISABLE=1
NCCL_DEBUG=INFO
# When train model with multi machines, use "--config_file accelerate.yaml" instead of "--mixed_precision='bf16'".
accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_config.json --deepspeed_multinode_launcher standard scripts/train_control.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATASET_NAME \
--train_data_meta=$DATASET_META_NAME \
--config_path "config/easyanimate_video_v5.1_magvit_qwen.yaml" \
--image_sample_size=1024 \
--video_sample_size=256 \
--token_sample_size=512 \
--video_sample_stride=3 \
--video_sample_n_frames=49 \
--train_batch_size=1 \
--video_repeat=1 \
--gradient_accumulation_steps=1 \
--dataloader_num_workers=8 \
--num_train_epochs=100 \
--checkpointing_steps=100 \
--learning_rate=2e-05 \
--lr_scheduler="constant_with_warmup" \
--lr_warmup_steps=100 \
--seed=42 \
--output_dir="output_dir" \
--gradient_checkpointing \
--mixed_precision="bf16" \
--adam_weight_decay=5e-3 \
--adam_epsilon=1e-10 \
--vae_mini_batch=1 \
--max_grad_norm=0.05 \
--random_hw_adapt \
--training_with_video_token_length \
--enable_bucket \
--uniform_sampling \
--loss_type="flow" \
--train_mode="control_ref" \
--control_ref_image="first_frame" \
--use_deepspeed \
--trainable_modules "."
EasyAnimateV5 without deepspeed:
export MODEL_NAME="models/Diffusion_Transformer/EasyAnimateV5-12b-zh-Control"
export DATASET_NAME="datasets/internal_datasets/"
export DATASET_META_NAME="datasets/internal_datasets/metadata.json"
export NCCL_IB_DISABLE=1
export NCCL_P2P_DISABLE=1
NCCL_DEBUG=INFO
# When train model with multi machines, use "--config_file accelerate.yaml" instead of "--mixed_precision='bf16'".
accelerate launch --mixed_precision="bf16" scripts/train_control.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATASET_NAME \
--train_data_meta=$DATASET_META_NAME \
--config_path "config/easyanimate_video_v5_magvit_multi_text_encoder.yaml" \
--image_sample_size=1024 \
--video_sample_size=256 \
--token_sample_size=512 \
--video_sample_stride=3 \
--video_sample_n_frames=49 \
--train_batch_size=1 \
--video_repeat=1 \
--gradient_accumulation_steps=1 \
--dataloader_num_workers=8 \
--num_train_epochs=100 \
--checkpointing_steps=100 \
--learning_rate=2e-05 \
--lr_scheduler="constant_with_warmup" \
--lr_warmup_steps=100 \
--seed=42 \
--output_dir="output_dir" \
--gradient_checkpointing \
--mixed_precision="bf16" \
--adam_weight_decay=5e-3 \
--adam_epsilon=1e-10 \
--vae_mini_batch=1 \
--max_grad_norm=0.05 \
--random_hw_adapt \
--training_with_video_token_length \
--loss_type="ddpm" \
--enable_bucket \
--uniform_sampling \
--trainable_modules "."
EasyAnimateV5 with deepspeed:
export MODEL_NAME="models/Diffusion_Transformer/EasyAnimateV5-12b-zh-Control"
export DATASET_NAME="datasets/internal_datasets/"
export DATASET_META_NAME="datasets/internal_datasets/metadata.json"
export NCCL_IB_DISABLE=1
export NCCL_P2P_DISABLE=1
NCCL_DEBUG=INFO
# When train model with multi machines, use "--config_file accelerate.yaml" instead of "--mixed_precision='bf16'".
accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_config.json --deepspeed_multinode_launcher standard scripts/train_control.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATASET_NAME \
--train_data_meta=$DATASET_META_NAME \
--config_path "config/easyanimate_video_v5_magvit_multi_text_encoder.yaml" \
--image_sample_size=1024 \
--video_sample_size=256 \
--token_sample_size=512 \
--video_sample_stride=3 \
--video_sample_n_frames=49 \
--train_batch_size=1 \
--video_repeat=1 \
--gradient_accumulation_steps=1 \
--dataloader_num_workers=8 \
--num_train_epochs=100 \
--checkpointing_steps=100 \
--learning_rate=2e-05 \
--lr_scheduler="constant_with_warmup" \
--lr_warmup_steps=100 \
--seed=42 \
--output_dir="output_dir" \
--gradient_checkpointing \
--mixed_precision="bf16" \
--adam_weight_decay=5e-3 \
--adam_epsilon=1e-10 \
--vae_mini_batch=1 \
--max_grad_norm=0.05 \
--random_hw_adapt \
--training_with_video_token_length \
--loss_type="ddpm" \
--enable_bucket \
--uniform_sampling \
--use_deepspeed \
--trainable_modules "."