diff --git a/scripts/cogvideox_fun/README_TRAIN.md b/scripts/cogvideox_fun/README_TRAIN.md index fec7675..b58198b 100755 --- a/scripts/cogvideox_fun/README_TRAIN.md +++ b/scripts/cogvideox_fun/README_TRAIN.md @@ -7,6 +7,15 @@ We can choose whether to use deepspeed and fsdp in CogVideoX-Fun, which can save Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -36,8 +45,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -79,8 +88,8 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -124,8 +133,8 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ diff --git a/scripts/cogvideox_fun/README_TRAIN_CONTROL.md b/scripts/cogvideox_fun/README_TRAIN_CONTROL.md index 9957cee..c63597c 100755 --- a/scripts/cogvideox_fun/README_TRAIN_CONTROL.md +++ b/scripts/cogvideox_fun/README_TRAIN_CONTROL.md @@ -27,6 +27,15 @@ The metadata_control.json is a little different from normal json in CogVideoX-Fu Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -55,8 +64,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train_control.p --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -97,8 +106,8 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -141,8 +150,8 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ diff --git a/scripts/cogvideox_fun/README_TRAIN_LORA.md b/scripts/cogvideox_fun/README_TRAIN_LORA.md index 1d05f9e..187401f 100755 --- a/scripts/cogvideox_fun/README_TRAIN_LORA.md +++ b/scripts/cogvideox_fun/README_TRAIN_LORA.md @@ -5,6 +5,15 @@ We can choose whether to use deepspeed and fsdp in CogVideoX-Fun, which can save Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -39,8 +48,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -84,8 +93,8 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -130,8 +139,8 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ diff --git a/scripts/cogvideox_fun/train.sh b/scripts/cogvideox_fun/train.sh index 154559e..d3e7508 100755 --- a/scripts/cogvideox_fun/train.sh +++ b/scripts/cogvideox_fun/train.sh @@ -10,8 +10,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -50,8 +50,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train.py \ # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ +# --image_sample_size=512 \ +# --video_sample_size=512 \ # --token_sample_size=512 \ # --video_sample_stride=3 \ # --video_sample_n_frames=85 \ diff --git a/scripts/cogvideox_fun/train_control.sh b/scripts/cogvideox_fun/train_control.sh index bf9eb52..1ee000e 100755 --- a/scripts/cogvideox_fun/train_control.sh +++ b/scripts/cogvideox_fun/train_control.sh @@ -10,8 +10,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train_control.p --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -49,8 +49,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train_control.p # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ +# --image_sample_size=512 \ +# --video_sample_size=512 \ # --token_sample_size=512 \ # --video_sample_stride=3 \ # --video_sample_n_frames=85 \ diff --git a/scripts/cogvideox_fun/train_lora.sh b/scripts/cogvideox_fun/train_lora.sh index 3243e53..19ec0ac 100755 --- a/scripts/cogvideox_fun/train_lora.sh +++ b/scripts/cogvideox_fun/train_lora.sh @@ -10,8 +10,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ + --image_sample_size=512 \ + --video_sample_size=512 \ --token_sample_size=512 \ --video_sample_stride=3 \ --video_sample_n_frames=49 \ @@ -52,8 +52,8 @@ accelerate launch --mixed_precision="bf16" scripts/cogvideox_fun/train_lora.py \ # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ +# --image_sample_size=512 \ +# --video_sample_size=512 \ # --token_sample_size=512 \ # --video_sample_stride=3 \ # --video_sample_n_frames=85 \ diff --git a/scripts/hunyuanvideo/README_TRAIN.md b/scripts/hunyuanvideo/README_TRAIN.md index 34e0ac7..efd9e5f 100644 --- a/scripts/hunyuanvideo/README_TRAIN.md +++ b/scripts/hunyuanvideo/README_TRAIN.md @@ -7,6 +7,15 @@ We can choose whether to use DeepSpeed and FSDP in HunyuanVideo, which can save Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -48,9 +57,9 @@ accelerate launch --mixed_precision="bf16" scripts/hunyuanvideo/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -94,9 +103,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -140,9 +149,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -186,9 +195,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/hunyuanvideo/README_TRAIN_LORA.md b/scripts/hunyuanvideo/README_TRAIN_LORA.md index 818c170..c969fdf 100644 --- a/scripts/hunyuanvideo/README_TRAIN_LORA.md +++ b/scripts/hunyuanvideo/README_TRAIN_LORA.md @@ -7,6 +7,15 @@ We can choose whether to use DeepSpeed and FSDP in HunyuanVideo, which can save Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -52,9 +61,9 @@ accelerate launch --mixed_precision="bf16" scripts/hunyuanvideo/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -99,9 +108,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -146,9 +155,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -193,9 +202,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/hunyuanvideo/train.sh b/scripts/hunyuanvideo/train.sh index 0a36ba3..2ba2234 100644 --- a/scripts/hunyuanvideo/train.sh +++ b/scripts/hunyuanvideo/train.sh @@ -10,9 +10,9 @@ accelerate launch --mixed_precision="bf16" scripts/hunyuanvideo/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/hunyuanvideo/train_lora.sh b/scripts/hunyuanvideo/train_lora.sh index ca06dea..6200750 100644 --- a/scripts/hunyuanvideo/train_lora.sh +++ b/scripts/hunyuanvideo/train_lora.sh @@ -10,9 +10,9 @@ accelerate launch --mixed_precision="bf16" scripts/hunyuanvideo/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/qwenimage/README_TRAIN.md b/scripts/qwenimage/README_TRAIN.md index 94ff338..6b2de6b 100755 --- a/scripts/qwenimage/README_TRAIN.md +++ b/scripts/qwenimage/README_TRAIN.md @@ -37,7 +37,7 @@ accelerate launch --mixed_precision="bf16" scripts/qwenimage/train.py \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -74,7 +74,7 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -119,7 +119,7 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -156,7 +156,7 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ diff --git a/scripts/qwenimage/README_TRAIN_EDIT.md b/scripts/qwenimage/README_TRAIN_EDIT.md index 98631cb..a5db00a 100755 --- a/scripts/qwenimage/README_TRAIN_EDIT.md +++ b/scripts/qwenimage/README_TRAIN_EDIT.md @@ -63,7 +63,7 @@ accelerate launch --mixed_precision="bf16" scripts/qwenimage/train_edit.py \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -101,7 +101,7 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -147,7 +147,7 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -185,7 +185,7 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ diff --git a/scripts/qwenimage/README_TRAIN_LORA.md b/scripts/qwenimage/README_TRAIN_LORA.md index 2804919..628ec41 100755 --- a/scripts/qwenimage/README_TRAIN_LORA.md +++ b/scripts/qwenimage/README_TRAIN_LORA.md @@ -41,7 +41,7 @@ accelerate launch --mixed_precision="bf16" scripts/qwenimage/train_lora.py \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -79,7 +79,7 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -127,7 +127,7 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ @@ -161,7 +161,7 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ diff --git a/scripts/qwenimage/train.sh b/scripts/qwenimage/train.sh index 2814055..2e81f31 100644 --- a/scripts/qwenimage/train.sh +++ b/scripts/qwenimage/train.sh @@ -11,7 +11,7 @@ accelerate launch --mixed_precision="bf16" scripts/qwenimage/train.py \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ diff --git a/scripts/qwenimage/train_edit.sh b/scripts/qwenimage/train_edit.sh index 14e8a9d..50c6f51 100644 --- a/scripts/qwenimage/train_edit.sh +++ b/scripts/qwenimage/train_edit.sh @@ -11,7 +11,7 @@ accelerate launch --mixed_precision="bf16" scripts/qwenimage/train_edit.py \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ diff --git a/scripts/qwenimage/train_edit_lora.sh b/scripts/qwenimage/train_edit_lora.sh index 6450449..ec17fe7 100644 --- a/scripts/qwenimage/train_edit_lora.sh +++ b/scripts/qwenimage/train_edit_lora.sh @@ -11,7 +11,7 @@ accelerate launch --mixed_precision="bf16" scripts/qwenimage/train_edit_lora.py --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ diff --git a/scripts/qwenimage/train_lora.sh b/scripts/qwenimage/train_lora.sh index c273ad3..7a35068 100644 --- a/scripts/qwenimage/train_lora.sh +++ b/scripts/qwenimage/train_lora.sh @@ -11,7 +11,7 @@ accelerate launch --mixed_precision="bf16" scripts/qwenimage/train_lora.py \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ --train_batch_size=1 \ - --image_sample_size=1024 \ + --image_sample_size=1328 \ --gradient_accumulation_steps=1 \ --dataloader_num_workers=8 \ --num_train_epochs=100 \ diff --git a/scripts/wan2.1/README_TRAIN.md b/scripts/wan2.1/README_TRAIN.md index d2c769d..63c82f8 100755 --- a/scripts/wan2.1/README_TRAIN.md +++ b/scripts/wan2.1/README_TRAIN.md @@ -7,6 +7,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan, which can save a lot of Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -50,9 +59,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -99,9 +108,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -155,9 +164,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -204,9 +213,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1/README_TRAIN_DISTILL.md b/scripts/wan2.1/README_TRAIN_DISTILL.md index 785a57e..4adb31d 100755 --- a/scripts/wan2.1/README_TRAIN_DISTILL.md +++ b/scripts/wan2.1/README_TRAIN_DISTILL.md @@ -7,6 +7,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan distill, which can save a Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. diff --git a/scripts/wan2.1/README_TRAIN_DISTILL_LORA.md b/scripts/wan2.1/README_TRAIN_DISTILL_LORA.md index 5da4d08..3c3c6ea 100755 --- a/scripts/wan2.1/README_TRAIN_DISTILL_LORA.md +++ b/scripts/wan2.1/README_TRAIN_DISTILL_LORA.md @@ -7,6 +7,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan distill, which can save a Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. diff --git a/scripts/wan2.1/README_TRAIN_LORA.md b/scripts/wan2.1/README_TRAIN_LORA.md index f04d036..345ffb9 100755 --- a/scripts/wan2.1/README_TRAIN_LORA.md +++ b/scripts/wan2.1/README_TRAIN_LORA.md @@ -5,6 +5,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan, which can save a lot of Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -52,9 +61,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -101,9 +110,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -158,9 +167,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -202,9 +211,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1/train.sh b/scripts/wan2.1/train.sh index 2756fe6..0afb828 100755 --- a/scripts/wan2.1/train.sh +++ b/scripts/wan2.1/train.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -54,9 +54,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1/train.py \ # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ -# --token_sample_size=512 \ +# --image_sample_size=640 \ +# --video_sample_size=640 \ +# --token_sample_size=640 \ # --video_sample_stride=2 \ # --video_sample_n_frames=81 \ # --train_batch_size=1 \ diff --git a/scripts/wan2.1/train_lora.sh b/scripts/wan2.1/train_lora.sh index 3738616..be0ece9 100755 --- a/scripts/wan2.1/train_lora.sh +++ b/scripts/wan2.1/train_lora.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -54,9 +54,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1/train_lora.py \ # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ -# --token_sample_size=512 \ +# --image_sample_size=640 \ +# --video_sample_size=640 \ +# --token_sample_size=640 \ # --video_sample_stride=2 \ # --video_sample_n_frames=81 \ # --train_batch_size=1 \ diff --git a/scripts/wan2.1_fun/README_TRAIN.md b/scripts/wan2.1_fun/README_TRAIN.md index 61eebd1..d3ddd7b 100755 --- a/scripts/wan2.1_fun/README_TRAIN.md +++ b/scripts/wan2.1_fun/README_TRAIN.md @@ -5,6 +5,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan, which can save a lot of Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -48,9 +57,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -97,9 +106,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -153,9 +162,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -202,9 +211,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1_fun/README_TRAIN_CONTROL.md b/scripts/wan2.1_fun/README_TRAIN_CONTROL.md index 45077dd..00fea70 100755 --- a/scripts/wan2.1_fun/README_TRAIN_CONTROL.md +++ b/scripts/wan2.1_fun/README_TRAIN_CONTROL.md @@ -25,6 +25,15 @@ The metadata_control.json is a little different from normal json in Wan-Fun, you Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -72,9 +81,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_control.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -120,9 +129,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -178,9 +187,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -229,9 +238,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -281,9 +290,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_control.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -328,9 +337,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -383,9 +392,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1_fun/README_TRAIN_CONTROL_LORA.md b/scripts/wan2.1_fun/README_TRAIN_CONTROL_LORA.md index cf82ff9..9e85ee7 100755 --- a/scripts/wan2.1_fun/README_TRAIN_CONTROL_LORA.md +++ b/scripts/wan2.1_fun/README_TRAIN_CONTROL_LORA.md @@ -25,6 +25,15 @@ The metadata_control.json is a little different from normal json in Wan-Fun, you Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -76,9 +85,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_control_lora --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -125,9 +134,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -186,9 +195,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -235,9 +244,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -289,9 +298,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_control_lora --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -333,9 +342,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -385,9 +394,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1_fun/README_TRAIN_LORA.md b/scripts/wan2.1_fun/README_TRAIN_LORA.md index 9673306..19c1385 100755 --- a/scripts/wan2.1_fun/README_TRAIN_LORA.md +++ b/scripts/wan2.1_fun/README_TRAIN_LORA.md @@ -5,6 +5,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan, which can save a lot of Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -52,9 +61,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -102,9 +111,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -161,9 +170,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -208,9 +217,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1_fun/train.sh b/scripts/wan2.1_fun/train.sh index a57e9f6..385f033 100755 --- a/scripts/wan2.1_fun/train.sh +++ b/scripts/wan2.1_fun/train.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -54,9 +54,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train.py \ # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ -# --token_sample_size=512 \ +# --image_sample_size=640 \ +# --video_sample_size=640 \ +# --token_sample_size=640 \ # --video_sample_stride=2 \ # --video_sample_n_frames=81 \ # --train_batch_size=1 \ diff --git a/scripts/wan2.1_fun/train_control.sh b/scripts/wan2.1_fun/train_control.sh index 5de5174..4cf7fbd 100755 --- a/scripts/wan2.1_fun/train_control.sh +++ b/scripts/wan2.1_fun/train_control.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_control.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1_fun/train_control_lora.sh b/scripts/wan2.1_fun/train_control_lora.sh index 24c9de5..a5df6dd 100755 --- a/scripts/wan2.1_fun/train_control_lora.sh +++ b/scripts/wan2.1_fun/train_control_lora.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_control_lora --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1_fun/train_lora.sh b/scripts/wan2.1_fun/train_lora.sh index 1445d3f..4d0aa0c 100755 --- a/scripts/wan2.1_fun/train_lora.sh +++ b/scripts/wan2.1_fun/train_lora.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -55,9 +55,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_fun/train_lora.py \ # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ -# --token_sample_size=512 \ +# --image_sample_size=640 \ +# --video_sample_size=640 \ +# --token_sample_size=640 \ # --video_sample_stride=2 \ # --video_sample_n_frames=81 \ # --train_batch_size=1 \ diff --git a/scripts/wan2.1_vace/README_TRAIN.md b/scripts/wan2.1_vace/README_TRAIN.md index 52aeaed..d04424f 100755 --- a/scripts/wan2.1_vace/README_TRAIN.md +++ b/scripts/wan2.1_vace/README_TRAIN.md @@ -27,6 +27,15 @@ The metadata_control.json is a little different from normal json in Wan-Fun, you Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -73,9 +82,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_vace/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -119,9 +128,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -174,9 +183,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -222,9 +231,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.1_vace/train.sh b/scripts/wan2.1_vace/train.sh index 0398836..867964f 100644 --- a/scripts/wan2.1_vace/train.sh +++ b/scripts/wan2.1_vace/train.sh @@ -10,9 +10,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.1_vace/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2/README_TRAIN.md b/scripts/wan2.2/README_TRAIN.md index b49ab6d..aa76feb 100755 --- a/scripts/wan2.2/README_TRAIN.md +++ b/scripts/wan2.2/README_TRAIN.md @@ -7,6 +7,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan, which can save a lot of Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -53,9 +62,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -103,9 +112,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -160,9 +169,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -210,9 +219,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -258,9 +267,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2/README_TRAIN_ANIMATE.md b/scripts/wan2.2/README_TRAIN_ANIMATE.md index e46b47e..08d7ea3 100755 --- a/scripts/wan2.2/README_TRAIN_ANIMATE.md +++ b/scripts/wan2.2/README_TRAIN_ANIMATE.md @@ -79,9 +79,8 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train_animate.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -126,9 +125,8 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -180,9 +178,8 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -228,9 +225,8 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2/README_TRAIN_DISTILL.md b/scripts/wan2.2/README_TRAIN_DISTILL.md index 22f7fab..bfffc59 100755 --- a/scripts/wan2.2/README_TRAIN_DISTILL.md +++ b/scripts/wan2.2/README_TRAIN_DISTILL.md @@ -7,6 +7,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan distill, which can save a Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. diff --git a/scripts/wan2.2/README_TRAIN_DISTILL_LORA.md b/scripts/wan2.2/README_TRAIN_DISTILL_LORA.md index 053a402..8dc6a19 100755 --- a/scripts/wan2.2/README_TRAIN_DISTILL_LORA.md +++ b/scripts/wan2.2/README_TRAIN_DISTILL_LORA.md @@ -7,6 +7,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan distill, which can save a Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. diff --git a/scripts/wan2.2/README_TRAIN_LORA.md b/scripts/wan2.2/README_TRAIN_LORA.md index 4013f70..69aa36e 100755 --- a/scripts/wan2.2/README_TRAIN_LORA.md +++ b/scripts/wan2.2/README_TRAIN_LORA.md @@ -5,6 +5,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan, which can save a lot of Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -55,9 +64,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -106,9 +115,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -165,9 +174,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -211,9 +220,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -260,9 +269,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2/train.sh b/scripts/wan2.2/train.sh index 7891013..1224fb6 100644 --- a/scripts/wan2.2/train.sh +++ b/scripts/wan2.2/train.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -58,9 +58,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train.py \ # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ -# --token_sample_size=512 \ +# --image_sample_size=640 \ +# --video_sample_size=640 \ +# --token_sample_size=640 \ # --video_sample_stride=2 \ # --video_sample_n_frames=81 \ # --train_batch_size=1 \ diff --git a/scripts/wan2.2/train_lora.sh b/scripts/wan2.2/train_lora.sh index d008bde..112b669 100755 --- a/scripts/wan2.2/train_lora.sh +++ b/scripts/wan2.2/train_lora.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -59,9 +59,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2/train_lora.py \ # --pretrained_model_name_or_path=$MODEL_NAME \ # --train_data_dir=$DATASET_NAME \ # --train_data_meta=$DATASET_META_NAME \ -# --image_sample_size=1024 \ -# --video_sample_size=256 \ -# --token_sample_size=512 \ +# --image_sample_size=640 \ +# --video_sample_size=640 \ +# --token_sample_size=640 \ # --video_sample_stride=2 \ # --video_sample_n_frames=81 \ # --train_batch_size=1 \ diff --git a/scripts/wan2.2_fun/README_TRAIN.md b/scripts/wan2.2_fun/README_TRAIN.md index ac33120..a938759 100755 --- a/scripts/wan2.2_fun/README_TRAIN.md +++ b/scripts/wan2.2_fun/README_TRAIN.md @@ -5,6 +5,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan, which can save a lot of Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -51,9 +60,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_fun/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -101,9 +110,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -158,9 +167,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -208,9 +217,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -256,9 +265,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_fun/README_TRAIN_CONTROL.md b/scripts/wan2.2_fun/README_TRAIN_CONTROL.md index dcd1595..ad85200 100755 --- a/scripts/wan2.2_fun/README_TRAIN_CONTROL.md +++ b/scripts/wan2.2_fun/README_TRAIN_CONTROL.md @@ -25,6 +25,15 @@ The metadata_control.json is a little different from normal json in Wan-Fun, you Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -76,9 +85,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_fun/train_control.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -126,9 +135,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -186,9 +195,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -239,9 +248,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -290,9 +299,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA.md b/scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA.md index f874f7a..2c28cf8 100755 --- a/scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA.md +++ b/scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA.md @@ -25,6 +25,15 @@ The metadata_control.json is a little different from normal json in Wan-Fun, you Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -80,9 +89,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_fun/train_control_lora --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -131,9 +140,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -194,9 +203,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -245,9 +254,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -299,9 +308,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_fun/README_TRAIN_LORA.md b/scripts/wan2.2_fun/README_TRAIN_LORA.md index 25b9421..018b768 100755 --- a/scripts/wan2.2_fun/README_TRAIN_LORA.md +++ b/scripts/wan2.2_fun/README_TRAIN_LORA.md @@ -5,6 +5,15 @@ We can choose whether to use DeepSpeed and FSDP in Wan, which can save a lot of Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -55,9 +64,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_fun/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -106,9 +115,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -165,9 +174,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -211,9 +220,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -260,9 +269,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_fun/train.sh b/scripts/wan2.2_fun/train.sh index 0f6d27c..dc1efb1 100644 --- a/scripts/wan2.2_fun/train.sh +++ b/scripts/wan2.2_fun/train.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_fun/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_fun/train_control.sh b/scripts/wan2.2_fun/train_control.sh index 5fd33c6..0ae5eab 100644 --- a/scripts/wan2.2_fun/train_control.sh +++ b/scripts/wan2.2_fun/train_control.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_fun/train_control.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_fun/train_control_lora.sh b/scripts/wan2.2_fun/train_control_lora.sh index d7cb822..ebafdef 100644 --- a/scripts/wan2.2_fun/train_control_lora.sh +++ b/scripts/wan2.2_fun/train_control_lora.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_fun/train_control_lora --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_fun/train_lora.sh b/scripts/wan2.2_fun/train_lora.sh index 5f13f97..3ca2d05 100644 --- a/scripts/wan2.2_fun/train_lora.sh +++ b/scripts/wan2.2_fun/train_lora.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_fun/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_vace_fun/README_TRAIN.md b/scripts/wan2.2_vace_fun/README_TRAIN.md index 258dbd2..221e035 100755 --- a/scripts/wan2.2_vace_fun/README_TRAIN.md +++ b/scripts/wan2.2_vace_fun/README_TRAIN.md @@ -27,6 +27,15 @@ The metadata_control.json is a little different from normal json in Wan-Fun, you Some parameters in the sh file can be confusing, and they are explained in this document: - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images and videos at the center, but instead, it trains the entire images and videos after grouping them into buckets based on resolution. +- Sample size Configuration Guide + - `video_sample_size` represents the resolution size of videos; when `random_hw_adapt` is True, it represents the minimum value between video and image resolutions. + - `image_sample_size` represents the resolution size of images; when `random_hw_adapt` is True, it represents the maximum value between video and image resolutions. + - `token_sample_size` represents the resolution corresponding to the maximum token length when `training_with_video_token_length` is True. + - Due to potential confusion in configuration, **if you don't require arbitrary resolution for finetuning**, it is recommended to set `video_sample_size`, `image_sample_size`, and `token_sample_size` to the same fixed value, such as **(320, 480, 512, 640, 960)**. + - **All set to 320** represents **240P**. + - **All set to 480** represents **320P**. + - **All set to 640** represents **480P**. + - **All set to 960** represents **720P**. - `random_frame_crop` is used for random cropping on video frames to simulate videos with different frame counts. - `random_hw_adapt` is used to enable automatic height and width scaling for images and videos. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. For training videos, the height and width will be set to `image_sample_size` as the maximum and `min(video_sample_size, 512)` as the minimum. - For example, when `random_hw_adapt` is enabled, with `video_sample_n_frames=49`, `video_sample_size=1024`, and `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024`, and the resolution of video inputs for training is `512x512x49` to `1024x1024x49`. @@ -74,9 +83,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_vace_fun/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -121,9 +130,9 @@ accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_con --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -177,9 +186,9 @@ accelerate launch --zero_stage 3 --zero3_save_16bit_model true --zero3_init_flag --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ @@ -226,9 +235,9 @@ accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TR --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \ diff --git a/scripts/wan2.2_vace_fun/train.sh b/scripts/wan2.2_vace_fun/train.sh index c207afb..26c59c7 100644 --- a/scripts/wan2.2_vace_fun/train.sh +++ b/scripts/wan2.2_vace_fun/train.sh @@ -11,9 +11,9 @@ accelerate launch --mixed_precision="bf16" scripts/wan2.2_vace_fun/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ - --image_sample_size=1024 \ - --video_sample_size=256 \ - --token_sample_size=512 \ + --image_sample_size=640 \ + --video_sample_size=640 \ + --token_sample_size=640 \ --video_sample_stride=2 \ --video_sample_n_frames=81 \ --train_batch_size=1 \