20 KiB
MOVA Full Training Guide
This document provides a complete workflow for MOVA (Audio-Video Generation Model) full parameter training, including environment setup, data preparation, distributed training, and inference testing.
Note
: MOVA is an audio-video generation model that can simultaneously generate video and corresponding audio. Training data requires both video and audio files.
Table of Contents
- 1. Environment Setup
- 2. Data Preparation
- 3. Full Parameter Training
- 4. Inference Testing
- 5. Additional Resources
1. Environment Setup
Option 1: Using requirements.txt
pip install -r requirements.txt
Option 2: Manual Installation
pip install Pillow einops safetensors timm tomesd librosa transformers accelerate diffusers peft decord imageio imageio-ffmpeg moviepy ftfy tensorboard sentencepiece modelscope
2. Data Preparation
2.1 Quick Test Dataset
We provide a test dataset containing several video-audio training samples.
# Download demo dataset
modelscope download --dataset PAI/X-Fun-Videos-Audios-Demo --local_dir ./datasets/X-Fun-Videos-Audios-Demo
2.2 Dataset Structure
📦 datasets/
├── 📂 my_dataset/
│ ├── 📂 train/
│ │ ├── 📄 video001.mp4
│ │ ├── 📄 video002.mp4
│ │ └── 📄 ...
│ ├── 📂 wav/
│ │ ├── 📄 audio001.wav
│ │ ├── 📄 audio002.wav
│ │ └── 📄 ...
│ └── 📄 metadata.json
2.3 metadata.json Format
⚠️ Important: MOVA is an audio-video generation model. Unlike normal video training, you must provide the
audio_pathfield in metadata.json.
Relative Path Format (Example):
[
{
"file_path": "train/video001.mp4",
"audio_path": "wav/audio001.wav",
"text": "A brown dog barks on a sofa, sitting on a light-colored couch in a cozy room",
"type": "video",
"width": 768,
"height": 512
},
{
"file_path": "train/video002.mp4",
"audio_path": "wav/audio002.wav",
"text": "A group of young men in suits and sunglasses are walking down a city street",
"type": "video",
"width": 640,
"height": 640
}
]
Absolute Path Format:
[
{
"file_path": "/mnt/data/videos/dog.mp4",
"audio_path": "/mnt/data/wavs/dog.wav",
"text": "A brown dog barks on a sofa",
"type": "video",
"width": 768,
"height": 512
}
]
Key Fields:
file_path: Video file path (relative or absolute)audio_path: Audio file path (MOVA-specific and required, the main difference from regular video training)- Audio files are typically in
.wavformat - The path should correspond to
file_path, e.g.,train/video001.mp4corresponds towav/audio001.wav
- Audio files are typically in
text: Video description (English prompt)type: Data type, fixed as"video"width/height: Video dimensions (recommended to enable bucket training; if not provided, they will be automatically read during training)
2.4 Using Relative and Absolute Paths
Relative Paths:
If your data uses relative paths, configure the training script as follows:
export DATASET_NAME="datasets/X-Fun-Videos-Audios-Demo/"
export DATASET_META_NAME="datasets/X-Fun-Videos-Audios-Demo/metadata_add_width_height.json"
Absolute Paths:
If your data uses absolute paths, configure the training script as follows:
export DATASET_NAME=""
export DATASET_META_NAME="/mnt/data/metadata_add_width_height.json"
💡 Tip: If the dataset is small and stored locally, use relative paths. If the dataset is stored on external storage (e.g., NAS, OSS) or shared across multiple machines, use absolute paths.
3. Full Parameter Training
3.1 Download Pretrained Model
# Create model directory
mkdir -p models/Diffusion_Transformer
hf download OpenMOSS-Team/MOVA-360p --local-dir models/Diffusion_Transformer/MOVA-360p
3.2 Quick Start (DeepSpeed-Zero-2)
If you have downloaded the data as per 2.1 Quick Test Dataset and the weights as per 3.1 Download Pretrained Model, you can directly copy and run the quick start command.
FSDP training is recommended as it can significantly save VRAM.
export MODEL_NAME="models/Diffusion_Transformer/MOVA-360p"
export DATASET_NAME="datasets/X-Fun-Videos-Audios-Demo/"
export DATASET_META_NAME="datasets/X-Fun-Videos-Audios-Demo/metadata_add_width_height.json"
# NCCL_IB_DISABLE=1 and NCCL_P2P_DISABLE=1 are used in multi-node environments without RDMA
# export NCCL_IB_DISABLE=1
# export NCCL_P2P_DISABLE=1
NCCL_DEBUG=INFO
accelerate launch --mixed_precision="bf16" scripts/mova/train.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATASET_NAME \
--train_data_meta=$DATASET_META_NAME \
--image_sample_size=480 \
--video_sample_size=480 \
--token_sample_size=480 \
--video_sample_stride=1 \
--video_sample_n_frames=193 \
--train_batch_size=1 \
--video_repeat=1 \
--gradient_accumulation_steps=1 \
--dataloader_num_workers=8 \
--num_train_epochs=100 \
--checkpointing_steps=100 \
--learning_rate=2e-05 \
--lr_scheduler="constant_with_warmup" \
--lr_warmup_steps=100 \
--seed=42 \
--output_dir="output_dir_mova" \
--gradient_checkpointing \
--mixed_precision="bf16" \
--adam_weight_decay=3e-2 \
--adam_epsilon=1e-10 \
--vae_mini_batch=1 \
--max_grad_norm=0.05 \
--random_hw_adapt \
--training_with_video_token_length \
--enable_bucket \
--uniform_sampling \
--low_vram \
--trainable_modules "." \
--boundary_type="low" \
--boundary_ratio=0.9 \
--train_components="transformer,transformer_2"
3.3 Common Training Parameters
Core Parameters:
| Parameter | Description | Example Value |
|---|---|---|
--pretrained_model_name_or_path |
Pretrained model path | models/Diffusion_Transformer/MOVA-360p |
--train_data_dir |
Training data directory | datasets/X-Fun-Videos-Audios-Demo/ |
--train_data_meta |
Training data metadata file | datasets/X-Fun-Videos-Audios-Demo/metadata_add_width_height.json |
--train_batch_size |
Number of samples per batch | 1 |
--image_sample_size |
Maximum training resolution for auto bucket | 480 |
--video_sample_size |
Maximum video training resolution | 480 |
--token_sample_size |
Token length sampling size | 480 |
--video_sample_stride |
Frame sampling stride | 1 |
--video_sample_n_frames |
Number of video frames to sample | 193 |
--num_train_epochs |
Number of training epochs | 100 |
--learning_rate |
Learning rate | 2e-05 |
--output_dir |
Output directory for checkpoints | output_dir_mova |
--checkpointing_steps |
Save checkpoint every N steps | 100 |
Some parameters in the sh file can be confusing, and they are explained in this document:
enable_bucket: Used to enable bucket training. When enabled, the model does not crop the videos at the center, but instead trains the videos after grouping them into buckets based on resolution.random_frame_crop: Used for random cropping on video frames to simulate videos with different frame counts.random_hw_adapt: Used to enable automatic height and width scaling for videos. Whenrandom_hw_adaptis enabled, for training videos, the height and width will be set tovideo_sample_sizeas the maximum and512as the minimum.- For example, when
random_hw_adaptis enabled, withvideo_sample_n_frames=49,video_sample_size=768, the resolution of video inputs for training is512x512x49,768x768x49.
- For example, when
training_with_video_token_length: Specifies training the model according to token length. For training videos, the height and width will be set tovideo_sample_sizeas the maximum and256as the minimum.- For example, when
training_with_video_token_lengthis enabled, withvideo_sample_n_frames=49,token_sample_size=512,video_sample_size=768, the resolution of video inputs for training is256x256x49,512x512x49,768x768x21. - The token length for a video with dimensions 512x512 and 49 frames is 13,312. We need to set the
token_sample_size = 512.- At 512x512 resolution, the number of video frames is 49 (~= 512 * 512 * 49 / 512 / 512).
- At 768x768 resolution, the number of video frames is 21 (~= 512 * 512 * 49 / 768 / 768).
- At 1024x1024 resolution, the number of video frames is 9 (~= 512 * 512 * 49 / 1024 / 1024).
- These resolutions combined with their corresponding lengths allow the model to generate videos of different sizes.
- For example, when
resume_from_checkpoint: Used to set whether training should be resumed from a previous checkpoint. Use a path or"latest"to automatically select the last available checkpoint.trainable_modules: Represents the modules to be trained, using "." means training all modules.boundary_type: Specifies which DiT to train: "low" = only low-noise DiT, "high" = only high-noise DiT, "full" = both DiTs.boundary_ratio: Is the boundary ratio for switching between high-noise and low-noise DiT. Timesteps below this ratio use low-noise DiT (default: 0.9).train_components: Specifies which components to train. Comma-separated list of: "transformer", "transformer_2", "transformer_audio", "dual_tower_bridge", or "all".i2v_ratio: Is the ratio of I2V samples in training. 0.0 = pure T2V, 1.0 = pure I2V, 0.5 = 50% T2V + 50% I2V (default).low_vram: Enables low VRAM mode to reduce memory usage.
3.4 Training Validation
During training, the model will automatically generate validation videos to monitor training progress. You can find these videos in the output_dir directory.
To manually validate a trained model:
python examples/mova/predict_i2v.py
Edit the script to load your trained checkpoint and generate test videos.
3.5 Training with FSDP
export MODEL_NAME="models/Diffusion_Transformer/MOVA-360p"
export DATASET_NAME="datasets/X-Fun-Videos-Audios-Demo/"
export DATASET_META_NAME="datasets/X-Fun-Videos-Audios-Demo/metadata_add_width_height.json"
# NCCL_IB_DISABLE=1 and NCCL_P2P_DISABLE=1 are used in multi nodes without RDMA.
# export NCCL_IB_DISABLE=1
# export NCCL_P2P_DISABLE=1
NCCL_DEBUG=INFO
accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TRANSFORMER_BASED_WRAP \
--fsdp_transformer_layer_cls_to_wrap=WanAttentionBlock,AudioWanAttentionBlock,ConditionalCrossAttentionBlock --fsdp_sharding_strategy "FULL_SHARD" \
--fsdp_state_dict_type=SHARDED_STATE_DICT --fsdp_backward_prefetch "BACKWARD_PRE" --fsdp_cpu_ram_efficient_loading False \
scripts/mova/train.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATASET_NAME \
--train_data_meta=$DATASET_META_NAME \
--image_sample_size=480 \
--video_sample_size=480 \
--token_sample_size=480 \
--video_sample_stride=1 \
--video_sample_n_frames=193 \
--train_batch_size=1 \
--video_repeat=1 \
--gradient_accumulation_steps=1 \
--dataloader_num_workers=8 \
--num_train_epochs=100 \
--checkpointing_steps=50 \
--learning_rate=2e-05 \
--lr_scheduler="constant_with_warmup" \
--lr_warmup_steps=100 \
--seed=42 \
--output_dir="output_dir_mova" \
--gradient_checkpointing \
--mixed_precision="bf16" \
--adam_weight_decay=3e-2 \
--adam_epsilon=1e-10 \
--vae_mini_batch=1 \
--max_grad_norm=0.05 \
--random_hw_adapt \
--training_with_video_token_length \
--enable_bucket \
--uniform_sampling \
--low_vram \
--trainable_modules "." \
--boundary_type="low" \
--boundary_ratio=0.9 \
--train_components="transformer,transformer_2"
3.6 Training without DeepSpeed or FSDP
export MODEL_NAME="models/Diffusion_Transformer/MOVA-360p"
export DATASET_NAME="datasets/X-Fun-Videos-Audios-Demo/"
export DATASET_META_NAME="datasets/X-Fun-Videos-Audios-Demo/metadata_add_width_height.json"
# NCCL_IB_DISABLE=1 and NCCL_P2P_DISABLE=1 are used in multi nodes without RDMA.
# export NCCL_IB_DISABLE=1
# export NCCL_P2P_DISABLE=1
NCCL_DEBUG=INFO
accelerate launch --mixed_precision="bf16" scripts/mova/train.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATASET_NAME \
--train_data_meta=$DATASET_META_NAME \
--image_sample_size=480 \
--video_sample_size=480 \
--token_sample_size=480 \
--video_sample_stride=1 \
--video_sample_n_frames=193 \
--train_batch_size=1 \
--video_repeat=1 \
--gradient_accumulation_steps=1 \
--dataloader_num_workers=8 \
--num_train_epochs=100 \
--checkpointing_steps=50 \
--learning_rate=2e-05 \
--lr_scheduler="constant_with_warmup" \
--lr_warmup_steps=100 \
--seed=42 \
--output_dir="output_dir_mova" \
--gradient_checkpointing \
--mixed_precision="bf16" \
--adam_weight_decay=3e-2 \
--adam_epsilon=1e-10 \
--vae_mini_batch=1 \
--max_grad_norm=0.05 \
--random_hw_adapt \
--training_with_video_token_length \
--enable_bucket \
--uniform_sampling \
--low_vram \
--trainable_modules "." \
--boundary_type="low" \
--boundary_ratio=0.9 \
--train_components="transformer,transformer_2"
3.7 Multi-Node Distributed Training
When training with multiple machines, set the parameters as follows:
export MASTER_ADDR="your master address"
export MASTER_PORT=10086
export WORLD_SIZE=1 # The number of machines
export NUM_PROCESS=8 # The number of processes, such as WORLD_SIZE * 8
export RANK=0 # The rank of this machine
accelerate launch --mixed_precision="bf16" --main_process_ip=$MASTER_ADDR --main_process_port=$MASTER_PORT --num_machines=$WORLD_SIZE --num_processes=$NUM_PROCESS --machine_rank=$RANK scripts/mova/train.py
4. Inference Testing
4.1 Inference Parameters
Core Parameters:
| Parameter | Description | Example Value |
|---|---|---|
GPU_memory_mode |
GPU memory mode, see table below for options | sequential_cpu_offload |
ulysses_degree |
Head dimension parallelism degree, 1 for single GPU | 1 |
ring_degree |
Sequence dimension parallelism degree, 1 for single GPU | 1 |
fsdp_dit |
Use FSDP for Transformer during multi-GPU inference to save VRAM | False |
fsdp_text_encoder |
Use FSDP for text encoder during multi-GPU inference | True |
compile_dit |
Compile Transformer for faster inference (effective for fixed resolution, not compatible with sequential_cpu_offload) | False |
model_name |
Model path | models/Diffusion_Transformer/MOVA-360p |
sampler_name |
Sampler type: Flow, Flow_Unipc, Flow_DPM++ |
Flow |
boundary_ratio |
Boundary ratio for switching between high-noise and low-noise DiT | 0.9 |
transformer_path |
Path to trained low-noise Transformer weights | None |
transformer_high_path |
Path to trained high-noise Transformer weights | None |
transformer_audio_path |
Path to trained audio Transformer weights | None |
bridge_path |
Path to trained dual-tower bridge weights | None |
vae_path |
Path to trained video VAE weights | None |
audio_vae_path |
Path to trained audio VAE weights | None |
lora_path |
Path to low-noise model LoRA weights | None |
lora_high_path |
Path to high-noise model LoRA weights | None |
validation_image |
Input image for I2V mode | asset/8.png |
sample_size |
Generated video resolution [height, width] |
[640, 352] |
video_length |
Number of frames to generate | 81 |
fps |
Frames per second | 24 |
weight_dtype |
Model weight precision, use torch.float16 for GPUs without bf16 (e.g., v100, 2080ti) |
torch.bfloat16 |
prompt |
Positive prompt describing what to generate | "Medium shot of a girl..." |
negative_prompt |
Negative prompt describing what to avoid | "oversaturated, overexposed..." |
guidance_scale |
Guidance strength | 5.0 |
seed |
Random seed for reproducibility | 43 |
num_inference_steps |
Number of inference steps | 50 |
lora_weight |
Low-noise model LoRA weight strength | 0.55 |
lora_high_weight |
High-noise model LoRA weight strength | 0.55 |
save_path |
Path to save generated videos | samples/mova-videos-i2v |
audio_sample_rate |
Audio sample rate (read from vocoder config) | 24000 |
GPU Memory Mode Description:
| Mode | Description | Memory Usage |
|---|---|---|
model_full_load |
Load entire model to GPU | Highest |
model_full_load_and_qfloat8 |
Full load + FP8 quantization | High |
model_cpu_offload |
Offload model to CPU after use | Medium |
model_cpu_offload_and_qfloat8 |
CPU offload + FP8 quantization | Medium-Low |
model_group_offload |
Transfer layer groups between CPU/CUDA | Low |
sequential_cpu_offload |
Offload each layer to CPU after use (slowest) | Lowest |
4.2 Single GPU Inference
Run single GPU inference:
python examples/mova/predict_i2v.py
Edit examples/mova/predict_i2v.py according to your needs. For first-time inference, focus on modifying the following parameters. For other parameters, see the inference parameter description above.
# Select based on your GPU memory
GPU_memory_mode = "sequential_cpu_offload"
# Your actual model path
model_name = "models/Diffusion_Transformer/MOVA-360p"
# I2V input image
validation_image = "asset/8.png"
# Paths to trained weights (if needed)
transformer_path = None
transformer_high_path = None
transformer_audio_path = None
bridge_path = None
# LoRA weight paths (if needed)
lora_path = None
lora_high_path = None
# Write according to what you want to generate
prompt = "Medium shot of a girl by the ocean. She starts with a bright smile, then gently nods her head while speaking. Her mouth moves naturally to say: \"Hi, nice to meet you.\" She maintains eye contact throughout. The background shows calm waves. Smooth motion, cinematic quality, realistic facial expressions."
# ...
4.3 Multi-GPU Parallel Inference
Applicable Scenarios: High-resolution generation, accelerated inference
Install Parallel Inference Dependencies
pip install xfuser==0.4.2 yunchang==0.6.2
Configure Parallel Strategy
Edit examples/mova/predict_i2v.py:
# Ensure ulysses_degree × ring_degree = number of GPUs used
# For example, using 8 GPUs:
ulysses_degree = 2 # Head dimension parallelism
ring_degree = 1 # Sequence dimension parallelism
Configuration Principles:
ulysses_degreemust be divisible by the model's number of headsring_degreesplits along the sequence dimension and affects communication overhead; avoid using it when heads can be evenly divided
Configuration Examples:
| GPU Count | ulysses_degree | ring_degree | Description |
|---|---|---|---|
| 1 | 1 | 1 | Single GPU |
| 4 | 4 | 1 | Head parallelism |
| 8 | 2 | 4 | Hybrid parallelism |
| 8 | 8 | 1 | Head parallelism |
Run Multi-GPU Inference
torchrun --nproc-per-node=2 examples/mova/predict_i2v.py
5. Additional Resources
- Official GitHub: https://github.com/aigc-apps/VideoX-Fun