* update train_lora && update deepspeed && update training with max token length * fix bug in train.py * fix bug in training_with_video_token_length * update v2v && update v2v api * add rope2d embedding precomputation; move text encoder to dataloader to reduce gpu memory consumpution * add cuda multi-stream to speedup vae encode * update new vae && new comfyui * fix some bug in training code * Add lcm lora (#89) Co-authored-by: xuanyuan.lb <xuanyuan.lb@alibaba-inc.com> * Update Training Code and fix bug in low vram mode * fix bug in low vram mode * update report * update cfg * actual text clip --------- Co-authored-by: mengli.cml <mengli.cml@alibaba-inc.com> Co-authored-by: liubo0902 <38622806+liubo0902@users.noreply.github.com> Co-authored-by: xuanyuan.lb <xuanyuan.lb@alibaba-inc.com>