2024-04-22 19:39:16 -07:00
2024-03-13 07:07:40 -07:00
2024-04-17 18:48:49 +02:00
2024-04-22 19:24:28 -07:00
2024-04-22 12:22:20 -07:00
2024-04-14 08:09:33 -07:00
2024-04-17 05:34:01 -07:00
2024-03-11 10:56:24 -07:00
2024-04-22 19:39:16 -07:00
2024-04-22 12:22:20 -07:00
2024-04-22 19:24:28 -07:00
2024-04-19 01:39:43 -07:00
2024-04-19 01:39:43 -07:00

Trainer

Code for finetuning and training LoRa modules on top of Stable Diffusion.

Setup

Install all dependencies manually and run: python main.py -c training_args.json

Adjust the arguments inside training_args.json accordingly.


You can also run this through cog as a docker image:

  1. Install Replicate 'cog':
sudo curl -o /usr/local/bin/cog -L "https://github.com/replicate/cog/releases/latest/download/cog_$(uname -s)_$(uname -m)"
sudo chmod +x /usr/local/bin/cog
  1. Build the image with sudo cog build
  2. Run a training run with sudo sh cog_test_train.sh

Evaluation

Download the aesthetic predictor model checkpoint first from google drive. This should give you a file named: aesthetic_score_best_model.pth (99.2 MB)

gdown 1thEIlXVc8lkULVUBY9Ab45tsOERxkjxns

Once the model is downloaded, you can run the eval script with the following CLI args:

  • output_folder: this is where the outputs of the model get saved as jpeg files
  • lora_path: path to your LoRA checkpoint (make sure you edit path_to_your_model_checkpoints to point to the correct folder. It generally ends with something like checkpoint-600 where 600 was the training step)
  • output_json: save all scores in this json file
  • config_filename: config file used for training
python3 evaluate.py \
--output_folder eval_images \
--lora_path path_to_your_model_checkpoint  \
--output_json eval_results.json \
--config_filename training_args.json

TODO's

Code / Cleanup:

  • cleanup optimizers / optimizeable params code into optimizer.py
  • Make sure the saved LoRa's are compatible with ComfyUI / AUTO1111
  • Figure out how to swap out a lora_adapter module onto a base model without reloading the entire model pipe...

Algo:

Bigger improvements:

  • add stronger token regularization (eg CelebBasis spanning basis):
    • remove the fix_embedding_std() hack with gradient based std-matching penalty
    • add covariance_loss to token_warmup phase
    • grid-search the new CovarianceLoss() strength
    • continue plotting the token_warmup_loss post warmup to visualize the distance to the chatgpt_description
  • Add multi-token training
  • implement perfusion ideas (key locking with superclass): https://research.nvidia.com/labs/par/Perfusion/
  • implement prompt-aligned: https://prompt-aligned.github.io/

Tuning Experiments once code is fully ready:

  • gridsearch over LoRa target_modules=["to_k", "to_q", "to_v", "to_out.0"] for both unet and txt-encoder
  • try-out conditioning noise injection during training to increase robustness
  • re-test / tweak the adaptive learning rates (also test Prodigy vs Adam)
  • right now it looks like the diffusion model gets partially "destroyed" in the beginning of training (outputs from steps 100-200 look terrible), but it then recovers. Can we avoid this collapse? Is the learning rate too high?
  • offset noise
  • AB test Dora vs Lora
  • sweep n_trainable_tokens to inject
S
Description
No description provided
Readme
6.8 MiB
Languages
Python 99.8%
Shell 0.2%