2024-04-08 23:47:44 -07:00
2024-03-13 07:07:40 -07:00
2024-04-08 19:36:54 -07:00
2024-04-08 02:07:52 -07:00
2024-04-08 01:24:07 -07:00
2024-04-06 00:28:12 -07:00
2024-04-05 01:10:02 -07:00
2024-03-11 10:56:24 -07:00
2024-04-08 23:43:33 -07:00
2024-04-08 01:30:16 -07:00
2024-04-08 23:43:33 -07:00
2024-04-03 23:53:37 -07:00

Trainer

Code for finetuning and training LoRa modules on top of Stable Diffusion.

Setup

Install all dependencies manually and run: python main.py -c training_args.json

Adjust the arguments inside training_args.json accordingly.


You can also run this through cog as a docker image:

  1. Install Replicate 'cog':
sudo curl -o /usr/local/bin/cog -L "https://github.com/replicate/cog/releases/latest/download/cog_$(uname -s)_$(uname -m)"
sudo chmod +x /usr/local/bin/cog
  1. Build the image with sudo cog build
  2. Run a training run with sudo sh cog_test_train.sh

TODO's

Code / Cleanup:

  • turn all/most of the args of the main() function in trainer_pti.py and the preprocess() function into a clean args_dict that makes it easy to add and distribute new parameters over the code and save these args to a .json file at the end.
  • Modularize the logic in train.py as much as possible, trying to minimize dev work that needs to happen when SD3 drops (in progress)
  • make a clean train.py entrypoint that can be run as a normal python command (instead of having to use cog)
  • make it so the textual_inversion optimizer only optimizes the actual trained token embeddings instead of all of them + resetting later
  • Figure out how to swap out a lora_adapter module onto a base model without reloading the entire model pipe...
  • Make sure the saved LoRa's are compatible with ComfyUI / AUTO1111
  • properly measure regularization target_values for conditioning and add pooled_prompt_embeds into regularizer for SDXL

Algo:

Bigger improvements:

Tuning Experiments once code is fully ready:

  • gridsearch over LoRa target_modules=["to_k", "to_q", "to_v", "to_out.0"]
  • try-out conditioning noise injection during training to increase robustness
  • re-test / tweak the adaptive learning rates (also test Prodigy vs Adam)
  • right now it looks like the diffusion model gets partially "destroyed" in the beginning of training (outputs from steps 100-200 look terrible), but it then recovers. Can we avoid this collapse? Is the learning rate too high?
  • offset noise
  • AB test Dora vs Lora
  • sweep n_trainable_tokens to inject
S
Description
No description provided
Readme
6.8 MiB
Languages
Python 99.8%
Shell 0.2%