From 46ce69a29ab8664e3d0b1d101ad05cee950efd03 Mon Sep 17 00:00:00 2001 From: xander Date: Tue, 28 May 2024 17:04:17 +0200 Subject: [PATCH] update readme --- README.md | 18 ++++++++++++++---- 1 file changed, 14 insertions(+), 4 deletions(-) diff --git a/README.md b/README.md index 9b75a65..897f8d6 100755 --- a/README.md +++ b/README.md @@ -3,10 +3,18 @@ Code for finetuning and training LoRa modules on top of Stable Diffusion. Uses a single training script and loss module that works for both **SDv15** and **SDXL**! +

+ Image 1 +

+

+ Image 2 +

+ ## Setup -Install all dependencies manually and run: -`python main.py -c training_args.json` +Install all dependencies using `pip install -r requirements.txt` +and run: +`python main.py -c training_args.json` to start a training job. Adjust the arguments inside `training_args.json` accordingly. @@ -25,6 +33,7 @@ sudo chmod +x /usr/local/bin/cog ## Automatic Checkpoint Evaluation +This script uses CLIP img/txt similarity scores to evaluate how good the LoRa is vs how overfit. Download the aesthetic predictor model checkpoint first from google drive. This should give you a file named: `aesthetic_score_best_model.pth` (99.2 MB) ```bash @@ -49,6 +58,9 @@ python3 evaluate.py \ ## TODO's +Bugs: +- pure textual inversion for SD15 does not seem to work well... (but it works amazingly well for SDXL...) ---> if anyone can figure this one out I'd be forever grateful! + Algo: - Improve some of the chatgpt functionality: - separate the "gpt_description" / "gpt_segmentation" prompt calls and make them run on a subset of prompts in case there's a lot of imgs / prompts (possibly use img_grids for some gpt4-v calls) @@ -74,5 +86,3 @@ but it then recovers. Can we avoid this collapse? Is the learning rate too high? - offset noise - AB test Dora vs Lora -Bugs: -- pure textual inversion for SD15 does not seem to work well... (but it works amazingly well for SDXL...)