diff --git a/README.md b/README.md index 18a3d8a..86a0dbd 100644 --- a/README.md +++ b/README.md @@ -4,14 +4,17 @@ A small neural network to provide interoperability between the latents generated I wanted to see if it was possible to pass latents generated by the new SDXL model directly into SDv1.5 models without decoding and re-encoding them using a VAE first. ## Installation -To install it, simply clone this repo to your custom_nodes folder using the following command: `git clone https://github.com/city96/SD-Latent-Interposer custom_nodes/SD-Latent-Interposer`. +To install it, simply clone this repo to your custom_nodes folder using the following command: +``` +git clone https://github.com/city96/SD-Latent-Interposer custom_nodes/SD-Latent-Interposer +``` Alternatively, you can download the [comfy_latent_interposer.py](https://github.com/city96/SD-Latent-Interposer/raw/main/comfy_latent_interposer.py) file to your `ComfyUI/custom_nodes` folder as well. You may need to install hfhub using the command `pip install huggingface-hub` inside your venv. -If you need the model weights for something else, they are [hosted on HF](https://huggingface.co/city96/SD-Latent-Interposer/tree/main) under the same Apache2 license as the rest of the repo. +If you need the model weights for something else, they are [hosted on HF](https://huggingface.co/city96/SD-Latent-Interposer/tree/main) under the same Apache2 license as the rest of the repo. The current files are in the **"v4.0"** subfolder. ## Usage -See the image below for an example on how to use it. xl=>v1 conversion is almost flawless, **v1=>xl seems to produce artifacts.** +Simply place it where you would normally place a VAE decode followed by a VAE encode. Set the denoise as appropirate to hide any artifacts while keeping the composition. See image below. ![LATENT_INTERPOSER_V3 1_TEST](https://github.com/city96/SD-Latent-Interposer/assets/125218114/849574b4-2565-4090-85d3-ae63ab425ee2) @@ -20,14 +23,53 @@ Without the interposer, the two latent spaces are incompatible: ![LATENT_INTERPOSER_V3 1](https://github.com/city96/SD-Latent-Interposer/assets/125218114/13e2c01f-580e-4ecb-af1f-b6b21699127b) ### Local models -The node pulls the required files from huggingface hub by default. You can create a `models` folder and place the modules there if you have a flaky connection or prefer to use it completely offline. The custom node will prefer local files over HF when available. The path should be: `ComfyUI/custom_nodes/SD-Latent-Interposer/models` +The node pulls the required files from huggingface hub by default. You can create a `models` folder and place the models there if you have a flaky connection or prefer to use it completely offline. The custom node will prefer local files over HF when available. The path should be: `ComfyUI/custom_nodes/SD-Latent-Interposer/models` -Alternatively, just clone the entire HF repo to it: `git clone https://huggingface.co/city96/SD-Latent-Interposer custom_nodes/SD-Latent-Interposer/models` +Alternatively, just clone the entire HF repo to it: +``` +git clone https://huggingface.co/city96/SD-Latent-Interposer custom_nodes/SD-Latent-Interposer/models +``` + +### Supported Models + +Model names: + +| code | name | +| ---- | -------------------------- | +| `v1` | SDXL | +| `xl` | Stable Diffusion v1.x | +| `ca` | Stable Cascade (Stage A/B) | + +Available models: + +| From | to `v1` | to `xl` | to `ca` | +|:----:|:-------:|:-------:|:-------:| +| `v1` | - | v4.0 | No | +| `xl` | v4.0 | - | No | +| `ca` | v4.0 | v4.0 | - | ## Training -Most of the training/preprocessing code is a 1:1 mirror from my latent upscaler. The folder layout it expects is also the same. + +The training code initializes most training parameters from the provided config file. The dataset should be a single .bin file saved with `torch.save` for each latent version. The format should be [batch, channels, height, width] with the "batch" being as large as the dataset, ie 88000. + +### Interposer v4.0 + +The training code currently initializes two copies of the model, one in the target direction and one in the opposite. The losses are defined based on this. + +- `p_loss` is the main criterion for the primary model. +- `b_loss` is the main criterion for the secondary one. +- `r_loss` is the output of the primary model back through the secondary model and checked against the source latent (basically a round trip through the two models). +- `h_loss` is the same as `r_loss` but for the secondary model. + +All models were trained for 50000 steps with either batch size 128 (xl/v1) or 48 (cascade). +The training was done locally on an RTX 3080 and a Tesla V100S. + +### Older versions + +
Interposer v3.1 ### Interposer v3.1 + This is basically a complete rewrite. Replaced the mediocre bunch of conv2d layers with something that looks more like a proper neural network. No VGG loss because I still don't have a better GPU. Training was done on combined Flickr2K + DIV2K, with each image being processed into 6 1024x1024 segments. Padded with some of my random images for a total of 22,000 source images in the dataset. @@ -38,7 +80,7 @@ v3.0 was 500k steps at a constant LR of 1e-4, v3.1 was 1M steps using a CosineAn ![INTERPOSER_V3 1](https://github.com/city96/SD-Latent-Interposer/assets/125218114/daff0ae2-4739-4cef-ba54-ac1d156d3388) -### Older versions +
Interposer v1.1 @@ -50,6 +92,7 @@ Overall, it seems to perform a lot better, especially for real life photos. I al
+
Interposer v1.0 ### Interposer v1.0