SD-Latent-Interposer
A small neural network to provide interoperability between the latents generated by the different Stable Diffusion models.
I wanted to see if it was possible to pass latents generated by the new SDXL model directly into SDv1.5 models without decoding and re-encoding them using a VAE first.
Installation
To install it, simply clone this repo to your custom_nodes folder using the following command: git clone https://github.com/city96/SD-Latent-Interposer custom_nodes/SD-Latent-Interposer.
Alternatively, you can download the comfy_latent_interposer.py file to your ComfyUI/custom_nodes folder as well. You may need to install hfhub using the command pip install huggingface-hub inside your venv.
If you need the model weights for something else, they are hosted on HF under the same Apache2 license as the rest of the repo.
Usage
See the image below for an example on how to use it. xl=>v1 conversion is almost flawless, v1=>xl seems to produce artifacts.
Without the interposer, the two latent spaces are incompatible:
Training
The training script should spit out a working model, nn layout is probably not optimal but I'm pretty short on VRAM to trial and error a better layout. PRs welcome.
Interposer v1.1
This is the second release using the "spaceship" architecture. It was trained on the Flickr2K dataset and was continued from the v1.0 checkpoint. Overall, it seems to perform a lot better, especially for real life photos. I also investigated the odd v1->xl artifacts but in the end it seems inherent to the VAE decoder stage.
Interposer v1.0
Not sure why the training loss is so different, it might be due to the """highly curated""" dataset of 1000 random images from my Downloads folder that I used to train it.
I probably should've just grabbed LAION.
I also trained a v1-to-v2 mode, before realizing v1 and v2 shared the same latent space. Oh well.