Update README.md
This commit is contained in:
@@ -20,7 +20,22 @@ Without the interposer, the two latent spaces are incompatible:
|
||||

|
||||
|
||||
## Training
|
||||
The training script should spit out a working model, nn layout is probably not optimal but I'm pretty short on VRAM to trial and error a better layout. PRs welcome.
|
||||
Most of the training/preprocessing code is a 1:1 mirror from my latent upscaler. The folder layout it expects is also the same.
|
||||
|
||||
### Interposer v3.1
|
||||
This is basically a complete rewrite. Replaced the mediocre bunch of conv2d layers with something that looks more like a proper neural network. No VGG loss because I still don't have a better GPU.
|
||||
|
||||
Training was done on combined Flickr2K + DIV2K, with each image being processed into 6 1024x1024 segments. Padded with some of my random images for a total of 22,000 source images in the dataset.
|
||||
|
||||
I think I got rid of most of the XL artifacts, but the color/hue/saturation shift issues are still there. I actually saved the optimizer state this time so I might be able to do 100K steps with visual loss on my P40s. Hopefully they won't burn up.
|
||||
|
||||
v3.0 was 500k steps at a constant LR of 1e-4, v3.1 was 1M steps using a CosineAnnealingLR to drop the learning rate towards the end. Both used AdamW.
|
||||
|
||||

|
||||
|
||||
### Older versions
|
||||
|
||||
<details><summary>Interposer v1.1</summary>
|
||||
|
||||
### Interposer v1.1
|
||||
This is the second release using the "spaceship" architecture. It was trained on the Flickr2K dataset and was continued from the v1.0 checkpoint.
|
||||
@@ -28,6 +43,10 @@ Overall, it seems to perform a lot better, especially for real life photos. I al
|
||||
|
||||

|
||||
|
||||
</details>
|
||||
|
||||
<details><summary>Interposer v1.0</summary>
|
||||
|
||||
### Interposer v1.0
|
||||
Not sure why the training loss is so different, it might be due to the """highly curated""" dataset of 1000 random images from my Downloads folder that I used to train it.
|
||||
|
||||
@@ -35,12 +54,10 @@ I probably should've just grabbed LAION.
|
||||
|
||||
I also trained a v1-to-v2 mode, before realizing v1 and v2 shared the same latent space. Oh well.
|
||||
|
||||
<details>
|
||||
<summary>Loss graphs for v1.0 models</summary>
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
</details>
|
||||
|
||||
</details>
|
||||
|
||||
Reference in New Issue
Block a user