diff --git a/README.md b/README.md index a1eeaa9..748ebca 100644 --- a/README.md +++ b/README.md @@ -1,15 +1,16 @@ # Tiled Diffusion & VAE for ComfyUI -See [this](https://github.com/pkuliyi2015/multidiffusion-upscaler-for-automatic1111/) for more info. +See [this](https://github.com/pkuliyi2015/multidiffusion-upscaler-for-automatic1111/) for more information. -The extension enables **large image drawing & upscaling with limited VRAM** via the following techniques: +This extension enables **large image drawing & upscaling with limited VRAM** via the following techniques: 1. Two SOTA diffusion tiling algorithms: [Mixture of Diffusers](https://github.com/albarji/mixture-of-diffusers) and [MultiDiffusion](https://github.com/omerbt/MultiDiffusion) -2. pkuliyi2015's Tiled VAE algorithm. -3. ~~pkuliyi2015's TIled Noise Inversion for better upscaling.~~ +2. pkuliyi2015 & Kahsolt's Tiled VAE algorithm. +3. ~~pkuliyi2015 & Kahsolt's TIled Noise Inversion for better upscaling.~~ > [!NOTE] -> Sizes are in pixel-space that then get converted into latent-space sizes. +> Sizes/dimensions are in pixels and then converted to latent-space sizes. + ## Features - [x] SDXL model support @@ -21,4 +22,75 @@ The extension enables **large image drawing & upscaling with limited VRAM** via - [x] Img2img upscale - [x] Ultra-Large image generation -Some conditioning nodes aren't working at the moment like SetArea or GLIGEN. \ No newline at end of file +Some conditioning nodes like SetArea or GLIGEN aren't working at the moment. + +## Tiled Diffusion + +
The image is split into tiles, which are then padded with 11/32 pixels' in the decoder/encoder.| +| `fast` |
| +| `color_fix` |When Fast Mode is disabled:
- The original VAE forward is decomposed into a task queue and a task worker, which starts to process each tile.
- When GroupNorm is needed, it suspends, stores current GroupNorm mean and var, send everything to RAM, and turns to the next tile.
- After all GroupNorm means and vars are summarized, it applies group norm to tiles and continues.
- A zigzag execution order is used to reduce unnecessary data transfer.
When Fast Mode is enabled:
- The original input is downsampled and passed to a separate task queue.
- Its group norm parameters are recorded and used by all tiles' task queues.
- Each tile is separately processed without any RAM-VRAM data transfer.
After all tiles are processed, tiles are written to a result buffer and returned.
Only estimate GroupNorm before downsampling, i.e., run in a semi-fast mode.
Only for the encoder. Can restore colors if tiles are too small.
| + + + +## Workflows + +The following images can be loaded in ComfyUI. + + +Simple upscale.
+4x upscale. 3 passes.
+