Merge pull request #26 from GiusTex/fix-loaded-keys

Fix loaded keys
This commit is contained in:
Gius
2025-05-15 20:19:11 +02:00
committed by GitHub
11 changed files with 981 additions and 1281 deletions
File diff suppressed because it is too large Load Diff
+21 -25
View File
@@ -1,8 +1,9 @@
ComfyUI nodes for outpainting images with diffusers, based on [diffusers-image-outpaint](https://huggingface.co/spaces/fffiloni/diffusers-image-outpaint/tree/main) by fffiloni.
![Extension-Overview](https://github.com/user-attachments/assets/b801698e-e666-4179-98bd-42dfb1f033ba)
![image](https://github.com/user-attachments/assets/1a02c2d1-f24e-4ad2-acdc-a2cbb15a1f14)
#### Updates:
- 15/05/2025: Fixed `missing 'loaded_keys'` error. More details below.
- 17/11/2024:
- Added more options to Pad Image node (resize image, custom resize image percentage, mask overlap percentage, overlap left/right/top/bottom).
- Side notes:
@@ -15,44 +16,39 @@ ComfyUI nodes for outpainting images with diffusers, based on [diffusers-image-o
- 20/10/2024: No more need to download tokenizers nor text encoders! Now comfyui clip loader works, and you can use your clip models. You can also use the Checkpoint Loader Simple node, to skip the clip selection part.
- 10/2024: You don't need any more the diffusers vae, and can use the extension in low vram mode using `sequential_cpu_offload` (also thanks to [zmwv823](https://github.com/GiusTex/ComfyUI-DiffusersImageOutpaint/pull/4)) that pushes the vram usage from *8,3 gb* down to **_6 gb_**.
#### To do list to [change model used](https://github.com/GiusTex/ComfyUI-DiffusersImageOutpaint/pull/14):
- - [x] ComfyUI Clip Loader Node
- ~[ ] ComfyUI Load Diffusion Model Node~ (more info [below](https://github.com/GiusTex/ComfyUI-DiffusersImageOutpaint#unet-and-controlnet-models-loader-using-comfyui-nodes-canceled))
- ~[ ] ComfyUI Load Conotrolnet Model Node~ (more info [below](https://github.com/GiusTex/ComfyUI-DiffusersImageOutpaint#unet-and-controlnet-models-loader-using-comfyui-nodes-canceled))
## Installation
- Download this extension or `git clone` it in comfyui/custom_nodes, then (if comfyui-manager didn't already install the requirements or you have missing modules), from comfyui virtual env write `cd your/path/to/this/extension` and `pip install -r requirements.txt`.
- Download models in comfyui/models/diffusion_models:
- model_name:
- unet:
- `diffusion_pytorch_model.fp16.safetensors` ([example](https://huggingface.co/SG161222/RealVisXL_V5.0_Lightning/blob/main/unet/diffusion_pytorch_model.fp16.safetensors))
- `config.json` ([example](https://huggingface.co/SG161222/RealVisXL_V5.0_Lightning/blob/main/unet/config.json))
- scheduler:
- `scheduler_config.json` ([example](https://huggingface.co/SG161222/RealVisXL_V5.0_Lightning/blob/main/scheduler/scheduler_config.json))
- `model_index.json` ([example](https://huggingface.co/SG161222/RealVisXL_V5.0_Lightning/blob/main/model_index.json))
- controlnet_name:
- `config_promax.json` ([example](https://huggingface.co/xinsir/controlnet-union-sdxl-1.0/blob/main/config_promax.json)), `diffusion_pytorch_model_promax.safetensors` ([example](https://huggingface.co/xinsir/controlnet-union-sdxl-1.0/blob/main/diffusion_pytorch_model_promax.safetensors))
- Download a sdxl model ([example](https://huggingface.co/SG161222/RealVisXL_V5.0_Lightning/resolve/main/unet/diffusion_pytorch_model.fp16.safetensors)) in comfyui/models/diffusion_models;
- Download a sdxl controlnet model ([example](https://huggingface.co/xinsir/controlnet-union-sdxl-1.0/blob/main/diffusion_pytorch_model_promax.safetensors)) in comfyui/models/controlnet.
**⚠ Choosing model and controlnet**: As of now, I only tried `RealVisXL_V5.0_Lightning` and `controlnet-union-promax_sdxl`. Mixing RealVisXL with controlnet-union (non promax version) gave error, so it could be that other models/controlnets give error as well, but I haven't tried much combinations so I can't tell.
<details>
<summary>Some considerations</summary>
Flux is still beyond me (even if I was quite there, I think). I haven't tried integrating other model types, and after my flux failure I don't think I'll try adding other model types.
Since for now only sdxl models work, the configs are hardcoded.
</details>
- (Dual) Clip Loader node: if you use the Clip Loader instead of Checkpoint Loader Simple, and want to use an `sdxl type` model like RealVisXL_V5.0_Lightning, you can download `clip_I` and `clip_g` from [here](https://huggingface.co/Comfy-Org/stable-diffusion-3.5-fp8/tree/main/text_encoders). You can use [this workflow](https://github.com/GiusTex/ComfyUI-DiffusersImageOutpaint/blob/New-Pad-Node-Options/Diffusers-Outpaint-DoubleWorkflow.json) (change model.fp16 with `clip_g`).
## Overview
- **Minimum VRAM**: 6 gb with 1280x720 image, rtx 3060, RealVisXL_V5.0_Lightning, sdxl-vae-fp16-fix, controlnet-union-sdxl-promax using `sequential_cpu_offload`, otherwise 8,3 gb;
- ~As seen in [this issue](https://github.com/GiusTex/ComfyUI-DiffusersImageOutpaint/issues/7#issuecomment-2410852908), images with **square corners** are required~.
The extension gives 4 nodes:
- **Load Diffusion Outpaint Models**: a simple node to load diffusion `models`. You can download them from Huggingface (the extension doesn't download them automatically);
The extension gives 5 nodes:
- **Load Diffuser Model**: a simple node to load diffusion `models`. You can download them from Huggingface (the extension doesn't download them automatically). Put them inside the `diffusion_models` folder;
- **Load Diffuser Controlnet**: a simple node to load diffusion `models`. You can download them from Huggingface (the extension doesn't download them automatically). Put them inside the `controlnet` folder;
- **Paid Image for Diffusers Outpaint**: this node resizes the image based on the specified `width` and `height`, then resizes it again based on the `resize_image` percentage, and if possible it will put the mask based on the `alignment` specified, otherwise it will revert back to the default "middle" `alignment`;
- **Encode Diffusers Outpaint Prompt**: self explanatory. Works as `clip text encode (prompt)`, and specifies what to add to the image;
- **Diffusers Image Outpaint**: This is the main node, that outpaints the image. Currently the generation process is based on fffiloni's one, so you can't reproduce a specific a specific outpaint, and the `seed` option you see is only used to update the UI and generate a new image. You can specify the amount of `steps` to generate the image.
You _can_ also pass image and mask to `vae encode (for inpainting)` node, then pass the latent to a `sampler`, but controlnets and ip-adapters won't always give good results like with diffusers outpaint, and they require a different workflow, not covered by this extension.
### Change model used
- **Main model**: On huggingface, choose a model from [text2image models](https://huggingface.co/models?pipeline_tag=text-to-image&sort=trending) (**sdxl and maybe sd1.5 model types should work, while flux doesn't**), then create a new folder named after it in `comfyui/models/diffusion_models`, then download in it the subfolders `unet` (if not available use `transformer`) and `scheduler`.
- Hint: sometimes in the `unet` or `transformer` folder there are more model files and not all are required. If you have `model.fp16` and `model`, I suggest you to use the fp16 variant; if you have `model-001-of-002`, `model-002-of-002`, `model`, choose model (instead of the fragmented version).
- **Controlnet model**: download `config.json` and the safetensors `model`.
#### Unet and Controlnet Models Loader using ComfYUI nodes canceled
I can load them but then they don't work in the inference code, since comfyui load diffusers models in a different format ([reddit post](https://www.reddit.com/r/comfyui/comments/17fvb49/comment/k6cz9yv/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button)).
## Missing 'loaded_keys' error
Recent versions of `transformers` and `diffusers` broke somethings, you need to revert back, command with some working versions (found [here](https://huggingface.co/spaces/fffiloni/diffusers-image-outpaint/blob/main/requirements.txt)) (do it inside your comfyui env): `pip install transformers==4.45.0 --upgrade diffusers==0.32.2 --upgrade`.
## Credits
diffusers-image-outpaint by [fffiloni](https://huggingface.co/spaces/fffiloni/diffusers-image-outpaint/tree/main)
+5 -3
View File
@@ -1,15 +1,17 @@
from .nodes import (PadImageForDiffusersOutpaint, LoadDiffusersOutpaintModels, EncodeDiffusersOutpaintPrompt, DiffusersImageOutpaint)
from .nodes import (PadImageForDiffusersOutpaint, LoadDiffuserModel, LoadDiffuserControlnet, EncodeDiffusersOutpaintPrompt, DiffusersImageOutpaint)
NODE_CLASS_MAPPINGS = {
"PadImageForDiffusersOutpaint": PadImageForDiffusersOutpaint,
"LoadDiffusersOutpaintModels": LoadDiffusersOutpaintModels,
"LoadDiffuserModel": LoadDiffuserModel,
"LoadDiffuserControlnet": LoadDiffuserControlnet,
"EncodeDiffusersOutpaintPrompt": EncodeDiffusersOutpaintPrompt,
"DiffusersImageOutpaint": DiffusersImageOutpaint
}
NODE_DISPLAY_NAME_MAPPINGS = {
"PadImageForDiffusersOutpaint": "Pad Image For Diffusers Outpaint",
"LoadDiffusersOutpaintModels": "Load Diffusers Outpaint Models",
"LoadDiffuserModel": "Load Diffuser Model",
"LoadDiffuserControlnet": "Load Diffuser Controlnet",
"EncodeDiffusersOutpaintPrompt": "Encode Diffusers Outpaint Prompt",
"DiffusersImageOutpaint": "Diffusers Image Outpaint"
}
@@ -0,0 +1,57 @@
{
"_class_name": "ControlNetModel",
"_diffusers_version": "0.20.0.dev0",
"act_fn": "silu",
"addition_embed_type": "text_time",
"addition_embed_type_num_heads": 64,
"addition_time_embed_dim": 256,
"attention_head_dim": [
5,
10,
20
],
"block_out_channels": [
320,
640,
1280
],
"class_embed_type": null,
"conditioning_channels": 3,
"conditioning_embedding_out_channels": [
16,
32,
96,
256
],
"controlnet_conditioning_channel_order": "rgb",
"cross_attention_dim": 2048,
"down_block_types": [
"DownBlock2D",
"CrossAttnDownBlock2D",
"CrossAttnDownBlock2D"
],
"downsample_padding": 1,
"encoder_hid_dim": null,
"encoder_hid_dim_type": null,
"flip_sin_to_cos": true,
"freq_shift": 0,
"global_pool_conditions": false,
"in_channels": 4,
"layers_per_block": 2,
"mid_block_scale_factor": 1,
"norm_eps": 1e-05,
"norm_num_groups": 32,
"num_attention_heads": null,
"num_class_embeds": null,
"only_cross_attention": false,
"projection_class_embeddings_input_dim": 2816,
"resnet_time_scale_shift": "default",
"transformer_layers_per_block": [
1,
2,
10
],
"upcast_attention": null,
"use_linear_projection": true,
"num_control_type": 8
}
@@ -0,0 +1,19 @@
{
"_class_name": "DDIMScheduler",
"_diffusers_version": "0.30.0.dev0",
"beta_end": 0.012,
"beta_schedule": "scaled_linear",
"beta_start": 0.00085,
"clip_sample": false,
"clip_sample_range": 1.0,
"dynamic_thresholding_ratio": 0.995,
"num_train_timesteps": 1000,
"prediction_type": "epsilon",
"rescale_betas_zero_snr": false,
"sample_max_value": 1.0,
"set_alpha_to_one": false,
"steps_offset": 1,
"thresholding": false,
"timestep_spacing": "leading",
"trained_betas": null
}
+71
View File
@@ -0,0 +1,71 @@
{
"_class_name": "UNet2DConditionModel",
"act_fn": "silu",
"addition_embed_type": "text_time",
"addition_embed_type_num_heads": 64,
"addition_time_embed_dim": 256,
"attention_head_dim": [
5,
10,
20
],
"attention_type": "default",
"block_out_channels": [
320,
640,
1280
],
"center_input_sample": false,
"class_embed_type": null,
"class_embeddings_concat": false,
"conv_in_kernel": 3,
"conv_out_kernel": 3,
"cross_attention_dim": 2048,
"cross_attention_norm": null,
"down_block_types": [
"DownBlock2D",
"CrossAttnDownBlock2D",
"CrossAttnDownBlock2D"
],
"downsample_padding": 1,
"dropout": 0.0,
"dual_cross_attention": false,
"encoder_hid_dim": null,
"encoder_hid_dim_type": null,
"flip_sin_to_cos": true,
"freq_shift": 0,
"in_channels": 4,
"layers_per_block": 2,
"mid_block_only_cross_attention": null,
"mid_block_scale_factor": 1,
"mid_block_type": "UNetMidBlock2DCrossAttn",
"norm_eps": 1e-05,
"norm_num_groups": 32,
"num_attention_heads": null,
"num_class_embeds": null,
"only_cross_attention": false,
"out_channels": 4,
"projection_class_embeddings_input_dim": 2816,
"resnet_out_scale_factor": 1.0,
"resnet_skip_time_act": false,
"resnet_time_scale_shift": "default",
"reverse_transformer_layers_per_block": null,
"sample_size": 128,
"time_cond_proj_dim": null,
"time_embedding_act_fn": null,
"time_embedding_dim": null,
"time_embedding_type": "positional",
"timestep_post_act": null,
"transformer_layers_per_block": [
1,
2,
10
],
"up_block_types": [
"CrossAttnUpBlock2D",
"CrossAttnUpBlock2D",
"UpBlock2D"
],
"upcast_attention": false,
"use_linear_projection": true
}
+629
View File
@@ -0,0 +1,629 @@
{
"id": "5e709e31-1e9f-475e-a837-14abe1d4f292",
"revision": 0,
"last_node_id": 586,
"last_link_id": 1133,
"nodes": [
{
"id": 499,
"type": "PadImageForDiffusersOutpaint",
"pos": [
-5320,
420
],
"size": [
290,
314
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 906
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": []
},
{
"name": "MASK",
"type": "MASK",
"links": []
},
{
"name": "diffuser_outpaint_cnet_image",
"type": "IMAGE",
"links": [
1100
]
}
],
"properties": {
"cnr_id": "ComfyUI-DiffusersImageOutpaint",
"ver": "6a51ce5d3baa2171a85f51d462ef9f30ff9b5d26",
"Node name for S&R": "PadImageForDiffusersOutpaint",
"aux_id": "GiusTex/ComfyUI-DiffusersImageOutpaint",
"widget_ue_connectable": {}
},
"widgets_values": [
720,
1280,
"Middle",
"Full",
50,
10,
true,
true,
true,
true
],
"color": "#233",
"bgcolor": "#355"
},
{
"id": 494,
"type": "VAEDecode",
"pos": [
-4710,
-70
],
"size": [
140,
46
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 1101
},
{
"name": "vae",
"type": "VAE",
"link": 908
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
899
]
}
],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.30",
"Node name for S&R": "VAEDecode",
"widget_ue_connectable": {}
},
"widgets_values": [],
"color": "#323",
"bgcolor": "#535"
},
{
"id": 495,
"type": "PreviewImage",
"pos": [
-4550,
-70
],
"size": [
250,
310
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 899
}
],
"outputs": [],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.30",
"Node name for S&R": "PreviewImage",
"widget_ue_connectable": {}
},
"widgets_values": []
},
{
"id": 497,
"type": "DualCLIPLoader",
"pos": [
-5570,
200
],
"size": [
270,
130
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "CLIP",
"type": "CLIP",
"links": [
1110,
1111
]
}
],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.30",
"Node name for S&R": "DualCLIPLoader",
"widget_ue_connectable": {}
},
"widgets_values": [
"clip_l.safetensors",
"clip_g.safetensors",
"sdxl",
"default"
],
"color": "#223",
"bgcolor": "#335"
},
{
"id": 569,
"type": "EncodeDiffusersOutpaintPrompt",
"pos": [
-5280,
200
],
"size": [
252.08065795898438,
136
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 1111
}
],
"outputs": [
{
"name": "diffusers_conditioning",
"type": "CONDITIONING",
"links": [
1099
]
}
],
"properties": {
"cnr_id": "ComfyUI-DiffusersImageOutpaint",
"ver": "6a51ce5d3baa2171a85f51d462ef9f30ff9b5d26",
"Node name for S&R": "EncodeDiffusersOutpaintPrompt",
"aux_id": "GiusTex/ComfyUI-DiffusersImageOutpaint",
"widget_ue_connectable": {}
},
"widgets_values": [
"auto",
"auto",
""
],
"color": "#322",
"bgcolor": "#533"
},
{
"id": 501,
"type": "VAELoader",
"pos": [
-5010,
260
],
"size": [
270,
58
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "VAE",
"type": "VAE",
"links": [
908
]
}
],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.30",
"Node name for S&R": "VAELoader",
"widget_ue_connectable": {}
},
"widgets_values": [
"sdxl_vae.safetensors"
],
"color": "#223",
"bgcolor": "#335"
},
{
"id": 573,
"type": "DiffusersImageOutpaint",
"pos": [
-4990,
-60
],
"size": [
247.341796875,
278
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 1131
},
{
"name": "scheduler_configs",
"type": "SCHEDULER",
"link": 1132
},
{
"name": "control_net",
"type": "CONTROL_NET",
"link": 1133
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 1098
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 1099
},
{
"name": "diffuser_outpaint_cnet_image",
"type": "IMAGE",
"link": 1100
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
1101
]
}
],
"properties": {
"cnr_id": "ComfyUI-DiffusersImageOutpaint",
"ver": "6a51ce5d3baa2171a85f51d462ef9f30ff9b5d26",
"Node name for S&R": "DiffusersImageOutpaint",
"aux_id": "GiusTex/ComfyUI-DiffusersImageOutpaint",
"widget_ue_connectable": {}
},
"widgets_values": [
1.5,
1,
8,
"auto",
"auto",
false
],
"color": "#232",
"bgcolor": "#353"
},
{
"id": 531,
"type": "LoadDiffuserModel",
"pos": [
-5610,
-190
],
"size": [
290,
150
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "model",
"type": "MODEL",
"links": [
1131
]
},
{
"name": "scheduler configs",
"type": "SCHEDULER",
"links": [
1132
]
}
],
"properties": {
"cnr_id": "ComfyUI-DiffusersImageOutpaint",
"ver": "6a51ce5d3baa2171a85f51d462ef9f30ff9b5d26",
"Node name for S&R": "LoadDiffuserModel",
"aux_id": "GiusTex/ComfyUI-DiffusersImageOutpaint",
"widget_ue_connectable": {}
},
"widgets_values": [
"RealVisXL_V5.0_Lightning_unet.safetensors",
"auto",
"auto",
"sdxl"
],
"color": "#223",
"bgcolor": "#335"
},
{
"id": 500,
"type": "LoadImage",
"pos": [
-5640,
430
],
"size": [
270,
314
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
906
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.30",
"Node name for S&R": "LoadImage",
"widget_ue_connectable": {}
},
"widgets_values": [
"20230403_183417.jpg",
"image"
]
},
{
"id": 570,
"type": "EncodeDiffusersOutpaintPrompt",
"pos": [
-5280,
10
],
"size": [
252.08065795898438,
136
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 1110
}
],
"outputs": [
{
"name": "diffusers_conditioning",
"type": "CONDITIONING",
"links": [
1098
]
}
],
"properties": {
"cnr_id": "ComfyUI-DiffusersImageOutpaint",
"ver": "6a51ce5d3baa2171a85f51d462ef9f30ff9b5d26",
"Node name for S&R": "EncodeDiffusersOutpaintPrompt",
"aux_id": "GiusTex/ComfyUI-DiffusersImageOutpaint",
"widget_ue_connectable": {}
},
"widgets_values": [
"auto",
"auto",
"a verdant valley with waterfalls, rainbow"
],
"color": "#232",
"bgcolor": "#353"
},
{
"id": 532,
"type": "LoadDiffuserControlnet",
"pos": [
-5650,
10
],
"size": [
330,
130
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "CONTROL_NET",
"type": "CONTROL_NET",
"links": [
1133
]
}
],
"properties": {
"cnr_id": "ComfyUI-DiffusersImageOutpaint",
"ver": "6a51ce5d3baa2171a85f51d462ef9f30ff9b5d26",
"Node name for S&R": "LoadDiffuserControlnet",
"aux_id": "GiusTex/ComfyUI-DiffusersImageOutpaint",
"widget_ue_connectable": {}
},
"widgets_values": [
"controlnet-union-promax_sdxl.safetensors",
"auto",
"auto",
"controlnet-sdxl-promax"
],
"color": "#432",
"bgcolor": "#653"
}
],
"links": [
[
899,
494,
0,
495,
0,
"IMAGE"
],
[
906,
500,
0,
499,
0,
"IMAGE"
],
[
908,
501,
0,
494,
1,
"VAE"
],
[
1098,
570,
0,
573,
3,
"CONDITIONING"
],
[
1099,
569,
0,
573,
4,
"CONDITIONING"
],
[
1100,
499,
2,
573,
5,
"IMAGE"
],
[
1101,
573,
0,
494,
0,
"LATENT"
],
[
1110,
497,
0,
570,
0,
"CLIP"
],
[
1111,
497,
0,
569,
0,
"CLIP"
],
[
1131,
531,
0,
573,
0,
"MODEL"
],
[
1132,
531,
1,
573,
1,
"SCHEDULER"
],
[
1133,
532,
0,
573,
2,
"CONTROL_NET"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 0.7972024500000006,
"offset": [
5835.209356481534,
220.1462025110432
]
},
"frontendVersion": "1.19.9",
"groupNodes": {},
"ue_links": [],
"links_added_by_ue": [],
"VHS_latentpreview": true,
"VHS_latentpreviewrate": 0,
"VHS_MetadataImage": true,
"VHS_KeepIntermediate": true
},
"version": 0.4
}
+110 -51
View File
@@ -1,12 +1,31 @@
import torch
import os
from PIL import Image, ImageDraw
from .utils import get_first_folder_list, tensor2pil, pil2tensor, diffuserOutpaintSamples, get_device_by_name, get_dtype_by_name, clearVram
from .utils import get_config_folder_list, tensor2pil, pil2tensor, diffuserOutpaintSamples, get_device_by_name, get_dtype_by_name, clearVram, test_scheduler_scale_model_input
import folder_paths
from diffusers.models import UNet2DConditionModel
from diffusers import TCDScheduler
from .controlnet_union import ControlNetModel_Union
from diffusers.models.model_loading_utils import load_state_dict
from safetensors.torch import load_file
import logging
# Get the absolute path of various directories
my_dir = os.path.dirname(os.path.abspath(__file__))
def update_folder_names_and_paths(key, targets=[]):
# check for existing key
base = folder_paths.folder_names_and_paths.get(key, ([], {}))
base = base[0] if isinstance(base[0], (list, set, tuple)) else []
# find base key & add w/ fallback, sanity check + warning
target = next((x for x in targets if x in folder_paths.folder_names_and_paths), targets[0])
orig, _ = folder_paths.folder_names_and_paths.get(target, ([], {}))
folder_paths.folder_names_and_paths[key] = (orig or base, {".gguf"})
if base and base != orig:
logging.warning(f"Unknown file list already present on key {key}: {base}")
def can_expand(source_width, source_height, target_width, target_height, alignment):
"""Checks if the image can be expanded based on the alignment."""
if alignment in ("Left", "Right") and source_width >= target_width:
@@ -173,44 +192,92 @@ class PadImageForDiffusersOutpaint:
return (new_image, tensor_mask, tensor_cnet_image,)
class LoadDiffusersOutpaintModels:
class LoadDiffuserModel:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"model": (get_first_folder_list("diffusion_models"), {"default": "RealVisXL_V5.0_Lightning", "tooltip": "The diffuser model used for denoising the input latent. (Put model files in a folder, in diffusion_models folder)."}),
"controlnet_model": (get_first_folder_list("diffusion_models"), {"default": "controlnet-union-sdxl-1.0", "tooltip": "The controlnet model used for denoising the input latent. (Put model files in a folder, in diffusion_models folder)."}),
"unet_name": (folder_paths.get_filename_list("diffusion_models"), {"tooltip": "The name of the unet (model) to load."}),
"device": (["auto", "cuda", "cpu", "mps", "xpu", "meta"],{"default": "auto", "tooltip": "Device for inference, default is auto checked by comfyui"}),
"dtype": (["auto","fp16","bf16","fp32", "fp8_e4m3fn", "fp8_e4m3fnuz", "fp8_e5m2", "fp8_e5m2fnuz"],{"default":"auto", "tooltip": "Model precision for inference, default is auto checked by comfyui"}),
"sequential_cpu_offload": ("BOOLEAN", {"default": False, "tooltip": "Inference by default needs around 8gb vram, if this option is on it will move controlnet and unet back and forth between cpu and vram, to have only one model loaded at a time (around 6 gb vram used), useful for gpus under 8gb but will impact inference speed."}),
"model_type": (get_config_folder_list("configs"), {"default": "sdxl", "tooltip": "The json configs used for the unet. (Put unet config in \"configs/your model type/unet\", and scheduler config in \"configs/your model type/scheduler\")."}),
},
}
RETURN_TYPES = ("PIPE",)
RETURN_NAMES = ("diffusers_outpaint_pipe",)
RETURN_TYPES = ("MODEL", "SCHEDULER")
RETURN_NAMES = ("model", "scheduler configs")
FUNCTION = "load"
CATEGORY = "DiffusersOutpaint"
def load(self, model, controlnet_model, device, dtype, sequential_cpu_offload):
def load(self, unet_name, device, dtype, model_type):
# Go 2 folders back
comfy_dir = os.path.dirname(os.path.dirname(my_dir))
model_path = f"{comfy_dir}/models/diffusion_models/{model}"
controlnet_path = f"{comfy_dir}/models/diffusion_models/{controlnet_model}"
unet_path = folder_paths.get_full_path_or_raise("diffusion_models", unet_name)
device = get_device_by_name(device)
dtype = get_dtype_by_name(dtype)
diffusers_outpaint_pipe = {
"model_path": model_path,
"controlnet_model": controlnet_model,
"controlnet_path": controlnet_path,
"device": device,
"dtype": dtype,
"keep_model_device": sequential_cpu_offload,
}
if model_type == "sdxl":
print("Loading sdxl unet...")
unet = UNet2DConditionModel.from_config(f"{comfy_dir}/custom_nodes/ComfyUI-DiffusersImageOutpaint/configs", subfolder=f"{model_type}/unet").to(device, dtype)
unet.load_state_dict(load_file(unet_path))
return (diffusers_outpaint_pipe,)
scheduler = TCDScheduler.from_config(f"{comfy_dir}/custom_nodes/ComfyUI-DiffusersImageOutpaint/configs", subfolder=f"{model_type}/scheduler")
scale_model_input_method = test_scheduler_scale_model_input(comfy_dir, model_type)
scheduler_configs = {
"scheduler": scheduler,
"scale_model_input_method": scale_model_input_method,
}
return (unet, scheduler_configs,)
class LoadDiffuserControlnet:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"controlnet_model": (folder_paths.get_filename_list("controlnet"), {"tooltip": "The controlnet model used for denoising the input latent."}),
"device": (["auto", "cuda", "cpu", "mps", "xpu", "meta"],{"default": "auto", "tooltip": "Device for inference, default is auto checked by comfyui"}),
"dtype": (["auto","fp16","bf16","fp32", "fp8_e4m3fn", "fp8_e4m3fnuz", "fp8_e5m2", "fp8_e5m2fnuz"],{"default":"auto", "tooltip": "Model precision for inference, default is auto checked by comfyui"}),
"controlnet_type": (get_config_folder_list("configs"), {"default": "controlnet-sdxl-promax", "tooltip": "The json configs used for controlnet. (Put config(s) in \"configs/your controlnet type\")."}),
},
}
RETURN_TYPES = ("CONTROL_NET",)
FUNCTION = "load"
CATEGORY = "DiffusersOutpaint"
def load(self, controlnet_model, device, dtype, controlnet_type):
# Go 2 folders back
comfy_dir = os.path.dirname(os.path.dirname(my_dir))
controlnet_path = folder_paths.get_full_path_or_raise("controlnet", controlnet_model)
device = get_device_by_name(device)
dtype = get_dtype_by_name(dtype)
if controlnet_type == "controlnet-sdxl-promax":
print("Loading controlnet-sdxl-promax...")
controlnet_model = ControlNetModel_Union.from_config(f"{comfy_dir}/custom_nodes/ComfyUI-DiffusersImageOutpaint/configs/{controlnet_type}/config_promax.json")
state_dict = load_state_dict(load_file(controlnet_path))
model, _, _, _, _ = ControlNetModel_Union._load_pretrained_model(
controlnet_model, state_dict, controlnet_path, controlnet_path
)
controlnet_model.to(device, dtype)
del model, state_dict, controlnet_path
clearVram(device)
return (controlnet_model,)
class EncodeDiffusersOutpaintPrompt:
@@ -218,21 +285,22 @@ class EncodeDiffusersOutpaintPrompt:
def INPUT_TYPES(s):
return {
"required": {
"diffusers_outpaint_pipe": ("PIPE", {"tooltip": "Load the diffusers outpaint models."}),
"device": (["auto", "cuda", "cpu", "mps", "xpu", "meta"],{"default": "auto", "tooltip": "Device for inference, default is auto checked by comfyui"}),
"dtype": (["auto","fp16","bf16","fp32", "fp8_e4m3fn", "fp8_e4m3fnuz", "fp8_e5m2", "fp8_e5m2fnuz"],{"default":"auto", "tooltip": "Model precision for inference, default is auto checked by comfyui"}),
"text": ("STRING", {"multiline": True, "dynamicPrompts": True, "tooltip": "The text to be encoded."}),
"clip": ("CLIP", {"tooltip": "The CLIP model used for encoding the text."})
}
}
RETURN_TYPES = ("PIPE","CONDITIONING",)
RETURN_NAMES = ("diffusers_outpaint_pipe","diffusers_conditioning",)
RETURN_TYPES = ("CONDITIONING",)
RETURN_NAMES = ("diffusers_conditioning",)
OUTPUT_TOOLTIPS = ("A conditioning containing the embedded text used to guide the diffusion model.",)
FUNCTION = "encode"
CATEGORY = "DiffusersOutpaint"
DESCRIPTION = "Encodes a text prompt using a CLIP model into an embedding that can be used to guide the diffusion model towards generating specific images."
def encode(self, diffusers_outpaint_pipe, text, clip):
dtype = diffusers_outpaint_pipe["dtype"]
device = diffusers_outpaint_pipe["device"]
def encode(self, device, dtype, text, clip):
device = get_device_by_name(device)
dtype = get_dtype_by_name(dtype)
text = f"{text}, high quality, 4k"
tokens = clip.tokenize(text)
@@ -255,7 +323,7 @@ class EncodeDiffusersOutpaintPrompt:
"pooled_prompt_embeds": pooled_prompt_embeds,
}
return (diffusers_outpaint_pipe,diffusers_conditioning,)
return (diffusers_conditioning,)
class DiffusersImageOutpaint:
@@ -263,45 +331,36 @@ class DiffusersImageOutpaint:
def INPUT_TYPES(s):
return {
"required": {
"diffusers_outpaint_pipe": ("PIPE", {"tooltip": "Load the diffusers outpaint models."}),
"model": ("MODEL", {"tooltip": "The model used for denoising the input latent."}),
"scheduler_configs": ("SCHEDULER",),
"control_net": ("CONTROL_NET",),
"positive": ("CONDITIONING", {"tooltip": "The prompt describing what you want."}),
"negative": ("CONDITIONING", {"tooltip": "The prompt describing what you don't want."}),
"diffuser_outpaint_cnet_image": ("IMAGE", {"tooltip": "The image to outpaint."}),
"guidance_scale": ("FLOAT", {"default": 1.50, "min": 1.01, "max": 10, "step": 0.01, "tooltip": "The Classifier-Free Guidance scale balances creativity and adherence to the prompt. Higher values result in images more closely matching the prompt, however too high values will negatively impact quality."}),
"controlnet_strength": ("FLOAT", {"default": 1.00, "min": 0.00, "max": 10, "step": 0.01}),
"seed": ("INT", {"default": 0, "min": 0, "max": 0xffffffffffffffff, "tooltip": "Fake seed, workaround used to keep generating different outpaints. Set to -1 to generate different images, or a fixed number to stop that."}),
"steps": ("INT", {"default": 8, "min": 4, "max": 20, "tooltip": "The number of steps used in the denoising process."}),
}
"device": (["auto", "cuda", "cpu", "mps", "xpu", "meta"],{"default": "auto", "tooltip": "Device for inference, default is auto checked by comfyui"}),
"dtype": (["auto","fp16","bf16","fp32", "fp8_e4m3fn", "fp8_e4m3fnuz", "fp8_e5m2", "fp8_e5m2fnuz"],{"default":"auto", "tooltip": "Model precision for inference, default is auto checked by comfyui"}),
"sequential_cpu_offload": ("BOOLEAN", {"default": False, "tooltip": "Inference by default needs around 8gb vram, if this option is on it will move controlnet and unet back and forth between cpu and vram, to have only one model loaded at a time (around 6 gb vram used), useful for gpus under 8gb but will impact inference speed."}),
},
}
RETURN_TYPES = ("LATENT",)
FUNCTION = "sample"
CATEGORY = "DiffusersOutpaint"
def sample(self, diffusers_outpaint_pipe, positive, negative, diffuser_outpaint_cnet_image, guidance_scale, controlnet_strength, seed, steps):
def sample(self, device, dtype, sequential_cpu_offload, scheduler_configs, model, control_net, positive, negative, diffuser_outpaint_cnet_image, guidance_scale, controlnet_strength, steps):
cnet_image = diffuser_outpaint_cnet_image
cnet_image=tensor2pil(cnet_image)
cnet_image=cnet_image.convert('RGB')
model_path = diffusers_outpaint_pipe["model_path"]
controlnet_model = diffusers_outpaint_pipe["controlnet_model"]
controlnet_path = diffusers_outpaint_pipe["controlnet_path"]
dtype = diffusers_outpaint_pipe["dtype"]
device = diffusers_outpaint_pipe["device"]
keep_model_device = diffusers_outpaint_pipe["keep_model_device"]
prompt_embeds = positive["prompt_embeds"]
pooled_prompt_embeds = positive["pooled_prompt_embeds"]
negative_prompt_embeds = negative["prompt_embeds"]
negative_pooled_prompt_embeds = negative["pooled_prompt_embeds"]
keep_model_device = sequential_cpu_offload
last_rgb_latent = diffuserOutpaintSamples(model_path, controlnet_model, diffuser_outpaint_cnet_image, dtype, controlnet_path,
prompt_embeds, negative_prompt_embeds, pooled_prompt_embeds, negative_pooled_prompt_embeds,
device, steps, controlnet_strength, guidance_scale,
keep_model_device)
scheduler = scheduler_configs["scheduler"]
scale_model_input_method = scheduler_configs["scale_model_input_method"]
last_rgb_latent = diffuserOutpaintSamples(device, dtype, keep_model_device, scheduler, scale_model_input_method, model, control_net, positive, negative,
cnet_image, controlnet_strength, guidance_scale, steps)
del prompt_embeds, pooled_prompt_embeds, negative_prompt_embeds, negative_pooled_prompt_embeds
clearVram(device)
return ({"samples":last_rgb_latent},)
+17 -25
View File
@@ -17,12 +17,10 @@ from typing import List, Optional, Union
import cv2
import PIL.Image
import torch
import gc
from diffusers.image_processor import PipelineImageInput, VaeImageProcessor
from diffusers.models import AutoencoderKL, UNet2DConditionModel
from diffusers.pipelines.pipeline_utils import DiffusionPipeline, StableDiffusionMixin
from diffusers.schedulers import KarrasDiffusionSchedulers
from diffusers.utils.torch_utils import randn_tensor
from tqdm import tqdm
from .controlnet_union import ControlNetModel_Union
from comfy.utils import ProgressBar
@@ -67,21 +65,13 @@ def retrieve_timesteps(
return timesteps, num_inference_steps
class StableDiffusionXLFillPipeline(DiffusionPipeline, StableDiffusionMixin):
class StableDiffusionXLFillPipeline:
def __init__(
self,
unet: UNet2DConditionModel,
scheduler: KarrasDiffusionSchedulers,
force_zeros_for_empty_prompt: bool = True,
):
super().__init__()
self.register_modules(
unet=unet,
scheduler=scheduler,
)
self.vae_scale_factor = 8
self.image_processor = VaeImageProcessor(
vae_scale_factor=self.vae_scale_factor, do_convert_rgb=True
@@ -91,10 +81,6 @@ class StableDiffusionXLFillPipeline(DiffusionPipeline, StableDiffusionMixin):
do_convert_rgb=True,
do_normalize=False,
)
self.register_to_config(
force_zeros_for_empty_prompt=force_zeros_for_empty_prompt
)
self.controlnet_model = None
def prepare_image(self, image, device, dtype, do_classifier_free_guidance=False):
image = self.control_image_processor.preprocess(image).to(dtype=torch.float32)
@@ -134,7 +120,10 @@ class StableDiffusionXLFillPipeline(DiffusionPipeline, StableDiffusionMixin):
# corresponds to doing no classifier free guidance.
@property
def do_classifier_free_guidance(self):
return self._guidance_scale > 1 and self.unet.config.time_cond_proj_dim is None
if hasattr(self.unet, 'config'):
return self._guidance_scale > 1 and self.unet.config.time_cond_proj_dim is None
else:
return self._guidance_scale > 1
@property
def num_timesteps(self):
@@ -147,6 +136,10 @@ class StableDiffusionXLFillPipeline(DiffusionPipeline, StableDiffusionMixin):
device,
dtype,
keep_model_device,
scheduler: KarrasDiffusionSchedulers,
unet: object,
timesteps,
scale_model_input_method,
prompt_embeds: torch.Tensor,
pooled_prompt_embeds: torch.Tensor,
negative_prompt_embeds: torch.Tensor,
@@ -158,6 +151,11 @@ class StableDiffusionXLFillPipeline(DiffusionPipeline, StableDiffusionMixin):
):
self.controlnet = controlnet_model
self._guidance_scale = guidance_scale
self.unet = unet
self.scheduler = scheduler
self.timesteps = timesteps
self.scale_model_input_method=scale_model_input_method
# 2. Define call parameters
batch_size = 1
@@ -228,7 +226,7 @@ class StableDiffusionXLFillPipeline(DiffusionPipeline, StableDiffusionMixin):
num_warmup_steps = len(timesteps) - num_inference_steps * self.scheduler.order
ComfyUI_ProgressBar = ProgressBar(int(num_inference_steps))
with self.progress_bar(total=num_inference_steps) as progress_bar:
with tqdm(total=num_inference_steps) as pbar:
for i, t in enumerate(timesteps):
# expand the latents if we are doing classifier free guidance
latent_model_input = (
@@ -314,14 +312,8 @@ class StableDiffusionXLFillPipeline(DiffusionPipeline, StableDiffusionMixin):
if i == len(timesteps) - 1 or (
(i + 1) > num_warmup_steps and (i + 1) % self.scheduler.order == 0
):
progress_bar.update()
pbar.update()
ComfyUI_ProgressBar.update(1)
#yield latents_to_rgb(latents)
del self.unet
del self.controlnet
gc.collect()
torch.cuda.empty_cache()
latents = latents / 0.13025
yield latents
+3 -3
View File
@@ -1,7 +1,7 @@
torch
numpy==1.26.4
transformers
transformers==4.45.0
accelerate
diffusers
diffusers==0.32.2
fastapi<0.113.0
opencv-python
opencv-python
+49 -74
View File
@@ -7,12 +7,7 @@ import comfy.model_management as mm
from PIL import Image
from folder_paths import map_legacy, folder_names_and_paths
from .controlnet_union import ControlNetModel_Union
from .pipeline_fill_sd_xl import StableDiffusionXLFillPipeline
from diffusers import AutoencoderKL, TCDScheduler
from diffusers.models.model_loading_utils import load_state_dict
from transformers import CLIPTextModel, CLIPTextModelWithProjection, CLIPTokenizer
from diffusers import UNet2DConditionModel
def get_first_folder_list(folder_name: str) -> tuple[list[str], dict[str, float], float]:
@@ -23,9 +18,17 @@ def get_first_folder_list(folder_name: str) -> tuple[list[str], dict[str, float]
root_folder = folders[0][0]
elif folder_name == "diffusion_models":
root_folder = folders[0][1]
elif folder_name == "controlnet":
root_folder = folders[0][0]
visible_folders = [name for name in os.listdir(root_folder) if os.path.isdir(os.path.join(root_folder, name))]
return visible_folders
def get_config_folder_list(folder_name: str) -> tuple[list[str], dict[str, float], float]:
my_dir = os.path.dirname(os.path.abspath(__file__))
configs_dir = f"{my_dir}/{folder_name}"
folders = [f for f in os.listdir(configs_dir) if os.path.isdir(os.path.join(configs_dir, f))]
return folders
# Tensor to PIL (grabbed from WAS Suite)
def tensor2pil(image: torch.Tensor) -> Image.Image:
@@ -67,16 +70,6 @@ def get_dtype_by_name(dtype):
return dtype
def loadDiffModels1(model_path, dtype, device):
tokenizer = CLIPTokenizer.from_pretrained(model_path, subfolder="tokenizer", use_fast=False)
tokenizer_2 = CLIPTokenizer.from_pretrained(model_path, subfolder="tokenizer_2", use_fast=False)
text_encoder = CLIPTextModel.from_pretrained(model_path, subfolder="text_encoder", torch_dtype=dtype).requires_grad_(False).to(device)
text_encoder_2 = CLIPTextModelWithProjection.from_pretrained(model_path, subfolder="text_encoder_2", torch_dtype=dtype).requires_grad_(False).to(device)
return tokenizer, tokenizer_2, text_encoder, text_encoder_2
def clearVram(device):
gc.collect()
@@ -91,73 +84,51 @@ def clearVram(device):
torch.xpu.empty_cache()
elif device.type == "meta":
torch.meta.empty_cache()
class TCDScheduler_Custom:
def __init__(self, **kwargs):
for key, value in kwargs.items():
setattr(self, key, value)
# torch.ipc_collect() not available, and ipc_collect seems available only for cuda
def loadControlnetModel(device, dtype, controlnet_path):
config_file = f"{controlnet_path}/config_promax.json"
config = ControlNetModel_Union.load_config(config_file)
controlnet_model = ControlNetModel_Union.from_config(config)
def scale_model_input(self, input, t):
scale_factor = getattr(self, 'scale_factor', 1)
return input * scale_factor
model_file = f"{controlnet_path}/diffusion_pytorch_model_promax.safetensors"
state_dict = load_state_dict(model_file)
model, _, _, _, _ = ControlNetModel_Union._load_pretrained_model(
controlnet_model, state_dict, model_file, f"{controlnet_path}"
)
controlnet_model.to(device, dtype)
del model, state_dict, model_file
def __repr__(self):
attrs = {key: value for key, value in self.__dict__.items()}
return f"TCDScheduler({attrs})"
clearVram(device)
return controlnet_model
def test_scheduler_scale_model_input(comfy_dir, model_type):
scheduler_config_path = f"{comfy_dir}/custom_nodes/ComfyUI-DiffusersImageOutpaint/configs/{model_type}/scheduler/scheduler_config.json"
def loadVaeModel(vae_path, device, dtype, enable_vae_slicing, enable_vae_tiling):
vae = AutoencoderKL.from_pretrained(f"{vae_path}").to(device, dtype)
if enable_vae_slicing:
vae.enable_slicing()
else:
vae.disable_slicing()
with open(scheduler_config_path, 'r') as f:
config = json.load(f)
if enable_vae_tiling:
vae.enable_tiling()
else:
vae.disable_tiling()
return vae
scheduler = TCDScheduler_Custom(**config)
scale_model_input_method = scheduler.scale_model_input
return scale_model_input_method
def loadUnetModel(model_path, device, dtype):
unet = UNet2DConditionModel.from_pretrained(model_path, subfolder="unet", use_safetensors=True)
unet.to(device, dtype)
return unet
def diffuserOutpaintSamples(model_path, controlnet_model, diffuser_outpaint_cnet_image, dtype, controlnet_path,
prompt_embeds, negative_prompt_embeds, pooled_prompt_embeds, negative_pooled_prompt_embeds,
device, steps, controlnet_strength, guidance_scale,
keep_model_device):
def diffuserOutpaintSamples(device, dtype, keep_model_device, scheduler, scale_model_input_method, model, control_net, positive, negative,
cnet_image, controlnet_strength, guidance_scale, steps):
controlnet_model = loadControlnetModel(device, dtype, controlnet_path)
unet = loadUnetModel(model_path, device, dtype)
with open(f"{model_path}/scheduler/scheduler_config.json", "r") as f:
scheduler_config = json.load(f)
scheduler = TCDScheduler.from_config(scheduler_config)
prompt_embeds = positive["prompt_embeds"]
pooled_prompt_embeds = positive["pooled_prompt_embeds"]
negative_prompt_embeds = negative["prompt_embeds"]
negative_pooled_prompt_embeds = negative["pooled_prompt_embeds"]
controlnet_model = control_net
pipe = StableDiffusionXLFillPipeline(
unet,
scheduler=scheduler,
)
if not keep_model_device:
pipe.to(device)
device = get_device_by_name(device)
dtype = get_dtype_by_name(dtype)
timesteps = None
unet = model
pipe = StableDiffusionXLFillPipeline()
cnet_image = diffuser_outpaint_cnet_image
cnet_image=tensor2pil(cnet_image)
cnet_image=cnet_image.convert('RGB')
rgb_latents = list(pipe(
prompt_embeds=prompt_embeds,
negative_prompt_embeds=negative_prompt_embeds,
@@ -170,13 +141,17 @@ def diffuserOutpaintSamples(model_path, controlnet_model, diffuser_outpaint_cnet
guidance_scale=guidance_scale,
device=device,
dtype=dtype,
unet=unet,
timesteps=timesteps,
scale_model_input_method=scale_model_input_method,
keep_model_device=keep_model_device,
))
scheduler=scheduler,
))
last_rgb_latent = rgb_latents[-1] # Access the last image
del pipe, controlnet_model, scheduler, prompt_embeds, negative_prompt_embeds, pooled_prompt_embeds, negative_pooled_prompt_embeds
del pipe, unet, controlnet_model, scheduler, prompt_embeds, negative_prompt_embeds, pooled_prompt_embeds, negative_pooled_prompt_embeds
clearVram(device)
return last_rgb_latent