Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
2d7912f9a5 | ||
|
|
93b90fad99 | ||
|
|
d82b218e1c | ||
|
|
a9238253e0 | ||
|
|
391ca9625f | ||
|
|
e7e8166e02 | ||
|
|
f564ee886b | ||
|
|
0168172781 | ||
|
|
be34cf3054 | ||
|
|
649d6ea465 | ||
|
|
c054966ae8 | ||
|
|
e513cf6778 | ||
|
|
32cf848e93 | ||
|
|
12c8035ffa |
@@ -13,17 +13,29 @@
|
||||
|
||||
Universally applicable inpainting ability for every model. LanPaint sampler lets the model "think" through multiple iterations before denoising, enabling you to invest more computation time for superior inpainting quality.
|
||||
|
||||
This is the official implementation of ["LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling"](https://arxiv.org/abs/2502.03491), accepted by TMLR.
|
||||
## What LanPaint Enables
|
||||
|
||||
The repository is for ComfyUI extension.
|
||||
- Training-free image inpainting
|
||||
- Outpainting and generative fill
|
||||
- Mask-constrained local image editing
|
||||
- Object / region replacement with text guidance
|
||||
- Character-consistent local generation
|
||||
- Video inpainting and local video editing
|
||||
- Video + audio masked generation
|
||||
|
||||
Diffusers Support: [LanPaint-Diffusers](https://github.com/charrywhite/LanPaint-diffusers) by [@charrywhite](https://github.com/charrywhite/)
|
||||
## Research & Benchmark
|
||||
|
||||
Benchmark code for paper reproduce: [LanPaintBench](https://github.com/scraed/LanPaintBench).
|
||||
LanPaint is a training-free partial conditional sampler that enables mask-constrained inpainting and local editing with pretrained diffusion and rectified-flow models, without fine-tuning or backpropagation.
|
||||
|
||||
## Citation
|
||||
* 📄 **Paper:** [LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling](https://openreview.net/forum?id=JPC8JyOUSW) — TMLR 2025
|
||||
* 🧩 **ComfyUI Implementation:** This repository
|
||||
* 🐍 **Diffusers Implementation:** [LanPaint-Diffusers](https://github.com/charrywhite/LanPaint-diffusers) by [@charrywhite](https://github.com/charrywhite/)
|
||||
* 🧪 **Benchmark & Reproduction:** [LanPaintBench](https://github.com/scraed/LanPaintBench)
|
||||
* 🌐 **Project Website:** [LanPaint Page](https://scraed.github.io/scraedBlog/lanpaint/)
|
||||
|
||||
```
|
||||
### Citation
|
||||
|
||||
```bibtex id="2r5ioa"
|
||||
@article{
|
||||
zheng2025lanpaint,
|
||||
title={LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling},
|
||||
@@ -31,17 +43,27 @@ author={Candi Zheng and Yuan Lan and Yang Wang},
|
||||
journal={Transactions on Machine Learning Research},
|
||||
issn={2835-8856},
|
||||
year={2025},
|
||||
url={https://openreview.net/forum?id=JPC8JyOUSW},
|
||||
note={}
|
||||
url={https://openreview.net/forum?id=JPC8JyOUSW}
|
||||
}
|
||||
```
|
||||
|
||||
**🎉 NEW 2026: Join our discord!**
|
||||
|
||||
[Join our Discord](https://discord.gg/yN5wYDE6W4) to share experiences, discuss features, and explore future development.
|
||||
|
||||
`v1.5.0` fixes an important hidden bug that reduced performance and could blur images (especially with `z-image-base`) and also boosts overall LanPaint performance across other models.
|
||||
`v2.1.0` significantly accelerates LanPaint with a new schedule mechanism and fixes MiniMax H3 support on the latest ComfyUI.
|
||||
If your inpainting results have wierd (glowing / broken) mask boundary, check this [issue](https://github.com/scraed/LanPaint/issues/80).
|
||||
|
||||
**🎨 NEW: LanPaint now supports Qwen-Image 2.1 - transparency, and masked image editing!**
|
||||
|
||||

|
||||
|
||||
Qwen 2.1's **image edit** model now works under a LanPaint mask: tell it what to change, paint over the part you want it to touch, and only that part changes. Hand it a second picture to borrow from if you want one. Check our latest [Qwen Image 2.1 Image Edit Example](#example-qwen-image-21-image-edit-masked-inpaintlanpaint-k-sampler-5-steps-of-thinking).
|
||||
|
||||

|
||||
|
||||
And if your picture carries transparency, it gets inpainted too - the rebuilt part comes back with a new outline, not just new colours. Check our latest [Qwen Image 2.1 Example](#example-qwen-image-21-inpaint-with-transparencylanpaint-k-sampler-5-steps-of-thinking).
|
||||
|
||||
**🎬 NEW: LanPaint now supports MiniMax H3 video + audio inpainting!**
|
||||
|
||||
| Masked Input (paint in the editor) | Mask (visible overlay) | Inpainted Result |
|
||||
@@ -103,10 +125,11 @@ Check our latest [Krea2 Example](#example-krea2-inpaintlanpaint-k-sampler-3-step
|
||||
- [Features](#features)
|
||||
- [Quickstart](#quickstart)
|
||||
- [How to Use Examples](#how-to-use-examples)
|
||||
- [Video Examples (Beta)](#video-examples-beta)
|
||||
- [Video Examples](#video-examples)
|
||||
- [Wan 2.2 Video Inpainting](#wan-22-video-inpainting)
|
||||
- [Wan 2.2 5B Video Inpainting](#wan-22-5b-video-inpainting)
|
||||
- [Wan 2.2 Video Outpainting](#wan-22-video-outpainting)
|
||||
- [MiniMax H3 Video + Audio Inpainting](#minimax-h3-video--audio-inpainting-av-pipeline)
|
||||
- [Resource Consumption](#resource-consumption)
|
||||
- [Image Examples](#image-examples)
|
||||
- [Flux.2.Dev](#example-flux2dev-inpaintlanpaint-k-sampler-5-steps-of-thinking)
|
||||
@@ -121,6 +144,8 @@ Check our latest [Krea2 Example](#example-krea2-inpaintlanpaint-k-sampler-3-step
|
||||
- [Wan 2.2 T2I with reference](#example-wan22-partial-inpaintlanpaint-k-sampler-5-steps-of-thinking)
|
||||
- [Qwen Image Edit 2511 2509](#example-qwen-edit-2509-inpaint)
|
||||
- [Qwen Image Edit 2508](#example-qwen-edit-2508-inpaint)
|
||||
- [Qwen Image 2.1 Image Edit](#example-qwen-image-21-image-edit-masked-inpaintlanpaint-k-sampler-5-steps-of-thinking)
|
||||
- [Qwen Image 2.1](#example-qwen-image-21-inpaint-with-transparencylanpaint-k-sampler-5-steps-of-thinking)
|
||||
- [Qwen Image](#example-qwen-image-inpaintlanpaint-k-sampler-5-steps-of-thinking)
|
||||
- [HiDream](#example-hidream-inpaint-lanpaint-k-sampler-5-steps-of-thinking)
|
||||
- [SD 3.5](#example-sd-35-inpaintlanpaint-k-sampler-5-steps-of-thinking)
|
||||
@@ -138,7 +163,7 @@ Check our latest [Krea2 Example](#example-krea2-inpaintlanpaint-k-sampler-3-step
|
||||
|
||||
## Features
|
||||
|
||||
- **Universal Compatibility** – Works instantly with almost any model (**Ideogram4, Krea2, Z-image, Z-image-base, Hunyuan, Wan 2.2, Qwen Image/Edit, Anima, HiDream, SD 3.5, Flux-series, SDXL, SD 1.5 or custom LoRAs**) and ControlNet.
|
||||
- **Universal Compatibility** – Works instantly with almost any model (**Ideogram4, Krea2, Z-image, Z-image-base, Hunyuan, Wan 2.2, Qwen Image 2.1/Image/Edit, Anima, HiDream, SD 3.5, Flux-series, SDXL, SD 1.5 or custom LoRAs**) and ControlNet.
|
||||

|
||||
- **No Training Needed** – Works out of the box with your existing model.
|
||||
- **Easy to Use** – Same workflow as standard ComfyUI KSampler.
|
||||
@@ -177,7 +202,7 @@ Once installed, you'll find the LanPaint nodes under the "sampling" category in
|
||||
- **[VAE Encode for Inpainting](https://comfyanonymous.github.io/ComfyUI_examples/inpaint/)**
|
||||
- **[Set Latent Noise Mask](https://comfyui-wiki.com/en/tutorial/basic/how-to-inpaint-an-image-in-comfyui)**
|
||||
|
||||
## Video Examples (Beta)
|
||||
## Video Examples
|
||||
|
||||
LanPaint now supports video inpainting with Wan 2.2, enabling you to seamlessly inpaint masked regions across video frames while maintaining temporal consistency.
|
||||
|
||||
@@ -435,6 +460,20 @@ Check [Mased Qwen Edit Workflow](https://github.com/scraed/LanPaint/tree/master/
|
||||
|
||||
|
||||
|
||||
### Example Qwen Image 2.1 Image Edit: Masked InPaint(LanPaint K Sampler, 5 steps of thinking)
|
||||
|
||||
Qwen-Image 2.1's image edit model now works under a LanPaint mask: write what you want changed, paint over the part it should touch, and only that part changes - everything else, transparency included, comes back exactly as it was. In this example a second picture supplies the material for the earcups, and the headband and stitching stay as they are. Workflow and images are in `examples/Example_32`; drag `InPainted_Drag_Me_to_ComfyUI.png` into ComfyUI to load it. Use your own pictures with the official [Qwen Image 2.1 Image Edit template](https://docs.comfy.org/tutorials/image/qwen/qwen-image-2-1).
|
||||
|
||||

|
||||
[View Workflow & Masks](https://github.com/scraed/LanPaint/tree/master/examples/Example_32) · [Workflow JSON](https://github.com/scraed/LanPaint/blob/master/example_workflows/Qwen_Image_2.1_Edit_Masked_Inpaint.json)
|
||||
|
||||
### Example Qwen Image 2.1: InPaint with Transparency(LanPaint K Sampler, 5 steps of thinking)
|
||||
|
||||
Qwen-Image 2.1 inpaints a picture's transparency along with its pixels, so the rebuilt part can come back with a new outline instead of merely new colours - here the boot's sole is replaced and the silhouette grows with it. Workflow and images are in `examples/Example_31`; drag `InPainted_Drag_Me_to_ComfyUI.png` into ComfyUI to load it.
|
||||
|
||||

|
||||
[View Workflow & Masks](https://github.com/scraed/LanPaint/tree/master/examples/Example_31) · [Workflow JSON](https://github.com/scraed/LanPaint/blob/master/example_workflows/Transparent_Edit_EncodeDecode_Inpaint.json)
|
||||
|
||||
### Example Qwen Image: InPaint(LanPaint K Sampler, 5 steps of thinking)
|
||||
|
||||

|
||||
@@ -619,6 +658,11 @@ Submit a PR to add your tutorial/video here, or open an [Issue](https://github.c
|
||||
[Working togather with crop&stitch](https://github.com/scraed/LanPaint/issues/46)
|
||||
|
||||
## Updates
|
||||
- 2026/09/28
|
||||
- Add Qwen-Image 2.1 image edit support: masked, instruction-driven editing (Example_32).
|
||||
- Add Qwen-Image 2.1 inpainting support with LanPaint KSampler (Example_31).
|
||||
- Inpainting a picture that carries transparency now works end to end: the 2.1 VAE is 4-in/4-out, so the alpha travels through the latent and is edited alongside the pixels. Keep the inpainting mask in its own greyscale file, since 2.1's alpha channel means image transparency.
|
||||
- `LanPaint_ImageDecode` now matches the decoded channel count to the source image, so an RGBA source comes back RGBA and an RGB source still comes back RGB.
|
||||
- 2026/08/12
|
||||
- `v2.1.0`: Significantly accelerated LanPaint using a new schedule mechanism.
|
||||
- Fix bugs for MiniMax H3 on the latest ComfyUI.
|
||||
|
||||
|
After Width: | Height: | Size: 199 KiB |
|
After Width: | Height: | Size: 219 KiB |
@@ -0,0 +1,976 @@
|
||||
{
|
||||
"id": "c4e8a1f7-2b93-4d6e-8a50-1f7e9c3b6d28",
|
||||
"revision": 0,
|
||||
"last_node_id": 14,
|
||||
"last_link_id": 19,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 1,
|
||||
"type": "UNETLoader",
|
||||
"pos": [
|
||||
-1180,
|
||||
40
|
||||
],
|
||||
"size": [
|
||||
390,
|
||||
82
|
||||
],
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "MODEL",
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
1
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "UNETLoader",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"qwen_image_2.1_int8_convrot.safetensors",
|
||||
"default"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"type": "CLIPLoader",
|
||||
"pos": [
|
||||
-1180,
|
||||
170
|
||||
],
|
||||
"size": [
|
||||
390,
|
||||
106
|
||||
],
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "CLIP",
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
3,
|
||||
4
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPLoader",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"qwen3vl_8b_int8_convrot.safetensors",
|
||||
"qwen_image",
|
||||
"default"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "VAELoader",
|
||||
"pos": [
|
||||
-1180,
|
||||
320
|
||||
],
|
||||
"size": [
|
||||
390,
|
||||
58
|
||||
],
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "VAE",
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
5,
|
||||
6,
|
||||
7
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAELoader",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"qwen_image_2.1_vae_bf16.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"type": "QwenImage21Cache",
|
||||
"pos": [
|
||||
-740,
|
||||
40
|
||||
],
|
||||
"size": [
|
||||
310,
|
||||
82
|
||||
],
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"localized_name": "model",
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 1
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "MODEL",
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
2
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "QwenImage21Cache",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"auto",
|
||||
"default"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
-1180,
|
||||
420
|
||||
],
|
||||
"size": [
|
||||
390,
|
||||
440
|
||||
],
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "IMAGE",
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
8
|
||||
]
|
||||
},
|
||||
{
|
||||
"localized_name": "MASK",
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"slot_index": 1,
|
||||
"links": [
|
||||
9
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"Image_Load_Me_in_Loader.png",
|
||||
"image"
|
||||
],
|
||||
"title": "The picture (RGBA: the alpha is the background)"
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "JoinImageWithAlpha",
|
||||
"pos": [
|
||||
-740,
|
||||
420
|
||||
],
|
||||
"size": [
|
||||
330,
|
||||
60
|
||||
],
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"localized_name": "image",
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 8
|
||||
},
|
||||
{
|
||||
"localized_name": "alpha",
|
||||
"name": "alpha",
|
||||
"type": "MASK",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "IMAGE",
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
10,
|
||||
11,
|
||||
12
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "JoinImageWithAlpha",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [],
|
||||
"title": "Re-attach the alpha as the 4th channel"
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "LoadImageMask",
|
||||
"pos": [
|
||||
-1180,
|
||||
900
|
||||
],
|
||||
"size": [
|
||||
390,
|
||||
480
|
||||
],
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "MASK",
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
13,
|
||||
14
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImageMask",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"Mask_Load_Me_in_Loader.png",
|
||||
"red"
|
||||
],
|
||||
"title": "The mask, on its own (white = rebuild)"
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "TextEncodeQwenImageEditPlus",
|
||||
"pos": [
|
||||
-740,
|
||||
520
|
||||
],
|
||||
"size": [
|
||||
440,
|
||||
250
|
||||
],
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"localized_name": "clip",
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
},
|
||||
{
|
||||
"localized_name": "prompt",
|
||||
"name": "prompt",
|
||||
"type": "STRING",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "prompt"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "vae",
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 5
|
||||
},
|
||||
{
|
||||
"localized_name": "image1",
|
||||
"name": "image1",
|
||||
"type": "IMAGE",
|
||||
"link": 10
|
||||
},
|
||||
{
|
||||
"localized_name": "image2",
|
||||
"name": "image2",
|
||||
"type": "IMAGE",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"localized_name": "image3",
|
||||
"name": "image3",
|
||||
"type": "IMAGE",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "CONDITIONING",
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
15
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "TextEncodeQwenImageEditPlus",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"the boot's sole and forefoot replaced by a sleek futuristic mechanical unit: segmented brushed titanium plating with fine panel lines, glowing cyan energy strips along the flanks, a sculpted dark carbon-fibre heel block, precise engineered details and tiny bolts, the brown leather upper above stays exactly as it is, high-end sci-fi product photography, transparent background"
|
||||
],
|
||||
"title": "Text Encode Qwen Image Edit Plus (reference + prompt)",
|
||||
"color": "#232",
|
||||
"bgcolor": "#353"
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
-740,
|
||||
800
|
||||
],
|
||||
"size": [
|
||||
440,
|
||||
150
|
||||
],
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"localized_name": "clip",
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"localized_name": "text",
|
||||
"name": "text",
|
||||
"type": "STRING",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "text"
|
||||
}
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "CONDITIONING",
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
16
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感,模糊,变形"
|
||||
],
|
||||
"title": "CLIP Text Encode (Negative Prompt)",
|
||||
"color": "#223",
|
||||
"bgcolor": "#335"
|
||||
},
|
||||
{
|
||||
"id": 10,
|
||||
"type": "LanPaint_ImageEncode",
|
||||
"pos": [
|
||||
-240,
|
||||
40
|
||||
],
|
||||
"size": [
|
||||
300,
|
||||
110
|
||||
],
|
||||
"flags": {},
|
||||
"order": 10,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"localized_name": "image",
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 11
|
||||
},
|
||||
{
|
||||
"localized_name": "vae",
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"localized_name": "mask",
|
||||
"name": "mask",
|
||||
"type": "MASK",
|
||||
"link": 13
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "LATENT",
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
17
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LanPaint_ImageEncode",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": []
|
||||
},
|
||||
{
|
||||
"id": 11,
|
||||
"type": "LanPaint_KSampler",
|
||||
"pos": [
|
||||
-240,
|
||||
190
|
||||
],
|
||||
"size": [
|
||||
330,
|
||||
320
|
||||
],
|
||||
"flags": {},
|
||||
"order": 11,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"localized_name": "model",
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 2
|
||||
},
|
||||
{
|
||||
"localized_name": "positive",
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 15
|
||||
},
|
||||
{
|
||||
"localized_name": "negative",
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 16
|
||||
},
|
||||
{
|
||||
"localized_name": "latent_image",
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 17
|
||||
},
|
||||
{
|
||||
"localized_name": "seed",
|
||||
"name": "seed",
|
||||
"type": "INT",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "seed"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "steps",
|
||||
"name": "steps",
|
||||
"type": "INT",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "steps"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "cfg",
|
||||
"name": "cfg",
|
||||
"type": "FLOAT",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "cfg"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "sampler_name",
|
||||
"name": "sampler_name",
|
||||
"type": "COMBO",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "sampler_name"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "scheduler",
|
||||
"name": "scheduler",
|
||||
"type": "COMBO",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "scheduler"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "denoise",
|
||||
"name": "denoise",
|
||||
"type": "FLOAT",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "denoise"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "LanPaint_NumSteps",
|
||||
"name": "LanPaint_NumSteps",
|
||||
"type": "INT",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "LanPaint_NumSteps"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "LanPaint_PromptMode",
|
||||
"name": "LanPaint_PromptMode",
|
||||
"type": "COMBO",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "LanPaint_PromptMode"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "LanPaint_Info",
|
||||
"name": "LanPaint_Info",
|
||||
"type": "STRING",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "LanPaint_Info"
|
||||
}
|
||||
},
|
||||
{
|
||||
"localized_name": "Inpainting_mode",
|
||||
"name": "Inpainting_mode",
|
||||
"type": "COMBO",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "Inpainting_mode"
|
||||
}
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "LATENT",
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
18
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LanPaint_KSampler",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
777,
|
||||
"fixed",
|
||||
20,
|
||||
4.0,
|
||||
"euler",
|
||||
"simple",
|
||||
1.0,
|
||||
5,
|
||||
"Image First",
|
||||
"LanPaint KSampler. For more info, visit https://github.com/scraed/LanPaint. If you find it useful, please give a star ⭐!",
|
||||
"🖼️ Image Inpainting",
|
||||
"lanpaint_star_button"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LanPaint_ImageDecode",
|
||||
"pos": [
|
||||
180,
|
||||
40
|
||||
],
|
||||
"size": [
|
||||
300,
|
||||
120
|
||||
],
|
||||
"flags": {},
|
||||
"order": 12,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"localized_name": "samples",
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 18
|
||||
},
|
||||
{
|
||||
"localized_name": "vae",
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"localized_name": "image",
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 12
|
||||
},
|
||||
{
|
||||
"localized_name": "mask",
|
||||
"name": "mask",
|
||||
"type": "MASK",
|
||||
"link": 14
|
||||
},
|
||||
{
|
||||
"localized_name": "blend_overlap",
|
||||
"name": "blend_overlap",
|
||||
"type": "INT",
|
||||
"link": null,
|
||||
"widget": {
|
||||
"name": "blend_overlap"
|
||||
}
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"localized_name": "IMAGE",
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"slot_index": 0,
|
||||
"links": [
|
||||
19
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LanPaint_ImageDecode",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
9
|
||||
],
|
||||
"title": "Decode and merge (keeps RGBA)"
|
||||
},
|
||||
{
|
||||
"id": 13,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
180,
|
||||
220
|
||||
],
|
||||
"size": [
|
||||
400,
|
||||
420
|
||||
],
|
||||
"flags": {},
|
||||
"order": 13,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"localized_name": "images",
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 19
|
||||
}
|
||||
],
|
||||
"outputs": [],
|
||||
"properties": {
|
||||
"Node name for S&R": "SaveImage",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"Qwen2.1_Transparent_Edit"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 14,
|
||||
"type": "MarkdownNote",
|
||||
"pos": [
|
||||
-1180,
|
||||
1420
|
||||
],
|
||||
"size": [
|
||||
900,
|
||||
640
|
||||
],
|
||||
"flags": {},
|
||||
"order": 14,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [],
|
||||
"properties": {
|
||||
"Node name for S&R": "MarkdownNote",
|
||||
"cnr_id": "comfy-core"
|
||||
},
|
||||
"widgets_values": [
|
||||
"## Qwen-Image 2.1 + LanPaint, inpainting a picture with transparency\n\nThe source is RGBA and the edit is allowed to change the alpha, so the rebuilt sole gets\na new silhouette instead of just new pixels.\n\n`Join Image With Alpha` re-attaches the picture's alpha - `Load Image` hands it out on its\n`MASK` output rather than keeping it on the image. That mask is `1 - alpha` (1 =\ntransparent), which is what `Join Image With Alpha` wants, so it connects with no invert.\n\nThe mask is a separate greyscale file, white = repaint, and it is kept modest here (about\n16% of the frame). LanPaint anchors everything outside the mask to the original on every\nstep, so a mask covering most of the picture leaves the model without context."
|
||||
],
|
||||
"title": "How this works (and the mask size trap)",
|
||||
"color": "#432",
|
||||
"bgcolor": "#653"
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
1,
|
||||
1,
|
||||
0,
|
||||
4,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
2,
|
||||
4,
|
||||
0,
|
||||
11,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
3,
|
||||
2,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
2,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
5,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
2,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
6,
|
||||
3,
|
||||
0,
|
||||
10,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
12,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
8,
|
||||
5,
|
||||
0,
|
||||
6,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
5,
|
||||
1,
|
||||
6,
|
||||
1,
|
||||
"MASK"
|
||||
],
|
||||
[
|
||||
10,
|
||||
6,
|
||||
0,
|
||||
8,
|
||||
3,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
11,
|
||||
6,
|
||||
0,
|
||||
10,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
12,
|
||||
6,
|
||||
0,
|
||||
12,
|
||||
2,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
13,
|
||||
7,
|
||||
0,
|
||||
10,
|
||||
2,
|
||||
"MASK"
|
||||
],
|
||||
[
|
||||
14,
|
||||
7,
|
||||
0,
|
||||
12,
|
||||
3,
|
||||
"MASK"
|
||||
],
|
||||
[
|
||||
15,
|
||||
8,
|
||||
0,
|
||||
11,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
16,
|
||||
9,
|
||||
0,
|
||||
11,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
17,
|
||||
10,
|
||||
0,
|
||||
11,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
18,
|
||||
11,
|
||||
0,
|
||||
12,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
19,
|
||||
12,
|
||||
0,
|
||||
13,
|
||||
0,
|
||||
"IMAGE"
|
||||
]
|
||||
],
|
||||
"groups": [
|
||||
{
|
||||
"id": 1,
|
||||
"title": "Step 1 - Load models",
|
||||
"bounding": [
|
||||
-1190,
|
||||
0,
|
||||
420,
|
||||
400
|
||||
],
|
||||
"color": "#3f789e",
|
||||
"font_size": 24,
|
||||
"flags": {}
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"title": "Step 2 - Picture + its alpha + the mask",
|
||||
"bounding": [
|
||||
-1190,
|
||||
380,
|
||||
880,
|
||||
1020
|
||||
],
|
||||
"color": "#3f789e",
|
||||
"font_size": 24,
|
||||
"flags": {}
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"title": "Step 3 - Texts",
|
||||
"bounding": [
|
||||
-750,
|
||||
490,
|
||||
460,
|
||||
480
|
||||
],
|
||||
"color": "#3f789e",
|
||||
"font_size": 24,
|
||||
"flags": {}
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"title": "Step 4 - LanPaint: encode, sample, decode",
|
||||
"bounding": [
|
||||
-250,
|
||||
0,
|
||||
840,
|
||||
520
|
||||
],
|
||||
"color": "#3f789e",
|
||||
"font_size": 24,
|
||||
"flags": {}
|
||||
}
|
||||
],
|
||||
"config": {},
|
||||
"extra": {
|
||||
"ds": {
|
||||
"scale": 0.55,
|
||||
"offset": [
|
||||
1210,
|
||||
180
|
||||
]
|
||||
},
|
||||
"workflowRendererVersion": "LG"
|
||||
},
|
||||
"version": 0.4
|
||||
}
|
||||
|
Before Width: | Height: | Size: 4.3 MiB After Width: | Height: | Size: 4.3 MiB |
|
After Width: | Height: | Size: 793 KiB |
|
After Width: | Height: | Size: 1.2 MiB |
|
After Width: | Height: | Size: 1.0 MiB |
|
After Width: | Height: | Size: 2.6 KiB |
|
After Width: | Height: | Size: 1.6 MiB |
|
After Width: | Height: | Size: 965 KiB |
|
After Width: | Height: | Size: 1018 KiB |
|
After Width: | Height: | Size: 5.0 KiB |
|
After Width: | Height: | Size: 2.8 MiB |
@@ -1,6 +1,6 @@
|
||||
import json
|
||||
import os
|
||||
from contextlib import contextmanager
|
||||
from contextlib import contextmanager, nullcontext
|
||||
import math
|
||||
# import nodes.py
|
||||
import comfy
|
||||
@@ -56,6 +56,39 @@ def _version_tuple(value):
|
||||
|
||||
COMFYUI_VERSION_060_OR_NEWER = _version_tuple(comfyui_version.__version__) >= (0, 6, 0)
|
||||
|
||||
# ComfyUI >= 0.34 runs MiniMax H3 with per-token denoise-mask row timesteps
|
||||
# (commit ff6c8a8a). LanPaint keeps its 0.33 single-schedule contract, so this
|
||||
# gates hiding the mask from the model during the paint loop.
|
||||
COMFYUI_H3_DENOISE_MASK_CONTRACT = _version_tuple(getattr(comfyui_version, "__version__", "0.0.0")) >= (0, 34, 0)
|
||||
|
||||
@contextmanager
|
||||
def _hide_h3_denoise_mask(model):
|
||||
"""ComfyUI >= 0.34 hands MiniMax H3 a per-token denoise mask, which
|
||||
switches the DiT to mask-driven row timesteps (commit ff6c8a8a). LanPaint
|
||||
implements the 0.33 single-schedule contract in its own replace step and
|
||||
inner dynamics, so during the paint loop the mask is hidden from the
|
||||
model's extra_conds and the DiT keeps the uniform row timesteps that
|
||||
LanPaint's math expects."""
|
||||
had_instance_attr = "extra_conds" in getattr(model, "__dict__", {})
|
||||
original_extra_conds = model.extra_conds
|
||||
|
||||
def _extra_conds_without_mask(*args, **kwargs):
|
||||
kwargs.pop("denoise_mask", None)
|
||||
return original_extra_conds(*args, **kwargs)
|
||||
|
||||
model.extra_conds = _extra_conds_without_mask
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
if had_instance_attr:
|
||||
model.extra_conds = original_extra_conds
|
||||
else:
|
||||
try:
|
||||
del model.extra_conds
|
||||
except AttributeError:
|
||||
model.extra_conds = original_extra_conds
|
||||
|
||||
|
||||
def reshape_mask(input_mask, output_shape,video_inpainting=False):
|
||||
dims = len(output_shape) - 2
|
||||
scale_mode = "nearest-exact"
|
||||
@@ -190,6 +223,13 @@ class CFGGuider_LanPaint:
|
||||
self.minimax_h3_audio = _detect_minimax_h3_audio(
|
||||
self.model_patcher, self.model_options, kwargs.get("latent_shapes", None))
|
||||
|
||||
# ComfyUI >= 0.34 switches MiniMax H3 to per-token row timesteps driven
|
||||
# by the denoise mask; LanPaint's replace step and inner dynamics follow
|
||||
# the 0.33 single-schedule contract, so hide the mask from the model for
|
||||
# the paint loop and keep the row timesteps uniform.
|
||||
h3_hide_mask = self.minimax_h3_audio is not None and COMFYUI_H3_DENOISE_MASK_CONTRACT
|
||||
hide_mask_ctx = _hide_h3_denoise_mask(self.inner_model) if h3_hide_mask else nullcontext()
|
||||
|
||||
if denoise_mask is not None:
|
||||
video_inpainting = self.model_options.get("video_inpainting", False)
|
||||
if tuple(denoise_mask.shape) != tuple(noise.shape):
|
||||
@@ -204,7 +244,8 @@ class CFGGuider_LanPaint:
|
||||
|
||||
try:
|
||||
self.model_patcher.pre_run()
|
||||
output = self.inner_sample(noise, latent_image, device, sampler, sigmas, denoise_mask, callback, disable_pbar, seed, **kwargs)
|
||||
with hide_mask_ctx:
|
||||
output = self.inner_sample(noise, latent_image, device, sampler, sigmas, denoise_mask, callback, disable_pbar, seed, **kwargs)
|
||||
finally:
|
||||
self.model_patcher.cleanup()
|
||||
|
||||
@@ -1329,6 +1370,18 @@ class LanPaint_ImageDecode:
|
||||
)
|
||||
if image is None:
|
||||
return (img,)
|
||||
# Some VAEs decode to more channels than the source image: the Qwen Image
|
||||
# 2.1 VAE always emits a 4th channel for RGB input. ComfyUI's IMAGE type is
|
||||
# RGB, so trim to the original's channel count - which is what ComfyUI's own
|
||||
# RGBA -> RGB conversion does, compositing over white - and the merge below
|
||||
# can broadcast. Without this it raises
|
||||
# "The size of tensor a (3) must match the size of tensor b (4)".
|
||||
img_channels, orig_channels = img.shape[-1], image.shape[-1]
|
||||
if img_channels > orig_channels:
|
||||
img = img[..., :orig_channels]
|
||||
elif img_channels < orig_channels:
|
||||
pad = img.new_ones(img.shape[:-1] + (orig_channels - img_channels,))
|
||||
img = torch.cat((img, pad), dim=-1)
|
||||
target_h, target_w = image.shape[1], image.shape[2]
|
||||
if tuple(img.shape[1:3]) != (target_h, target_w):
|
||||
img = torch.nn.functional.interpolate(
|
||||
@@ -1339,6 +1392,10 @@ class LanPaint_ImageDecode:
|
||||
).movedim(1, -1)
|
||||
if mask is None:
|
||||
return (img,)
|
||||
# Keep the merge channel-agnostic: an RGBA image (a source that carries its
|
||||
# own alpha, e.g. a Qwen Image 2.1 transparent-background render) keeps its
|
||||
# 4th channel all the way through, so transparency can be inpainted alongside
|
||||
# the content. An RGB image still comes back as RGB.
|
||||
return (merge_video_with_mask(image, img, mask, blend_overlap),)
|
||||
|
||||
|
||||
|
||||
@@ -322,3 +322,26 @@ def test_prepare_step_size_handles_per_row_parameters() -> None:
|
||||
adt = (A_x * dtx).flatten()
|
||||
assert adt[0] == pytest.approx(0.2) # 1/(1-0.5) * 0.1
|
||||
assert adt[-1] == pytest.approx(0.2) # 1/(1-0.9) * 0.02 -- bounded invariant
|
||||
|
||||
|
||||
def test_hide_h3_denoise_mask_strips_and_restores(monkeypatch) -> None:
|
||||
# ComfyUI >= 0.34: the per-token denoise mask switches the MiniMax H3 DiT
|
||||
# to mask-driven row timesteps. LanPaint hides it from extra_conds during
|
||||
# the paint loop so the DiT keeps the uniform 0.33 row timesteps.
|
||||
nodes = _import_nodes(monkeypatch)
|
||||
captured = {}
|
||||
|
||||
class FakeH3Model:
|
||||
def extra_conds(self, **kwargs): # type: ignore[no-untyped-def]
|
||||
captured.update(kwargs)
|
||||
return {}
|
||||
|
||||
model = FakeH3Model()
|
||||
with nodes._hide_h3_denoise_mask(model):
|
||||
model.extra_conds(denoise_mask="MASK", latent_shapes=[(1, 2), (1, 2)])
|
||||
assert "denoise_mask" not in captured # hidden from the model
|
||||
assert captured["latent_shapes"] is not None # other conds still pass
|
||||
|
||||
captured.clear()
|
||||
model.extra_conds(denoise_mask="MASK", latent_shapes=[(1, 2), (1, 2)])
|
||||
assert captured["denoise_mask"] == "MASK" # restored after the loop
|
||||
|
||||