Compare commits

..
14 Commits
Author SHA1 Message Date
scraed 2d7912f9a5 Describe the Qwen 2.1 examples for readers, not for nodes
The two example sections were written from the inside of the graph: node class
names, the `<image1>` / `<image2>` notation, which loader to point where, and an
aside about which of the two 2.1 encoders keeps alpha. None of that helps
someone deciding whether to try the example.

Both sections now say what a user does and what comes back, in two sentences
each, and the news lines and the Updates entry follow. Example 31 also gains the
Workflow JSON link it was missing.

The MarkdownNote embedded in the Example 32 workflow gets the same treatment:
it keeps `<image1>` / `<image2>` (the instruction needs those) and drops the
`Join Image With Alpha` / `Load Image` MASK-output explanation.

The result image is byte-identical to the committed one; only the embedded
workflow metadata changed, which is why nothing was re-run.
2026-09-28 15:16:49 +08:00
scraed 93b90fad99 Enlarge the Example 32 mask so the new shape can grow
The first version swapped a material onto the earcups without changing their
outline, so nothing in the result showed that the alpha channel is being
inpainted too - the silhouette moved by 47 pixels.

The mask now covers both earcups plus a ring of the transparent background
around them (25.1% of the frame, 70.9% of it on the headphones), and the
instruction asks for the earcups to be rebuilt as oversized turbine cups that
flare out past their old outline. The rebuilt region grows 11735 pixels of new
silhouette over what used to be empty background, 100% of it inside the mask,
while the headband, stitching, yokes and hinges stay pixel-identical (0.27/255
mean channel difference outside the mask).

A 25% mask is still clean here; the earlier 16% guidance came from a picture
whose mask had far less context left around it.

Also verified that the alpha is a real channel rather than a trimmed-off fourth
one: re-running with the identical RGB and a fully opaque alpha changes the
masked region by 61.9/255 on average (max 254).
2026-09-28 15:06:42 +08:00
scraed d82b218e1c Add a Qwen-Image 2.1 image edit example with a LanPaint mask
This is the official `Qwen Image 2.1 Image Edit` graph - a `TextEncodeQwenImage21`
fed through its autogrow `images` input, with the prompt naming the references as
`<image1>` / `<image2>` - carrying a LanPaint mask.

Three changes against the stock template: `LanPaint_ImageEncode` takes the mask,
`LanPaint_KSampler` does the sampling, and `LanPaint_ImageDecode` merges the result
back inside the mask and keeps RGBA. No new node inputs or outputs.

The demo re-surfaces a pair of headphones' earcups with the iridescent titanium of a
second reference image. The mask is 14.8% of the frame and 89% of it sits on the
headphones; outside the mask the output is pixel-identical to the source (mean
channel difference 0.098/255), so the headband, the stitching and the hinges are
untouched rather than merely similar. The source and the result are both RGBA.

`Text Encode Qwen Image 2.1` is the right reference encoder for transparent pictures:
it hands the vision tower the alpha composited over white and encodes all four
channels into the reference latent, where the older `Text Encode Qwen Image Edit Plus`
trims to RGB.

Verified by replaying the API prompt embedded in the packaged PNG against the shipped
input files: bit-identical output. The README gets the news block, a TOC entry, the
example section and an Updates line; the previous 2.1 heading now says "with
Transparency" so the two 2.1 examples are distinguishable.
2026-09-28 14:38:33 +08:00
scraed a9238253e0 Announce Qwen-Image 2.1 support in the README news block 2026-09-28 00:34:37 +08:00
scraed 391ca9625f Trim the Qwen 2.1 section to a paragraph
The section explained the VAE's channel count, the autogrow mismatch and the mask-size
measurements before saying where the example was. Keep what a reader needs: the model is
supported, transparency is inpainted too, and here is how to point it at your own picture.
The workflow's own note gets the same treatment.
2026-09-28 00:31:59 +08:00
scraed e7e8166e02 List Qwen-Image 2.1 in the features and announce it, with a before/masked/after sheet
The Features model list and the Updates log both predate the Qwen 2.1 example, so the
support was only discoverable from the table of contents. Example_31 also gets a
Comparison.png showing the original, the covered region and the result side by side,
since the new sole growing past the old outline is the part worth seeing.
2026-09-28 00:25:23 +08:00
scraed f564ee886b Add the Qwen-Image 2.1 example with transparent-background inpainting
LanPaint runs on Qwen-Image 2.1 unchanged: it is a rectified-flow model, so the
existing Flux/Qwen-Image conversions apply. Example_31 uses a text-to-image render
of its own model as the before/after, generated with a transparent background, and
keeps the inpainting mask in a separate greyscale file rather than in the picture's
alpha - 2.1's alpha channel means transparency, so overloading it would make the
two indistinguishable.

The transparency is carried into the latent and edited along with the pixels: the
2.1 VAE is 4-in/4-out (encoder.conv1 takes 4 channels, the decoder head emits 4),
and LoadImage's MASK output re-attaches through Join Image With Alpha, since that
mask is already 1 - alpha. The rebuilt sole grows past the old outline, so the
alpha in that region is generated rather than copied.

LanPaint_ImageDecode now matches the decoded channels to the source image: the 2.1
VAE always emits a 4th channel, which the merge could not broadcast. It keeps the
input's channel count, so an RGBA source comes back RGBA and an RGB source still
comes back RGB.
2026-09-28 00:25:23 +08:00
scraed 0168172781 Adapt MiniMax H3 to the ComfyUI 0.34 per-token denoise-mask contract
ComfyUI 0.34 (commit ff6c8a8a) hands MiniMax H3 a per-token denoise mask: the DiT derives per-row timesteps from it, presenting preserved rows at the video cond timestep (0.999) and the audio stream rescaled. LanPaint implements the 0.33 single-schedule contract in its own replace step and inner dynamics, so on 0.34 the mask-driven row timesteps contradicted the injected latents and both streams came out wrong.

Detect the contract from the ComfyUI version (>= 0.34) and hide the mask from the model's extra_conds for the paint loop, so the DiT keeps the uniform row timesteps LanPaint's math expects. Image models and ComfyUI <= 0.33 are unaffected.
2026-09-28 00:25:23 +08:00
Yuan Lan be34cf3054 Add project website link to README
Added project website link to the README.
2026-09-25 18:41:32 +08:00
Yuan Lan 649d6ea465 Revise README with new research and benchmark information
Updated README to reflect new structure and added research and benchmark section with citation details.
2026-09-25 09:29:12 +08:00
Yuan Lan c054966ae8 Update README with LanPaint capabilities
Added features enabled by LanPaint to README.
2026-09-25 09:18:27 +08:00
Yuan Lan e513cf6778 Update README with LanPaint features
Added description of LanPaint as a training-free sampler for local editing.
2026-09-25 09:14:35 +08:00
scraed 32cf848e93 Drop the Beta label from the video examples section and untrack the combined showcase
Only the InPainted mp4/gif pair is tracked for Example_29; the combined
H3_Input_Mask_Result media stays on disk but out of git.
2026-08-12 17:36:12 +08:00
scraed 12c8035ffa Update the README news and index for v2.1.0 and refresh the Example_29 GIFs 2026-08-12 17:35:28 +08:00
18 changed files with 2119 additions and 14 deletions
+56 -12
View File
@@ -13,17 +13,29 @@
Universally applicable inpainting ability for every model. LanPaint sampler lets the model "think" through multiple iterations before denoising, enabling you to invest more computation time for superior inpainting quality.
This is the official implementation of ["LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling"](https://arxiv.org/abs/2502.03491), accepted by TMLR.
## What LanPaint Enables
The repository is for ComfyUI extension.
- Training-free image inpainting
- Outpainting and generative fill
- Mask-constrained local image editing
- Object / region replacement with text guidance
- Character-consistent local generation
- Video inpainting and local video editing
- Video + audio masked generation
Diffusers Support: [LanPaint-Diffusers](https://github.com/charrywhite/LanPaint-diffusers) by [@charrywhite](https://github.com/charrywhite/)
## Research & Benchmark
Benchmark code for paper reproduce: [LanPaintBench](https://github.com/scraed/LanPaintBench).
LanPaint is a training-free partial conditional sampler that enables mask-constrained inpainting and local editing with pretrained diffusion and rectified-flow models, without fine-tuning or backpropagation.
## Citation
* 📄 **Paper:** [LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling](https://openreview.net/forum?id=JPC8JyOUSW) — TMLR 2025
* 🧩 **ComfyUI Implementation:** This repository
* 🐍 **Diffusers Implementation:** [LanPaint-Diffusers](https://github.com/charrywhite/LanPaint-diffusers) by [@charrywhite](https://github.com/charrywhite/)
* 🧪 **Benchmark & Reproduction:** [LanPaintBench](https://github.com/scraed/LanPaintBench)
* 🌐 **Project Website:** [LanPaint Page](https://scraed.github.io/scraedBlog/lanpaint/)
```
### Citation
```bibtex id="2r5ioa"
@article{
zheng2025lanpaint,
title={LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling},
@@ -31,17 +43,27 @@ author={Candi Zheng and Yuan Lan and Yang Wang},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2025},
url={https://openreview.net/forum?id=JPC8JyOUSW},
note={}
url={https://openreview.net/forum?id=JPC8JyOUSW}
}
```
**🎉 NEW 2026: Join our discord!**
[Join our Discord](https://discord.gg/yN5wYDE6W4) to share experiences, discuss features, and explore future development.
`v1.5.0` fixes an important hidden bug that reduced performance and could blur images (especially with `z-image-base`) and also boosts overall LanPaint performance across other models.
`v2.1.0` significantly accelerates LanPaint with a new schedule mechanism and fixes MiniMax H3 support on the latest ComfyUI.
If your inpainting results have wierd (glowing / broken) mask boundary, check this [issue](https://github.com/scraed/LanPaint/issues/80).
**🎨 NEW: LanPaint now supports Qwen-Image 2.1 - transparency, and masked image editing!**
![Qwen 2.1 image edit: the canvas, the mask, the second reference and the result](https://github.com/scraed/LanPaint/blob/master/examples/Example_32/Comparison.png)
Qwen 2.1's **image edit** model now works under a LanPaint mask: tell it what to change, paint over the part you want it to touch, and only that part changes. Hand it a second picture to borrow from if you want one. Check our latest [Qwen Image 2.1 Image Edit Example](#example-qwen-image-21-image-edit-masked-inpaintlanpaint-k-sampler-5-steps-of-thinking).
![Qwen 2.1 before / masked / after](https://github.com/scraed/LanPaint/blob/master/examples/Example_31/Comparison.png)
And if your picture carries transparency, it gets inpainted too - the rebuilt part comes back with a new outline, not just new colours. Check our latest [Qwen Image 2.1 Example](#example-qwen-image-21-inpaint-with-transparencylanpaint-k-sampler-5-steps-of-thinking).
**🎬 NEW: LanPaint now supports MiniMax H3 video + audio inpainting!**
| Masked Input (paint in the editor) | Mask (visible overlay) | Inpainted Result |
@@ -103,10 +125,11 @@ Check our latest [Krea2 Example](#example-krea2-inpaintlanpaint-k-sampler-3-step
- [Features](#features)
- [Quickstart](#quickstart)
- [How to Use Examples](#how-to-use-examples)
- [Video Examples (Beta)](#video-examples-beta)
- [Video Examples](#video-examples)
- [Wan 2.2 Video Inpainting](#wan-22-video-inpainting)
- [Wan 2.2 5B Video Inpainting](#wan-22-5b-video-inpainting)
- [Wan 2.2 Video Outpainting](#wan-22-video-outpainting)
- [MiniMax H3 Video + Audio Inpainting](#minimax-h3-video--audio-inpainting-av-pipeline)
- [Resource Consumption](#resource-consumption)
- [Image Examples](#image-examples)
- [Flux.2.Dev](#example-flux2dev-inpaintlanpaint-k-sampler-5-steps-of-thinking)
@@ -121,6 +144,8 @@ Check our latest [Krea2 Example](#example-krea2-inpaintlanpaint-k-sampler-3-step
- [Wan 2.2 T2I with reference](#example-wan22-partial-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [Qwen Image Edit 2511 2509](#example-qwen-edit-2509-inpaint)
- [Qwen Image Edit 2508](#example-qwen-edit-2508-inpaint)
- [Qwen Image 2.1 Image Edit](#example-qwen-image-21-image-edit-masked-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [Qwen Image 2.1](#example-qwen-image-21-inpaint-with-transparencylanpaint-k-sampler-5-steps-of-thinking)
- [Qwen Image](#example-qwen-image-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [HiDream](#example-hidream-inpaint-lanpaint-k-sampler-5-steps-of-thinking)
- [SD 3.5](#example-sd-35-inpaintlanpaint-k-sampler-5-steps-of-thinking)
@@ -138,7 +163,7 @@ Check our latest [Krea2 Example](#example-krea2-inpaintlanpaint-k-sampler-3-step
## Features
- **Universal Compatibility** – Works instantly with almost any model (**Ideogram4, Krea2, Z-image, Z-image-base, Hunyuan, Wan 2.2, Qwen Image/Edit, Anima, HiDream, SD 3.5, Flux-series, SDXL, SD 1.5 or custom LoRAs**) and ControlNet.
- **Universal Compatibility** – Works instantly with almost any model (**Ideogram4, Krea2, Z-image, Z-image-base, Hunyuan, Wan 2.2, Qwen Image 2.1/Image/Edit, Anima, HiDream, SD 3.5, Flux-series, SDXL, SD 1.5 or custom LoRAs**) and ControlNet.
![Inpainting Result 13](https://github.com/scraed/LanPaint/blob/master/examples/InpaintChara_13.jpg)
- **No Training Needed** – Works out of the box with your existing model.
- **Easy to Use** – Same workflow as standard ComfyUI KSampler.
@@ -177,7 +202,7 @@ Once installed, you'll find the LanPaint nodes under the "sampling" category in
- **[VAE Encode for Inpainting](https://comfyanonymous.github.io/ComfyUI_examples/inpaint/)**
- **[Set Latent Noise Mask](https://comfyui-wiki.com/en/tutorial/basic/how-to-inpaint-an-image-in-comfyui)**
## Video Examples (Beta)
## Video Examples
LanPaint now supports video inpainting with Wan 2.2, enabling you to seamlessly inpaint masked regions across video frames while maintaining temporal consistency.
@@ -435,6 +460,20 @@ Check [Mased Qwen Edit Workflow](https://github.com/scraed/LanPaint/tree/master/
### Example Qwen Image 2.1 Image Edit: Masked InPaint(LanPaint K Sampler, 5 steps of thinking)
Qwen-Image 2.1's image edit model now works under a LanPaint mask: write what you want changed, paint over the part it should touch, and only that part changes - everything else, transparency included, comes back exactly as it was. In this example a second picture supplies the material for the earcups, and the headband and stitching stay as they are. Workflow and images are in `examples/Example_32`; drag `InPainted_Drag_Me_to_ComfyUI.png` into ComfyUI to load it. Use your own pictures with the official [Qwen Image 2.1 Image Edit template](https://docs.comfy.org/tutorials/image/qwen/qwen-image-2-1).
![Qwen 2.1 image edit: canvas, mask, material, result](https://github.com/scraed/LanPaint/blob/master/examples/Example_32/Comparison.png)
[View Workflow & Masks](https://github.com/scraed/LanPaint/tree/master/examples/Example_32) · [Workflow JSON](https://github.com/scraed/LanPaint/blob/master/example_workflows/Qwen_Image_2.1_Edit_Masked_Inpaint.json)
### Example Qwen Image 2.1: InPaint with Transparency(LanPaint K Sampler, 5 steps of thinking)
Qwen-Image 2.1 inpaints a picture's transparency along with its pixels, so the rebuilt part can come back with a new outline instead of merely new colours - here the boot's sole is replaced and the silhouette grows with it. Workflow and images are in `examples/Example_31`; drag `InPainted_Drag_Me_to_ComfyUI.png` into ComfyUI to load it.
![Qwen 2.1: original, mask, result](https://github.com/scraed/LanPaint/blob/master/examples/Example_31/Comparison.png)
[View Workflow & Masks](https://github.com/scraed/LanPaint/tree/master/examples/Example_31) · [Workflow JSON](https://github.com/scraed/LanPaint/blob/master/example_workflows/Transparent_Edit_EncodeDecode_Inpaint.json)
### Example Qwen Image: InPaint(LanPaint K Sampler, 5 steps of thinking)
![Inpainting Result 14](https://github.com/scraed/LanPaint/blob/master/examples/InpaintChara_14.jpg)
@@ -619,6 +658,11 @@ Submit a PR to add your tutorial/video here, or open an [Issue](https://github.c
[Working togather with crop&stitch](https://github.com/scraed/LanPaint/issues/46)
## Updates
- 2026/09/28
- Add Qwen-Image 2.1 image edit support: masked, instruction-driven editing (Example_32).
- Add Qwen-Image 2.1 inpainting support with LanPaint KSampler (Example_31).
- Inpainting a picture that carries transparency now works end to end: the 2.1 VAE is 4-in/4-out, so the alpha travels through the latent and is edited alongside the pixels. Keep the inpainting mask in its own greyscale file, since 2.1's alpha channel means image transparency.
- `LanPaint_ImageDecode` now matches the decoded channel count to the source image, so an RGBA source comes back RGBA and an RGB source still comes back RGB.
- 2026/08/12
- `v2.1.0`: Significantly accelerated LanPaint using a new schedule mechanism.
- Fix bugs for MiniMax H3 on the latest ComfyUI.
Binary file not shown.

After

Width:  |  Height:  |  Size: 199 KiB

File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 219 KiB

@@ -0,0 +1,976 @@
{
"id": "c4e8a1f7-2b93-4d6e-8a50-1f7e9c3b6d28",
"revision": 0,
"last_node_id": 14,
"last_link_id": 19,
"nodes": [
{
"id": 1,
"type": "UNETLoader",
"pos": [
-1180,
40
],
"size": [
390,
82
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"localized_name": "MODEL",
"name": "MODEL",
"type": "MODEL",
"slot_index": 0,
"links": [
1
]
}
],
"properties": {
"Node name for S&R": "UNETLoader",
"cnr_id": "comfy-core"
},
"widgets_values": [
"qwen_image_2.1_int8_convrot.safetensors",
"default"
]
},
{
"id": 2,
"type": "CLIPLoader",
"pos": [
-1180,
170
],
"size": [
390,
106
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"localized_name": "CLIP",
"name": "CLIP",
"type": "CLIP",
"slot_index": 0,
"links": [
3,
4
]
}
],
"properties": {
"Node name for S&R": "CLIPLoader",
"cnr_id": "comfy-core"
},
"widgets_values": [
"qwen3vl_8b_int8_convrot.safetensors",
"qwen_image",
"default"
]
},
{
"id": 3,
"type": "VAELoader",
"pos": [
-1180,
320
],
"size": [
390,
58
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"localized_name": "VAE",
"name": "VAE",
"type": "VAE",
"slot_index": 0,
"links": [
5,
6,
7
]
}
],
"properties": {
"Node name for S&R": "VAELoader",
"cnr_id": "comfy-core"
},
"widgets_values": [
"qwen_image_2.1_vae_bf16.safetensors"
]
},
{
"id": 4,
"type": "QwenImage21Cache",
"pos": [
-740,
40
],
"size": [
310,
82
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"localized_name": "model",
"name": "model",
"type": "MODEL",
"link": 1
}
],
"outputs": [
{
"localized_name": "MODEL",
"name": "MODEL",
"type": "MODEL",
"slot_index": 0,
"links": [
2
]
}
],
"properties": {
"Node name for S&R": "QwenImage21Cache",
"cnr_id": "comfy-core"
},
"widgets_values": [
"auto",
"default"
]
},
{
"id": 5,
"type": "LoadImage",
"pos": [
-1180,
420
],
"size": [
390,
440
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"localized_name": "IMAGE",
"name": "IMAGE",
"type": "IMAGE",
"slot_index": 0,
"links": [
8
]
},
{
"localized_name": "MASK",
"name": "MASK",
"type": "MASK",
"slot_index": 1,
"links": [
9
]
}
],
"properties": {
"Node name for S&R": "LoadImage",
"cnr_id": "comfy-core"
},
"widgets_values": [
"Image_Load_Me_in_Loader.png",
"image"
],
"title": "The picture (RGBA: the alpha is the background)"
},
{
"id": 6,
"type": "JoinImageWithAlpha",
"pos": [
-740,
420
],
"size": [
330,
60
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"localized_name": "image",
"name": "image",
"type": "IMAGE",
"link": 8
},
{
"localized_name": "alpha",
"name": "alpha",
"type": "MASK",
"link": 9
}
],
"outputs": [
{
"localized_name": "IMAGE",
"name": "IMAGE",
"type": "IMAGE",
"slot_index": 0,
"links": [
10,
11,
12
]
}
],
"properties": {
"Node name for S&R": "JoinImageWithAlpha",
"cnr_id": "comfy-core"
},
"widgets_values": [],
"title": "Re-attach the alpha as the 4th channel"
},
{
"id": 7,
"type": "LoadImageMask",
"pos": [
-1180,
900
],
"size": [
390,
480
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"localized_name": "MASK",
"name": "MASK",
"type": "MASK",
"slot_index": 0,
"links": [
13,
14
]
}
],
"properties": {
"Node name for S&R": "LoadImageMask",
"cnr_id": "comfy-core"
},
"widgets_values": [
"Mask_Load_Me_in_Loader.png",
"red"
],
"title": "The mask, on its own (white = rebuild)"
},
{
"id": 8,
"type": "TextEncodeQwenImageEditPlus",
"pos": [
-740,
520
],
"size": [
440,
250
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"localized_name": "clip",
"name": "clip",
"type": "CLIP",
"link": 3
},
{
"localized_name": "prompt",
"name": "prompt",
"type": "STRING",
"link": null,
"widget": {
"name": "prompt"
}
},
{
"localized_name": "vae",
"name": "vae",
"type": "VAE",
"link": 5
},
{
"localized_name": "image1",
"name": "image1",
"type": "IMAGE",
"link": 10
},
{
"localized_name": "image2",
"name": "image2",
"type": "IMAGE",
"link": null
},
{
"localized_name": "image3",
"name": "image3",
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"localized_name": "CONDITIONING",
"name": "CONDITIONING",
"type": "CONDITIONING",
"slot_index": 0,
"links": [
15
]
}
],
"properties": {
"Node name for S&R": "TextEncodeQwenImageEditPlus",
"cnr_id": "comfy-core"
},
"widgets_values": [
"the boot's sole and forefoot replaced by a sleek futuristic mechanical unit: segmented brushed titanium plating with fine panel lines, glowing cyan energy strips along the flanks, a sculpted dark carbon-fibre heel block, precise engineered details and tiny bolts, the brown leather upper above stays exactly as it is, high-end sci-fi product photography, transparent background"
],
"title": "Text Encode Qwen Image Edit Plus (reference + prompt)",
"color": "#232",
"bgcolor": "#353"
},
{
"id": 9,
"type": "CLIPTextEncode",
"pos": [
-740,
800
],
"size": [
440,
150
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"localized_name": "clip",
"name": "clip",
"type": "CLIP",
"link": 4
},
{
"localized_name": "text",
"name": "text",
"type": "STRING",
"link": null,
"widget": {
"name": "text"
}
}
],
"outputs": [
{
"localized_name": "CONDITIONING",
"name": "CONDITIONING",
"type": "CONDITIONING",
"slot_index": 0,
"links": [
16
]
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode",
"cnr_id": "comfy-core"
},
"widgets_values": [
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感,模糊,变形"
],
"title": "CLIP Text Encode (Negative Prompt)",
"color": "#223",
"bgcolor": "#335"
},
{
"id": 10,
"type": "LanPaint_ImageEncode",
"pos": [
-240,
40
],
"size": [
300,
110
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"localized_name": "image",
"name": "image",
"type": "IMAGE",
"link": 11
},
{
"localized_name": "vae",
"name": "vae",
"type": "VAE",
"link": 6
},
{
"localized_name": "mask",
"name": "mask",
"type": "MASK",
"link": 13
}
],
"outputs": [
{
"localized_name": "LATENT",
"name": "LATENT",
"type": "LATENT",
"slot_index": 0,
"links": [
17
]
}
],
"properties": {
"Node name for S&R": "LanPaint_ImageEncode",
"cnr_id": "comfy-core"
},
"widgets_values": []
},
{
"id": 11,
"type": "LanPaint_KSampler",
"pos": [
-240,
190
],
"size": [
330,
320
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"localized_name": "model",
"name": "model",
"type": "MODEL",
"link": 2
},
{
"localized_name": "positive",
"name": "positive",
"type": "CONDITIONING",
"link": 15
},
{
"localized_name": "negative",
"name": "negative",
"type": "CONDITIONING",
"link": 16
},
{
"localized_name": "latent_image",
"name": "latent_image",
"type": "LATENT",
"link": 17
},
{
"localized_name": "seed",
"name": "seed",
"type": "INT",
"link": null,
"widget": {
"name": "seed"
}
},
{
"localized_name": "steps",
"name": "steps",
"type": "INT",
"link": null,
"widget": {
"name": "steps"
}
},
{
"localized_name": "cfg",
"name": "cfg",
"type": "FLOAT",
"link": null,
"widget": {
"name": "cfg"
}
},
{
"localized_name": "sampler_name",
"name": "sampler_name",
"type": "COMBO",
"link": null,
"widget": {
"name": "sampler_name"
}
},
{
"localized_name": "scheduler",
"name": "scheduler",
"type": "COMBO",
"link": null,
"widget": {
"name": "scheduler"
}
},
{
"localized_name": "denoise",
"name": "denoise",
"type": "FLOAT",
"link": null,
"widget": {
"name": "denoise"
}
},
{
"localized_name": "LanPaint_NumSteps",
"name": "LanPaint_NumSteps",
"type": "INT",
"link": null,
"widget": {
"name": "LanPaint_NumSteps"
}
},
{
"localized_name": "LanPaint_PromptMode",
"name": "LanPaint_PromptMode",
"type": "COMBO",
"link": null,
"widget": {
"name": "LanPaint_PromptMode"
}
},
{
"localized_name": "LanPaint_Info",
"name": "LanPaint_Info",
"type": "STRING",
"link": null,
"widget": {
"name": "LanPaint_Info"
}
},
{
"localized_name": "Inpainting_mode",
"name": "Inpainting_mode",
"type": "COMBO",
"link": null,
"widget": {
"name": "Inpainting_mode"
}
}
],
"outputs": [
{
"localized_name": "LATENT",
"name": "LATENT",
"type": "LATENT",
"slot_index": 0,
"links": [
18
]
}
],
"properties": {
"Node name for S&R": "LanPaint_KSampler",
"cnr_id": "comfy-core"
},
"widgets_values": [
777,
"fixed",
20,
4.0,
"euler",
"simple",
1.0,
5,
"Image First",
"LanPaint KSampler. For more info, visit https://github.com/scraed/LanPaint. If you find it useful, please give a star ⭐!",
"🖼️ Image Inpainting",
"lanpaint_star_button"
]
},
{
"id": 12,
"type": "LanPaint_ImageDecode",
"pos": [
180,
40
],
"size": [
300,
120
],
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"localized_name": "samples",
"name": "samples",
"type": "LATENT",
"link": 18
},
{
"localized_name": "vae",
"name": "vae",
"type": "VAE",
"link": 7
},
{
"localized_name": "image",
"name": "image",
"type": "IMAGE",
"link": 12
},
{
"localized_name": "mask",
"name": "mask",
"type": "MASK",
"link": 14
},
{
"localized_name": "blend_overlap",
"name": "blend_overlap",
"type": "INT",
"link": null,
"widget": {
"name": "blend_overlap"
}
}
],
"outputs": [
{
"localized_name": "IMAGE",
"name": "IMAGE",
"type": "IMAGE",
"slot_index": 0,
"links": [
19
]
}
],
"properties": {
"Node name for S&R": "LanPaint_ImageDecode",
"cnr_id": "comfy-core"
},
"widgets_values": [
9
],
"title": "Decode and merge (keeps RGBA)"
},
{
"id": 13,
"type": "SaveImage",
"pos": [
180,
220
],
"size": [
400,
420
],
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"localized_name": "images",
"name": "images",
"type": "IMAGE",
"link": 19
}
],
"outputs": [],
"properties": {
"Node name for S&R": "SaveImage",
"cnr_id": "comfy-core"
},
"widgets_values": [
"Qwen2.1_Transparent_Edit"
]
},
{
"id": 14,
"type": "MarkdownNote",
"pos": [
-1180,
1420
],
"size": [
900,
640
],
"flags": {},
"order": 14,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"Node name for S&R": "MarkdownNote",
"cnr_id": "comfy-core"
},
"widgets_values": [
"## Qwen-Image 2.1 + LanPaint, inpainting a picture with transparency\n\nThe source is RGBA and the edit is allowed to change the alpha, so the rebuilt sole gets\na new silhouette instead of just new pixels.\n\n`Join Image With Alpha` re-attaches the picture's alpha - `Load Image` hands it out on its\n`MASK` output rather than keeping it on the image. That mask is `1 - alpha` (1 =\ntransparent), which is what `Join Image With Alpha` wants, so it connects with no invert.\n\nThe mask is a separate greyscale file, white = repaint, and it is kept modest here (about\n16% of the frame). LanPaint anchors everything outside the mask to the original on every\nstep, so a mask covering most of the picture leaves the model without context."
],
"title": "How this works (and the mask size trap)",
"color": "#432",
"bgcolor": "#653"
}
],
"links": [
[
1,
1,
0,
4,
0,
"MODEL"
],
[
2,
4,
0,
11,
0,
"MODEL"
],
[
3,
2,
0,
8,
0,
"CLIP"
],
[
4,
2,
0,
9,
0,
"CLIP"
],
[
5,
3,
0,
8,
2,
"VAE"
],
[
6,
3,
0,
10,
1,
"VAE"
],
[
7,
3,
0,
12,
1,
"VAE"
],
[
8,
5,
0,
6,
0,
"IMAGE"
],
[
9,
5,
1,
6,
1,
"MASK"
],
[
10,
6,
0,
8,
3,
"IMAGE"
],
[
11,
6,
0,
10,
0,
"IMAGE"
],
[
12,
6,
0,
12,
2,
"IMAGE"
],
[
13,
7,
0,
10,
2,
"MASK"
],
[
14,
7,
0,
12,
3,
"MASK"
],
[
15,
8,
0,
11,
1,
"CONDITIONING"
],
[
16,
9,
0,
11,
2,
"CONDITIONING"
],
[
17,
10,
0,
11,
3,
"LATENT"
],
[
18,
11,
0,
12,
0,
"LATENT"
],
[
19,
12,
0,
13,
0,
"IMAGE"
]
],
"groups": [
{
"id": 1,
"title": "Step 1 - Load models",
"bounding": [
-1190,
0,
420,
400
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Step 2 - Picture + its alpha + the mask",
"bounding": [
-1190,
380,
880,
1020
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
},
{
"id": 3,
"title": "Step 3 - Texts",
"bounding": [
-750,
490,
460,
480
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
},
{
"id": 4,
"title": "Step 4 - LanPaint: encode, sample, decode",
"bounding": [
-250,
0,
840,
520
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.55,
"offset": [
1210,
180
]
},
"workflowRendererVersion": "LG"
},
"version": 0.4
}
Binary file not shown.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 4.3 MiB

After

Width:  |  Height:  |  Size: 4.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 793 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.0 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.6 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.6 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 965 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1018 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 5.0 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.8 MiB

+59 -2
View File
@@ -1,6 +1,6 @@
import json
import os
from contextlib import contextmanager
from contextlib import contextmanager, nullcontext
import math
# import nodes.py
import comfy
@@ -56,6 +56,39 @@ def _version_tuple(value):
COMFYUI_VERSION_060_OR_NEWER = _version_tuple(comfyui_version.__version__) >= (0, 6, 0)
# ComfyUI >= 0.34 runs MiniMax H3 with per-token denoise-mask row timesteps
# (commit ff6c8a8a). LanPaint keeps its 0.33 single-schedule contract, so this
# gates hiding the mask from the model during the paint loop.
COMFYUI_H3_DENOISE_MASK_CONTRACT = _version_tuple(getattr(comfyui_version, "__version__", "0.0.0")) >= (0, 34, 0)
@contextmanager
def _hide_h3_denoise_mask(model):
"""ComfyUI >= 0.34 hands MiniMax H3 a per-token denoise mask, which
switches the DiT to mask-driven row timesteps (commit ff6c8a8a). LanPaint
implements the 0.33 single-schedule contract in its own replace step and
inner dynamics, so during the paint loop the mask is hidden from the
model's extra_conds and the DiT keeps the uniform row timesteps that
LanPaint's math expects."""
had_instance_attr = "extra_conds" in getattr(model, "__dict__", {})
original_extra_conds = model.extra_conds
def _extra_conds_without_mask(*args, **kwargs):
kwargs.pop("denoise_mask", None)
return original_extra_conds(*args, **kwargs)
model.extra_conds = _extra_conds_without_mask
try:
yield
finally:
if had_instance_attr:
model.extra_conds = original_extra_conds
else:
try:
del model.extra_conds
except AttributeError:
model.extra_conds = original_extra_conds
def reshape_mask(input_mask, output_shape,video_inpainting=False):
dims = len(output_shape) - 2
scale_mode = "nearest-exact"
@@ -190,6 +223,13 @@ class CFGGuider_LanPaint:
self.minimax_h3_audio = _detect_minimax_h3_audio(
self.model_patcher, self.model_options, kwargs.get("latent_shapes", None))
# ComfyUI >= 0.34 switches MiniMax H3 to per-token row timesteps driven
# by the denoise mask; LanPaint's replace step and inner dynamics follow
# the 0.33 single-schedule contract, so hide the mask from the model for
# the paint loop and keep the row timesteps uniform.
h3_hide_mask = self.minimax_h3_audio is not None and COMFYUI_H3_DENOISE_MASK_CONTRACT
hide_mask_ctx = _hide_h3_denoise_mask(self.inner_model) if h3_hide_mask else nullcontext()
if denoise_mask is not None:
video_inpainting = self.model_options.get("video_inpainting", False)
if tuple(denoise_mask.shape) != tuple(noise.shape):
@@ -204,7 +244,8 @@ class CFGGuider_LanPaint:
try:
self.model_patcher.pre_run()
output = self.inner_sample(noise, latent_image, device, sampler, sigmas, denoise_mask, callback, disable_pbar, seed, **kwargs)
with hide_mask_ctx:
output = self.inner_sample(noise, latent_image, device, sampler, sigmas, denoise_mask, callback, disable_pbar, seed, **kwargs)
finally:
self.model_patcher.cleanup()
@@ -1329,6 +1370,18 @@ class LanPaint_ImageDecode:
)
if image is None:
return (img,)
# Some VAEs decode to more channels than the source image: the Qwen Image
# 2.1 VAE always emits a 4th channel for RGB input. ComfyUI's IMAGE type is
# RGB, so trim to the original's channel count - which is what ComfyUI's own
# RGBA -> RGB conversion does, compositing over white - and the merge below
# can broadcast. Without this it raises
# "The size of tensor a (3) must match the size of tensor b (4)".
img_channels, orig_channels = img.shape[-1], image.shape[-1]
if img_channels > orig_channels:
img = img[..., :orig_channels]
elif img_channels < orig_channels:
pad = img.new_ones(img.shape[:-1] + (orig_channels - img_channels,))
img = torch.cat((img, pad), dim=-1)
target_h, target_w = image.shape[1], image.shape[2]
if tuple(img.shape[1:3]) != (target_h, target_w):
img = torch.nn.functional.interpolate(
@@ -1339,6 +1392,10 @@ class LanPaint_ImageDecode:
).movedim(1, -1)
if mask is None:
return (img,)
# Keep the merge channel-agnostic: an RGBA image (a source that carries its
# own alpha, e.g. a Qwen Image 2.1 transparent-background render) keeps its
# 4th channel all the way through, so transparency can be inpainted alongside
# the content. An RGB image still comes back as RGB.
return (merge_video_with_mask(image, img, mask, blend_overlap),)
+23
View File
@@ -322,3 +322,26 @@ def test_prepare_step_size_handles_per_row_parameters() -> None:
adt = (A_x * dtx).flatten()
assert adt[0] == pytest.approx(0.2) # 1/(1-0.5) * 0.1
assert adt[-1] == pytest.approx(0.2) # 1/(1-0.9) * 0.02 -- bounded invariant
def test_hide_h3_denoise_mask_strips_and_restores(monkeypatch) -> None:
# ComfyUI >= 0.34: the per-token denoise mask switches the MiniMax H3 DiT
# to mask-driven row timesteps. LanPaint hides it from extra_conds during
# the paint loop so the DiT keeps the uniform 0.33 row timesteps.
nodes = _import_nodes(monkeypatch)
captured = {}
class FakeH3Model:
def extra_conds(self, **kwargs): # type: ignore[no-untyped-def]
captured.update(kwargs)
return {}
model = FakeH3Model()
with nodes._hide_h3_denoise_mask(model):
model.extra_conds(denoise_mask="MASK", latent_shapes=[(1, 2), (1, 2)])
assert "denoise_mask" not in captured # hidden from the model
assert captured["latent_shapes"] is not None # other conds still pass
captured.clear()
model.extra_conds(denoise_mask="MASK", latent_shapes=[(1, 2), (1, 2)])
assert captured["denoise_mask"] == "MASK" # restored after the loop