Compare commits

...
19 Commits
Author SHA1 Message Date
charrywhite 111f9fdffb 666 version code 2026-02-02 21:36:01 +08:00
charrywhite 9479106c1d Fix Flux2 Dev Inpaint workflow size bug 2026-01-10 12:43:40 +08:00
charrywhite db4bbbe9f2 Update Flux.2.Dev workflow Preview pic 2026-01-10 11:47:29 +08:00
charrywhite 1a3e15964c Merge branch 'master' of https://github.com/scraed/LanPaint 2026-01-10 11:41:58 +08:00
charrywhite 4cc30f8a1d Fix Flux2.Dev Inpainting Workflow MaskBlend Bug 2026-01-10 11:41:53 +08:00
scraed d1e609192f Bump version from 1.4.8 to 1.4.9 2026-01-06 23:13:59 +08:00
charrywhite bf46831196 Update README.md 2026-01-06 23:01:04 +08:00
charrywhite c9d9a79b18 Update README.md 2026-01-06 22:59:21 +08:00
charrywhite 175156af86 Add Flux.2 Dev support 2026-01-06 22:57:42 +08:00
charrywhite 3281946b32 Merge branch 'master' of https://github.com/scraed/LanPaint 2026-01-06 22:42:16 +08:00
charrywhite 83f557ff4c Support Flux.2 Dev Inpainting 2026-01-06 22:42:07 +08:00
charrywhite 5e7fe4d5e4 Enhance README with new Discord info and features
Updated README to highlight new Discord channel and features.
2026-01-06 21:06:25 +08:00
scraed 69f2aded4d Update README with Discord link and new features
Added Discord link for community engagement and announced new inpainting and outpainting features.
2026-01-06 20:02:15 +08:00
scraed 5f1b6d8989 add discord link 2026-01-06 20:01:41 +08:00
scraed 7b0d144db9 Update Qwen Image Edit 2511 support
Updated references for Qwen Image Edit to include version 2511 alongside 2509.
2025-12-28 18:40:41 +08:00
scraed 27ebb6e7af fix qwen edit shape error 2025-12-25 23:18:51 +08:00
charrywhite f148e4b631 Update paper title in README.md 2025-12-17 22:33:46 +08:00
charrywhite cc9ec8873a Revise citation and add TMLR link
Updated citation format and added TMLR link.
2025-12-17 22:30:06 +08:00
scraed 565087e8f8 fix z image resize error for inpaint 2025-12-15 18:15:15 +08:00
10 changed files with 2381 additions and 283 deletions
+53 -14
View File
@@ -1,18 +1,38 @@
<div align="center">
# LanPaint: Universal Inpainting Sampler with "Think Mode"
[![arXiv](https://img.shields.io/badge/Arxiv-2502.03491-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2502.03491)
[![TMLR PDF](https://img.shields.io/badge/TMLR-PDF-8A2BE2?logo=openreview&logoColor=white)](https://openreview.net/pdf?id=JPC8JyOUSW)
[![Python Benchmark](https://img.shields.io/badge/🐍-Python_Benchmark-3776AB?logo=python)](https://github.com/scraed/LanPaintBench)
[![ComfyUI Extension](https://img.shields.io/badge/ComfyUI-Extension-7B5DFF)](https://github.com/comfyanonymous/ComfyUI)
[![Hugging Face](https://img.shields.io/badge/Hugging%20Face-yellow?logo=huggingface&logoColor=white)](https://huggingface.co/charrywhite/LanPaint)
[![Blog](https://img.shields.io/badge/📝-Blog-9cf)](https://scraed.github.io/scraedBlog/)
[![GitHub stars](https://img.shields.io/github/stars/scraed/LanPaint)](https://github.com/scraed/LanPaint/stargazers)
[![Discord](https://img.shields.io/badge/Discord-5865F2?style=for-the-badge&logo=discord&logoColor=white)](https://discord.gg/aCGZutBV)
</div>
Universally applicable inpainting ability for every model. LanPaint sampler lets the model "think" through multiple iterations before denoising, enabling you to invest more computation time for superior inpainting quality.
This is the official implementation of ["Lanpaint: Training-Free Diffusion Inpainting with Exact and Fast Conditional Inference"](https://arxiv.org/abs/2502.03491), accepted by TMLR. The repository is for ComfyUI extension. Local Python benchmark code is published here: [LanPaintBench](https://github.com/scraed/LanPaintBench).
This is the official implementation of ["LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling"](https://arxiv.org/abs/2502.03491), accepted by TMLR. The repository is for ComfyUI extension. Local Python benchmark code is published here: [LanPaintBench](https://github.com/scraed/LanPaintBench).
## Citation
```
@article{
zheng2025lanpaint,
title={LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling},
author={Candi Zheng and Yuan Lan and Yang Wang},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2025},
url={https://openreview.net/forum?id=JPC8JyOUSW},
note={}
}
```
**🎉 NEW 2026: Join our discord!**
[Join our Discord](https://discord.gg/aCGZutBV) to share experiences, discuss features, and explore future development.
**🎬 NEW: LanPaint now supports inpainting and outpainting based on Z-Image!**
| Original | Masked | Inpainted |
@@ -46,11 +66,12 @@ Check our latest [Wan 2.2 Video Examples](#video-examples-beta), [Wan 2.2 Image
- [Wan 2.2 Video Outpainting](#wan-22-video-outpainting)
- [Resource Consumption](#resource-consumption)
- [Image Examples](#image-examples)
- [Flux.2.Dev](#example-flux2dev-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [Z-image](#example-z-image-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [Hunyuan T2I](#example-hunyuan-t2i-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [Wan 2.2 T2I](#example-wan22-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [Wan 2.2 T2I with reference](#example-wan22-partial-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [Qwen Image Edit 2509](#example-qwen-edit-2509-inpaint)
- [Qwen Image Edit 2511 2509](#example-qwen-edit-2509-inpaint)
- [Qwen Image Edit 2508](#example-qwen-edit-2508-inpaint)
- [Qwen Image](#example-qwen-image-inpaintlanpaint-k-sampler-5-steps-of-thinking)
- [HiDream](#example-hidream-inpaint-lanpaint-k-sampler-5-steps-of-thinking)
@@ -69,7 +90,7 @@ Check our latest [Wan 2.2 Video Examples](#video-examples-beta), [Wan 2.2 Image
## Features
- **Universal Compatibility** – Works instantly with almost any model (**SD 1.5, XL, 3.5, Flux, HiDream, Qwen-Image, Wan2.2 or custom LoRAs**) and ControlNet.
- **Universal Compatibility** – Works instantly with almost any model (**Z-image, Hunyuan, Wan 2.2, Qwen Image/Edit, HiDream, SD 3.5, Flux-series, SDXL, SD 1.5 or custom LoRAs**) and ControlNet.
![Inpainting Result 13](https://github.com/scraed/LanPaint/blob/master/examples/InpaintChara_13.jpg)
- **No Training Needed** – Works out of the box with your existing model.
- **Easy to Use** – Same workflow as standard ComfyUI KSampler.
@@ -263,7 +284,7 @@ You need to follow the ComfyUI version of [Wan2.2 T2V workflow](https://docs.com
### Example Qwen Edit 2509: InPaint
Check our latest updated [Mased Qwen Edit Workflow](https://github.com/scraed/LanPaint/tree/master/examples/Example_14) for Qwen Image Edit 2509. Download the model at [Qwen Image Edit 2509 Comfy](https://huggingface.co/Comfy-Org/Qwen-Image-Edit_ComfyUI/tree/main/split_files/diffusion_models).
Check our latest updated [Mased Qwen Edit Workflow](https://github.com/scraed/LanPaint/tree/master/examples/Example_14) for Qwen Image Edit 2509. Download the model at [Qwen Image Edit 2509 Comfy](https://huggingface.co/Comfy-Org/Qwen-Image-Edit_ComfyUI/tree/main/split_files/diffusion_models). This workflow also supports Qwen Image Edit 2511.
![Qwen Result 3](https://github.com/scraed/LanPaint/blob/master/examples/LanPaintQwen_04.jpg)
@@ -304,6 +325,23 @@ You need to follow the ComfyUI version of [HiDream workflow](https://docs.comfy.
You need to follow the ComfyUI version of [SD 3.5 workflow](https://comfyui-wiki.com/en/tutorial/advanced/stable-diffusion-3-5-comfyui-workflow) to download and install the model.
### Example Flux.2.Dev: InPaint(LanPaint K Sampler, 5 steps of thinking)
<details open>
<summary>View Original / Masked / Inpainted Comparison</summary>
| Original | Masked | Inpainted |
|:--------:|:------:|:---------:|
| ![Original Flux.2.Dev](https://github.com/scraed/LanPaint/blob/master/examples/Example_23/Original_No_Mask.png) | ![Masked Flux.2.Dev](https://github.com/scraed/LanPaint/blob/master/examples/Example_23/Masked_Load_Me_in_Loader.png) | ![Inpainted Flux.2.Dev](https://github.com/scraed/LanPaint/blob/master/examples/Example_23/InPainted_Drag_Me_to_ComfyUI.png) |
</details>
[View Workflow & Masks](https://github.com/scraed/LanPaint/tree/master/examples/Example_23)
[Model Used in This Example](https://huggingface.co/Comfy-Org/flux2-dev)
(Note: Prompt First mode is disabled on Flux.2.Dev. As it does not use CFG guidance.)
### Example Flux: InPaint(LanPaint K Sampler, 5 steps of thinking)
![Inpainting Result 7](https://github.com/scraed/LanPaint/blob/master/examples/InpaintChara_10.jpg)
[View Workflow & Masks](https://github.com/scraed/LanPaint/tree/master/examples/Example_7)
@@ -451,17 +489,19 @@ Submit a PR to add your tutorial/video here, or open an [Issue](https://github.c
- Try Implement Detailer
- ~~Provide inference code on without GUI.~~ Check our local Python benchmark code [LanPaintBench](https://github.com/scraed/LanPaintBench).
## Citation
```
@misc{zheng2025lanpainttrainingfreediffusioninpainting,
title={Lanpaint: Training-Free Diffusion Inpainting with Exact and Fast Conditional Inference},
author={Candi Zheng and Yuan Lan and Yang Wang},
year={2025},
eprint={2502.03491},
archivePrefix={arXiv},
primaryClass={eess.IV},
url={https://arxiv.org/abs/2502.03491},
@article{
zheng2025lanpaint,
title={LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling},
author={Candi Zheng and Yuan Lan and Yang Wang},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2025},
url={https://openreview.net/forum?id=JPC8JyOUSW},
note={}
}
```
@@ -469,4 +509,3 @@ Submit a PR to add your tutorial/video here, or open an [Issue](https://github.c
Binary file not shown.

After

Width:  |  Height:  |  Size: 1.6 MiB

File diff suppressed because it is too large Load Diff
+236 -234
View File
@@ -2,7 +2,7 @@
"id": "9ae6082b-c7f4-433c-9971-7a8f65a3ea65",
"revision": 0,
"last_node_id": 73,
"last_link_id": 93,
"last_link_id": 95,
"nodes": [
{
"id": 39,
@@ -29,9 +29,9 @@
}
],
"properties": {
"Node name for S&R": "CLIPLoader",
"cnr_id": "comfy-core",
"ver": "0.3.73",
"Node name for S&R": "CLIPLoader",
"models": [
{
"name": "qwen_3_4b.safetensors",
@@ -81,9 +81,9 @@
}
],
"properties": {
"Node name for S&R": "VAELoader",
"cnr_id": "comfy-core",
"ver": "0.3.73",
"Node name for S&R": "VAELoader",
"models": [
{
"name": "ae.safetensors",
@@ -134,9 +134,9 @@
}
],
"properties": {
"Node name for S&R": "ConditioningZeroOut",
"cnr_id": "comfy-core",
"ver": "0.3.73",
"Node name for S&R": "ConditioningZeroOut",
"enableTabs": false,
"tabWidth": 65,
"tabXOffset": 10,
@@ -147,54 +147,6 @@
},
"widgets_values": []
},
{
"id": 46,
"type": "UNETLoader",
"pos": [
158.31605577680375,
332.3053621604172
],
"size": [
323.734375,
145.390625
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
54
]
}
],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.73",
"Node name for S&R": "UNETLoader",
"models": [
{
"name": "z_image_turbo_bf16.safetensors",
"url": "https://huggingface.co/Comfy-Org/z_image_turbo/resolve/main/split_files/diffusion_models/z_image_turbo_bf16.safetensors",
"directory": "diffusion_models"
}
],
"enableTabs": false,
"tabWidth": 65,
"tabXOffset": 10,
"hasSecondTab": false,
"secondTabText": "Send Back",
"secondTabOffset": 80,
"secondTabWidth": 65
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"default"
]
},
{
"id": 41,
"type": "EmptySD3LatentImage",
@@ -207,7 +159,7 @@
179.390625
],
"flags": {},
"order": 3,
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
@@ -219,9 +171,9 @@
}
],
"properties": {
"Node name for S&R": "EmptySD3LatentImage",
"cnr_id": "comfy-core",
"ver": "0.3.64",
"Node name for S&R": "EmptySD3LatentImage",
"enableTabs": false,
"tabWidth": 65,
"tabXOffset": 10,
@@ -248,7 +200,7 @@
111.390625
],
"flags": {},
"order": 12,
"order": 13,
"mode": 0,
"inputs": [
{
@@ -268,9 +220,9 @@
}
],
"properties": {
"Node name for S&R": "ModelSamplingAuraFlow",
"cnr_id": "comfy-core",
"ver": "0.3.64",
"Node name for S&R": "ModelSamplingAuraFlow",
"enableTabs": false,
"tabWidth": 65,
"tabXOffset": 10,
@@ -321,9 +273,9 @@
}
],
"properties": {
"Node name for S&R": "VAEDecode",
"cnr_id": "comfy-core",
"ver": "0.3.64",
"Node name for S&R": "VAEDecode",
"enableTabs": false,
"tabWidth": 65,
"tabXOffset": 10,
@@ -366,9 +318,9 @@
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode",
"cnr_id": "comfy-core",
"ver": "0.3.73",
"Node name for S&R": "CLIPTextEncode",
"enableTabs": false,
"tabWidth": 65,
"tabXOffset": 10,
@@ -395,7 +347,7 @@
145.390625
],
"flags": {},
"order": 8,
"order": 10,
"mode": 4,
"inputs": [
{
@@ -414,9 +366,9 @@
}
],
"properties": {
"Node name for S&R": "LoraLoaderModelOnly",
"cnr_id": "comfy-core",
"ver": "0.3.75",
"Node name for S&R": "LoraLoaderModelOnly",
"models": [
{
"name": "pixel_art_style_z_image_turbo.safetensors",
@@ -451,7 +403,7 @@
"flags": {
"collapsed": false
},
"order": 4,
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [],
@@ -509,12 +461,12 @@
}
],
"properties": {
"Node name for S&R": "LanPaint_KSampler",
"cnr_id": "LanPaint",
"ver": "6109df6591a4cf2bc9d3b113d03f7297fa9248e9",
"Node name for S&R": "LanPaint_KSampler"
"ver": "6109df6591a4cf2bc9d3b113d03f7297fa9248e9"
},
"widgets_values": [
454543748915702,
880311146947153,
"randomize",
9,
1,
@@ -539,7 +491,7 @@
686.171875
],
"flags": {},
"order": 5,
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [
@@ -547,7 +499,6 @@
"name": "IMAGE",
"type": "IMAGE",
"links": [
74,
85,
88
]
@@ -556,15 +507,14 @@
"name": "MASK",
"type": "MASK",
"links": [
75,
89
]
}
],
"properties": {
"Node name for S&R": "LoadImage",
"cnr_id": "comfy-core",
"ver": "0.3.59",
"Node name for S&R": "LoadImage",
"ue_properties": {
"widget_ue_connectable": {},
"version": "7.1",
@@ -613,9 +563,9 @@
}
],
"properties": {
"Node name for S&R": "VAEEncode",
"cnr_id": "comfy-core",
"ver": "0.3.59",
"Node name for S&R": "VAEEncode",
"ue_properties": {
"widget_ue_connectable": {},
"version": "7.1",
@@ -660,9 +610,9 @@
}
],
"properties": {
"Node name for S&R": "SetLatentNoiseMask",
"cnr_id": "comfy-core",
"ver": "0.3.59",
"Node name for S&R": "SetLatentNoiseMask",
"ue_properties": {
"widget_ue_connectable": {},
"version": "7.1",
@@ -689,7 +639,7 @@
{
"name": "image1",
"type": "IMAGE",
"link": 74
"link": 94
},
{
"name": "image2",
@@ -699,7 +649,7 @@
{
"name": "mask",
"type": "MASK",
"link": 75
"link": 95
}
],
"outputs": [
@@ -712,9 +662,9 @@
}
],
"properties": {
"Node name for S&R": "LanPaint_MaskBlend",
"cnr_id": "LanPaint",
"ver": "4d3d5d17f0105b673df92da5b084cce567c9c712",
"Node name for S&R": "LanPaint_MaskBlend"
"ver": "4d3d5d17f0105b673df92da5b084cce567c9c712"
},
"widgets_values": [
9
@@ -762,7 +712,7 @@
101.421875
],
"flags": {},
"order": 13,
"order": 12,
"mode": 0,
"inputs": [
{
@@ -787,70 +737,12 @@
}
],
"properties": {
"Node name for S&R": "VAEDecode",
"cnr_id": "comfy-core",
"ver": "0.3.23",
"Node name for S&R": "VAEDecode"
"ver": "0.3.23"
},
"widgets_values": []
},
{
"id": 63,
"type": "ImageScale",
"pos": [
1283.837511214967,
1657.4367336671903
],
"size": [
323.90625,
213.375
],
"flags": {},
"order": 15,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 88
},
{
"name": "width",
"type": "INT",
"widget": {
"name": "width"
},
"link": 77
},
{
"name": "height",
"type": "INT",
"widget": {
"name": "height"
},
"link": 78
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.52",
"Node name for S&R": "ImageScale"
},
"widgets_values": [
"area",
512,
512,
"center"
]
},
{
"id": 64,
"type": "MaskToImage",
@@ -863,7 +755,7 @@
77.421875
],
"flags": {},
"order": 10,
"order": 9,
"mode": 0,
"inputs": [
{
@@ -882,51 +774,12 @@
}
],
"properties": {
"Node name for S&R": "MaskToImage",
"cnr_id": "comfy-core",
"ver": "0.3.51",
"Node name for S&R": "MaskToImage"
"ver": "0.3.51"
},
"widgets_values": []
},
{
"id": 65,
"type": "ImageToMask",
"pos": [
1274.556170930292,
2196.94755423526
],
"size": [
323.90625,
111.390625
],
"flags": {},
"order": 18,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 79
}
],
"outputs": [
{
"name": "MASK",
"type": "MASK",
"links": [
91
]
}
],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.51",
"Node name for S&R": "ImageToMask"
},
"widgets_values": [
"red"
]
},
{
"id": 67,
"type": "VAEEncode",
@@ -939,7 +792,7 @@
101.421875
],
"flags": {},
"order": 9,
"order": 8,
"mode": 0,
"inputs": [
{
@@ -963,9 +816,9 @@
}
],
"properties": {
"Node name for S&R": "VAEEncode",
"cnr_id": "comfy-core",
"ver": "0.3.50",
"Node name for S&R": "VAEEncode",
"enableTabs": false,
"tabWidth": 65,
"tabXOffset": 10,
@@ -1026,9 +879,9 @@
}
],
"properties": {
"Node name for S&R": "ImageScale",
"cnr_id": "comfy-core",
"ver": "0.3.52",
"Node name for S&R": "ImageScale"
"ver": "0.3.52"
},
"widgets_values": [
"nearest-exact",
@@ -1082,18 +935,196 @@
}
],
"properties": {
"Node name for S&R": "GetImageSize",
"cnr_id": "comfy-core",
"ver": "0.3.52",
"Node name for S&R": "GetImageSize"
"ver": "0.3.52"
},
"widgets_values": [
"width: 1024, height: 1024\n batch size: 1"
]
},
{
"id": 73,
"type": "PreviewImage",
"pos": [
2226.506787053337,
107.57278250350078
],
"size": [
225,
246
],
"flags": {},
"order": 23,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 93
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage",
"cnr_id": "comfy-core",
"ver": "0.3.76"
},
"widgets_values": []
},
{
"id": 63,
"type": "ImageScale",
"pos": [
1283.837511214967,
1657.4367336671903
],
"size": [
323.90625,
213.375
],
"flags": {},
"order": 15,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 88
},
{
"name": "width",
"type": "INT",
"widget": {
"name": "width"
},
"link": 77
},
{
"name": "height",
"type": "INT",
"widget": {
"name": "height"
},
"link": 78
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
90,
94
]
}
],
"properties": {
"Node name for S&R": "ImageScale",
"cnr_id": "comfy-core",
"ver": "0.3.52"
},
"widgets_values": [
"area",
512,
512,
"center"
]
},
{
"id": 65,
"type": "ImageToMask",
"pos": [
1274.556170930292,
2196.94755423526
],
"size": [
323.90625,
111.390625
],
"flags": {},
"order": 18,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 79
}
],
"outputs": [
{
"name": "MASK",
"type": "MASK",
"links": [
91,
95
]
}
],
"properties": {
"Node name for S&R": "ImageToMask",
"cnr_id": "comfy-core",
"ver": "0.3.51"
},
"widgets_values": [
"red"
]
},
{
"id": 46,
"type": "UNETLoader",
"pos": [
159.05646022112083,
288.62105108441466
],
"size": [
323.734375,
145.390625
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
54
]
}
],
"properties": {
"Node name for S&R": "UNETLoader",
"cnr_id": "comfy-core",
"ver": "0.3.73",
"models": [
{
"name": "z_image_turbo_bf16.safetensors",
"url": "https://huggingface.co/Comfy-Org/z_image_turbo/resolve/main/split_files/diffusion_models/z_image_turbo_bf16.safetensors",
"directory": "diffusion_models"
}
],
"enableTabs": false,
"tabWidth": 65,
"tabXOffset": 10,
"hasSecondTab": false,
"secondTabText": "Send Back",
"secondTabOffset": 80,
"secondTabWidth": 65
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"default"
]
},
{
"id": 72,
"type": "Note",
"pos": [
1266.1247223180148,
1534.1298096819087
1288.7909707861209,
1503.5796522859714
],
"size": [
225,
@@ -1110,35 +1141,6 @@
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 73,
"type": "PreviewImage",
"pos": [
2226.506787053337,
107.57278250350078
],
"size": [
225,
77.421875
],
"flags": {},
"order": 23,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 93
}
],
"outputs": [],
"properties": {
"cnr_id": "comfy-core",
"ver": "0.3.76",
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
}
],
"links": [
@@ -1254,22 +1256,6 @@
1,
"IMAGE"
],
[
74,
57,
0,
60,
0,
"IMAGE"
],
[
75,
57,
1,
60,
2,
"MASK"
],
[
76,
67,
@@ -1397,6 +1383,22 @@
73,
0,
"IMAGE"
],
[
94,
63,
0,
60,
0,
"IMAGE"
],
[
95,
65,
0,
60,
2,
"MASK"
]
],
"groups": [
@@ -1456,13 +1458,13 @@
"config": {},
"extra": {
"ds": {
"scale": 0.17731177475370102,
"scale": 0.3800835362432756,
"offset": [
2305.975484371013,
1284.903135448736
-902.897494105397,
-396.4805064253962
]
},
"frontendVersion": "1.32.10",
"frontendVersion": "1.33.14",
"VHS_latentpreview": false,
"VHS_latentpreviewrate": 0,
"VHS_MetadataImage": true,
Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.1 MiB

After

Width:  |  Height:  |  Size: 1.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.6 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.6 MiB

+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "LanPaint"
version = "1.4.6"
version = "1.4.9"
description = "Achieve seamless inpainting results without needing a specialized inpainting model."
authors = [
{name = "LanPaint", email = "czhengac@connect.ust.hk"}
+205 -34
View File
@@ -12,8 +12,28 @@ from comfy.model_base import ModelType
from .utils import *
from .lanpaint import LanPaint
from comfy.model_base import WAN22
import comfyui_version
import comfy.nested_tensor
def reshape_mask(input_mask, output_shape,video_inpainting=False):
import comfy.nested_tensor
# 修改这里的判断条件,不能只用 hasattr("unbind")
if isinstance(input_mask, comfy.nested_tensor.NestedTensor):
masks = input_mask.unbind()
# 如果 output_shape 也是嵌套的(通常 noise.shape 在 NestedTensor 下返回 tuple of shapes)
if isinstance(output_shape, (list, tuple)) and len(output_shape) > 0 and not isinstance(output_shape[0], int):
reshaped_parts = []
for i in range(len(masks)):
# 递归处理每一个子部分,并传入对应的子 shape
reshaped_parts.append(reshape_mask(masks[i], output_shape[i], video_inpainting))
return comfy.nested_tensor.NestedTensor(tuple(reshaped_parts))
else:
# 如果 output_shape 是单一形状(降级处理)
return comfy.nested_tensor.NestedTensor(tuple(reshape_mask(m, output_shape, video_inpainting) for m in masks))
dims = len(output_shape) - 2
print('output shape',output_shape)
scale_mode = "nearest-exact"
@@ -22,43 +42,81 @@ def reshape_mask(input_mask, output_shape,video_inpainting=False):
print('input_mask.ndim:', input_mask.ndim, 'output_shape len:', len(output_shape))
# Handle video case with temporal dimension
if video_inpainting: # Video case: (batch, channels, frames, height, width)
target_frames = output_shape[2]
target_height, target_width = output_shape[-2:]
# if video_inpainting: # Video case: (batch, channels, frames, height, width)
# target_frames = output_shape[2]
# target_height, target_width = output_shape[-2:]
print('Video case - input_mask initial shape:', input_mask.shape)
# print('Video case - input_mask initial shape:', input_mask.shape)
# First reshape input_mask to have proper dimensions for video processing
# Assume input is (frames, channels, height, width) -> (1, channels, frames, height, width)
input_mask = input_mask.permute(1, 0, 2, 3).unsqueeze(0)
print('Video case - input_mask after reshaping:', input_mask.shape)
# Ensure we have the correct 5D shape: (batch, channels, frames, height, width)
batch_size, channels, frames, height, width = input_mask.shape
print('Video case - dimensions: batch_size={}, channels={}, frames={}, height={}, width={}'.format(batch_size, channels, frames, height, width))
print('Video case - target size:', (target_frames, target_height, target_width))
# # First reshape input_mask to have proper dimensions for video processing
# # Assume input is (frames, channels, height, width) -> (1, channels, frames, height, width)
# ## if comfy version < 0.6.0
# if comfyui_version.__version__ < "0.6.0":
# input_mask = input_mask.permute(1, 0, 2, 3).unsqueeze(0)
# print('Video case - input_mask after reshaping:', input_mask.shape)
# # Ensure we have the correct 5D shape: (batch, channels, frames, height, width)
# batch_size, channels, frames, height, width = input_mask.shape
# print('Video case - dimensions: batch_size={}, channels={}, frames={}, height={}, width={}'.format(batch_size, channels, frames, height, width))
# print('Video case - target size:', (target_frames, target_height, target_width))
# 3D nearest-exact interpolation: (batch, channels, frames, height, width) -> (batch, channels, target_frames, target_height, target_width)
temp_mask = torch.nn.functional.interpolate(
input_mask,
size=(target_frames, target_height, target_width),
mode=scale_mode,
)
# # 3D nearest-exact interpolation: (batch, channels, frames, height, width) -> (batch, channels, target_frames, target_height, target_width)
# temp_mask = torch.nn.functional.interpolate(
# input_mask,
# size=(target_frames, target_height, target_width),
# mode=scale_mode,
# )
# temp_mask is already 5D: (batch, channels, target_frames, target_height, target_width)
mask = temp_mask
print('after mask',mask.shape)
# Handle channel dimension expansion if needed
if mask.shape[1] < output_shape[1]:
mask = mask.repeat(1, output_shape[1], 1, 1, 1)[:, :output_shape[1]]
# Handle batch dimension
mask = repeat_to_batch_size(mask, output_shape[0])
# # temp_mask is already 5D: (batch, channels, target_frames, target_height, target_width)
# mask = temp_mask
# print('after mask',mask.shape)
# # Handle channel dimension expansion if needed
# if mask.shape[1] < output_shape[1]:
# mask = mask.repeat(1, output_shape[1], 1, 1, 1)[:, :output_shape[1]]
# # Handle batch dimension
# mask = repeat_to_batch_size(mask, output_shape[0])
if video_inpainting:
# 如果是 3D Token 序列 (LTXV 压平后的情况)
if input_mask.ndim == 3 and len(output_shape) == 3:
mask = torch.nn.functional.interpolate(
input_mask,
size=output_shape[2],
mode=scale_mode
)
return mask
# 只有在确认为 5D 视频张量时才执行原有逻辑
if input_mask.ndim == 5:
target_frames = output_shape[2]
target_height, target_width = output_shape[-2:]
# (这里保留你原有的 permute 和 unsqueeze 逻辑,但要确保它是针对非 5D 输入的补救)
if input_mask.ndim < 5:
# 假设输入是 (F, C, H, W) -> (1, C, F, H, W)
if hasattr(comfyui_version, "__version__") and comfyui_version.__version__ < "0.6.0":
input_mask = input_mask.permute(1, 0, 2, 3).unsqueeze(0)
# 现在可以安全地解包 5D 形状了
batch_size, channels, frames, height, width = input_mask.shape
mask = torch.nn.functional.interpolate(
input_mask,
size=(target_frames, target_height, target_width),
mode=scale_mode,
)
if mask.shape[1] < output_shape[1]:
mask = mask.repeat(1, output_shape[1], 1, 1, 1)[:, :output_shape[1]]
mask = repeat_to_batch_size(mask, output_shape[0])
return mask
else: # Original 2D image case
mask = torch.nn.functional.interpolate(input_mask, size=output_shape[-2:], mode=scale_mode)
if comfyui_version.__version__ < "0.6.0":
mask = torch.nn.functional.interpolate(input_mask, size=output_shape[-2:], mode=scale_mode)
else:
mask = torch.nn.functional.interpolate(input_mask, size=output_shape[2:], mode=scale_mode)
if mask.shape[1] < output_shape[1]:
mask = mask.repeat((1, output_shape[1]) + (1,) * dims)[:,:output_shape[1]]
mask = repeat_to_batch_size(mask, output_shape[0])
print('resize mask',mask.shape,type(mask),torch.max(mask),torch.min(mask))
return mask
def prepare_mask(noise_mask, shape, device,video_inpainting=False):
return reshape_mask(noise_mask, shape,video_inpainting).to(device)
@@ -88,9 +146,9 @@ class CFGGuider_LanPaint:
if isinstance(self.inner_model, WAN22):
print("WAN22 detected")
self.inner_model.extra_conds = super(WAN22, self.inner_model).extra_conds
if denoise_mask is not None:
video_inpainting = self.model_options.get("video_inpainting", False)
print('denoise_mask',denoise_mask.shape,type(denoise_mask))
denoise_mask = prepare_mask(denoise_mask, noise.shape, device, video_inpainting)
noise = noise.to(device)
@@ -138,8 +196,6 @@ class KSamplerX0Inpaint:
abt = (1 - Flow_t)**2 / ((1 - Flow_t)**2 + Flow_t**2 )
VE_Sigma = Flow_t / (1 - Flow_t)
#print("t", torch.mean( sigma ).item(), "VE_Sigma", torch.mean( VE_Sigma ).item())
else:
VE_Sigma = sigma
abt = 1/( 1+VE_Sigma**2 )
@@ -149,6 +205,31 @@ class KSamplerX0Inpaint:
if "denoise_mask_function" in model_options:
denoise_mask = model_options["denoise_mask_function"](sigma, denoise_mask, extra_options={"model": self.inner_model, "sigmas": self.sigmas})
if isinstance(denoise_mask, comfy.nested_tensor.NestedTensor):
masks = denoise_mask.unbind()
xs = x.unbind()
latent_imgs = self.latent_image.unbind()
noises = self.noise.unbind()
outs = []
# 针对 LTXV,通常 i=0 是视频,i=1 是音频
for i in range(len(xs)):
m = (masks[i] > 0.5).float()
lm = 1 - m
# 这里的 PaintMethod 通常只支持普通 Tensor,所以我们分块处理
# 注意:如果音频部分不需要 Inpaint,可以增加判断
current_times = (VE_Sigma, abt, Flow_t)
# 只有视频部分 (i=0) 应用 LanPaint 逻辑,音频部分通常直接 pass 或原样返回
if i == 0:
out_part = self.PaintMethod(xs[i], latent_imgs[i], noises[i], sigma, lm, current_times, model_options, seed)
else:
# 音频部分如果没有对应的 Inpaint 逻辑,通常直接调用 inner_model
out_part, _ = self.inner_model(xs[i], sigma, model_options=model_options, seed=seed)
outs.append(out_part)
return comfy.nested_tensor.NestedTensor(tuple(outs))
denoise_mask = (denoise_mask > 0.5).float()
latent_mask = 1 - denoise_mask
@@ -183,6 +264,7 @@ class KSAMPLER(comfy.samplers.KSAMPLER):
#noise here is a randn noise from comfy.sample.prepare_noise
#latent_image is the latent image as input of the KSampler node. For inpainting, it is the masked latent image. Otherwise it is zero tensor.
extra_args["denoise_mask"] = denoise_mask
print("LanPaint KSampler start sampler_function",denoise_mask.shape if denoise_mask is not None else None)
model_k = KSamplerX0Inpaint(model_wrap, sigmas)
model_k.latent_image = latent_image
if self.inpaint_options.get("random", False): #TODO: Should this be the default?
@@ -447,6 +529,77 @@ class MaskBlend:
return kernel
class MaskBlendAlpha:
"""
Create an RGBA image by writing the mask into the PNG alpha channel.
Requirement:
- inpaint region: alpha = 0 (transparent)
- other region: alpha = 1 (opaque)
This node writes the mask into the PNG alpha channel.
Current default behavior matches the previous `invert_mask=True` behavior:
alpha = mask.
"""
def __init__(self):
pass
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"image": ("IMAGE", {"tooltip": "VAE-decoded image (RGB)."}),
"mask": ("MASK", {"tooltip": "Mask used as alpha channel (alpha = mask)."}),
},
}
RETURN_TYPES = ("IMAGE",)
FUNCTION = "to_rgba"
CATEGORY = "image/postprocessing"
def to_rgba(self, image: torch.Tensor, mask: torch.Tensor):
"""
image: [B,H,W,3] float in [0,1]
mask: [B,H,W] (or [H,W]) float in [0,1] used as alpha
returns RGBA image: [B,H,W,4] float in [0,1]
"""
if image.ndim != 4 or image.shape[-1] != 3:
raise ValueError(f"Expected IMAGE tensor [B,H,W,3], got {tuple(image.shape)}")
# Normalize mask shape to [B,H,W]
if mask.ndim == 2:
mask = mask.unsqueeze(0)
elif mask.ndim == 3:
pass
else:
# Some pipelines may carry mask as [B,1,H,W]
if mask.ndim == 4 and mask.shape[1] == 1:
mask = mask[:, 0, :, :]
else:
raise ValueError(f"Expected MASK tensor [B,H,W] or [H,W], got {tuple(mask.shape)}")
b, h, w, _ = image.shape
# Batch align
if mask.shape[0] != b:
if mask.shape[0] == 1:
mask = mask.repeat(b, 1, 1)
else:
raise ValueError(f"Batch mismatch: image batch={b}, mask batch={mask.shape[0]}")
# Spatial align (resize mask to image resolution if needed)
if mask.shape[1] != h or mask.shape[2] != w:
mask_4d = mask.unsqueeze(1) # [B,1,H,W]
mask_4d = torch.nn.functional.interpolate(mask_4d, size=(h, w), mode="nearest")
mask = mask_4d[:, 0, :, :]
mask = mask.float().clamp(0.0, 1.0)
# Default behavior (matches previous invert_mask=True path):
# alpha = mask
rgba = torch.cat([image, mask.unsqueeze(-1)], dim=-1)
return (rgba,)
class Noise_EmptyNoise:
def generate_noise(self, latent):
return torch.zeros_like(latent["samples"])
@@ -475,6 +628,7 @@ class LanPaint_SamplerCustom:
"LanPaint_NumSteps": ("INT", {"default": 5, "min": 0, "max": 100, "tooltip": "Number of steps for Langevin dynamics, representing turns of thinking per step."}),
"LanPaint_PromptMode": (["Image First", "Prompt First"], {"tooltip": "Image First: prioritizes image quality; Prompt First: prioritizes prompt adherence."}),
"LanPaint_Info": ("STRING", {"default": "LanPaint Custom Sampler. For more info, visit https://github.com/scraed/LanPaint. If you find it useful, please give a star ⭐️!", "multiline": True}),
"Inpainting_mode": (["🖼️ Image Inpainting", "🎬 Video Inpainting"], {"default": "🖼️ Image Inpainting", "tooltip": "Choose Image mode for photos or Video mode for video frames with temporal consistency"}),
}
}
@@ -483,7 +637,7 @@ class LanPaint_SamplerCustom:
FUNCTION = "sample"
CATEGORY = "sampling/custom_sampling"
def sample(self, model, sampler, sigmas, add_noise, noise_seed, cfg, positive, negative, latent_image, LanPaint_NumSteps, LanPaint_PromptMode, LanPaint_Info=""):
def sample(self, model, sampler, sigmas, add_noise, noise_seed, cfg, positive, negative, latent_image, LanPaint_NumSteps, LanPaint_PromptMode, LanPaint_Info="",Inpainting_mode="🖼️ Image Inpainting"):
model.LanPaint_StepSize = 0.2
model.LanPaint_Lambda = 16.0
model.LanPaint_Beta = 1.
@@ -494,6 +648,10 @@ class LanPaint_SamplerCustom:
model.LanPaint_cfg_BIG = cfg
else:
model.LanPaint_cfg_BIG = 0 * cfg - 0.5
video_inpainting = (Inpainting_mode == "🎬 Video Inpainting")
if not hasattr(model, 'model_options') or model.model_options is None:
model.model_options = {}
model.model_options["video_inpainting"] = video_inpainting
with override_sample_function():
latent = latent_image.copy()
latent_image = latent["samples"]
@@ -541,6 +699,7 @@ class LanPaint_SamplerCustomAdvanced:
"LanPaint_PromptMode": (["Image First", "Prompt First"], {"tooltip": "Image First: prioritizes image quality; Prompt First: prioritizes prompt adherence."}),
"LanPaint_EarlyStop": ("INT", {"default": 1, "min": 0, "max": 10000, "tooltip": "Steps to stop LanPaint early, preventing irregular patterns."}),
"LanPaint_Info": ("STRING", {"default": "LanPaint Custom Sampler Adv. For more info, visit https://github.com/scraed/LanPaint. If you find it useful, please give a star ⭐️!", "multiline": True}),
"Inpainting_mode": (["🖼️ Image Inpainting", "🎬 Video Inpainting"], {"default": "🖼️ Image Inpainting", "tooltip": "Choose Image mode for photos or Video mode for video frames with temporal consistency"}),
}
}
@@ -551,7 +710,7 @@ class LanPaint_SamplerCustomAdvanced:
CATEGORY = "sampling/custom_sampling"
def sample(self, noise, guider, sampler, sigmas, latent_image, LanPaint_NumSteps, LanPaint_Lambda, LanPaint_StepSize, LanPaint_Beta, LanPaint_Friction, LanPaint_PromptMode, LanPaint_EarlyStop, LanPaint_Info=""):
def sample(self, noise, guider, sampler, sigmas, latent_image, LanPaint_NumSteps, LanPaint_Lambda, LanPaint_StepSize, LanPaint_Beta, LanPaint_Friction, LanPaint_PromptMode, LanPaint_EarlyStop, LanPaint_Info="",Inpainting_mode="🖼️ Image Inpainting"):
model = guider.model_patcher
model.LanPaint_StepSize = LanPaint_StepSize
model.LanPaint_Lambda = LanPaint_Lambda
@@ -563,16 +722,25 @@ class LanPaint_SamplerCustomAdvanced:
model.LanPaint_cfg_BIG = guider.cfg
else:
model.LanPaint_cfg_BIG = 0 * guider.cfg - 0.5
video_inpainting = (Inpainting_mode == "🎬 Video Inpainting")
if not hasattr(model, 'model_options') or model.model_options is None:
model.model_options = {}
model.model_options["video_inpainting"] = video_inpainting
with override_sample_function():
latent = latent_image
latent_image = latent["samples"]
print('before fix_empty_latent_channels latent_image shape',latent_image.shape)
latent = latent.copy()
latent_image = comfy.sample.fix_empty_latent_channels(guider.model_patcher, latent_image)
latent["samples"] = latent_image
print('latent_image shape',latent_image.shape)
print('outside noise_mask',latent["noise_mask"].shape if "noise_mask" in latent else 'no noise_mask')
print('latent keys',latent.keys())
noise_mask = None
if "noise_mask" in latent:
noise_mask = latent["noise_mask"]
print('inside noise_mask shape',noise_mask.shape)
x0_output = {}
callback = latent_preview.prepare_callback(guider.model_patcher, sigmas.shape[-1] - 1, x0_output)
@@ -588,6 +756,7 @@ class LanPaint_SamplerCustomAdvanced:
out_denoised["samples"] = guider.model_patcher.model.process_latent_out(x0_output["x0"].cpu())
else:
out_denoised = out
# print('output',out.keys(),out["samples"].shape,out['noise_mask'].shape)
return (out, out_denoised)
@@ -599,6 +768,7 @@ NODE_CLASS_MAPPINGS = {
"LanPaint_SamplerCustom" : LanPaint_SamplerCustom,
"LanPaint_SamplerCustomAdvanced" : LanPaint_SamplerCustomAdvanced,
"LanPaint_MaskBlend": MaskBlend,
"LanPaint_MaskBlendAlpha": MaskBlendAlpha,
# "LanPaint_UpSale_LatentNoiseMask": LanPaint_UpSale_LatentNoiseMask,
}
@@ -609,5 +779,6 @@ NODE_DISPLAY_NAME_MAPPINGS = {
"LanPaint_SamplerCustom" : "LanPaint Sampler Custom",
"LanPaint_SamplerCustomAdvanced" : "LanPaint Sampler Custom (Advanced)",
"LanPaint_MaskBlend": "LanPaint Mask Blend",
"LanPaint_MaskBlendAlpha": "MaskBlend (alpha)",
# "LanPaint_UpSale_LatentNoiseMask": "LanPaint UpSale Latent Noise Mask"
}