LanPaint runs on Qwen-Image 2.1 unchanged: it is a rectified-flow model, so the existing Flux/Qwen-Image conversions apply. Example_31 uses a text-to-image render of its own model as the before/after, generated with a transparent background, and keeps the inpainting mask in a separate greyscale file rather than in the picture's alpha - 2.1's alpha channel means transparency, so overloading it would make the two indistinguishable. The transparency is carried into the latent and edited along with the pixels: the 2.1 VAE is 4-in/4-out (encoder.conv1 takes 4 channels, the decoder head emits 4), and LoadImage's MASK output re-attaches through Join Image With Alpha, since that mask is already 1 - alpha. The rebuilt sole grows past the old outline, so the alpha in that region is generated rather than copied. LanPaint_ImageDecode now matches the decoded channels to the source image: the 2.1 VAE always emits a 4th channel, which the merge could not broadcast. It keeps the input's channel count, so an RGBA source comes back RGBA and an RGB source still comes back RGB.
219 KiB
1024x1024px
219 KiB
1024x1024px