[Added] Foreground estimation (2 nodes) and Empty Image (from ref)

This commit is contained in:
Salvador E. Tropea
2025-09-30 10:05:32 -03:00
parent 87ee4aebb7
commit 2ddd61ee87
2 changed files with 285 additions and 3 deletions
+60 -1
View File
@@ -24,6 +24,9 @@ Currently we just have a few nodes used by other nodes I maintain.
- [Normalize Image to [-0.5, 0.5]](#5-normalize-image-to-05-05)
- [Normalize Image to [-1, 1]](#6-normalize-image-to-1-1)
- [Apply Mask using AFFCE](#7-apply-mask-using-affce)
- [Estimate foreground (AFFCE)](#8-estimate-foreground-affce)
- [Estimate foreground (FMLFE)](#9-estimate-foreground-fmlfe)
- [Create Empty Image](#10-create-empty-image)
- 📝 [Usage Notes](#-usage-notes)
- 📜 [Project History](#-project-history)
- ⚖️ [License](#️-license)
@@ -134,7 +137,7 @@ Currently we just have a few nodes used by other nodes I maintain.
- **Display Name:** `Apply Mask using AFFCE`
- **Internal Name:** `SET_ApplyMaskAFFCE`
- **Category:** `image/manipulation
- **Category:** `image/manipulation`
- **Description:** Applies a mask to an image using [Approximate Fast Foreground Colour Estimation](https://github.com/Photoroom/fast-foreground-estimation). This blends the image contour in a better way.
- **Purpose:** Used to apply the mask of a background removal model.
- **Inputs:**
@@ -150,6 +153,62 @@ Currently we just have a few nodes used by other nodes I maintain.
- `mask` (`MASK`): The input mask
### 8. Estimate foreground (AFFCE)
- **Display Name:** `Estimate foreground (AFFCE)`
- **Internal Name:** `SET_AFFCE`
- **Category:** `image/foreground`
- **Description:** Estimates the foreground of an image using [Approximate Fast Foreground Colour Estimation](https://github.com/Photoroom/fast-foreground-estimation). The result is suitable for background replacement.
- **Purpose:** Used to get a better foreground for background removal.
- **Inputs:**
- `images` (`IMAGE`): One ore more ComfyUI images
- `masks` (`MASK`): Masks to apply
- `blur_size` (`INT`): Diameter for the coarse gaussian blur
- `blur_size_two` (`INT`): Diameter for the fine gaussian blur
- `batched` (`BOOLEAN`): Process all the images at once. Otherwise do it one at a time.
- **Output:**
- `image` (`IMAGE`): The image after foreground estimation.
- `mask` (`MASK`): The input mask
### 9. Estimate foreground (FMLFE)
- **Display Name:** `Estimate foreground (FMLFE)`
- **Internal Name:** `SET_FMLFE`
- **Category:** `image/foreground`
- **Description:** Estimates the foreground of an image using [Fast Multi-Level Foreground Estimation](https://arxiv.org/abs/2006.14970). The result is suitable for background replacement.
- **Purpose:** Used to get a better foreground for background removal.
- **Inputs:**
- `images` (`IMAGE`): The source image(s) from which to estimate the foreground and background.
- `masks` (`MASK`): The alpha matte that guides the estimation. White areas are treated as known foreground, black as known background, and gray areas are the semi-transparent regions the algorithm will solve for.
- `implementation" (`auto`, `cupy`, `opencl`, `numba`, `torch`): Which implementation to use. The `auto` will use the fastest available. The `torch` implementation is slow and approximated, but doesn't need extra dependencies. The `numba` implementation is good and just needs [Numba](https://numba.pydata.org/). The `cupy` implementation is the fastest, but needs [CuPy](https://cupy.dev/) and a full CUDA environment.
- `regularization` (`FLOAT`): The regularization strength (epsilon). This acts as a smoothness prior. Higher values result in smoother, more blended foreground and background colors, but may lose very fine details. Lower values preserve more detail but can be noisier.
- `n_small_iterations` (`INT`): The number of solver iterations to perform on the lower-resolution levels of the image pyramid. More iterations can improve quality at the cost of speed.
- `n_big_iterations` (`INT`): The number of solver iterations to perform on the higher-resolution (larger) levels of the image pyramid. Fewer iterations are typically needed at high resolution as the details are propagated up from the smaller levels.
- `small_size` (`INT`): The pixel dimension threshold. Image pyramid levels smaller than this size will use the higher 'n_small_iterations' count, while larger levels will use 'n_big_iterations'.
- `gradient_weight` (`FLOAT`): Controls how strongly the edges in the alpha matte influence color blending. A higher value makes the algorithm respect the mask's edges more, leading to sharper color boundaries. A lower value allows more color bleeding, an effect similar to increasing regularization.
- **Output:**
- `image` (`IMAGE`): The estimated foreground image (F). This is a full-color image where the algorithm has estimated the true, un-blended color of the foreground object.
- `mask` (`MASK`): The estimated background image (B). The algorithm has effectively "inpainted" the area behind the foreground object, creating a clean background plate.
### 10. Create Empty Image
- **Display Name:** `Create Empty Image`
- **Internal Name:** `SET_CreateEmptyImage`
- **Category:** `image/generation`
- **Description:** Creates an image filled with a solid color. Similar to standard `EmptyImage`, but you can use an image as reference and you have more options to select the color.
- **Purpose:** Create an empty image
- **Inputs:**
- `width` (`INT`): The width of the new image in pixels. This value is ignored if a `reference` is provided.
- `height` (`INT`): The height of the new image in pixels. This value is ignored if a `reference` is provided.
- `batch_size` (`INT`): The number of images to create in the batch. This value is ignored if a `reference` is provided.
- `color` (`STRING`): The solid color to fill the image with. Can be a named color (e.g., "black", "red"), a hex string (e.g., "#FF0000", "00ff00"), or comma-separated components in the [0, 255] or [0.0, 1.0] range.
- `reference` (`IMAGE`, optional): If an image is connected here, its dimensions (batch size, height, and width) will be used for the new image, overriding the manual width, height, and batch_size inputs.
- **Output:**
- `image` (`IMAGE`): A new image tensor of the specified size and color.
## 🚀 Installation
You can install the nodes from the ComfyUI nodes manager, the name is *Image Misc*, or just do it manually:
+225 -2
View File
@@ -8,6 +8,9 @@ import numpy as np
import os
from PIL import Image # Import the Python Imaging Library
from seconohe.apply_mask import apply_mask
from seconohe.foreground_estimation.affce import affce
from seconohe.foreground_estimation.fmlfe import fmlfe, IMPL_PRIORITY
from seconohe.color import color_to_rgb_float
from seconohe.downloader import download_file
# We are the main source, so we use the main_logger
from . import main_logger
@@ -33,6 +36,7 @@ BASE_CATEGORY = "image"
IO_CATEGORY = "io"
MANIPULATION_CATEGORY = "manipulation"
NORMALIZATION = "normalization"
FOREGROUND = "foreground"
BLUR_SIZE_OPT = ("INT", {"default": 90, "min": 1, "max": 255, "step": 1, })
BLUR_SIZE_TWO_OPT = ("INT", {"default": 6, "min": 1, "max": 255, "step": 1, })
COLOR_OPT = ("STRING", {
@@ -138,10 +142,10 @@ if has_load_image:
# by the ComfyUI widget, which is just the filename. It internally
# resolves the path using folder_paths.
logger.debug(f"Calling built-in LoadImage.load_image() with filename: '{filename}'")
logger.debug(f"Calling built-in LoadImage.load_image() with filename: '{dest_fname}'")
# Call the method and return its result directly
result = loader_instance.load_image(filename)
result = loader_instance.load_image(dest_fname)
# This information is for the preview, as we are an output node and we return images
# they will be displayed in our node. Quite simple.
downloaded_file = {
@@ -440,3 +444,222 @@ class ApplyMaskAFFCE:
out_images = apply_mask(logger, images, masks, model_management.get_torch_device(), blur_size, blur_size_two,
fill_color, color, batched)
return out_images.cpu(), masks.cpu()
class AFFCE:
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"images": ("IMAGE",),
"masks": ("MASK",),
"blur_size": BLUR_SIZE_OPT,
"blur_size_two": BLUR_SIZE_TWO_OPT,
"batched": ("BOOLEAN", {
"default": True,
"tooltip": ("Process the images at once.\n"
"Faster, needs more memory")
}),
}
}
RETURN_TYPES = ("IMAGE", "MASK",)
RETURN_NAMES = ("foreground", "mask",)
FUNCTION = "get_foreground"
CATEGORY = BASE_CATEGORY + "/" + FOREGROUND
DESCRIPTION = ("Estimate the foreground image using\n"
"Approximate Fast Foreground Colour Estimation.\n"
"https://github.com/Photoroom/fast-foreground-estimation")
UNIQUE_NAME = "SET_AFFCE"
DISPLAY_NAME = "Estimate foreground (AFFCE)"
def get_foreground(self, images, masks, blur_size=91, blur_size_two=7, batched=True):
device = model_management.get_torch_device()
images_on_device = images.to(device)
masks_on_device = masks.to(device)
out_images = affce(images_on_device, masks_on_device, r1=blur_size, r2=blur_size_two, batched=batched)
return out_images.cpu(), masks.cpu()
class FMLFE:
"""
A ComfyUI node that uses the Fast Multi-Level Foreground Estimation algorithm
to produce a high-quality foreground and background separation. It can
intelligently select the best available backend (CuPy, OpenCL, Numba, or PyTorch).
"""
@classmethod
def INPUT_TYPES(cls):
# Create the dropdown list for the implementation choice
impl_list = ['auto'] + IMPL_PRIORITY
return {
"required": {
"images": ("IMAGE", {
"tooltip": "The source image(s) from which to estimate the foreground and background."
}),
"masks": ("MASK", {
"tooltip": "The alpha matte that guides the estimation. White areas are treated as known "
"foreground, black as known background, and gray areas are the semi-transparent "
"regions the algorithm will solve for."
}),
"implementation": (impl_list, {
"default": "auto",
"tootip": "Select the computation backend. 'auto' mode will automatically try to use the "
"fastest available implementation, in order of priority: CuPy (NVIDIA GPU), "
"OpenCL (GPU), Numba (CPU/GPU), and finally the pure PyTorch version."
}),
},
"optional": {
"regularization": ("FLOAT", {
"default": 1e-5,
"min": 0.0,
"max": 0.1,
"step": 1e-5,
"display": "number",
"tooltip": "The regularization strength (epsilon). This acts as a smoothness prior. "
"Higher values result in smoother, more blended foreground and background colors, "
"but may lose very fine details. Lower values preserve more detail but can be noisier."
}),
"n_small_iterations": ("INT", {
"default": 10,
"min": 1,
"max": 100,
"tooltip": "The number of solver iterations to perform on the lower-resolution levels of the "
"image pyramid. More iterations can improve quality at the cost of speed."
}),
"n_big_iterations": ("INT", {
"default": 2,
"min": 1,
"max": 100,
"tooltip": "The number of solver iterations to perform on the higher-resolution (larger) levels "
"of the image pyramid. Fewer iterations are typically needed at high resolution as the "
"details are propagated up from the smaller levels."
}),
"small_size": ("INT", {
"default": 32,
"min": 8,
"max": 256,
"tooltip": "The pixel dimension threshold. Image pyramid levels smaller than this size will use "
"the higher 'n_small_iterations' count, while larger levels will use 'n_big_iterations'."
}),
"gradient_weight": ("FLOAT", {
"default": 1.0,
"min": 0.0,
"max": 10.0,
"step": 0.1,
"tooltip": "Controls how strongly the edges in the alpha matte influence color blending. "
"A higher value makes the algorithm respect the mask's edges more, leading to sharper "
"color boundaries. A lower value allows more color bleeding, an effect similar to "
"increasing regularization."
}),
}
}
RETURN_TYPES = ("IMAGE", "IMAGE", "MASK",)
RETURN_NAMES = ("foreground", "background", "mask")
FUNCTION = "estimate"
CATEGORY = BASE_CATEGORY + "/" + FOREGROUND
DESCRIPTION = ("Estimate the foreground image using\n"
"Fast Multi-Level Foreground Estimation.")
UNIQUE_NAME = "SET_FMLFE"
DISPLAY_NAME = "Estimate foreground (FMLFE)"
def estimate(self, images: torch.Tensor, masks: torch.Tensor, implementation: str,
regularization: float, n_small_iterations: int, n_big_iterations: int,
small_size: int, gradient_weight: float):
try:
foregrounds, backgrounds = fmlfe(
images=images,
masks=masks,
logger=logger,
implementation=implementation,
regularization=regularization,
n_small_iterations=n_small_iterations,
n_big_iterations=n_big_iterations,
small_size=small_size,
gradient_weight=gradient_weight
)
return (foregrounds, backgrounds, masks,)
except Exception as e:
# This ensures that if all backends fail, the error is clearly visible in the ComfyUI console.
logger.error("Failed to execute ML Foreground Estimation. All backends failed.")
logger.error(f"Last error: {e}")
# Raising the exception will stop the workflow and show the error to the user.
raise e
class CreateEmptyImage:
"""
A ComfyUI node to create a solid-color image tensor.
The output dimensions can be specified manually or inherited from an optional input image.
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"width": ("INT", {
"default": 1024,
"min": 1,
"max": 8192,
"step": 8,
"tooltip": "The width of the new image in pixels. This value is ignored if a `reference` is provided."
}),
"height": ("INT", {
"default": 1024,
"min": 1,
"max": 8192,
"step": 8,
"tooltip": "The height of the new image in pixels. This value is ignored if a `reference` is provided."
}),
"batch_size": ("INT", {
"default": 1,
"min": 1,
"max": 64,
"tooltip": "The number of images to create in the batch. This value is ignored if a `reference` "
"is provided."
}),
"color": COLOR_OPT,
},
"optional": {
"reference": ("IMAGE", {
"tooltip": "If an image is connected here, its dimensions (batch size, height, and width) will be "
"used for the new image, overriding the manual width, height, and batch_size inputs."
}),
}
}
RETURN_TYPES = ("IMAGE",)
RETURN_NAMES = ("image",)
FUNCTION = "create_image"
CATEGORY = BASE_CATEGORY + "/generation"
DESCRIPTION = ("Create a solid-color image.\n"
"If the optional image is provides uses its shape.")
UNIQUE_NAME = "SET_CreateEmptyImage"
DISPLAY_NAME = "Create Empty Image"
def create_image(self, width: int, height: int, batch_size: int, color: str,
reference: Optional[torch.Tensor] = None):
# --- 1. Determine the final shape of the output tensor ---
if reference is not None:
# If an image is provided, its shape overrides the manual inputs
b, h, w, _ = reference.shape
else:
b, h, w = batch_size, height, width
# --- 2. Parse the color string ---
# The function returns a tuple of floats in the [0, 1] range
rgb_color = color_to_rgb_float(logger, color)
# --- 3. Create the tensor efficiently ---
# Create a small color tensor and then expand it to the final size.
# This is highly memory-efficient as it creates a view, not a full-size copy.
# Tensors should be created on the CPU by default in generator nodes.
color_tensor = torch.tensor(rgb_color, dtype=torch.float32, device="cpu").view(1, 1, 1, 3)
final_image = color_tensor.expand(b, h, w, 3)
return (final_image,)