diff --git a/README.md b/README.md index 7465fc2..9789dce 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ # ComfyUI-ExternalAPI-Helpers -A collection of powerful custom nodes for ComfyUI that connect your local workflows to closed-source AI models via their APIs. Use Google's Gemini, OpenAI's GPT-Image-1, and Black Forest Labs' FLUX models directly within ComfyUI. +A collection of powerful custom nodes for ComfyUI that connect your local workflows to closed-source AI models via their APIs. Use Google's Gemini, Imagen, Veo, OpenAI's GPT-Image-1, and Black Forest Labs' FLUX models directly within ComfyUI. @@ -8,7 +8,11 @@ A collection of powerful custom nodes for ComfyUI that connect your local workfl * **FLUX Kontext Pro & Max:** Image-to-image transformations using the FLUX models via the Replicate API. * **Gemini Chat:** Google's powerful multimodal AI. Ask questions about an image, generate detailed descriptions or create prompts for other models. Supports thinking budget controls for applicable models. +* **Gemini Segmentation:** Generate segmentation masks for objects in an image using Gemini. * **GPT Image Edit:** OpenAI's `gpt-image-1` for prompt-based image editing and inpainting. Simply mask an area and describe the change you want to see. +* **Google Imagen Generator & Edit:** Create and edit images with Google's Imagen models, with support for Vertex AI. +* **Nano Banana:** A creative image generation node using a specialized Gemini model. +* **Veo Text-to-Video:** Generate high-quality video clips from text prompts using Google's Veo model via Vertex AI. * **Seamless Integration:** All nodes are designed to work seamlessly with standard ComfyUI inputs (IMAGE, MASK, STRING) and outputs, allowing you to chain them into complex and creative workflows. * **Secure & Simple:** Simply provide your API key in the node's input field to get started. @@ -40,10 +44,11 @@ A collection of powerful custom nodes for ComfyUI that connect your local workfl All nodes in this collection require API keys to function. * **FLUX Nodes (Replicate):** You will need a [Replicate API Token](https://replicate.com/account/api-tokens). -* **Gemini Chat Node:** You will need a [Google AI Studio API Key](https://aistudio.google.com/app/api-keys). +* **Gemini, Imagen, and Nano Banana Nodes:** You will need a [Google AI Studio API Key](https://aistudio.google.com/app/api-keys). * **GPT Image Edit Node:** You will need an [OpenAI API Key](https://platform.openai.com/api-keys). +* **Vertex AI Nodes (Imagen Edit, Veo):** You will need a Google Cloud Project ID, a service account with appropriate permissions, and the location for the resources. -You can paste your key directly into the `api_key` or `replicate_api_token` field on the corresponding node. +You can paste your key directly into the `api_key` field on the corresponding node. For Vertex AI nodes, you will need to provide the project ID, location, and path to your service account JSON file. --- @@ -53,7 +58,7 @@ You can paste your key directly into the `api_key` or `replicate_api_token` fiel These nodes allow you to transform an input image based on a text prompt. They are ideal for applying artistic styles or making significant conceptual changes to an existing image. -* **Category:** `image/generation` +* **Category:** `image/edit` * **Inputs:** * `image`: The source image to transform. * `prompt`: A text description of the desired output (e.g., "A vibrant Van Gogh painting", "Make this a 90s cartoon"). @@ -68,7 +73,7 @@ These nodes allow you to transform an input image based on a text prompt. They a A versatile node for text generation and image analysis. Use it to understand an image's content or to generate creative text for other nodes. -* **Category:** `AI/Gemini` +* **Category:** `text/generation` * **Inputs:** * `prompt`: The text prompt or question you want to ask the model. * `image` (Optional): An input image for the model to analyze. @@ -80,12 +85,25 @@ A versatile node for text generation and image analysis. Use it to understand an * **Output:** * `response`: The text generated by the Gemini model. +### Gemini Segmentation + +This node uses a Gemini model to generate segmentation masks for specified objects within an image. + +* **Category:** `image/generation` +* **Inputs:** + * `image`: The source image for segmentation. + * `segment_prompt`: A text description of the objects to segment (e.g., "the car", "all people"). + * `api_key`: Your API key from Google AI Studio. + * `model`: The Gemini model to use. + * `...other_params`: Controls for temperature, thinking, and seed. +* **Output:** + * `mask`: A black and white mask of the segmented objects. ### GPT Image Edit This node uses OpenAI's API to perform powerful, prompt-based inpainting and editing. -* **Category:** `image/ai` +* **Category:** `image/edit` * **Inputs:** * `image`: The source image to edit. * `mask` (Optional): A black and white mask. The model will edit the **white area** of the mask. @@ -97,6 +115,63 @@ This node uses OpenAI's API to perform powerful, prompt-based inpainting and edi **Note:** If a mask is provided, the edits will be constrained to the masked region. If no mask is provided, the model will attempt to edit the entire image based on the prompt. +### Google Imagen Generator + +Generate images from a text prompt using Google's Imagen models. + +* **Category:** `image/generation` +* **Inputs:** + * `prompt`: A text description of the image to generate. + * `api_key`: Your API key from Google AI Studio. + * `model`: The Imagen model to use. + * `...other_params`: Options for number of images, aspect ratio, and image size. +* **Output:** + * `images`: The generated image(s). + +### Google Imagen Edit (Vertex AI only) + +Perform advanced image editing, inpainting, outpainting, and background swapping using Imagen on Google's Vertex AI platform. + +* **Category:** `image/edit` +* **Inputs:** + * `image`: The source image to edit. + * `mask`: A mask defining the area to edit. + * `prompt`: A description of the desired edit. + * `project_id`: Your Google Cloud Project ID. + * `location`: The Google Cloud location for the model. + * `service_account`: Path to your Google Cloud service account JSON file. + * `edit_mode`: The type of edit to perform (e.g., inpainting, outpainting). + * `...other_params`: Controls for negative prompt, seed, and steps. +* **Output:** + * `edited_images`: The edited image(s). + +### Nano Banana + +A creative image generation node that can take a combination of text and up to five images as input. + +* **Category:** `image/generation` +* **Inputs:** + * `api_key`: Your API key from Google AI Studio. + * `prompt` (Optional): A text prompt. + * `image_1` to `image_5` (Optional): Up to five source images. + * `...other_params`: Controls for aspect ratio, temperature, top_p, and seed. +* **Output:** + * `image`: The generated image. + +### Veo Text-to-Video (Vertex AI) + +Generate short, high-quality video clips from a text description using Google's Veo model on Vertex AI. + +* **Category:** `video/generation` +* **Inputs:** + * `prompt`: A text description of the video to generate. + * `project_id`: Your Google Cloud Project ID. + * `location`: The Google Cloud location for the model. + * `service_account`: Path to your Google Cloud service account JSON file. + * `...other_params`: Controls for negative prompt, aspect ratio, audio generation, and seed. +* **Output:** + * `frames`: The generated video frames, output as an image batch. + --- diff --git a/__init__.py b/__init__.py index ac92ac3..c501b35 100644 --- a/__init__.py +++ b/__init__.py @@ -6,9 +6,10 @@ from .imagen import NODE_CLASS_MAPPINGS as IMAGEN_IMAGE_MAPPINGS, NODE_DISPLAY_N from .imagen_edit import NODE_CLASS_MAPPINGS as IMAGEN_EDIT_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as IMAGEN_EDIT_DISPLAY from .veo import NODE_CLASS_MAPPINGS as VEO_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as VEO_DISPLAY from .gemini_segment import NODE_CLASS_MAPPINGS as GEMINI_SEGMENT_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as GEMINI_SEGMENT_DISPLAY +from .nano_banana import NODE_CLASS_MAPPINGS as NANO_BANANA_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as NANO_BANANA_DISPLAY # Combine both mappings -NODE_CLASS_MAPPINGS = {**PRO_MAPPINGS, **MAX_MAPPINGS, **GEMINI_MAPPINGS, **GPT_IMAGE_MAPPINGS, **IMAGEN_IMAGE_MAPPINGS, **IMAGEN_EDIT_MAPPINGS, **VEO_MAPPINGS, **GEMINI_SEGMENT_MAPPINGS} -NODE_DISPLAY_NAME_MAPPINGS = {**PRO_DISPLAY, **MAX_DISPLAY, **GEMINI_DISPLAY, **GPT_IMAGE_DISPLAY, **IMAGEN_IMAGE_DISPLAY, **IMAGEN_EDIT_DISPLAY, **VEO_DISPLAY, **GEMINI_SEGMENT_DISPLAY} +NODE_CLASS_MAPPINGS = {**PRO_MAPPINGS, **MAX_MAPPINGS, **GEMINI_MAPPINGS, **GPT_IMAGE_MAPPINGS, **IMAGEN_IMAGE_MAPPINGS, **IMAGEN_EDIT_MAPPINGS, **VEO_MAPPINGS, **GEMINI_SEGMENT_MAPPINGS, **NANO_BANANA_MAPPINGS} +NODE_DISPLAY_NAME_MAPPINGS = {**PRO_DISPLAY, **MAX_DISPLAY, **GEMINI_DISPLAY, **GPT_IMAGE_DISPLAY, **IMAGEN_IMAGE_DISPLAY, **IMAGEN_EDIT_DISPLAY, **VEO_DISPLAY, **GEMINI_SEGMENT_DISPLAY, **NANO_BANANA_DISPLAY} __all__ = ['NODE_CLASS_MAPPINGS', 'NODE_DISPLAY_NAME_MAPPINGS'] \ No newline at end of file diff --git a/flux_kontext_max_node.py b/flux_kontext_max_node.py index 7600368..76d1a65 100644 --- a/flux_kontext_max_node.py +++ b/flux_kontext_max_node.py @@ -16,7 +16,7 @@ class FluxKontextMaxNode: "multiline": True, "default": "Make this a 90s cartoon" }), - "replicate_api_token": ("STRING", { + "api_key": ("STRING", { "default": "" }), "aspect_ratio": (["1:1", "16:9", "9:16", "4:3", "3:4", "3:2", "2:3", "5:4", "4:5", "21:9", "9:21", "2:1", "1:2", "match_input_image"], { @@ -37,11 +37,11 @@ class FluxKontextMaxNode: RETURN_TYPES = ("IMAGE",) RETURN_NAMES = ("image",) FUNCTION = "generate_image" - CATEGORY = "image/generation" + CATEGORY = "image/edit" - def generate_image(self, image, prompt, replicate_api_token, aspect_ratio, output_format, safety_tolerance): + def generate_image(self, image, prompt, api_key, aspect_ratio, output_format, safety_tolerance): try: - os.environ["REPLICATE_API_TOKEN"] = replicate_api_token + os.environ["REPLICATE_API_TOKEN"] = api_key # Convert tensor to PIL and save to buffer tensor = image.squeeze(0) if len(image.shape) == 4 else image diff --git a/flux_kontext_pro_node.py b/flux_kontext_pro_node.py index b47ac60..faaaee3 100644 --- a/flux_kontext_pro_node.py +++ b/flux_kontext_pro_node.py @@ -16,7 +16,7 @@ class FluxKontextProNode: "multiline": True, "default": "Make this a 90s cartoon" }), - "replicate_api_token": ("STRING", { + "api_key": ("STRING", { "default": "" }), "aspect_ratio": (["1:1", "16:9", "9:16", "4:3", "3:4", "3:2", "2:3", "5:4", "4:5", "21:9", "9:21", "2:1", "1:2", "match_input_image"], { @@ -37,11 +37,11 @@ class FluxKontextProNode: RETURN_TYPES = ("IMAGE",) RETURN_NAMES = ("image",) FUNCTION = "generate_image" - CATEGORY = "image/generation" + CATEGORY = "image/edit" - def generate_image(self, image, prompt, replicate_api_token, aspect_ratio, output_format, safety_tolerance): + def generate_image(self, image, prompt, api_key, aspect_ratio, output_format, safety_tolerance): try: - os.environ["REPLICATE_API_TOKEN"] = replicate_api_token + os.environ["REPLICATE_API_TOKEN"] = api_key # Convert tensor to PIL and save to buffer tensor = image.squeeze(0) if len(image.shape) == 4 else image diff --git a/gemini_node.py b/gemini_node.py index 0f08caf..24cabdc 100644 --- a/gemini_node.py +++ b/gemini_node.py @@ -32,7 +32,7 @@ class GeminiChatNode: RETURN_TYPES = ("STRING",) RETURN_NAMES = ("response",) FUNCTION = "generate" - CATEGORY = "AI/Gemini" + CATEGORY = "text/generation" def generate(self, prompt: str, model: str, temperature: float, thinking: bool, seed: int, api_key: str, system_instruction: Optional[str] = None, thinking_budget: int = -1, diff --git a/gemini_segment.py b/gemini_segment.py index e70764c..59f3414 100644 --- a/gemini_segment.py +++ b/gemini_segment.py @@ -31,7 +31,7 @@ class GeminiSegmentationNode: RETURN_TYPES = ("MASK",) RETURN_NAMES = ("mask",) FUNCTION = "generate_segmentation" - CATEGORY = "AI/Gemini" + CATEGORY = "image/generation" def generate_segmentation(self, image: torch.Tensor, segment_prompt: str, model: str, temperature: float, thinking: bool, seed: int, api_key: str, diff --git a/gpt_image1.py b/gpt_image1.py index e06cae7..cff3ee6 100644 --- a/gpt_image1.py +++ b/gpt_image1.py @@ -53,7 +53,7 @@ class GPTImageEditNode: RETURN_TYPES = ("IMAGE",) RETURN_NAMES = ("image",) FUNCTION = "edit_image" - CATEGORY = "image/ai" + CATEGORY = "image/edit" def tensor_to_pil(self, tensor): if len(tensor.shape) == 3: diff --git a/imagen.py b/imagen.py index 8a83552..95eb46c 100644 --- a/imagen.py +++ b/imagen.py @@ -23,7 +23,7 @@ class GoogleImagenNode: RETURN_TYPES = ("IMAGE",) RETURN_NAMES = ("images",) FUNCTION = "generate_images" - CATEGORY = "image/ai" + CATEGORY = "image/generation" def pil_to_tensor(self, images): if not isinstance(images, list): diff --git a/imagen_edit.py b/imagen_edit.py index 8bbb4ab..67c208a 100644 --- a/imagen_edit.py +++ b/imagen_edit.py @@ -31,7 +31,7 @@ class GoogleImagenEditNode: RETURN_TYPES = ("IMAGE",) RETURN_NAMES = ("edited_images",) FUNCTION = "edit_image" - CATEGORY = "image/ai" + CATEGORY = "image/edit" def tensor_to_pil(self, tensor): array = (tensor.cpu().numpy() * 255).astype(np.uint8) diff --git a/nano_banana.py b/nano_banana.py new file mode 100644 index 0000000..0ec5757 --- /dev/null +++ b/nano_banana.py @@ -0,0 +1,119 @@ +import os +import io +import torch +import numpy as np +from PIL import Image +from google import genai +from google.genai import types + +class NanoBananaNode: + + @classmethod + def INPUT_TYPES(cls): + return { + "required": { + "api_key": ("STRING", {"multiline": False, "default": ""}), + "aspect_ratio": (["1:1", "2:3", "3:2", "3:4", "4:3", "9:16", "16:9", "21:9"],), + "temperature": ("FLOAT", {"default": 0.5, "min": 0.0, "max": 1.0, "step": 0.01}), + "top_p": ("FLOAT", {"default": 0.85, "min": 0.0, "max": 1.0, "step": 0.01}), + "seed": ("INT", {"default": 69, "min": -1, "max": 2147483646, "step": 1}), + }, + "optional": { + "prompt": ("STRING", {"multiline": True, "default": ""}), + "system_instruction": ("STRING", {"multiline": True, "default": ""}), + "image_1": ("IMAGE",), + "image_2": ("IMAGE",), + "image_3": ("IMAGE",), + "image_4": ("IMAGE",), + "image_5": ("IMAGE",), + } + } + + RETURN_TYPES = ("IMAGE",) + RETURN_NAMES = ("image",) + FUNCTION = "generate" + CATEGORY = "image/generation" + + def tensor_to_pil(self, tensor): + if tensor.dim() == 4: + tensor = tensor[0] + array = (tensor.cpu().numpy() * 255).astype(np.uint8) + return Image.fromarray(array) + + def pil_to_tensor(self, image): + if image.mode != 'RGB': + image = image.convert('RGB') + array = np.array(image).astype(np.float32) / 255.0 + tensor = torch.from_numpy(array) + return tensor.unsqueeze(0) + + def generate(self, api_key, aspect_ratio, temperature, top_p, seed, + prompt="", system_instruction="", + image_1=None, image_2=None, image_3=None, image_4=None, image_5=None): + + key = api_key.strip() or os.environ.get("GEMINI_API_KEY") + if not key: + raise ValueError("No API key provided.") + + client = genai.Client(api_key=key) + + # Build parts list + parts = [] + + # Add images + for img_tensor in [image_1, image_2, image_3, image_4, image_5]: + if img_tensor is not None: + pil_img = self.tensor_to_pil(img_tensor) + buffer = io.BytesIO() + pil_img.save(buffer, format='PNG') + parts.append(types.Part.from_bytes( + mime_type="image/png", + data=buffer.getvalue() + )) + + # Add prompt if provided + if prompt.strip(): + parts.append(types.Part.from_text(text=prompt)) + + if not parts: + raise ValueError("At least one image or prompt must be provided.") + + contents = [types.Content(role="user", parts=parts)] + + # Build config + config = types.GenerateContentConfig( + temperature=temperature, + seed=seed, + top_p=top_p, + response_modalities=["IMAGE"], + image_config=types.ImageConfig(aspect_ratio=aspect_ratio), + ) + + if system_instruction.strip(): + config.system_instruction = [types.Part.from_text(text=system_instruction)] + + # Generate + response = client.models.generate_content( + model="gemini-2.5-flash-image", + contents=contents, + config=config, + ) + + result_image = None + if response.candidates and response.candidates[0].content and response.candidates[0].content.parts: + for part in response.candidates[0].content.parts: + if part.inline_data and part.inline_data.data: + result_image = Image.open(io.BytesIO(part.inline_data.data)) + break + + if result_image is None: + raise ValueError("No image generated by the API.") + + return (self.pil_to_tensor(result_image),) + + @classmethod + def IS_CHANGED(cls, **kwargs): + return f"{kwargs.get('prompt', '')}-{kwargs.get('temperature', 0.5)}-{kwargs.get('top_p', 0.85)}-{kwargs.get('seed', 69)}-{kwargs.get('aspect_ratio', '1:1')}-{kwargs.get('image_1')}-{kwargs.get('image_2')}-{kwargs.get('image_3')}-{kwargs.get('image_4')}-{kwargs.get('image_5')}" + +NODE_CLASS_MAPPINGS = {"NanoBananaNode": NanoBananaNode} +NODE_DISPLAY_NAME_MAPPINGS = {"NanoBananaNode": "Nano Banana"} \ No newline at end of file