feat: add Gemini Prompt Engineer node

- Add GeminiPromptNode for AI-powered prompt engineering
- Integrates with Google's Gemini API for prompt generation
- Includes various prompt templates and generation modes
- Add comprehensive tests and documentation
- Register node in ComfyAssets category
This commit is contained in:
Vito Sansevero
2025-08-01 09:41:21 -07:00
parent 932e30ade0
commit f559fe220e
8 changed files with 1075 additions and 0 deletions
+162
View File
@@ -0,0 +1,162 @@
# Gemini Prompt Engineer
The Gemini Prompt Engineer node uses Google's Gemini AI to analyze images and generate optimized prompts for various AI image generation models.
## Features
- **Multi-Model Support**: Generate prompts optimized for FLUX, SDXL, Danbooru, and Video generation
- **Custom Prompts**: Override templates with your own system prompts
- **Visual Feedback**: UI shows processing status and error states
- **Flexible API Key Management**: Multiple ways to provide API credentials
## Setup
### 1. Get API Key
Get your free Gemini API key from [Google AI Studio](https://makersuite.google.com/app/apikey)
### 2. Install Dependencies
```bash
pip install google-generativeai
```
### 3. Configure API Key
Choose one of these methods:
1. **Environment Variable** (Recommended):
```bash
export GEMINI_API_KEY="your-api-key-here"
```
2. **Config File**:
Create `gemini_config.json` in your ComfyUI root directory:
```json
{
"api_key": "your-api-key-here"
}
```
3. **Node Input**:
Enter the API key directly in the node's `api_key` field
## Inputs
- **image** (IMAGE): The image to analyze
- **prompt_type** (DROPDOWN): Type of prompt to generate
- `flux`: Detailed artistic prompts with quality markers
- `sdxl`: Positive/negative prompt pairs with weight emphasis
- `danbooru`: Anime-style booru tags with underscores
- `video`: Motion and temporal descriptions for video generation
- **api_key** (STRING, optional): Gemini API key if not set elsewhere
- **custom_prompt** (STRING, optional): Override template with custom system prompt
## Outputs
- **prompt** (STRING): Generated prompt text
- **negative_prompt** (STRING): Negative prompt (only populated for SDXL format)
## Prompt Type Details
### FLUX Format
Generates detailed prompts optimized for FLUX models:
- Starts with main subject and action
- Includes style and medium descriptors
- Adds lighting and atmosphere details
- Uses quality markers like "4K", "highly detailed", "award-winning"
Example output:
```
majestic mountain landscape at golden hour, oil painting style, dramatic lighting with sun rays piercing through clouds, wide angle composition, warm color palette with orange and purple hues, highly detailed, 4K resolution, trending on ArtStation, photorealistic rendering
```
### SDXL Format
Generates positive and negative prompt pairs:
- Detailed positive prompts with weight emphasis
- Comprehensive negative prompts to avoid common issues
- Uses parentheses for emphasis: `(detailed eyes:1.2)`
Example output:
```
Positive: beautiful woman, (detailed eyes:1.2), flowing red dress, golden hour lighting, professional photography, 85mm lens, shallow depth of field, bokeh, high resolution, masterpiece
Negative: low quality, blurry, distorted features, bad anatomy, poorly drawn, amateur, oversaturated, jpeg artifacts
```
### Danbooru Format
Generates booru-style tags for anime artwork:
- Uses underscores for multi-word concepts
- Includes character count descriptors (1girl, 2boys)
- Orders tags from most to least important
Example output:
```
1girl, solo, long_hair, blue_eyes, blonde_hair, school_uniform, serafuku, pleated_skirt, thighhighs, smile, looking_at_viewer, classroom, sitting, desk, window, sunlight, highres, masterpiece
```
### Video Format
Generates prompts for video generation models:
- Describes motion and camera movements
- Includes temporal markers and transitions
- Specifies technical details like fps and duration
Example output:
```
Aerial shot slowly descending toward a misty forest at dawn, camera smoothly transitions to tracking shot following a deer through the trees, photorealistic style, soft golden hour lighting with fog, 10 second duration, 4K resolution 24fps, ending with close-up of deer looking at camera
```
## Custom System Prompts
You can override any template by providing your own system prompt. This is useful for:
- Specialized use cases
- Different language outputs
- Custom formatting requirements
- Integration with specific workflows
Example custom prompt:
```
You are an expert at analyzing images and creating simple, concise descriptions.
Focus only on the main subject and primary colors.
Keep your response under 50 words.
```
## Error Handling
The node provides clear error messages for common issues:
- Missing API key
- API request failures
- Invalid image inputs
- Rate limiting
Errors are displayed in the prompt output for easy debugging.
## Tips
1. **API Usage**: Gemini has generous free tier limits, but be mindful of rate limits
2. **Image Quality**: Higher resolution images provide better analysis results
3. **Prompt Refinement**: You can chain multiple Gemini nodes with different custom prompts
4. **Caching**: Results are not cached, so identical images will make new API calls
## Example Workflow
1. Load an image using Load Image node
2. Connect to Gemini Prompt Engineer
3. Select appropriate prompt_type for your target model
4. Connect prompt output to your generation model
5. For SDXL, connect both prompt and negative_prompt outputs
## Troubleshooting
**"API key not found" error**:
- Check environment variable is set correctly
- Verify config file path and JSON format
- Try entering key directly in node
**"No response generated" error**:
- Check internet connection
- Verify API key is valid
- Image might be too large (resize if needed)
**Import error for google-generativeai**:
- Run `pip install google-generativeai` in your ComfyUI environment
- Restart ComfyUI after installation
+3
View File
@@ -10,6 +10,7 @@ from .tools.sampler_combo import SamplerComboNode, SamplerComboCompactNode
from .tools.empty_latent_batch import EmptyLatentBatchNode
from .tools.kiko_save_image import KikoSaveImageNode
from .tools.image_to_multiple_of import ImageToMultipleOfNode
from .tools.gemini_prompt import GeminiPromptNode
# ComfyUI node registration mappings
NODE_CLASS_MAPPINGS = {
@@ -21,6 +22,7 @@ NODE_CLASS_MAPPINGS = {
"EmptyLatentBatch": EmptyLatentBatchNode,
"KikoSaveImage": KikoSaveImageNode,
"ImageToMultipleOf": ImageToMultipleOfNode,
"GeminiPrompt": GeminiPromptNode,
}
NODE_DISPLAY_NAME_MAPPINGS = {
@@ -32,6 +34,7 @@ NODE_DISPLAY_NAME_MAPPINGS = {
"EmptyLatentBatch": "Empty Latent Batch",
"KikoSaveImage": "Kiko Save Image",
"ImageToMultipleOf": "Image to Multiple of",
"GeminiPrompt": "Gemini Prompt Engineer",
}
__all__ = ["NODE_CLASS_MAPPINGS", "NODE_DISPLAY_NAME_MAPPINGS"]
@@ -0,0 +1,5 @@
"""Gemini Prompt Engineer node for ComfyUI."""
from .node import GeminiPromptNode
__all__ = ["GeminiPromptNode"]
+153
View File
@@ -0,0 +1,153 @@
"""Logic for Gemini API integration and prompt generation."""
import base64
import io
import json
import os
from typing import Optional, Tuple
import numpy as np
from PIL import Image
from .prompts import PROMPT_TEMPLATES
def tensor_to_pil(tensor: np.ndarray) -> Image.Image:
"""Convert ComfyUI tensor to PIL Image.
Args:
tensor: Input tensor in ComfyUI format (B, H, W, C)
Returns:
PIL Image object
"""
# ComfyUI tensors are in [0, 1] range
if tensor.ndim == 4:
# Take first image from batch
tensor = tensor[0]
# Convert to uint8
image_array = (tensor * 255).astype(np.uint8)
# Convert to PIL
return Image.fromarray(image_array, mode='RGB')
def image_to_base64(image: Image.Image, format: str = "PNG") -> str:
"""Convert PIL Image to base64 string.
Args:
image: PIL Image object
format: Image format (PNG or JPEG)
Returns:
Base64 encoded string
"""
buffer = io.BytesIO()
image.save(buffer, format=format)
buffer.seek(0)
return base64.b64encode(buffer.read()).decode('utf-8')
def get_api_key() -> Optional[str]:
"""Get Gemini API key from environment or config.
Returns:
API key string or None if not found
"""
# Check environment variable first
api_key = os.environ.get("GEMINI_API_KEY")
if not api_key:
# Check for config file in ComfyUI directory
try:
config_path = os.path.join(os.path.dirname(__file__), "..", "..", "..", "gemini_config.json")
if os.path.exists(config_path):
with open(config_path, 'r') as f:
config = json.load(f)
api_key = config.get("api_key")
except Exception:
pass
return api_key
def analyze_image_with_gemini(
image: np.ndarray,
prompt_type: str,
api_key: Optional[str] = None,
custom_prompt: Optional[str] = None,
model_name: str = "gemini-1.5-flash"
) -> Tuple[str, Optional[str]]:
"""Analyze image using Gemini API and generate appropriate prompt.
Args:
image: Input image tensor
prompt_type: Type of prompt to generate (flux, sdxl, danbooru, video)
api_key: Gemini API key (optional, will try to get from env/config)
custom_prompt: Custom system prompt to use instead of templates
model_name: Gemini model to use (default: gemini-1.5-flash)
Returns:
Tuple of (generated_prompt, error_message)
"""
# Get API key
if not api_key:
api_key = get_api_key()
if not api_key:
return "", "Gemini API key not found. Please set GEMINI_API_KEY environment variable or provide it in the node."
# Convert tensor to PIL image
try:
pil_image = tensor_to_pil(image)
except Exception as e:
return "", f"Failed to convert image: {str(e)}"
# Get system prompt
if custom_prompt:
system_prompt = custom_prompt
else:
system_prompt = PROMPT_TEMPLATES.get(prompt_type, PROMPT_TEMPLATES["flux"])
# Here we would normally make the API call to Gemini
# For now, we'll import the google-generativeai library
try:
import google.generativeai as genai
except ImportError:
return "", "google-generativeai library not installed. Please run: pip install google-generativeai"
try:
# Configure Gemini
genai.configure(api_key=api_key)
# Create model
model = genai.GenerativeModel(model_name)
# Generate content
response = model.generate_content([
system_prompt,
pil_image,
"Analyze this image and generate an appropriate prompt according to the instructions."
])
# Extract text from response
if response.text:
return response.text.strip(), None
else:
return "", "No response generated from Gemini"
except Exception as e:
return "", f"Gemini API error: {str(e)}"
def validate_prompt_type(prompt_type: str) -> bool:
"""Validate if prompt type is supported.
Args:
prompt_type: Type of prompt to validate
Returns:
True if valid, False otherwise
"""
return prompt_type in PROMPT_TEMPLATES
+111
View File
@@ -0,0 +1,111 @@
"""Gemini Prompt Engineer node implementation."""
import torch
from ...base import ComfyAssetsBaseNode
from .logic import analyze_image_with_gemini, validate_prompt_type
from .prompts import PROMPT_OPTIONS, GEMINI_MODELS
class GeminiPromptNode(ComfyAssetsBaseNode):
"""Analyzes images using Gemini AI to generate optimized prompts for various AI models."""
@classmethod
def INPUT_TYPES(cls):
"""Define input types for the node."""
return {
"required": {
"image": ("IMAGE",),
"prompt_type": (PROMPT_OPTIONS, {"default": "flux"}),
"model": (GEMINI_MODELS, {"default": "gemini-1.5-flash"}),
},
"optional": {
"api_key": ("STRING", {"default": "", "multiline": False}),
"custom_prompt": (
"STRING",
{
"default": "",
"multiline": True,
"placeholder": "Optional: Enter custom system prompt instead of using templates",
},
),
},
}
RETURN_TYPES = ("STRING", "STRING")
RETURN_NAMES = ("prompt", "negative_prompt")
FUNCTION = "generate_prompt"
CATEGORY = "ComfyAssets"
DESCRIPTION = """
Analyzes images using Google's Gemini AI to generate optimized prompts.
Supports multiple prompt formats:
- FLUX: Detailed artistic prompts with quality markers
- SDXL: Positive/negative prompt pairs with weight emphasis
- Danbooru: Anime-style booru tags with underscores
- Video: Motion and temporal descriptions for video generation
Requires Gemini API key (set GEMINI_API_KEY env var or provide in node).
Install: pip install google-generativeai
"""
def generate_prompt(self, image, prompt_type, model, api_key="", custom_prompt=""):
"""Generate prompt from image using Gemini.
Args:
image: Input image tensor
prompt_type: Type of prompt to generate
model: Gemini model to use
api_key: Optional API key
custom_prompt: Optional custom system prompt
Returns:
Tuple of (prompt, negative_prompt)
"""
# Validate prompt type
if not validate_prompt_type(prompt_type):
raise ValueError(f"Invalid prompt type: {prompt_type}")
# Convert torch tensor to numpy if needed
if isinstance(image, torch.Tensor):
image_np = image.cpu().numpy()
else:
image_np = image
# Analyze image with Gemini
prompt, error = analyze_image_with_gemini(
image_np, prompt_type, api_key=api_key or None, custom_prompt=custom_prompt or None, model_name=model
)
if error:
# Return error as prompt for visibility
return (f"Error: {error}", "")
# Handle different prompt types
if prompt_type == "sdxl":
# SDXL returns positive and negative prompts
lines = prompt.split("\n")
positive_prompt = ""
negative_prompt = ""
for line in lines:
if line.startswith("Positive:"):
positive_prompt = line.replace("Positive:", "").strip()
elif line.startswith("Negative:"):
negative_prompt = line.replace("Negative:", "").strip()
# If format not found, assume entire response is positive prompt
if not positive_prompt:
positive_prompt = prompt
return (positive_prompt, negative_prompt)
else:
# Other formats don't use negative prompts
return (prompt, "")
# Node display name
NODE_DISPLAY_NAME = "Gemini Prompt Engineer"
+200
View File
@@ -0,0 +1,200 @@
"""System prompts for different AI model types."""
FLUX_PROMPT = """You are an expert visual analyst and FLUX prompt engineer. Your role is to examine images in detail and create precise, effective prompts that can recreate similar images using the FLUX image generation model.
When analyzing an image, systematically observe and document:
1. **Subject & Composition**
- Primary subjects and their positions
- Background elements and environment
- Overall composition and framing
- Perspective and camera angle
2. **Visual Style & Technique**
- Art style (photorealistic, illustration, painting, etc.)
- Rendering technique (digital art, oil painting, watercolor, etc.)
- Level of detail and texture quality
- Any specific artistic influences or movements
3. **Lighting & Atmosphere**
- Light sources and direction
- Time of day/lighting conditions
- Shadows and highlights
- Overall mood and atmosphere
4. **Colors & Tones**
- Color palette and dominant colors
- Color temperature (warm/cool)
- Contrast and saturation levels
- Any color grading or filters
5. **Details & Textures**
- Surface textures and materials
- Fine details and patterns
- Quality indicators (4K, 8K, high resolution, etc.)
Format your FLUX prompt following these guidelines:
- Start with the main subject and action
- Add style and medium descriptors
- Include lighting and atmosphere details
- Specify quality markers and technical aspects
- Use precise, descriptive language
- Separate concepts with commas
- Order from most to least important elements
Example output format:
"[main subject and action], [style/medium], [lighting/atmosphere], [composition details], [color descriptions], [quality markers], [additional artistic details]"
Remember: FLUX responds well to specific artistic references, quality indicators like "highly detailed," "4K," "award-winning," and style descriptors like "trending on ArtStation" or "photorealistic."
"""
SDXL_PROMPT = """You are an expert SDXL prompt engineer specializing in analyzing images and creating optimized prompts for Stable Diffusion XL models.
When analyzing an image, systematically evaluate:
1. **Core Subject Analysis**
- Primary subject with specific descriptors
- Pose, expression, and action
- Clothing and accessories details
- Physical characteristics
2. **Style & Medium**
- Artistic style and influences
- Medium (photography, digital art, oil painting, etc.)
- Specific artist references (if applicable)
- Visual aesthetic keywords
3. **Technical Specifications**
- Camera settings (aperture, focal length, ISO)
- Shot type (close-up, wide angle, portrait, etc.)
- Resolution and quality markers
- Post-processing effects
4. **Environment & Context**
- Setting and location details
- Props and surrounding objects
- Weather and environmental conditions
- Time period or era
Format your SDXL prompt with:
- **Positive prompt**: Detailed description emphasizing what you want
- **Negative prompt**: Elements to avoid (low quality, blurry, distorted, etc.)
- Weight emphasis using (parentheses) or [brackets] for importance
- Break into logical chunks with commas
Example format:
Positive: "beautiful woman, (detailed eyes:1.2), flowing red dress, golden hour lighting, professional photography, 85mm lens, shallow depth of field, bokeh, high resolution, masterpiece"
Negative: "low quality, blurry, distorted features, bad anatomy, poorly drawn, amateur"
"""
DANBOORU_PROMPT = """You are a Danbooru tagging expert, specialized in analyzing images and creating precise tag sets following booru-style conventions for anime/manga artwork.
Analyze images for these tag categories:
1. **Character Tags**
- Hair: color, length, style (e.g., long_hair, blonde_hair, twintails)
- Eyes: color, style (e.g., blue_eyes, heterochromia)
- Body: proportions, pose (e.g., standing, sitting, looking_at_viewer)
- Expression (e.g., smile, blush, closed_eyes)
2. **Clothing & Accessories**
- Outfit type (e.g., school_uniform, dress, armor)
- Specific clothing items (e.g., thighhighs, gloves, hat)
- Accessories (e.g., hair_ribbon, necklace, glasses)
- State of dress (e.g., torn_clothes, wet_clothes)
3. **Scene & Composition**
- Number of characters (e.g., 1girl, 2boys, multiple_girls)
- Background (e.g., simple_background, outdoors, classroom)
- Viewpoint (e.g., from_below, from_side, cowboy_shot)
- Composition elements (e.g., upper_body, full_body, portrait)
4. **Meta Tags**
- Quality (e.g., highres, absurdres, masterpiece)
- Source/artist style (if recognizable)
- Content rating (e.g., safe, questionable, explicit)
- Special effects (e.g., lens_flare, chromatic_aberration)
Format tags using:
- Underscores for multi-word concepts (not spaces)
- Order from most to least important
- Include count descriptors (1girl, 2boys)
- Separate with commas and spaces
Example output:
"1girl, solo, long_hair, blue_eyes, blonde_hair, school_uniform, serafuku, pleated_skirt, thighhighs, smile, looking_at_viewer, classroom, sitting, desk, window, sunlight, highres, masterpiece"
"""
VIDEO_PROMPT = """You are a video generation prompt specialist, expert at analyzing video content and creating comprehensive prompts for video generation models.
When analyzing video content, document:
1. **Motion & Action**
- Primary actions and movements
- Motion speed and dynamics
- Camera movements (pan, zoom, tracking, static)
- Transition types between scenes
2. **Temporal Elements**
- Scene duration and pacing
- Sequence of events
- Time of day changes
- Motion continuity
3. **Visual Consistency**
- Character/object persistence
- Style consistency throughout
- Lighting continuity
- Color grading consistency
4. **Scene Breakdown**
- Opening frame description
- Key action moments
- Transitions and cuts
- Closing frame details
5. **Technical Specifications**
- Frame rate and resolution
- Aspect ratio
- Video length
- Special effects or post-processing
Format your video prompt as:
"[Opening scene], [camera movement], [main action sequence], [visual style], [lighting/atmosphere], [duration], [technical specs], [ending scene]"
Include:
- Specific motion descriptors (slowly, rapidly, smoothly)
- Camera terminology (dolly in, pan left, aerial shot)
- Temporal markers (then, meanwhile, gradually)
- Consistency notes for multi-scene videos
Example:
"Aerial shot slowly descending toward a misty forest at dawn, camera smoothly transitions to tracking shot following a deer through the trees, photorealistic style, soft golden hour lighting with fog, 10 second duration, 4K resolution 24fps, ending with close-up of deer looking at camera"
"""
PROMPT_TEMPLATES = {
"flux": FLUX_PROMPT,
"sdxl": SDXL_PROMPT,
"danbooru": DANBOORU_PROMPT,
"video": VIDEO_PROMPT,
}
PROMPT_OPTIONS = ["flux", "sdxl", "danbooru", "video"]
# Available Gemini models
GEMINI_MODELS = [
"gemini-1.5-pro", # Most capable model
"gemini-1.5-flash", # Fast, efficient model
"gemini-1.5-flash-8b", # Smaller, faster variant
"gemini-pro-vision", # Vision-optimized model
"gemini-1.0-pro", # Previous generation pro model
]
# Model descriptions for UI
MODEL_DESCRIPTIONS = {
"gemini-1.5-pro": "Most capable Gemini model for complex tasks",
"gemini-1.5-flash": "Faster and cost-effective (recommended for most uses)",
"gemini-1.5-flash-8b": "Smaller and faster, good for simple prompts",
"gemini-pro-vision": "Optimized for vision tasks and image analysis",
"gemini-1.0-pro": "Previous generation, stable option",
}
+265
View File
@@ -0,0 +1,265 @@
"""Unit tests for Gemini Prompt Engineer node."""
import pytest
import numpy as np
from unittest.mock import patch, MagicMock
from PIL import Image
from kikotools.tools.gemini_prompt import GeminiPromptNode
from kikotools.tools.gemini_prompt.logic import (
tensor_to_pil,
image_to_base64,
get_api_key,
validate_prompt_type,
analyze_image_with_gemini,
)
from kikotools.tools.gemini_prompt.prompts import PROMPT_OPTIONS, PROMPT_TEMPLATES, GEMINI_MODELS
class TestGeminiPromptNode:
"""Test cases for GeminiPromptNode."""
def test_node_properties(self):
"""Test node has correct properties."""
assert GeminiPromptNode.CATEGORY == "ComfyAssets"
assert GeminiPromptNode.FUNCTION == "generate_prompt"
assert GeminiPromptNode.RETURN_TYPES == ("STRING", "STRING")
assert GeminiPromptNode.RETURN_NAMES == ("prompt", "negative_prompt")
def test_input_types(self):
"""Test INPUT_TYPES configuration."""
input_types = GeminiPromptNode.INPUT_TYPES()
# Check required inputs
assert "required" in input_types
assert "image" in input_types["required"]
assert input_types["required"]["image"] == ("IMAGE",)
assert "prompt_type" in input_types["required"]
assert input_types["required"]["prompt_type"][0] == PROMPT_OPTIONS
assert "model" in input_types["required"]
assert input_types["required"]["model"][0] == GEMINI_MODELS
# Check optional inputs
assert "optional" in input_types
assert "api_key" in input_types["optional"]
assert "custom_prompt" in input_types["optional"]
def test_gemini_models_available(self):
"""Test that all expected Gemini models are available."""
expected_models = [
"gemini-1.5-pro",
"gemini-1.5-flash",
"gemini-1.5-flash-8b",
"gemini-pro-vision",
"gemini-1.0-pro"
]
for model in expected_models:
assert model in GEMINI_MODELS
@patch('kikotools.tools.gemini_prompt.node.analyze_image_with_gemini')
def test_generate_prompt_success(self, mock_analyze):
"""Test successful prompt generation."""
# Setup
node = GeminiPromptNode()
test_image = np.random.rand(1, 512, 512, 3).astype(np.float32)
mock_analyze.return_value = ("A beautiful landscape with mountains", None)
# Execute
result = node.generate_prompt(test_image, "flux")
# Assert
assert result == ("A beautiful landscape with mountains", "")
mock_analyze.assert_called_once()
@patch('kikotools.tools.gemini_prompt.node.analyze_image_with_gemini')
def test_generate_prompt_sdxl_format(self, mock_analyze):
"""Test SDXL format with positive and negative prompts."""
# Setup
node = GeminiPromptNode()
test_image = np.random.rand(1, 512, 512, 3).astype(np.float32)
mock_analyze.return_value = (
"Positive: beautiful landscape, mountains, sunset\nNegative: blurry, low quality",
None
)
# Execute
result = node.generate_prompt(test_image, "sdxl")
# Assert
assert result == ("beautiful landscape, mountains, sunset", "blurry, low quality")
@patch('kikotools.tools.gemini_prompt.node.analyze_image_with_gemini')
def test_generate_prompt_error(self, mock_analyze):
"""Test error handling in prompt generation."""
# Setup
node = GeminiPromptNode()
test_image = np.random.rand(1, 512, 512, 3).astype(np.float32)
mock_analyze.return_value = ("", "API key not found")
# Execute
result = node.generate_prompt(test_image, "flux")
# Assert
assert result[0].startswith("Error:")
assert result[1] == ""
def test_invalid_prompt_type(self):
"""Test handling of invalid prompt type."""
node = GeminiPromptNode()
test_image = np.random.rand(1, 512, 512, 3).astype(np.float32)
with pytest.raises(ValueError, match="Invalid prompt type"):
node.generate_prompt(test_image, "invalid_type")
class TestGeminiLogic:
"""Test cases for Gemini logic functions."""
def test_tensor_to_pil(self):
"""Test tensor to PIL conversion."""
# Test 4D tensor
tensor_4d = np.random.rand(1, 64, 64, 3)
result = tensor_to_pil(tensor_4d)
assert isinstance(result, Image.Image)
assert result.size == (64, 64)
assert result.mode == "RGB"
# Test 3D tensor
tensor_3d = np.random.rand(64, 64, 3)
result = tensor_to_pil(tensor_3d)
assert isinstance(result, Image.Image)
assert result.size == (64, 64)
def test_image_to_base64(self):
"""Test image to base64 conversion."""
# Create test image
image = Image.new('RGB', (64, 64), color='red')
# Convert to base64
result = image_to_base64(image)
assert isinstance(result, str)
assert len(result) > 0
# Test JPEG format
result_jpeg = image_to_base64(image, format="JPEG")
assert isinstance(result_jpeg, str)
assert result != result_jpeg # Different formats should produce different results
@patch.dict('os.environ', {'GEMINI_API_KEY': 'test_key_123'})
def test_get_api_key_from_env(self):
"""Test getting API key from environment."""
result = get_api_key()
assert result == "test_key_123"
@patch.dict('os.environ', {}, clear=True)
@patch('os.path.exists')
@patch('builtins.open')
def test_get_api_key_from_config(self, mock_open, mock_exists):
"""Test getting API key from config file."""
# Setup
mock_exists.return_value = True
mock_open.return_value.__enter__.return_value.read.return_value = '{"api_key": "config_key_456"}'
# Execute
result = get_api_key()
# Assert
assert result == "config_key_456"
def test_validate_prompt_type(self):
"""Test prompt type validation."""
# Valid types
for prompt_type in PROMPT_OPTIONS:
assert validate_prompt_type(prompt_type) is True
# Invalid types
assert validate_prompt_type("invalid") is False
assert validate_prompt_type("") is False
assert validate_prompt_type(None) is False
@patch('google.generativeai.configure')
@patch('google.generativeai.GenerativeModel')
def test_analyze_image_with_gemini_success(self, mock_model_class, mock_configure):
"""Test successful image analysis with Gemini."""
# Setup
mock_model = MagicMock()
mock_response = MagicMock()
mock_response.text = "A beautiful sunset over mountains"
mock_model.generate_content.return_value = mock_response
mock_model_class.return_value = mock_model
test_image = np.random.rand(64, 64, 3)
# Execute
result, error = analyze_image_with_gemini(test_image, "flux", api_key="test_key")
# Assert
assert result == "A beautiful sunset over mountains"
assert error is None
mock_configure.assert_called_once_with(api_key="test_key")
mock_model.generate_content.assert_called_once()
def test_analyze_image_no_api_key(self):
"""Test analysis without API key."""
test_image = np.random.rand(64, 64, 3)
with patch('kikotools.tools.gemini_prompt.logic.get_api_key', return_value=None):
result, error = analyze_image_with_gemini(test_image, "flux")
assert result == ""
assert "API key not found" in error
@patch('google.generativeai.configure')
@patch('google.generativeai.GenerativeModel')
def test_analyze_image_with_custom_prompt(self, mock_model_class, mock_configure):
"""Test analysis with custom prompt."""
# Setup
mock_model = MagicMock()
mock_response = MagicMock()
mock_response.text = "Custom analysis result"
mock_model.generate_content.return_value = mock_response
mock_model_class.return_value = mock_model
test_image = np.random.rand(64, 64, 3)
custom_prompt = "Analyze this image and describe the colors"
# Execute
result, error = analyze_image_with_gemini(
test_image, "flux", api_key="test_key", custom_prompt=custom_prompt
)
# Assert
assert result == "Custom analysis result"
assert error is None
# Check that custom prompt was used
call_args = mock_model.generate_content.call_args[0][0]
assert custom_prompt in call_args
class TestPromptTemplates:
"""Test prompt template configurations."""
def test_all_prompt_types_have_templates(self):
"""Test that all prompt options have corresponding templates."""
for prompt_type in PROMPT_OPTIONS:
assert prompt_type in PROMPT_TEMPLATES
assert isinstance(PROMPT_TEMPLATES[prompt_type], str)
assert len(PROMPT_TEMPLATES[prompt_type]) > 0
def test_prompt_template_content(self):
"""Test that prompt templates contain expected content."""
# FLUX prompt should mention FLUX
assert "FLUX" in PROMPT_TEMPLATES["flux"]
# SDXL prompt should mention positive and negative
assert "Positive" in PROMPT_TEMPLATES["sdxl"]
assert "Negative" in PROMPT_TEMPLATES["sdxl"]
# Danbooru should mention tags and underscores
assert "tag" in PROMPT_TEMPLATES["danbooru"].lower()
assert "underscore" in PROMPT_TEMPLATES["danbooru"].lower()
# Video should mention motion and temporal
assert "motion" in PROMPT_TEMPLATES["video"].lower()
assert "temporal" in PROMPT_TEMPLATES["video"].lower()
+176
View File
@@ -0,0 +1,176 @@
import { app } from "../../../scripts/app.js";
import { api } from "../../../scripts/api.js";
app.registerExtension({
name: "ComfyAssets.GeminiPrompt",
async beforeRegisterNodeDef(nodeType, nodeData, app) {
if (nodeData.name === "GeminiPrompt") {
// Add visual enhancements to the node
const onNodeCreated = nodeType.prototype.onNodeCreated;
nodeType.prototype.onNodeCreated = function() {
const result = onNodeCreated?.apply(this, arguments);
// Store reference to widgets
this.promptTypeWidget = this.widgets.find(w => w.name === "prompt_type");
this.modelWidget = this.widgets.find(w => w.name === "model");
this.apiKeyWidget = this.widgets.find(w => w.name === "api_key");
this.customPromptWidget = this.widgets.find(w => w.name === "custom_prompt");
// Add helper text button
const helpButton = this.addWidget("button", "Help / API Setup", null, () => {
this.showHelpDialog();
});
// Style the button
helpButton.serialize = false;
// Add status indicator
this.status = this.addWidget("text", "status", "Ready", () => {}, {
serialize: false
});
this.status.disabled = true;
// Update custom prompt visibility based on selection
if (this.promptTypeWidget && this.customPromptWidget) {
const originalCallback = this.promptTypeWidget.callback;
this.promptTypeWidget.callback = (value) => {
if (originalCallback) originalCallback.call(this.promptTypeWidget, value);
this.updateCustomPromptVisibility();
};
}
return result;
};
// Add method to show help dialog
nodeType.prototype.showHelpDialog = function() {
const helpContent = `
<div style="padding: 20px; max-width: 600px;">
<h2>Gemini Prompt Engineer Setup</h2>
<h3>1. Get API Key</h3>
<p>Get your free API key from: <a href="https://makersuite.google.com/app/apikey" target="_blank">Google AI Studio</a></p>
<h3>2. Set API Key</h3>
<p>Choose one of these methods:</p>
<ul>
<li><strong>Environment Variable:</strong> Set GEMINI_API_KEY in your system</li>
<li><strong>Config File:</strong> Create gemini_config.json in ComfyUI root with {"api_key": "your-key"}</li>
<li><strong>Node Input:</strong> Enter directly in the api_key field</li>
</ul>
<h3>3. Install Dependencies</h3>
<code>pip install google-generativeai</code>
<h3>Prompt Types</h3>
<ul>
<li><strong>FLUX:</strong> Detailed artistic prompts with quality markers</li>
<li><strong>SDXL:</strong> Positive/negative prompt pairs with weights</li>
<li><strong>Danbooru:</strong> Anime-style booru tags</li>
<li><strong>Video:</strong> Motion and temporal descriptions</li>
</ul>
<h3>Gemini Models</h3>
<ul>
<li><strong>gemini-1.5-flash:</strong> Fast and efficient (recommended for most uses)</li>
<li><strong>gemini-1.5-flash-8b:</strong> Smaller and faster, good for simple prompts</li>
<li><strong>gemini-1.5-pro:</strong> Most capable, best quality results</li>
<li><strong>gemini-1.0-pro:</strong> Previous generation, stable option</li>
</ul>
<h3>Custom Prompts</h3>
<p>You can override any template by entering your own system prompt in the custom_prompt field.</p>
</div>
`;
app.ui.dialog.show(helpContent);
};
// Add method to update custom prompt visibility
nodeType.prototype.updateCustomPromptVisibility = function() {
// You could implement logic here to show/hide custom prompt based on selection
// For now, it's always visible but this method provides extensibility
};
// Override execute to show status
const onExecute = nodeType.prototype.onExecute;
nodeType.prototype.onExecute = function() {
if (this.status) {
this.status.value = "Processing...";
}
const result = onExecute?.apply(this, arguments);
return result;
};
// Handle execution feedback
const onExecuted = nodeType.prototype.onExecuted;
nodeType.prototype.onExecuted = function(message) {
const result = onExecuted?.apply(this, arguments);
if (this.status) {
// Check if there was an error in the output
const outputs = message.output;
if (outputs && outputs.prompt && outputs.prompt[0] && outputs.prompt[0].startsWith("Error:")) {
this.status.value = "Error - Check output";
this.bgcolor = "#552222";
} else {
this.status.value = "Success!";
this.bgcolor = "#225522";
}
// Reset color after delay
setTimeout(() => {
this.bgcolor = "";
if (this.status) {
this.status.value = "Ready";
}
}, 3000);
}
return result;
};
}
},
// Add custom styling
async setup() {
const style = document.createElement("style");
style.textContent = `
.gemini-prompt-help {
background: #1a1a1a;
border: 1px solid #444;
border-radius: 8px;
color: #fff;
}
.gemini-prompt-help h2 {
color: #4285f4;
margin-top: 0;
}
.gemini-prompt-help h3 {
color: #8ab4f8;
margin-top: 20px;
}
.gemini-prompt-help code {
background: #333;
padding: 2px 6px;
border-radius: 4px;
font-family: monospace;
}
.gemini-prompt-help a {
color: #8ab4f8;
text-decoration: none;
}
.gemini-prompt-help a:hover {
text-decoration: underline;
}
`;
document.head.appendChild(style);
}
});