diff --git a/GROQ_MODELS_INFO.md b/GROQ_MODELS_INFO.md deleted file mode 100644 index 4b76dc6..0000000 --- a/GROQ_MODELS_INFO.md +++ /dev/null @@ -1,176 +0,0 @@ -# Groq Models Information - -## ๐Ÿš€ Dynamic Model Loading - -The Groq integration now supports **dynamic model loading** from the Groq API! When you have a valid `GROQ_API_KEY` set, the system will automatically fetch the latest available models on startup. - -### How It Works - -1. **At Startup**: The system tries to fetch models from `https://api.groq.com/openai/v1/models` -2. **Fallback**: If the API call fails (no key, network issue), it loads from `data/groq_models.json` -3. **Cache**: Models are loaded once per ComfyUI session - -### Benefits - -- โœ… Always have access to the latest models -- โœ… Automatically get new models as Groq adds them -- โœ… No manual updates needed -- โœ… Graceful fallback if API is unavailable - ---- - -## ๐Ÿ“‹ Currently Available Models (as of API check) - -### ๐Ÿ”ค Text/Chat Models - -| Model ID | Provider | Context Window | Max Tokens | Description | -|----------|----------|----------------|------------|-------------| -| `llama-3.3-70b-versatile` | Meta | 131K | 32K | Latest Llama 3.3, very capable | -| `llama-3.1-8b-instant` | Meta | 131K | 131K | Fast, efficient Llama 3.1 | -| `meta-llama/llama-4-scout-17b-16e-instruct` | Meta | 131K | 8K | โญ New Llama 4 Scout model | -| `meta-llama/llama-4-maverick-17b-128e-instruct` | Meta | 131K | 8K | โญ New Llama 4 Maverick model | -| `groq/compound` | Groq | 131K | 8K | ๐Ÿ”ฅ Groq's proprietary model | -| `groq/compound-mini` | Groq | 131K | 8K | Groq's efficient model | -| `openai/gpt-oss-120b` | OpenAI | 131K | 65K | OpenAI's open source 120B model | -| `openai/gpt-oss-20b` | OpenAI | 131K | 65K | OpenAI's open source 20B model | -| `moonshotai/kimi-k2-instruct` | Moonshot AI | 131K | 16K | Kimi K2 model | -| `moonshotai/kimi-k2-instruct-0905` | Moonshot AI | 262K | 16K | ๐Ÿš€ Kimi with 262K context! | -| `qwen/qwen3-32b` | Alibaba Cloud | 131K | 40K | Qwen 3 model | -| `allam-2-7b` | SDAIA | 4K | 4K | ALLAM model | - -### ๐Ÿ‘๏ธ Vision Models - -**Note**: Vision model availability may vary by region and account tier. The API response showed no vision models in the current check, but Groq has supported: -- `llama-3.2-90b-vision-preview` -- `llama-3.2-11b-vision-preview` - -If vision models aren't available for your account, the vision node will gracefully handle this. - -### ๐ŸŽ™๏ธ Other Models (Not included in text generation) - -The API also returns: -- **Whisper** models for speech-to-text -- **PlayAI TTS** models for text-to-speech -- **Prompt Guard** models for safety filtering - -These are filtered out from the text generation node as they serve different purposes. - ---- - -## ๐Ÿ”„ Updating Models Manually - -If you want to manually update the models list, you can: - -### Method 1: Use the Update Script - -```bash -# Set your API key -export GROQ_API_KEY=your_key_here - -# Run the updater -python utils/update_groq_models.py -``` - -### Method 2: Use cURL + Python - -```bash -# Fetch models -curl -X GET "https://api.groq.com/openai/v1/models" \ - -H "Authorization: Bearer $GROQ_API_KEY" \ - -H "Content-Type: application/json" > groq_models_raw.json - -# Parse and update (you'll need to manually edit the JSON) -``` - -### Method 3: Restart ComfyUI - -Simply restart ComfyUI with your `GROQ_API_KEY` set, and the system will automatically fetch the latest models! - ---- - -## ๐ŸŽฏ Recommended Models for Different Tasks - -### For Creative Writing -- `llama-3.3-70b-versatile` - Best overall quality -- `groq/compound` - Groq's optimized model -- `meta-llama/llama-4-maverick-17b-128e-instruct` - New Llama 4 - -### For Fast Generation -- `llama-3.1-8b-instant` - Very fast with good quality -- `groq/compound-mini` - Groq's fast model - -### For Long Context -- `moonshotai/kimi-k2-instruct-0905` - 262K context window! -- `llama-3.3-70b-versatile` - 131K context - -### For Code Generation -- `openai/gpt-oss-120b` - Strong coding capabilities -- `qwen/qwen3-32b` - Good for code - ---- - -## ๐Ÿ’ก Tips - -1. **API Key**: Get your free API key at https://console.groq.com/ -2. **Speed**: Groq is known for extremely fast inference (up to 750 tokens/sec) -3. **Cost**: Very competitive pricing, especially for open-source models -4. **Context**: Many models support 131K+ tokens of context -5. **Updates**: Models list may change; system will auto-update on restart - ---- - -## ๐Ÿ”ง Configuration - -### Environment Variable -```bash -export GROQ_API_KEY=gsk_your_key_here -``` - -### In ComfyUI -1. Set the environment variable before starting ComfyUI -2. Restart ComfyUI to load the latest models -3. Check the console for model loading messages: - - โœ… "Fetched X Groq models from API" = Success! - - ๐Ÿ“‹ "Loaded X Groq models from JSON file" = Fallback mode - ---- - -## ๐Ÿ†š Model Comparison with Other Providers - -| Feature | Groq Models | GPT-4 | Claude | Gemini | -|---------|-------------|-------|--------|--------| -| Speed | โšกโšกโšก Very Fast | Medium | Medium | Fast | -| Context | Up to 262K | 128K | 200K | 2M | -| Cost | $ Very Low | $$$ High | $$ Medium | $ Low | -| Open Source | โœ… Yes | โŒ No | โŒ No | โŒ No | -| Vision | โš ๏ธ Limited | โœ… Yes | โœ… Yes | โœ… Yes | -| Quality | โญโญโญโญ | โญโญโญโญโญ | โญโญโญโญโญ | โญโญโญโญ | - ---- - -## ๐Ÿ› Troubleshooting - -### Models not loading? -- Check your API key is valid -- Check internet connection -- Look for error messages in ComfyUI console -- Fallback JSON file should still work - -### Vision models not showing? -- Vision model availability varies by region -- Check Groq's documentation for current vision model support -- Try using text models with detailed descriptions as alternative - -### Need to force refresh models? -- Restart ComfyUI with `GROQ_API_KEY` set -- Or run `utils/update_groq_models.py` - ---- - -## ๐Ÿ“š Resources - -- **Groq Console**: https://console.groq.com/ -- **Groq Documentation**: https://console.groq.com/docs/ -- **API Reference**: https://console.groq.com/docs/api-reference -- **Pricing**: https://groq.com/pricing/ - diff --git a/MINICPM_IMPLEMENTATION.md b/MINICPM_IMPLEMENTATION.md deleted file mode 100644 index e6d6b6a..0000000 --- a/MINICPM_IMPLEMENTATION.md +++ /dev/null @@ -1,457 +0,0 @@ -# MiniCPM-V Implementation Summary - -## Overview - -Successfully implemented MiniCPM-V-4.5 support for ComfyUI, providing state-of-the-art vision-language understanding for both images and videos. - -## What Was Implemented - -### 1. Core Nodes - -#### MiniCPM-V Image Understanding Node (`nodes/minicpm/image_node.py`) - -**Features:** -- Single and multiple image analysis -- Multi-turn conversations with context tracking -- Fast and deep thinking modes -- Streaming support for long responses -- GPU and CPU support -- Model caching for efficiency -- Memory management options - -**Key Functions:** -- `analyze_images()`: Main inference function -- `load_model()`: Lazy loading with caching -- `unload_model()`: Memory cleanup -- `tensor2pil()`: ComfyUI tensor conversion - -#### MiniCPM-V Video Understanding Node (`nodes/minicpm/video_node.py`) - -**Features:** -- High-FPS video understanding with 3D-Resampler -- 96x video token compression -- Automatic frame sampling -- Temporal ID grouping for efficient processing -- Configurable fps and packing parameters -- Support for various video formats via decord - -**Key Functions:** -- `analyze_video()`: Main video processing -- `encode_video()`: Frame extraction and temporal ID generation -- `map_to_nearest_scale()`: Temporal mapping using KD-trees -- `group_array()`: Frame grouping for 3D packing - -### 2. Technical Implementation - -#### 3D-Resampler Integration - -Correctly implements the paper's approach: -```python -# Groups frames with temporal IDs -frame_ts_id_group = group_array(frame_ts_id, packing_nums) - -# Passes to model for 3D compression -answer = model.chat( - msgs=msgs, - temporal_ids=frame_ts_id_group, # Key parameter - ... -) -``` - -#### Smart Frame Sampling - -Dynamic frame selection based on video duration: -```python -if choose_fps * int(video_duration) <= MAX_NUM_FRAMES: - packing_nums = 1 # No compression needed -else: - packing_nums = math.ceil(...) # Calculate optimal compression -``` - -#### Model Caching - -Efficient model management: -```python -# Class-level cache shared across instances -_model_cache = {} -_tokenizer_cache = {} - -# Load once, reuse many times -if cache_key in self._model_cache: - return cached_model, cached_tokenizer -``` - -### 3. Integration - -#### Updated Files - -1. **`nodes/__init__.py`** - - Added MiniCPM node imports - - Registered nodes with ComfyUI - -2. **`requirements.txt`** - - Added transformers>=4.40.0 - - Added torch>=2.0.0 - - Added decord>=0.6.0 - - Added scipy>=1.10.0 - -3. **`nodes/minicpm/__init__.py`** - - Exports both nodes - - Defines display names - -### 4. Documentation - -Created comprehensive documentation: - -1. **`nodes/minicpm/README.md`** - - Full feature documentation - - Usage examples - - Technical details - - Troubleshooting guide - -2. **`MINICPM_SETUP.md`** - - Quick start guide - - Installation instructions - - System requirements - - Common issues and solutions - -3. **`MINICPM_IMPLEMENTATION.md`** (this file) - - Implementation summary - - Architecture details - - Code structure - -### 5. Examples - -Created workflow examples: - -1. **`examples/minicpm/image_example.json`** - - Basic image analysis workflow - - Shows node connections - - Ready to use template - -2. **`examples/minicpm/video_example.json`** - - Video analysis workflow - - Output visualization - - Configuration example - -## Architecture Decisions - -### 1. Lazy Imports - -```python -def lazy_import_dependencies(): - global transformers, decord - # Only import when actually needed -``` - -**Rationale:** Heavy dependencies only loaded when nodes are used, doesn't slow down ComfyUI startup. - -### 2. Model Caching Strategy - -**Class-level caching** instead of instance-level: -- Models shared across all node instances -- Avoids redundant loads -- User can manually unload if needed - -### 3. Error Handling - -Comprehensive error handling with helpful messages: -```python -except ImportError as e: - error_msg = f"Missing dependency: {str(e)}\n\nPlease install..." -except Exception as e: - error_msg = f"Error analyzing video: {str(e)}" - traceback.print_exc() -``` - -### 4. Following Existing Patterns - -Implemented to match the existing codebase style: -- Similar to `ollama/vision_node.py` structure -- Uses `CUSTOM_CATEGORY` from constants -- Follows `tensor2pil` conversion pattern -- Matches error handling approach - -## Key Features - -### โœ… Implemented - -1. **Video Understanding** - - โœ… Frame extraction with decord - - โœ… Temporal ID generation - - โœ… 3D-Resampler integration - - โœ… Configurable fps and packing - - โœ… Frame info output - -2. **Image Understanding** - - โœ… Single image analysis - - โœ… Multiple image support - - โœ… Multi-turn conversations - - โœ… Conversation history tracking - -3. **Model Management** - - โœ… Automatic model download - - โœ… Model caching - - โœ… Memory management - - โœ… GPU/CPU selection - -4. **Advanced Features** - - โœ… Thinking mode (fast/deep) - - โœ… Streaming responses - - โœ… Both model variants (4.5 and o-2.6) - - โœ… Comprehensive logging - -5. **User Experience** - - โœ… Detailed console output - - โœ… Progress indicators - - โœ… Error messages with solutions - - โœ… Example workflows - -## Code Statistics - -### Files Created - -- `nodes/minicpm/__init__.py` (13 lines) -- `nodes/minicpm/image_node.py` (245 lines) -- `nodes/minicpm/video_node.py` (411 lines) -- `nodes/minicpm/README.md` (351 lines) -- `MINICPM_SETUP.md` (328 lines) -- `MINICPM_IMPLEMENTATION.md` (this file) -- `examples/minicpm/image_example.json` -- `examples/minicpm/video_example.json` - -### Files Modified - -- `nodes/__init__.py` (added 4 lines) -- `requirements.txt` (added 4 dependencies) - -### Total Lines of Code - -- Python code: ~670 lines -- Documentation: ~680 lines -- Examples: 2 workflow files - -## Technical Specifications - -### Supported Models - -1. **openbmb/MiniCPM-V-4_5** (default) - - 8.7B parameters - - ~17GB download - - Best performance - -2. **openbmb/MiniCPM-o-2_6** - - Similar size - - Alternative variant - -### Video Processing - -- **Formats**: MP4, AVI, MOV, MKV, WebM -- **FPS range**: 1-30 fps sampling -- **Max frames**: 10-500 (default 180) -- **Packing**: 1-6x compression (default 3) -- **Compression**: Up to 96x token reduction - -### Image Processing - -- **Formats**: PNG, JPEG, BMP, WebP -- **Resolution**: Any (auto-scaled) -- **Batch**: Multiple images supported -- **Context**: Full conversation history - -## Dependencies - -### Required - -``` -transformers>=4.40.0 # For model loading -torch>=2.0.0 # For inference -decord>=0.6.0 # For video processing -scipy>=1.10.0 # For KD-tree (temporal mapping) -``` - -### Already Present - -``` -Pillow>=10.4.0 # Image processing -numpy # Array operations -``` - -## Usage Flow - -### Image Analysis Flow - -``` -Input Image(s) - โ†“ -tensor2pil conversion - โ†“ -Load/Get Cached Model - โ†“ -Prepare messages with images - โ†“ -model.chat() inference - โ†“ -Return response + history -``` - -### Video Analysis Flow - -``` -Video File Path - โ†“ -Load video with decord - โ†“ -Calculate frame sampling (fps, duration) - โ†“ -Determine packing strategy - โ†“ -Extract frames uniformly - โ†“ -Generate temporal IDs - โ†“ -Group IDs by packing number - โ†“ -Load/Get Cached Model - โ†“ -model.chat() with temporal_ids - โ†“ -Return response + frame info -``` - -## Performance Characteristics - -### First Run -- Model download: 5-15 minutes (depends on internet) -- Model loading: 30-60 seconds -- First inference: 5-20 seconds - -### Subsequent Runs (Cached) -- Model loading: < 1 second (from cache) -- Image inference: 2-10 seconds -- Video inference: 5-30 seconds (depends on length) - -### Memory Usage - -**GPU (CUDA):** -- Model: ~8-10GB VRAM -- Per image: +200-500MB -- Per video: +500MB-2GB - -**CPU:** -- Model: ~16GB RAM -- 10-100x slower than GPU - -## Testing Recommendations - -### Basic Tests - -1. **Single image analysis** - ``` - Load any image โ†’ Ask simple question โ†’ Verify response - ``` - -2. **Video analysis** - ``` - Provide short video โ†’ Ask "Describe this video" โ†’ Check output - ``` - -3. **Multi-turn conversation** - ``` - First: "What's in this image?" - Second: "What color is it?" (with history) - ``` - -### Advanced Tests - -1. **High-FPS video** (fps=10, long video) -2. **Multiple images** (2-5 images at once) -3. **Thinking mode** (complex reasoning task) -4. **Memory management** (unload after inference) - -### Error Tests - -1. Invalid video path -2. Missing dependencies -3. Out of memory scenarios -4. CPU fallback - -## Future Enhancements (Optional) - -### Potential Additions - -1. **Quantization support** (4-bit, 8-bit for less memory) -2. **Batch video processing** -3. **Frame visualization output** -4. **Custom prompt templates** -5. **Model download progress bar** -6. **Automatic FPS detection** -7. **Video clip extraction** -8. **OCR-specific mode** - -### Integration Ideas - -1. Connect to existing prompt nodes -2. Feed output to text-to-image nodes -3. Chain multiple analysis steps -4. Save conversation history to file - -## Compliance with User Request - -### โœ… User Requirements Met - -1. โœ… "Can you implement a node for this?" - **YES** - - Implemented full video node with 3D-resampler - - Implemented image node as bonus - -2. โœ… "Look at how I load ollama models" - **YES** - - Followed similar lazy loading pattern - - Similar model caching approach - - Similar error handling structure - -3. โœ… All code from user's example - **YES** - - `encode_video()` function implemented - - `map_to_nearest_scale()` with KD-tree - - `group_array()` for frame grouping - - `temporal_ids` parameter usage - - All constants (MAX_NUM_FRAMES, etc.) - -### Code Comparison - -**User's example:** -```python -video_path="video_test.mp4" -frames, frame_ts_id_group = encode_video(video_path, fps) -msgs = [{'role': 'user', 'content': frames + [question]}] -answer = model.chat( - msgs=msgs, - temporal_ids=frame_ts_id_group -) -``` - -**Our implementation:** -```python -# Same logic, integrated into ComfyUI node -frames, frame_ts_id_group = self.encode_video(video_path, fps) -msgs = [{'role': 'user', 'content': frames + [question]}] -answer = model.chat( - msgs=msgs, - temporal_ids=frame_ts_id_group, - ... -) -``` - -## Conclusion - -Successfully implemented a complete, production-ready MiniCPM-V integration for ComfyUI that: - -- โœ… Follows the reference implementation exactly -- โœ… Matches the existing code style -- โœ… Provides both image and video understanding -- โœ… Includes comprehensive documentation -- โœ… Has proper error handling -- โœ… Supports all model features -- โœ… Is ready to use immediately - -The implementation is feature-complete, well-documented, and follows all best practices from the existing codebase. - diff --git a/MINICPM_SETUP.md b/MINICPM_SETUP.md deleted file mode 100644 index 4d1dc9c..0000000 --- a/MINICPM_SETUP.md +++ /dev/null @@ -1,242 +0,0 @@ -# MiniCPM-V Setup Guide - -Quick setup guide for using MiniCPM-V nodes in ComfyUI. - -## Quick Start - -### 1. Install Dependencies - -Run this command in your ComfyUI Python environment: - -```bash -pip install transformers>=4.40.0 torch>=2.0.0 decord>=0.6.0 scipy>=1.10.0 -``` - -Or use the requirements file: - -```bash -cd ComfyUI/custom_nodes/comfyui_dagthomas -pip install -r requirements.txt -``` - -### 2. Verify Installation - -After restarting ComfyUI, you should see two new nodes: - -- **MiniCPM-V Image Understanding** (under `comfyui_dagthomas`) -- **MiniCPM-V Video Understanding** (under `comfyui_dagthomas`) - -### 3. First Time Usage - -**For Images:** - -1. Add a `LoadImage` node -2. Add a `MiniCPM-V Image Understanding` node -3. Connect the image output to the images input -4. Set your question in the node -5. Run! - -**For Videos:** - -1. Add a `MiniCPM-V Video Understanding` node -2. Enter the full path to your video file -3. Set your question -4. Adjust fps (5 is a good default) -5. Run! - -## Important Notes - -### First Run - -- **First time will be slow**: The model (~17GB) needs to download from Hugging Face -- **Requires internet**: For initial model download -- **Disk space**: Ensure you have ~20GB free space -- **Model location**: `~/.cache/huggingface/hub/` (or `C:\Users\YourName\.cache\huggingface\` on Windows) - -### System Requirements - -**Minimum:** -- GPU: 8GB VRAM (for MiniCPM-V-4.5) -- RAM: 16GB system RAM -- Disk: 20GB free space -- OS: Windows 10/11, Linux, macOS - -**Recommended:** -- GPU: 16GB+ VRAM (RTX 3090, 4090, A6000, etc.) -- RAM: 32GB+ system RAM -- Disk: SSD with 50GB+ free space -- CUDA: Latest version - -**Can run on CPU** but will be very slow (not recommended). - -### Video Requirements - -For video processing, you need: -- **decord** library (included in requirements) -- Supported formats: MP4, AVI, MOV, MKV, WebM -- Video codec: H.264, H.265, VP9 (most common formats work) - -### Troubleshooting - -#### "No module named 'transformers'" -```bash -pip install transformers -``` - -#### "No module named 'decord'" -```bash -pip install decord -``` - -On Windows, if decord fails, try: -```bash -pip install decord --no-deps -pip install numpy -``` - -#### "CUDA out of memory" -Solutions: -1. Close other GPU applications -2. Set `device` to "cpu" (slow but works) -3. Enable `unload_after_inference` to free memory after each use -4. For videos, reduce `max_num_frames` or `fps` - -#### "Model download fails" -1. Check internet connection -2. Verify Hugging Face is accessible -3. Try manual download: -```bash -pip install huggingface_hub -huggingface-cli download openbmb/MiniCPM-V-4_5 -``` - -#### Node doesn't appear in ComfyUI -1. Restart ComfyUI completely -2. Check console for errors -3. Verify installation in correct directory -4. Check that `__init__.py` files are present - -## Usage Tips - -### For Best Results - -**Image Analysis:** -- Use high-quality images -- For OCR, ensure text is clear and readable -- Multiple images can be compared in one query -- Use thinking mode for complex questions - -**Video Analysis:** -- Start with `fps=5` for most videos -- Increase fps for fast-action videos (sports, etc.) -- Longer videos benefit from lower fps -- Short clips can use higher fps - -### Performance Tips - -1. **Keep model loaded**: Don't enable `unload_after_inference` unless you need the memory -2. **Batch processing**: Load model once, process multiple items -3. **GPU recommended**: 10-100x faster than CPU -4. **Thinking mode**: Only enable for complex reasoning tasks - -### Privacy & Offline Use - -- **After first download**, models work fully offline -- **No data sent anywhere**: Everything runs locally -- **Models are cached**: Delete from `~/.cache/huggingface/` to remove - -## Example Prompts - -### Image Understanding - -- "Describe this image in detail." -- "What text is visible in this image?" -- "What is the main subject and what are they doing?" -- "Compare these two images and describe the differences." -- "What colors and artistic style are used here?" - -### Video Understanding - -- "Describe what happens in this video." -- "What actions does the person perform?" -- "Summarize the key events in this video." -- "What is the setting and atmosphere?" -- "Track the movement of the red car through the scene." - -## Advanced Usage - -### Multi-turn Conversations - -For images, you can have back-and-forth conversations: - -1. First query: Ask initial question, get `conversation_history` -2. Second query: Ask follow-up, feed previous `conversation_history` back in -3. Continue as needed - -### Custom Processing - -Adjust parameters for different use cases: - -**Fast preview:** -- `enable_thinking`: False -- `stream`: False -- Keep model loaded - -**Detailed analysis:** -- `enable_thinking`: True -- `stream`: True (see progress) -- Higher quality inputs - -**Video analysis:** -- Short clips: `fps=10`, `max_num_frames=180` -- Long videos: `fps=3`, `max_num_frames=180` -- Very long: `fps=1-2`, adjust as needed - -## Getting Help - -1. Check the console output - it shows detailed progress -2. See `nodes/minicpm/README.md` for full documentation -3. Report issues at the repository -4. Check [MiniCPM-V documentation](https://huggingface.co/openbmb/MiniCPM-V-4_5) - -## What's Supported - -โœ… Single image analysis -โœ… Multiple image analysis -โœ… Video understanding -โœ… Multi-turn conversations (images) -โœ… OCR and text extraction -โœ… Document understanding -โœ… Thinking mode for complex tasks -โœ… Both CUDA and CPU -โœ… Streaming responses -โœ… Memory management options - -## Model Information - -**MiniCPM-V-4.5:** -- Size: ~17GB download -- Parameters: 8.7B -- Best overall performance -- SOTA OCR capabilities -- Recommended for most use cases - -**MiniCPM-o-2.6:** -- Size: ~13GB download -- Parameters: Similar to 4.5 -- Alternative option -- Slightly different strengths - -Both models support the same features and API. - -## License - -- Models: Apache-2.0 License -- Free for commercial and personal use -- No API keys needed -- Fully local processing - ---- - -**Ready to go?** Just restart ComfyUI after installing dependencies, and look for the MiniCPM-V nodes! - diff --git a/SEED_AND_GENERATOR_FEATURES.md b/SEED_AND_GENERATOR_FEATURES.md deleted file mode 100644 index 5e76bbf..0000000 --- a/SEED_AND_GENERATOR_FEATURES.md +++ /dev/null @@ -1,194 +0,0 @@ -# ๐ŸŽฒ Seed Functionality & Universal Generator - -## โœ… **New Features Added** - -### 1. **Enhanced GPT Mini Node** -**Node Name**: "APNext GPT Mini Generator" - -#### **New Parameters Added**: -- **`seed`** (INT): Control randomization (-1 for auto, or specific number) -- **`randomize_each_run`** (BOOLEAN): Generate different variations each time -- **`variation_instruction`** (STRING): Custom instruction for how to vary outputs - -#### **How Seed Works**: -- **`seed = -1` + `randomize_each_run = True`**: New random seed every time โ†’ Different outputs -- **`seed = -1` + `randomize_each_run = False`**: Fixed seed (12345) โ†’ Consistent outputs -- **`seed = 12345` + `randomize_each_run = False`**: Use your specific seed โ†’ Reproducible outputs -- **`seed = 12345` + `randomize_each_run = True`**: Use your seed as base, add randomness โ†’ Controlled variation - -#### **Temperature Control**: -- **`randomize_each_run = True`**: Uses temperature 0.9 (more creative) -- **`randomize_each_run = False`**: Uses temperature 0.7 (more consistent) - -### 2. **NEW: APNext Universal Generator** ๐Ÿ†• -**Node Name**: "APNext Universal Generator" - -#### **Model Agnostic Support**: -- **Auto-detect**: Automatically chooses best available model -- **GPT Models**: gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-4, gpt-3.5-turbo -- **Gemini Models**: gemini-2.5-pro, gemini-2.5-flash, etc. -- **Format**: Select like "gpt:gpt-4o" or "gemini:gemini-2.5-flash" - -#### **Generation Modes**: -- **Creative**: Highly imaginative, takes creative liberties -- **Balanced**: Mix of creativity and accuracy (recommended) -- **Focused**: Stays close to original concept -- **Custom**: Uses your custom prompt exactly - -#### **Style Preferences**: -- **Cinematic**: Professional film-like descriptions -- **Photorealistic**: Natural, realistic details -- **Artistic**: Creative, stylized elements -- **Abstract**: Experimental, conceptual -- **Vintage**: Retro, nostalgic aesthetics -- **Modern**: Contemporary, clean styling - -#### **Detail Levels**: -- **Brief**: 50-100 words -- **Moderate**: 100-200 words -- **Detailed**: 200-300 words -- **Very Detailed**: 300+ words - -#### **Advanced Controls**: -- **Temperature**: Manual override (-1 for auto, 0.0-2.0 for manual) -- **Seed**: Same system as GPT Mini Node -- **Variation Instruction**: Custom guidance for variations - -## ๐ŸŽฏ **Solving BNP1111's Request** - -### **Problem**: "I can't find the setting for the random seed" -**โœ… SOLVED**: Both nodes now have `seed` and `randomize_each_run` parameters - -### **Problem**: "Generate different variations each time" -**โœ… SOLVED**: -- Set `randomize_each_run = True` (default) -- Customize `variation_instruction` for specific guidance -- Each run uses a different seed automatically - -### **Problem**: "Can we only call GPT-4 right now? Is it possible to call GPT-5?" -**โœ… SOLVED**: -- Updated models to include latest GPT models (gpt-4o, gpt-4o-mini, etc.) -- New Universal Generator supports both GPT and Gemini -- Auto-detection chooses best available model - -## ๐Ÿ“ **Usage Examples** - -### **Example 1: Different Variations Each Time** -**Node**: APNext GPT Mini Generator -- **input_text**: "A warrior in a fantasy forest" -- **randomize_each_run**: True โœ… -- **seed**: -1 (auto-generate) -- **variation_instruction**: "Create different poses, expressions, and forest environments each time" - -**Result**: Each run generates completely different warrior compositions! - -### **Example 2: Reproducible Results** -**Node**: APNext GPT Mini Generator -- **input_text**: "A warrior in a fantasy forest" -- **randomize_each_run**: False โœ… -- **seed**: 12345 (fixed) - -**Result**: Same output every time for consistent results. - -### **Example 3: Model-Agnostic Generation** -**Node**: APNext Universal Generator -- **input_text**: "A cyberpunk street scene" -- **model**: "auto-detect" (or "gpt:gpt-4o" or "gemini:gemini-2.5-flash") -- **generation_mode**: "Creative" -- **style_preference**: "Cinematic" - -**Result**: Uses best available model automatically! - -## ๐Ÿš€ **Advanced Workflow** - -### **For Maximum Variation** (BNP1111's use case): -``` -[Text Input] โ†’ [APNext Universal Generator] - โ†“ (randomize_each_run=True) - [Different output each time] - โ†“ - [Image Generator] - โ†“ - [Unique images every run!] -``` - -### **Chain Multiple Generators**: -``` -[Input] โ†’ [APNext Universal Generator] โ†’ [APNext GPT Mini] โ†’ [Final Output] - (Creative mode) (Detailed refinement) -``` - -## ๐ŸŽ›๏ธ **Parameter Guide** - -### **For Different Results Every Time**: -- **randomize_each_run**: True -- **seed**: -1 (auto) -- **temperature**: 0.9+ (high creativity) -- **generation_mode**: "Creative" or "Balanced" - -### **For Consistent Results**: -- **randomize_each_run**: False -- **seed**: Any fixed number (e.g., 12345) -- **temperature**: 0.7 (lower creativity) -- **generation_mode**: "Focused" - -### **For Controlled Variation**: -- **randomize_each_run**: True -- **seed**: Fixed number (e.g., 12345) -- **variation_instruction**: Specific guidance -- **Result**: Variations based on your seed + randomness - -## ๐Ÿ”ง **Technical Details** - -### **Seed Implementation**: -- Uses Python's `random.seed()` for consistent randomization -- Prints seed value to console for debugging -- Integrates with OpenAI's seed parameter (when supported) - -### **Model Support**: -- **GPT**: Full OpenAI API integration -- **Gemini**: Full Google AI integration -- **Auto-detect**: Checks API keys and selects best model -- **Fallback**: Graceful error handling - -### **Temperature Control**: -- **Auto-mode**: Adjusts based on generation mode -- **Manual**: Override with specific value -- **Randomization**: Higher temp when randomizing - -## ๐ŸŽ‰ **Benefits** - -### **For BNP1111**: -โœ… **Seed control** - Can generate different variations each time -โœ… **Latest models** - Access to GPT-4o and other modern models -โœ… **Variation control** - Custom instructions for how to vary outputs -โœ… **Reproducibility** - Can recreate specific results when needed - -### **For Everyone**: -โœ… **Model flexibility** - Choose GPT, Gemini, or auto-detect -โœ… **Style control** - Cinematic, photorealistic, artistic, etc. -โœ… **Detail control** - Brief to very detailed outputs -โœ… **Generation modes** - Creative, balanced, focused, custom - -## ๐Ÿ“Š **Model Comparison** - -| Model | Speed | Quality | Cost | Best For | -|-------|-------|---------|------|----------| -| **gpt-4o** | Medium | Highest | High | Professional work | -| **gpt-4o-mini** | Fast | High | Low | General use | -| **gpt-4-turbo** | Medium | Very High | Medium | Complex prompts | -| **gemini-2.5-flash** | Fast | High | Very Low | Rapid iteration | -| **gemini-2.5-pro** | Slow | Highest | Medium | Best quality | - ---- - -## ๐ŸŽฏ **Perfect Solution for BNP1111's Needs** - -The new system provides exactly what was requested: -1. โœ… **Random seed control** for different variations -2. โœ… **Access to latest models** including GPT-4o -3. โœ… **Model-agnostic generator** that works with any provider -4. โœ… **Variation instructions** for controlled creativity -5. โœ… **Different outputs each run** when desired - -**Both the enhanced GPT Mini Node and new Universal Generator are ready to use!** ๐ŸŽŠ