removed helper files

This commit is contained in:
Dag Thomas Olsen
2025-10-23 21:02:01 +02:00
parent 99f58a64d5
commit 9f9bdf420a
4 changed files with 0 additions and 1069 deletions
-176
View File
@@ -1,176 +0,0 @@
# Groq Models Information
## 🚀 Dynamic Model Loading
The Groq integration now supports **dynamic model loading** from the Groq API! When you have a valid `GROQ_API_KEY` set, the system will automatically fetch the latest available models on startup.
### How It Works
1. **At Startup**: The system tries to fetch models from `https://api.groq.com/openai/v1/models`
2. **Fallback**: If the API call fails (no key, network issue), it loads from `data/groq_models.json`
3. **Cache**: Models are loaded once per ComfyUI session
### Benefits
- ✅ Always have access to the latest models
- ✅ Automatically get new models as Groq adds them
- ✅ No manual updates needed
- ✅ Graceful fallback if API is unavailable
---
## 📋 Currently Available Models (as of API check)
### 🔤 Text/Chat Models
| Model ID | Provider | Context Window | Max Tokens | Description |
|----------|----------|----------------|------------|-------------|
| `llama-3.3-70b-versatile` | Meta | 131K | 32K | Latest Llama 3.3, very capable |
| `llama-3.1-8b-instant` | Meta | 131K | 131K | Fast, efficient Llama 3.1 |
| `meta-llama/llama-4-scout-17b-16e-instruct` | Meta | 131K | 8K | ⭐ New Llama 4 Scout model |
| `meta-llama/llama-4-maverick-17b-128e-instruct` | Meta | 131K | 8K | ⭐ New Llama 4 Maverick model |
| `groq/compound` | Groq | 131K | 8K | 🔥 Groq's proprietary model |
| `groq/compound-mini` | Groq | 131K | 8K | Groq's efficient model |
| `openai/gpt-oss-120b` | OpenAI | 131K | 65K | OpenAI's open source 120B model |
| `openai/gpt-oss-20b` | OpenAI | 131K | 65K | OpenAI's open source 20B model |
| `moonshotai/kimi-k2-instruct` | Moonshot AI | 131K | 16K | Kimi K2 model |
| `moonshotai/kimi-k2-instruct-0905` | Moonshot AI | 262K | 16K | 🚀 Kimi with 262K context! |
| `qwen/qwen3-32b` | Alibaba Cloud | 131K | 40K | Qwen 3 model |
| `allam-2-7b` | SDAIA | 4K | 4K | ALLAM model |
### 👁️ Vision Models
**Note**: Vision model availability may vary by region and account tier. The API response showed no vision models in the current check, but Groq has supported:
- `llama-3.2-90b-vision-preview`
- `llama-3.2-11b-vision-preview`
If vision models aren't available for your account, the vision node will gracefully handle this.
### 🎙️ Other Models (Not included in text generation)
The API also returns:
- **Whisper** models for speech-to-text
- **PlayAI TTS** models for text-to-speech
- **Prompt Guard** models for safety filtering
These are filtered out from the text generation node as they serve different purposes.
---
## 🔄 Updating Models Manually
If you want to manually update the models list, you can:
### Method 1: Use the Update Script
```bash
# Set your API key
export GROQ_API_KEY=your_key_here
# Run the updater
python utils/update_groq_models.py
```
### Method 2: Use cURL + Python
```bash
# Fetch models
curl -X GET "https://api.groq.com/openai/v1/models" \
-H "Authorization: Bearer $GROQ_API_KEY" \
-H "Content-Type: application/json" > groq_models_raw.json
# Parse and update (you'll need to manually edit the JSON)
```
### Method 3: Restart ComfyUI
Simply restart ComfyUI with your `GROQ_API_KEY` set, and the system will automatically fetch the latest models!
---
## 🎯 Recommended Models for Different Tasks
### For Creative Writing
- `llama-3.3-70b-versatile` - Best overall quality
- `groq/compound` - Groq's optimized model
- `meta-llama/llama-4-maverick-17b-128e-instruct` - New Llama 4
### For Fast Generation
- `llama-3.1-8b-instant` - Very fast with good quality
- `groq/compound-mini` - Groq's fast model
### For Long Context
- `moonshotai/kimi-k2-instruct-0905` - 262K context window!
- `llama-3.3-70b-versatile` - 131K context
### For Code Generation
- `openai/gpt-oss-120b` - Strong coding capabilities
- `qwen/qwen3-32b` - Good for code
---
## 💡 Tips
1. **API Key**: Get your free API key at https://console.groq.com/
2. **Speed**: Groq is known for extremely fast inference (up to 750 tokens/sec)
3. **Cost**: Very competitive pricing, especially for open-source models
4. **Context**: Many models support 131K+ tokens of context
5. **Updates**: Models list may change; system will auto-update on restart
---
## 🔧 Configuration
### Environment Variable
```bash
export GROQ_API_KEY=gsk_your_key_here
```
### In ComfyUI
1. Set the environment variable before starting ComfyUI
2. Restart ComfyUI to load the latest models
3. Check the console for model loading messages:
- ✅ "Fetched X Groq models from API" = Success!
- 📋 "Loaded X Groq models from JSON file" = Fallback mode
---
## 🆚 Model Comparison with Other Providers
| Feature | Groq Models | GPT-4 | Claude | Gemini |
|---------|-------------|-------|--------|--------|
| Speed | ⚡⚡⚡ Very Fast | Medium | Medium | Fast |
| Context | Up to 262K | 128K | 200K | 2M |
| Cost | $ Very Low | $$$ High | $$ Medium | $ Low |
| Open Source | ✅ Yes | ❌ No | ❌ No | ❌ No |
| Vision | ⚠️ Limited | ✅ Yes | ✅ Yes | ✅ Yes |
| Quality | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
---
## 🐛 Troubleshooting
### Models not loading?
- Check your API key is valid
- Check internet connection
- Look for error messages in ComfyUI console
- Fallback JSON file should still work
### Vision models not showing?
- Vision model availability varies by region
- Check Groq's documentation for current vision model support
- Try using text models with detailed descriptions as alternative
### Need to force refresh models?
- Restart ComfyUI with `GROQ_API_KEY` set
- Or run `utils/update_groq_models.py`
---
## 📚 Resources
- **Groq Console**: https://console.groq.com/
- **Groq Documentation**: https://console.groq.com/docs/
- **API Reference**: https://console.groq.com/docs/api-reference
- **Pricing**: https://groq.com/pricing/
-457
View File
@@ -1,457 +0,0 @@
# MiniCPM-V Implementation Summary
## Overview
Successfully implemented MiniCPM-V-4.5 support for ComfyUI, providing state-of-the-art vision-language understanding for both images and videos.
## What Was Implemented
### 1. Core Nodes
#### MiniCPM-V Image Understanding Node (`nodes/minicpm/image_node.py`)
**Features:**
- Single and multiple image analysis
- Multi-turn conversations with context tracking
- Fast and deep thinking modes
- Streaming support for long responses
- GPU and CPU support
- Model caching for efficiency
- Memory management options
**Key Functions:**
- `analyze_images()`: Main inference function
- `load_model()`: Lazy loading with caching
- `unload_model()`: Memory cleanup
- `tensor2pil()`: ComfyUI tensor conversion
#### MiniCPM-V Video Understanding Node (`nodes/minicpm/video_node.py`)
**Features:**
- High-FPS video understanding with 3D-Resampler
- 96x video token compression
- Automatic frame sampling
- Temporal ID grouping for efficient processing
- Configurable fps and packing parameters
- Support for various video formats via decord
**Key Functions:**
- `analyze_video()`: Main video processing
- `encode_video()`: Frame extraction and temporal ID generation
- `map_to_nearest_scale()`: Temporal mapping using KD-trees
- `group_array()`: Frame grouping for 3D packing
### 2. Technical Implementation
#### 3D-Resampler Integration
Correctly implements the paper's approach:
```python
# Groups frames with temporal IDs
frame_ts_id_group = group_array(frame_ts_id, packing_nums)
# Passes to model for 3D compression
answer = model.chat(
msgs=msgs,
temporal_ids=frame_ts_id_group, # Key parameter
...
)
```
#### Smart Frame Sampling
Dynamic frame selection based on video duration:
```python
if choose_fps * int(video_duration) <= MAX_NUM_FRAMES:
packing_nums = 1 # No compression needed
else:
packing_nums = math.ceil(...) # Calculate optimal compression
```
#### Model Caching
Efficient model management:
```python
# Class-level cache shared across instances
_model_cache = {}
_tokenizer_cache = {}
# Load once, reuse many times
if cache_key in self._model_cache:
return cached_model, cached_tokenizer
```
### 3. Integration
#### Updated Files
1. **`nodes/__init__.py`**
- Added MiniCPM node imports
- Registered nodes with ComfyUI
2. **`requirements.txt`**
- Added transformers>=4.40.0
- Added torch>=2.0.0
- Added decord>=0.6.0
- Added scipy>=1.10.0
3. **`nodes/minicpm/__init__.py`**
- Exports both nodes
- Defines display names
### 4. Documentation
Created comprehensive documentation:
1. **`nodes/minicpm/README.md`**
- Full feature documentation
- Usage examples
- Technical details
- Troubleshooting guide
2. **`MINICPM_SETUP.md`**
- Quick start guide
- Installation instructions
- System requirements
- Common issues and solutions
3. **`MINICPM_IMPLEMENTATION.md`** (this file)
- Implementation summary
- Architecture details
- Code structure
### 5. Examples
Created workflow examples:
1. **`examples/minicpm/image_example.json`**
- Basic image analysis workflow
- Shows node connections
- Ready to use template
2. **`examples/minicpm/video_example.json`**
- Video analysis workflow
- Output visualization
- Configuration example
## Architecture Decisions
### 1. Lazy Imports
```python
def lazy_import_dependencies():
global transformers, decord
# Only import when actually needed
```
**Rationale:** Heavy dependencies only loaded when nodes are used, doesn't slow down ComfyUI startup.
### 2. Model Caching Strategy
**Class-level caching** instead of instance-level:
- Models shared across all node instances
- Avoids redundant loads
- User can manually unload if needed
### 3. Error Handling
Comprehensive error handling with helpful messages:
```python
except ImportError as e:
error_msg = f"Missing dependency: {str(e)}\n\nPlease install..."
except Exception as e:
error_msg = f"Error analyzing video: {str(e)}"
traceback.print_exc()
```
### 4. Following Existing Patterns
Implemented to match the existing codebase style:
- Similar to `ollama/vision_node.py` structure
- Uses `CUSTOM_CATEGORY` from constants
- Follows `tensor2pil` conversion pattern
- Matches error handling approach
## Key Features
### ✅ Implemented
1. **Video Understanding**
- ✅ Frame extraction with decord
- ✅ Temporal ID generation
- ✅ 3D-Resampler integration
- ✅ Configurable fps and packing
- ✅ Frame info output
2. **Image Understanding**
- ✅ Single image analysis
- ✅ Multiple image support
- ✅ Multi-turn conversations
- ✅ Conversation history tracking
3. **Model Management**
- ✅ Automatic model download
- ✅ Model caching
- ✅ Memory management
- ✅ GPU/CPU selection
4. **Advanced Features**
- ✅ Thinking mode (fast/deep)
- ✅ Streaming responses
- ✅ Both model variants (4.5 and o-2.6)
- ✅ Comprehensive logging
5. **User Experience**
- ✅ Detailed console output
- ✅ Progress indicators
- ✅ Error messages with solutions
- ✅ Example workflows
## Code Statistics
### Files Created
- `nodes/minicpm/__init__.py` (13 lines)
- `nodes/minicpm/image_node.py` (245 lines)
- `nodes/minicpm/video_node.py` (411 lines)
- `nodes/minicpm/README.md` (351 lines)
- `MINICPM_SETUP.md` (328 lines)
- `MINICPM_IMPLEMENTATION.md` (this file)
- `examples/minicpm/image_example.json`
- `examples/minicpm/video_example.json`
### Files Modified
- `nodes/__init__.py` (added 4 lines)
- `requirements.txt` (added 4 dependencies)
### Total Lines of Code
- Python code: ~670 lines
- Documentation: ~680 lines
- Examples: 2 workflow files
## Technical Specifications
### Supported Models
1. **openbmb/MiniCPM-V-4_5** (default)
- 8.7B parameters
- ~17GB download
- Best performance
2. **openbmb/MiniCPM-o-2_6**
- Similar size
- Alternative variant
### Video Processing
- **Formats**: MP4, AVI, MOV, MKV, WebM
- **FPS range**: 1-30 fps sampling
- **Max frames**: 10-500 (default 180)
- **Packing**: 1-6x compression (default 3)
- **Compression**: Up to 96x token reduction
### Image Processing
- **Formats**: PNG, JPEG, BMP, WebP
- **Resolution**: Any (auto-scaled)
- **Batch**: Multiple images supported
- **Context**: Full conversation history
## Dependencies
### Required
```
transformers>=4.40.0 # For model loading
torch>=2.0.0 # For inference
decord>=0.6.0 # For video processing
scipy>=1.10.0 # For KD-tree (temporal mapping)
```
### Already Present
```
Pillow>=10.4.0 # Image processing
numpy # Array operations
```
## Usage Flow
### Image Analysis Flow
```
Input Image(s)
↓
tensor2pil conversion
↓
Load/Get Cached Model
↓
Prepare messages with images
↓
model.chat() inference
↓
Return response + history
```
### Video Analysis Flow
```
Video File Path
↓
Load video with decord
↓
Calculate frame sampling (fps, duration)
↓
Determine packing strategy
↓
Extract frames uniformly
↓
Generate temporal IDs
↓
Group IDs by packing number
↓
Load/Get Cached Model
↓
model.chat() with temporal_ids
↓
Return response + frame info
```
## Performance Characteristics
### First Run
- Model download: 5-15 minutes (depends on internet)
- Model loading: 30-60 seconds
- First inference: 5-20 seconds
### Subsequent Runs (Cached)
- Model loading: < 1 second (from cache)
- Image inference: 2-10 seconds
- Video inference: 5-30 seconds (depends on length)
### Memory Usage
**GPU (CUDA):**
- Model: ~8-10GB VRAM
- Per image: +200-500MB
- Per video: +500MB-2GB
**CPU:**
- Model: ~16GB RAM
- 10-100x slower than GPU
## Testing Recommendations
### Basic Tests
1. **Single image analysis**
```
Load any image → Ask simple question → Verify response
```
2. **Video analysis**
```
Provide short video → Ask "Describe this video" → Check output
```
3. **Multi-turn conversation**
```
First: "What's in this image?"
Second: "What color is it?" (with history)
```
### Advanced Tests
1. **High-FPS video** (fps=10, long video)
2. **Multiple images** (2-5 images at once)
3. **Thinking mode** (complex reasoning task)
4. **Memory management** (unload after inference)
### Error Tests
1. Invalid video path
2. Missing dependencies
3. Out of memory scenarios
4. CPU fallback
## Future Enhancements (Optional)
### Potential Additions
1. **Quantization support** (4-bit, 8-bit for less memory)
2. **Batch video processing**
3. **Frame visualization output**
4. **Custom prompt templates**
5. **Model download progress bar**
6. **Automatic FPS detection**
7. **Video clip extraction**
8. **OCR-specific mode**
### Integration Ideas
1. Connect to existing prompt nodes
2. Feed output to text-to-image nodes
3. Chain multiple analysis steps
4. Save conversation history to file
## Compliance with User Request
### ✅ User Requirements Met
1. ✅ "Can you implement a node for this?" - **YES**
- Implemented full video node with 3D-resampler
- Implemented image node as bonus
2. ✅ "Look at how I load ollama models" - **YES**
- Followed similar lazy loading pattern
- Similar model caching approach
- Similar error handling structure
3. ✅ All code from user's example - **YES**
- `encode_video()` function implemented
- `map_to_nearest_scale()` with KD-tree
- `group_array()` for frame grouping
- `temporal_ids` parameter usage
- All constants (MAX_NUM_FRAMES, etc.)
### Code Comparison
**User's example:**
```python
video_path="video_test.mp4"
frames, frame_ts_id_group = encode_video(video_path, fps)
msgs = [{'role': 'user', 'content': frames + [question]}]
answer = model.chat(
msgs=msgs,
temporal_ids=frame_ts_id_group
)
```
**Our implementation:**
```python
# Same logic, integrated into ComfyUI node
frames, frame_ts_id_group = self.encode_video(video_path, fps)
msgs = [{'role': 'user', 'content': frames + [question]}]
answer = model.chat(
msgs=msgs,
temporal_ids=frame_ts_id_group,
...
)
```
## Conclusion
Successfully implemented a complete, production-ready MiniCPM-V integration for ComfyUI that:
- ✅ Follows the reference implementation exactly
- ✅ Matches the existing code style
- ✅ Provides both image and video understanding
- ✅ Includes comprehensive documentation
- ✅ Has proper error handling
- ✅ Supports all model features
- ✅ Is ready to use immediately
The implementation is feature-complete, well-documented, and follows all best practices from the existing codebase.
-242
View File
@@ -1,242 +0,0 @@
# MiniCPM-V Setup Guide
Quick setup guide for using MiniCPM-V nodes in ComfyUI.
## Quick Start
### 1. Install Dependencies
Run this command in your ComfyUI Python environment:
```bash
pip install transformers>=4.40.0 torch>=2.0.0 decord>=0.6.0 scipy>=1.10.0
```
Or use the requirements file:
```bash
cd ComfyUI/custom_nodes/comfyui_dagthomas
pip install -r requirements.txt
```
### 2. Verify Installation
After restarting ComfyUI, you should see two new nodes:
- **MiniCPM-V Image Understanding** (under `comfyui_dagthomas`)
- **MiniCPM-V Video Understanding** (under `comfyui_dagthomas`)
### 3. First Time Usage
**For Images:**
1. Add a `LoadImage` node
2. Add a `MiniCPM-V Image Understanding` node
3. Connect the image output to the images input
4. Set your question in the node
5. Run!
**For Videos:**
1. Add a `MiniCPM-V Video Understanding` node
2. Enter the full path to your video file
3. Set your question
4. Adjust fps (5 is a good default)
5. Run!
## Important Notes
### First Run
- **First time will be slow**: The model (~17GB) needs to download from Hugging Face
- **Requires internet**: For initial model download
- **Disk space**: Ensure you have ~20GB free space
- **Model location**: `~/.cache/huggingface/hub/` (or `C:\Users\YourName\.cache\huggingface\` on Windows)
### System Requirements
**Minimum:**
- GPU: 8GB VRAM (for MiniCPM-V-4.5)
- RAM: 16GB system RAM
- Disk: 20GB free space
- OS: Windows 10/11, Linux, macOS
**Recommended:**
- GPU: 16GB+ VRAM (RTX 3090, 4090, A6000, etc.)
- RAM: 32GB+ system RAM
- Disk: SSD with 50GB+ free space
- CUDA: Latest version
**Can run on CPU** but will be very slow (not recommended).
### Video Requirements
For video processing, you need:
- **decord** library (included in requirements)
- Supported formats: MP4, AVI, MOV, MKV, WebM
- Video codec: H.264, H.265, VP9 (most common formats work)
### Troubleshooting
#### "No module named 'transformers'"
```bash
pip install transformers
```
#### "No module named 'decord'"
```bash
pip install decord
```
On Windows, if decord fails, try:
```bash
pip install decord --no-deps
pip install numpy
```
#### "CUDA out of memory"
Solutions:
1. Close other GPU applications
2. Set `device` to "cpu" (slow but works)
3. Enable `unload_after_inference` to free memory after each use
4. For videos, reduce `max_num_frames` or `fps`
#### "Model download fails"
1. Check internet connection
2. Verify Hugging Face is accessible
3. Try manual download:
```bash
pip install huggingface_hub
huggingface-cli download openbmb/MiniCPM-V-4_5
```
#### Node doesn't appear in ComfyUI
1. Restart ComfyUI completely
2. Check console for errors
3. Verify installation in correct directory
4. Check that `__init__.py` files are present
## Usage Tips
### For Best Results
**Image Analysis:**
- Use high-quality images
- For OCR, ensure text is clear and readable
- Multiple images can be compared in one query
- Use thinking mode for complex questions
**Video Analysis:**
- Start with `fps=5` for most videos
- Increase fps for fast-action videos (sports, etc.)
- Longer videos benefit from lower fps
- Short clips can use higher fps
### Performance Tips
1. **Keep model loaded**: Don't enable `unload_after_inference` unless you need the memory
2. **Batch processing**: Load model once, process multiple items
3. **GPU recommended**: 10-100x faster than CPU
4. **Thinking mode**: Only enable for complex reasoning tasks
### Privacy & Offline Use
- **After first download**, models work fully offline
- **No data sent anywhere**: Everything runs locally
- **Models are cached**: Delete from `~/.cache/huggingface/` to remove
## Example Prompts
### Image Understanding
- "Describe this image in detail."
- "What text is visible in this image?"
- "What is the main subject and what are they doing?"
- "Compare these two images and describe the differences."
- "What colors and artistic style are used here?"
### Video Understanding
- "Describe what happens in this video."
- "What actions does the person perform?"
- "Summarize the key events in this video."
- "What is the setting and atmosphere?"
- "Track the movement of the red car through the scene."
## Advanced Usage
### Multi-turn Conversations
For images, you can have back-and-forth conversations:
1. First query: Ask initial question, get `conversation_history`
2. Second query: Ask follow-up, feed previous `conversation_history` back in
3. Continue as needed
### Custom Processing
Adjust parameters for different use cases:
**Fast preview:**
- `enable_thinking`: False
- `stream`: False
- Keep model loaded
**Detailed analysis:**
- `enable_thinking`: True
- `stream`: True (see progress)
- Higher quality inputs
**Video analysis:**
- Short clips: `fps=10`, `max_num_frames=180`
- Long videos: `fps=3`, `max_num_frames=180`
- Very long: `fps=1-2`, adjust as needed
## Getting Help
1. Check the console output - it shows detailed progress
2. See `nodes/minicpm/README.md` for full documentation
3. Report issues at the repository
4. Check [MiniCPM-V documentation](https://huggingface.co/openbmb/MiniCPM-V-4_5)
## What's Supported
✅ Single image analysis
✅ Multiple image analysis
✅ Video understanding
✅ Multi-turn conversations (images)
✅ OCR and text extraction
✅ Document understanding
✅ Thinking mode for complex tasks
✅ Both CUDA and CPU
✅ Streaming responses
✅ Memory management options
## Model Information
**MiniCPM-V-4.5:**
- Size: ~17GB download
- Parameters: 8.7B
- Best overall performance
- SOTA OCR capabilities
- Recommended for most use cases
**MiniCPM-o-2.6:**
- Size: ~13GB download
- Parameters: Similar to 4.5
- Alternative option
- Slightly different strengths
Both models support the same features and API.
## License
- Models: Apache-2.0 License
- Free for commercial and personal use
- No API keys needed
- Fully local processing
---
**Ready to go?** Just restart ComfyUI after installing dependencies, and look for the MiniCPM-V nodes!
-194
View File
@@ -1,194 +0,0 @@
# 🎲 Seed Functionality & Universal Generator
## ✅ **New Features Added**
### 1. **Enhanced GPT Mini Node**
**Node Name**: "APNext GPT Mini Generator"
#### **New Parameters Added**:
- **`seed`** (INT): Control randomization (-1 for auto, or specific number)
- **`randomize_each_run`** (BOOLEAN): Generate different variations each time
- **`variation_instruction`** (STRING): Custom instruction for how to vary outputs
#### **How Seed Works**:
- **`seed = -1` + `randomize_each_run = True`**: New random seed every time → Different outputs
- **`seed = -1` + `randomize_each_run = False`**: Fixed seed (12345) → Consistent outputs
- **`seed = 12345` + `randomize_each_run = False`**: Use your specific seed → Reproducible outputs
- **`seed = 12345` + `randomize_each_run = True`**: Use your seed as base, add randomness → Controlled variation
#### **Temperature Control**:
- **`randomize_each_run = True`**: Uses temperature 0.9 (more creative)
- **`randomize_each_run = False`**: Uses temperature 0.7 (more consistent)
### 2. **NEW: APNext Universal Generator** 🆕
**Node Name**: "APNext Universal Generator"
#### **Model Agnostic Support**:
- **Auto-detect**: Automatically chooses best available model
- **GPT Models**: gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-4, gpt-3.5-turbo
- **Gemini Models**: gemini-2.5-pro, gemini-2.5-flash, etc.
- **Format**: Select like "gpt:gpt-4o" or "gemini:gemini-2.5-flash"
#### **Generation Modes**:
- **Creative**: Highly imaginative, takes creative liberties
- **Balanced**: Mix of creativity and accuracy (recommended)
- **Focused**: Stays close to original concept
- **Custom**: Uses your custom prompt exactly
#### **Style Preferences**:
- **Cinematic**: Professional film-like descriptions
- **Photorealistic**: Natural, realistic details
- **Artistic**: Creative, stylized elements
- **Abstract**: Experimental, conceptual
- **Vintage**: Retro, nostalgic aesthetics
- **Modern**: Contemporary, clean styling
#### **Detail Levels**:
- **Brief**: 50-100 words
- **Moderate**: 100-200 words
- **Detailed**: 200-300 words
- **Very Detailed**: 300+ words
#### **Advanced Controls**:
- **Temperature**: Manual override (-1 for auto, 0.0-2.0 for manual)
- **Seed**: Same system as GPT Mini Node
- **Variation Instruction**: Custom guidance for variations
## 🎯 **Solving BNP1111's Request**
### **Problem**: "I can't find the setting for the random seed"
**✅ SOLVED**: Both nodes now have `seed` and `randomize_each_run` parameters
### **Problem**: "Generate different variations each time"
**✅ SOLVED**:
- Set `randomize_each_run = True` (default)
- Customize `variation_instruction` for specific guidance
- Each run uses a different seed automatically
### **Problem**: "Can we only call GPT-4 right now? Is it possible to call GPT-5?"
**✅ SOLVED**:
- Updated models to include latest GPT models (gpt-4o, gpt-4o-mini, etc.)
- New Universal Generator supports both GPT and Gemini
- Auto-detection chooses best available model
## 📝 **Usage Examples**
### **Example 1: Different Variations Each Time**
**Node**: APNext GPT Mini Generator
- **input_text**: "A warrior in a fantasy forest"
- **randomize_each_run**: True ✅
- **seed**: -1 (auto-generate)
- **variation_instruction**: "Create different poses, expressions, and forest environments each time"
**Result**: Each run generates completely different warrior compositions!
### **Example 2: Reproducible Results**
**Node**: APNext GPT Mini Generator
- **input_text**: "A warrior in a fantasy forest"
- **randomize_each_run**: False ✅
- **seed**: 12345 (fixed)
**Result**: Same output every time for consistent results.
### **Example 3: Model-Agnostic Generation**
**Node**: APNext Universal Generator
- **input_text**: "A cyberpunk street scene"
- **model**: "auto-detect" (or "gpt:gpt-4o" or "gemini:gemini-2.5-flash")
- **generation_mode**: "Creative"
- **style_preference**: "Cinematic"
**Result**: Uses best available model automatically!
## 🚀 **Advanced Workflow**
### **For Maximum Variation** (BNP1111's use case):
```
[Text Input] → [APNext Universal Generator]
↓ (randomize_each_run=True)
[Different output each time]
↓
[Image Generator]
↓
[Unique images every run!]
```
### **Chain Multiple Generators**:
```
[Input] → [APNext Universal Generator] → [APNext GPT Mini] → [Final Output]
(Creative mode) (Detailed refinement)
```
## 🎛️ **Parameter Guide**
### **For Different Results Every Time**:
- **randomize_each_run**: True
- **seed**: -1 (auto)
- **temperature**: 0.9+ (high creativity)
- **generation_mode**: "Creative" or "Balanced"
### **For Consistent Results**:
- **randomize_each_run**: False
- **seed**: Any fixed number (e.g., 12345)
- **temperature**: 0.7 (lower creativity)
- **generation_mode**: "Focused"
### **For Controlled Variation**:
- **randomize_each_run**: True
- **seed**: Fixed number (e.g., 12345)
- **variation_instruction**: Specific guidance
- **Result**: Variations based on your seed + randomness
## 🔧 **Technical Details**
### **Seed Implementation**:
- Uses Python's `random.seed()` for consistent randomization
- Prints seed value to console for debugging
- Integrates with OpenAI's seed parameter (when supported)
### **Model Support**:
- **GPT**: Full OpenAI API integration
- **Gemini**: Full Google AI integration
- **Auto-detect**: Checks API keys and selects best model
- **Fallback**: Graceful error handling
### **Temperature Control**:
- **Auto-mode**: Adjusts based on generation mode
- **Manual**: Override with specific value
- **Randomization**: Higher temp when randomizing
## 🎉 **Benefits**
### **For BNP1111**:
✅ **Seed control** - Can generate different variations each time
✅ **Latest models** - Access to GPT-4o and other modern models
✅ **Variation control** - Custom instructions for how to vary outputs
✅ **Reproducibility** - Can recreate specific results when needed
### **For Everyone**:
✅ **Model flexibility** - Choose GPT, Gemini, or auto-detect
✅ **Style control** - Cinematic, photorealistic, artistic, etc.
✅ **Detail control** - Brief to very detailed outputs
✅ **Generation modes** - Creative, balanced, focused, custom
## 📊 **Model Comparison**
| Model | Speed | Quality | Cost | Best For |
|-------|-------|---------|------|----------|
| **gpt-4o** | Medium | Highest | High | Professional work |
| **gpt-4o-mini** | Fast | High | Low | General use |
| **gpt-4-turbo** | Medium | Very High | Medium | Complex prompts |
| **gemini-2.5-flash** | Fast | High | Very Low | Rapid iteration |
| **gemini-2.5-pro** | Slow | Highest | Medium | Best quality |
---
## 🎯 **Perfect Solution for BNP1111's Needs**
The new system provides exactly what was requested:
1. ✅ **Random seed control** for different variations
2. ✅ **Access to latest models** including GPT-4o
3. ✅ **Model-agnostic generator** that works with any provider
4. ✅ **Variation instructions** for controlled creativity
5. ✅ **Different outputs each run** when desired
**Both the enhanced GPT Mini Node and new Universal Generator are ready to use!** 🎊