- Add QwenGPUInference node for AI photo prompt optimization
- Implement smart GPU memory management with automatic detection
- Support CPU offload when GPU memory is insufficient
- Auto-download model config files from HuggingFace
- Remove <think> tags from model output
- Add bilingual (Chinese/English) support
- Remove deprecated GGUF inference node and related files
- Update README with comprehensive documentation
Features:
- Automatic model detection (qwen_3_4b.safetensors)
- Three loading strategies: Full GPU / CPU Offload / CPU-only
- Memory conflict prevention with ComfyUI models
- Professional photography prompt generation
- Default max_tokens: 2048 for detailed prompts
- Custom system prompt for photo optimization
Performance:
- Full GPU: ~26-30 tokens/second
- CPU Offload: ~1-2 tokens/second (reliable fallback)
- First load: 7-130 seconds depending on hardware
- Subsequent loads: Near-instant (model cached)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>