Emre 238a46fa2f Refactor DSD model loading and image processing functionalities
- Updated `load_model` and `download_and_load_model` methods in `dsd_nodes.py` to include new parameters for CPU memory management and offloading options.
- Introduced `DSDResizeSelector` class for flexible image resizing options, including methods for center cropping, padding, and fitting images.
- Enhanced `center_crop_and_resize` and added new utility functions in `utils.py` for improved image processing.
- Updated `README.md` to document the new resizing options available in the DSD Image Generator.

These changes improve memory efficiency and provide users with more control over image processing during model inference.
2025-03-14 00:26:32 +03:00
2025-03-12 06:14:58 +03:00
2025-03-12 06:14:58 +03:00
2025-03-12 06:14:58 +03:00
2025-03-12 06:14:58 +03:00
2025-03-12 06:14:58 +03:00

ComfyUI-DSD

An Unofficial ComfyUI custom node package that integrates Diffusion Self-Distillation (DSD) for zero-shot customized image generation.

DSD is a model for subject-preserving image generation that allows you to create images of a specific subject in novel contexts without per-instance tuning.

Features

  • Subject-preserving image generation using DSD model
  • Gemini API prompt enhancement
  • Direct model download from Hugging Face
  • Fine-grained control over generation parameters

Installation

  1. Clone this repository into your ComfyUI custom_nodes folder:
cd ComfyUI/custom_nodes
git clone https://github.com/irreveloper/ComfyUI-DSD.git
  1. Install the required dependencies:
pip install -r requirements.txt
  1. Get the model files (two options):

    • Option 1: Use the DSD Model Downloader node in ComfyUI to automatically download the model
    • Option 2: Download manually from Hugging Face or Google Drive

    The model files will be stored in:

    • ComfyUI/models/dsd_model/transformer/ (for transformer files)
    • ComfyUI/models/dsd_model/pytorch_lora_weights.safetensors (for LoRA file)
  2. Restart ComfyUI

Available Nodes

  1. DSD Model Downloader: Automatically downloads the model from Hugging Face

  2. DSD Model Loader: Loads a pre-downloaded model

  3. DSD Model Selector: Helps select models from local directories

  4. DSD Gemini Prompt Enhancer: Uses Google's Gemini API to enhance prompts for better image generation results. The API key can be provided in two ways:

    • As an input parameter to the node (not recommended for sharing workflows)
    • Through the GEMINI_API_KEY environment variable (strongly recommended)

    Note: To use the enhanced prompts, make sure to enable the use_gemini_prompt option on the DSD Image Generator node. If you don't enter a API Key it will be skipped automatically.

  5. DSD Image Generator: Generates images with the DSD model

  6. DSD Resize Selector: Provides flexible image resizing options for the DSD Image Generator:

    • center_crop: Center crops the image and resizes (default behavior)
    • crop: Simple center crop and resize
    • pad: Preserves aspect ratio and adds padding to reach target size
    • fit: Resizes to target dimensions without preserving aspect ratio
    • Additional options for interpolation method and padding color

Basic Workflow

Sample1

Sample2

Troubleshooting

  • Memory Issues: Try reducing precision (use bfloat16), lower resolution, or fewer steps
  • Gemini API: Ensure you have a valid API key (can be set via GEMINI_API_KEY environment variable)
  • Model Loading: If you see errors, try using the Model Downloader node to re-download files

Examples

Check the examples directory for sample workflows.

S
Description
No description provided
Readme GPL-3.0
10 MiB
Languages
Python 97.1%
JavaScript 2.9%