2023-09-28 11:26:17 +02:00
2023-09-11 13:18:47 +02:00
2023-09-16 15:27:33 +02:00
2023-09-28 11:20:39 +02:00
2023-09-28 11:26:17 +02:00
2023-09-26 10:12:34 +02:00

ComfyUI ExLlama Nodes

A simple prompt generator for ComfyUI utilizing ExLlama.

Installation

Clone the repository to custom_nodes in your ComfyUI directory and install dependencies:

git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes
pip install -r requirements.txt

If you see any ExLlama errors while loading, install it manually from here.
For example, on Windows with Python 3.10 and CUDA 11.7:

pip install https://github.com/turboderp/exllamav2/releases/download/v0.0.4/exllamav2-0.0.4+cu117-cp310-cp310-win_amd64.whl

Nodes

Name Description
Loader Loads GPTQ/EXL2 Llama models. You can find a lot of them on Hugging Face.
Clone the model repository or download all the files and place them in an empty directory, then specify the path in model_dir. The model.safetensors file won't work on its own.

ExLlama allocates memory based on max_seq_len. Lowering it is a good way to save on VRAM.
It's currently not possible to offload the model to RAM.
Generator Generates a string based on the given prompt for use with other nodes.
Default values correspond to the simple-1 preset from text-generation-webui.

ExLlama isn't deterministic, so the outputs may differ even with the same seed.
Previewer Displays generated outputs in the UI and appends them to workflow metadata.

Workflow

The workflow below can be opened in ComfyUI. Peak VRAM usage with SDXL around 10GB.
Model: MythoLogic-Mini-7B-GPTQ.

workflow

S
Description
No description provided
Readme MIT
214 KiB
Languages
Python 89.9%
JavaScript 10.1%