Files
Zuellni-ComfyUI-ExLlama-Nodes/README.md
T
2023-09-20 10:45:19 +02:00

2.1 KiB

ComfyUI ExLlama Nodes

A simple prompt generator for ComfyUI utilizing ExLlama.

Installation

Clone the repository to custom_nodes in your ComfyUI directory and install the dependencies:

git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes
python -m pip install -r requirements.txt

Next, install the latest pre-built ExLlama wheel from https://github.com/jllllll/exllama/releases/latest.
Choose the version matching your platform, Python, and PyTorch CUDA/ROCm.

Example for Windows with Python 3.10 and CUDA 11.8, which should match the portable ComfyUI build:

python -m pip install https://github.com/jllllll/exllama/releases/download/0.0.17/exllama-0.0.17+cu118-cp310-cp310-win_amd64.whl

Nodes

Name Description
Loader Loads 4-bit GPTQ Llama/2 models. You can find a lot of them on Hugging Face.
Clone the model repository or download all the files in it and place them in an empty directory, then specify the path in model_dir. The model.safetensors file won't work on its own.

ExLlama allocates memory based on max_seq_len. Lowering it is a good way to save on VRAM. It's currently not possible to offload the model to RAM.
Generator Returns a string based on the given prompt for use with other nodes. Default values correspond to the simple-1 preset from text-generation-webui. ExLlama isn't deterministic, so the outputs may differ slightly even with the same seed.

To load a LoRA specify the path to its directory in lora_dir. It should contain adapter_model.bin and adapter_config.json.
Previewer Displays generated outputs in the UI.

Workflow

The workflow below can be loaded directly in ComfyUI. Model used: MythoLogic-Mini-7B.

workflow