2023-11-24 17:42:03 +01:00
2023-09-11 13:18:47 +02:00
2023-11-23 20:04:14 +01:00
2023-09-16 15:27:33 +02:00
2023-11-24 17:42:03 +01:00
2023-11-22 12:38:08 +01:00
2023-11-22 16:14:28 +01:00
2023-11-07 12:55:39 +01:00

ComfyUI ExLlama Nodes

A simple text generator for ComfyUI utilizing ExLlamaV2.

Installation

While in the root ComfyUI directory, clone the repository to custom_nodes and install dependencies:

git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes custom_nodes/ComfyUI-ExLlama-Nodes
pip install -r custom_nodes/ComfyUI-ExLlama-Nodes/requirements.txt

Optionally, you can install flash-attention by uncommenting the relevant lines in the requirements file. It should lower VRAM usage but your mileage may vary.

Important

If you see any ExLlama-related errors while loading the nodes, try to install it manually following the official instructions.

Usage

Only EXL2 and 4-bit GPTQ models are supported. You can find a lot of them on Hugging Face. Refer to the model card in each repository for details about quant differences and instruction formats.

To use a model with the nodes, you should clone its repository with git or manually download all the files and place them in models/llm. For example, if you'd like download the 4-bit 32g version of Zephyr 7B Beta, use the following command:

git clone https://huggingface.co/TheBloke/zephyr-7B-beta-GPTQ -b gptq-4bit-32g-actorder_True models/llm/zephyr-7b-gptq-32g

Tip

You can add your own llm path to the extra_model_paths.yaml file and place the models there instead.

Nodes

Name Description
Loader Loads models from the llm directory.
gpu_split - comma-separated VRAM in MB per GPU, if using more than one.
cache_8bit - lower VRAM usage but also lower speed if set to True.
max_seq_len - max context length, higher number equals higher VRAM usage. Setting it to 0 will make the model use default context length from its config file.
Generator Generates text based on the given prompt. Refer to text-generation-webui for parameter explanations.
unload - unloads the model after each generation if set to True, freeing all the VRAM used.
single_line - stops generation on new line.
max_tokens - max new tokens to generate, setting it to 0 will make the model use all available context.
Preview Displays generated text in the UI.
Replace Replaces variable names enclosed in brackets, such as [a], with their values.

Workflow

The example workflow is embedded in the image below and can be opened in ComfyUI.

workflow

S
Description
No description provided
Readme MIT
214 KiB
Languages
Python 89.9%
JavaScript 10.1%