ComfyUI ExLlamaV2 Nodes

A simple local text generator for ComfyUI using ExLlamaV2.

Installation

Clone the repository to custom_nodes:

git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes custom_nodes/ComfyUI-ExLlamaV2-Nodes

Install requirements, use wheels for ExLlamaV2 and Flash Attention on Windows:

pip install -r custom_nodes/ComfyUI-ExLlamaV2-Nodes/requirements.txt

Usage

Only EXL2, 4-bit GPTQ and FP16 HF models are supported. You can find them on Hugging Face.

To use a model with the nodes, you should clone its repository with git or manually download all the files and place them in models/llm. For example, if you want to download the 6-bit Llama-3-8B-Instruct, use the following command:

git install lfs
git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw models/llm/Llama-3-8B-Instruct-exl2-6.0bpw

Tip

You can add your own llm path to the extra_model_paths.yaml file and put the models there instead.

Nodes

Loader Loads models from the llm directory.
cache_bits Lower value equals lower VRAM usage but also impacts generation speed and quality.
max_seq_len Max context, higher value equals higher VRAM usage. 0 will default to config.
Generator Generates text based on the given prompt. Refer to text-generation-webui for parameters.
unload Unloads the model after each generation.
single_line Stops the generation on newline.
max_tokens Max new tokens, 0 will use available context.
Previewer Displays generated text in the UI.
Replacer Replaces variable names enclosed in brackets, eg [a], with their values.

Workflow

The example workflow is embedded in the image below and can be opened in ComfyUI.

workflow

S
Description
No description provided
Readme MIT
214 KiB
Languages
Python 89.9%
JavaScript 10.1%