ComfyUI ExLlama Nodes
A simple text generator for ComfyUI utilizing ExLlamaV2.
Installation
Clone the repository to custom_nodes and install dependencies:
git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes
pip install -r requirements.txt
Optionally, you can install flash-attention by uncommenting the relevant lines in the requirements file.
If you see any ExLlama-related errors while loading the nodes, install it manually following the instructions here.
Usage
ExLlamaV2 supports EXL2 and 4-bit GPTQ models. You can find a lot of them on Hugging Face.
Refer to the model card in each repository for details about quant differences and instruction formats.
To use a model with the nodes, you should clone its repository with git or manually download all the files and place them in models/llm.
You can also add your own llm path to extra_model_paths.yaml and place the models there instead.
For instance, if you want to download the 4-bit 32g branch of Zephyr 7B Beta, use the following command:
git clone https://huggingface.co/TheBloke/zephyr-7B-beta-GPTQ -b gptq-4bit-32g-actorder_True models/llm/zephyr-7b-gptq-32g
Nodes
| Name | Description |
|---|---|
| Loader | Used to load Llama models.gpu_split - comma-separated VRAM in MB per GPU, if using more than one.cache_8bit - lower VRAM usage and lower speed if set to True.max_seq_len - max context length, higher number equals higher VRAM usage. Setting it to 0 will make the model use default context length based on its config file. |
| Generator | Generates text based on the given prompt. Refer to text-generation-webui for parameters.unload - unloads the model after each generation if set to True.single_line - stops generation on new line.max_tokens - max new tokens to generate, setting it to 0 will make it use all available context length. |
| Preview | Displays generated text in the UI. |
| Replace | Replaces variables enclosed in brackets, such as [a], with their values. |
Workflow
The image below can be opened in ComfyUI.