f54093078ee61563922a34b2a718aa45ff54a4bd
ComfyUI ExLlamaV2 Nodes
A simple local text generator for ComfyUI using ExLlamaV2.
Installation
Clone the repository to custom_nodes:
git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes custom_nodes/ComfyUI-ExLlamaV2-Nodes
Install requirements, use wheels for ExLlamaV2 and Flash Attention on Windows:
pip install -r custom_nodes/ComfyUI-ExLlamaV2-Nodes/requirements.txt
Usage
Only EXL2, 4-bit GPTQ, and unquantized HF models are supported. You can find them on Hugging Face.
To use a model with the nodes, you should clone its repository with git or manually download all the files and place them in models/llm.
For example, if you want to download the 6-bit Llama-3-8B-Instruct, use the following command:
git install lfs
git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw models/llm/Llama-3-8B-Instruct-exl2-6.0bpw
Tip
You can add your own
llmpath to the extra_model_paths.yaml file and put the models there instead.
Nodes
| Loader | Loads models from the llm directory. |
|
| cache_bits | Lower value equals lower VRAM usage but also impacts generation speed and quality. | |
| max_seq_len | Max context, higher value equals higher VRAM usage. 0 will default to config. |
|
| Generator | Generates text based on the given prompt. Refer to text-generation-webui for parameters. | |
| unload | Unloads the model after each generation. | |
| single_line | Stops the generation on newline. | |
| max_tokens | Max new tokens, 0 will use available context. |
|
| Previewer | Displays generated text in the UI. | |
| Replacer | Replaces variable names enclosed in brackets, eg [a], with their values. |
|
Workflow
The example workflow is embedded in the image below and can be opened in ComfyUI.
Languages
Python
89.9%
JavaScript
10.1%