From 5a90464b6ceb445e06e380e82623932046346091 Mon Sep 17 00:00:00 2001 From: Zuellni <123005779+Zuellni@users.noreply.github.com> Date: Fri, 14 Jun 2024 18:48:05 +0200 Subject: [PATCH] Update README.md --- README.md | 40 ++++++++++++++++++++++++++-------------- 1 file changed, 26 insertions(+), 14 deletions(-) diff --git a/README.md b/README.md index efa8079..a4577ec 100644 --- a/README.md +++ b/README.md @@ -2,18 +2,20 @@ A simple local text generator for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) using [ExLlamaV2](https://github.com/turboderp/exllamav2). ## Installation -Clone the repository to `custom_nodes`: +Clone the repository to `custom_nodes` and install the requirements: ``` git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes custom_nodes/ComfyUI-ExLlamaV2-Nodes -``` - -Install requirements, use wheels for [ExLlamaV2](https://github.com/turboderp/exllamav2/releases/latest) and [Flash Attention](https://github.com/bdashore3/flash-attention/releases/latest) on Windows: -``` pip install -r custom_nodes/ComfyUI-ExLlamaV2-Nodes/requirements.txt ``` +Use wheels for [ExLlamaV2](https://github.com/turboderp/exllamav2/releases/latest) and [Flash Attention](https://github.com/bdashore3/flash-attention/releases/latest) on Windows: +``` +pip install exllamav2-X.X.X+cuXXX.torch2.X.X-cp3XX-cp3XX-win_amd64.whl +pip install flash_attn-X.X.X+cuXXX.torch2.X.X-cp3XX-cp3XX-win_amd64.whl +``` + ## Usage -Only EXL2, 4-bit GPTQ and FP16 HF models are supported. You can find them on [Hugging Face](https://huggingface.co). +Only EXL2, 4-bit GPTQ and unquantized models are supported. You can find them on [Hugging Face](https://huggingface.co). To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in `models/llm`. For example, if you want to download the 6-bit [Llama-3-8B-Instruct](https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2), use the following command: @@ -34,26 +36,36 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo cache_bits - Lower value equals lower VRAM usage but also impacts generation speed and quality. + A lower value reduces VRAM usage, but also affects generation speed and quality. + + + + fast_tensors + Enabling reduces RAM usage and speeds up model loading. + + + + flash_attention + Enabling reduces VRAM usage, not supported on cards with compute capability below 8.0. max_seq_len - Max context, higher value equals higher VRAM usage. 0 will default to config. + Max context, higher value equals higher VRAM usage. 0 will default to model config. Generator - Generates text based on the given prompt. Refer to text-generation-webui for parameters. + Generates text based on the given prompt. Refer to SillyTavern for sampler parameters. unload - Unloads the model after each generation. + Unloads the model after each generation to reduce VRAM usage. - single_line - Stops the generation on newline. + stop_conditions + List of strings to stop generation on, e.g. ["\n"] to stop on newline. Leave empty to only stop on eos token. @@ -66,11 +78,11 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo Replacer - Replaces variable names enclosed in brackets, eg [a], with their values. + Replaces variable names in brackets, e.g. [a], with their values. ## Workflow -The example workflow is embedded in the image below and can be opened in ComfyUI. +An example workflow is embedded in the image below and can be opened in ComfyUI. ![workflow](https://github.com/Zuellni/ComfyUI-ExLlama-Nodes/assets/123005779/bf688acb-6f7a-4410-98ff-cf22b6937ae7)