diff --git a/README.md b/README.md index eadcaa9..522fec7 100644 --- a/README.md +++ b/README.md @@ -1,4 +1,4 @@ -# ComfyUI ExLlamaV2 Nodes +# ComfyUI ExLlama Nodes A simple local text generator for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) using [ExLlamaV2](https://github.com/turboderp/exllamav2). ## Installation @@ -15,9 +15,9 @@ pip install flash_attn-X.X.X+cuXXX.torch2.X.X-cp3XX-cp3XX-win_amd64.whl ``` ## Usage -Only EXL2, 4-bit GPTQ and unquantized models are supported. You can find them on [Hugging Face](https://huggingface.co). +Only EXL2, 4-bit GPTQ and FP16 models are supported. You can find them on [Hugging Face](https://huggingface.co). -To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in `models/llm`. +To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in a folder in `models/llm`. For example, if you want to download the 6-bit [Llama-3-8B-Instruct](https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2), use the following command: ``` git install lfs @@ -29,6 +29,9 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo ## Nodes
| ExLlama Nodes | +||||
| Loader | Loads models from the llm directory. |
@@ -45,17 +48,44 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo
|||
| - | flash_attention | -Enabling reduces VRAM usage, not supported on cards with compute capability below 8.0. |
+ flash_attention | +Enabling reduces VRAM usage, not supported on cards with compute capability lower than 8.0. |
| max_seq_len | Max context, higher value equals higher VRAM usage. 0 will default to model config. |
|||
| Formatter | +Formats messages using the model's chat template. | +|||
| + | add_assistant_role | +Appends an assistant role to the formatted output. | +||
| Tokenizer | +Tokenizes input text using the model's tokenizer. | +|||
| + | add_bos_token | +Prepends the input with a bos token if enabled. |
+ ||
| + | encode_special_tokens | +Encodes special tokens such as bos and eos if enabled, otherwise treats them as normal strings. |
+ ||
| Settings | +Optional sampler settings node. Refer to SillyTavern for parameters. | +|||
| Generator | -Generates text based on the given prompt. Refer to SillyTavern for sampler parameters. | +Generates text based on the given input. | ||
| @@ -65,24 +95,29 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo | ||||
| stop_conditions | -List of strings to stop generation on, e.g. ["\n"] to stop on newline. Leave empty to only stop on eos token. |
+ A list of strings to stop generation on, e.g. "\n" to stop on newline. Leave empty to only stop on eos. |
||
| max_tokens | -Max new tokens, 0 will use available context. |
+ Max new tokens to generate. 0 will use available context. |
||
| Previewer | +Text Nodes | +|||
| Message | +A message for the Formatter node. Can be chained to create a conversation. |
+ |||
| Preview | Displays generated text in the UI. | |||
| Replacer | +Replace | Replaces variable names in brackets, e.g. [a], with their values. |
||