diff --git a/README.md b/README.md index eadcaa9..522fec7 100644 --- a/README.md +++ b/README.md @@ -1,4 +1,4 @@ -# ComfyUI ExLlamaV2 Nodes +# ComfyUI ExLlama Nodes A simple local text generator for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) using [ExLlamaV2](https://github.com/turboderp/exllamav2). ## Installation @@ -15,9 +15,9 @@ pip install flash_attn-X.X.X+cuXXX.torch2.X.X-cp3XX-cp3XX-win_amd64.whl ``` ## Usage -Only EXL2, 4-bit GPTQ and unquantized models are supported. You can find them on [Hugging Face](https://huggingface.co). +Only EXL2, 4-bit GPTQ and FP16 models are supported. You can find them on [Hugging Face](https://huggingface.co). -To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in `models/llm`. +To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in a folder in `models/llm`. For example, if you want to download the 6-bit [Llama-3-8B-Instruct](https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2), use the following command: ``` git install lfs @@ -29,6 +29,9 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo ## Nodes + + + @@ -45,17 +48,44 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo - - + + + + + + + + + + + + + + + + + + + + + + + + + + + + + - + @@ -65,24 +95,29 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo - + - + - + + + + + + + + - +
ExLlama Nodes
Loader Loads models from the llm directory.
flash_attentionEnabling reduces VRAM usage, not supported on cards with compute capability below 8.0.flash_attentionEnabling reduces VRAM usage, not supported on cards with compute capability lower than 8.0.
max_seq_len Max context, higher value equals higher VRAM usage. 0 will default to model config.
FormatterFormats messages using the model's chat template.
add_assistant_roleAppends an assistant role to the formatted output.
TokenizerTokenizes input text using the model's tokenizer.
add_bos_tokenPrepends the input with a bos token if enabled.
encode_special_tokensEncodes special tokens such as bos and eos if enabled, otherwise treats them as normal strings.
SettingsOptional sampler settings node. Refer to SillyTavern for parameters.
GeneratorGenerates text based on the given prompt. Refer to SillyTavern for sampler parameters.Generates text based on the given input.
stop_conditionsList of strings to stop generation on, e.g. ["\n"] to stop on newline. Leave empty to only stop on eos token.A list of strings to stop generation on, e.g. "\n" to stop on newline. Leave empty to only stop on eos.
max_tokensMax new tokens, 0 will use available context.Max new tokens to generate. 0 will use available context.
PreviewerText Nodes
MessageA message for the Formatter node. Can be chained to create a conversation.
Preview Displays generated text in the UI.
ReplacerReplace Replaces variable names in brackets, e.g. [a], with their values.
## Workflow An example workflow is embedded in the image below and can be opened in ComfyUI. - -![workflow](https://github.com/Zuellni/ComfyUI-ExLlama-Nodes/assets/123005779/bf688acb-6f7a-4410-98ff-cf22b6937ae7)