Update README.md

This commit is contained in:
Zuellni
2024-07-28 21:31:21 +02:00
committed by GitHub
parent 7ae7ef9a24
commit 2acc730ad3
+46 -11
View File
@@ -1,4 +1,4 @@
# ComfyUI ExLlamaV2 Nodes # ComfyUI ExLlama Nodes
A simple local text generator for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) using [ExLlamaV2](https://github.com/turboderp/exllamav2). A simple local text generator for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) using [ExLlamaV2](https://github.com/turboderp/exllamav2).
## Installation ## Installation
@@ -15,9 +15,9 @@ pip install flash_attn-X.X.X+cuXXX.torch2.X.X-cp3XX-cp3XX-win_amd64.whl
``` ```
## Usage ## Usage
Only EXL2, 4-bit GPTQ and unquantized models are supported. You can find them on [Hugging Face](https://huggingface.co). Only EXL2, 4-bit GPTQ and FP16 models are supported. You can find them on [Hugging Face](https://huggingface.co).
To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in `models/llm`. To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in a folder in `models/llm`.
For example, if you want to download the 6-bit [Llama-3-8B-Instruct](https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2), use the following command: For example, if you want to download the 6-bit [Llama-3-8B-Instruct](https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2), use the following command:
``` ```
git install lfs git install lfs
@@ -29,6 +29,9 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo
## Nodes ## Nodes
<table> <table>
<tr>
<td colspan="3" align="center"><b>ExLlama Nodes</b></td>
</tr>
<tr> <tr>
<td><b>Loader</b></td> <td><b>Loader</b></td>
<td colspan="2">Loads models from the <code>llm</code> directory.</td> <td colspan="2">Loads models from the <code>llm</code> directory.</td>
@@ -46,16 +49,43 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo
<tr> <tr>
<td></td> <td></td>
<td ><i>flash_attention</i></td> <td ><i>flash_attention</i></td>
<td>Enabling reduces VRAM usage, not supported on cards with compute capability below <code>8.0</code>.</td> <td>Enabling reduces VRAM usage, not supported on cards with compute capability lower than <code>8.0</code>.</td>
</tr> </tr>
<tr> <tr>
<td></td> <td></td>
<td><i>max_seq_len</i></td> <td><i>max_seq_len</i></td>
<td>Max context, higher value equals higher VRAM usage. <code>0</code> will default to model config.</td> <td>Max context, higher value equals higher VRAM usage. <code>0</code> will default to model config.</td>
</tr> </tr>
<tr>
<td><b>Formatter</b></td>
<td colspan="2">Formats messages using the model's chat template.</td>
</tr>
<tr>
<td></td>
<td><i>add_assistant_role</i></td>
<td>Appends an assistant role to the formatted output.</td>
</tr>
<tr>
<td><b>Tokenizer</b></td>
<td colspan="2">Tokenizes input text using the model's tokenizer.</td>
</tr>
<tr>
<td></td>
<td><i>add_bos_token</i></td>
<td>Prepends the input with a <code>bos</code> token if enabled.</td>
</tr>
<tr>
<td></td>
<td><i>encode_special_tokens</i></td>
<td>Encodes special tokens such as <code>bos</code> and <code>eos</code> if enabled, otherwise treats them as normal strings.</td>
</tr>
<tr>
<td><b>Settings</b></td>
<td colspan="2">Optional sampler settings node. Refer to <a href="https://docs.sillytavern.app/usage/common-settings/#sampler-parameters">SillyTavern</a> for parameters.</td>
</tr>
<tr> <tr>
<td><b>Generator</b></td> <td><b>Generator</b></td>
<td colspan="2">Generates text based on the given prompt. Refer to <a href="https://docs.sillytavern.app/usage/common-settings/#sampler-parameters">SillyTavern</a> for sampler parameters.</td> <td colspan="2">Generates text based on the given input.</td>
</tr> </tr>
<tr> <tr>
<td></td> <td></td>
@@ -65,24 +95,29 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo
<tr> <tr>
<td></td> <td></td>
<td><i>stop_conditions</i></td> <td><i>stop_conditions</i></td>
<td>List of strings to stop generation on, e.g. <code>["\n"]</code> to stop on newline. Leave empty to only stop on <code>eos</code> token.</td> <td>A list of strings to stop generation on, e.g. <code>"\n"</code> to stop on newline. Leave empty to only stop on <code>eos</code>.</td>
</tr> </tr>
<tr> <tr>
<td></td> <td></td>
<td><i>max_tokens</i></td> <td><i>max_tokens</i></td>
<td>Max new tokens, <code>0</code> will use available context.</td> <td>Max new tokens to generate. <code>0</code> will use available context.</td>
</tr> </tr>
<tr> <tr>
<td><b>Previewer</b></td> <td colspan="3" align="center"><b>Text Nodes</b></td>
</tr>
<tr>
<td><b>Message</b></td>
<td colspan="2">A message for the <code>Formatter</code> node. Can be chained to create a conversation.</td>
</tr>
<tr>
<td><b>Preview</b></td>
<td colspan="2">Displays generated text in the UI.</td> <td colspan="2">Displays generated text in the UI.</td>
</tr> </tr>
<tr> <tr>
<td><b>Replacer</b></td> <td><b>Replace</b></td>
<td colspan="2">Replaces variable names in brackets, e.g. <code>[a]</code>, with their values.</td> <td colspan="2">Replaces variable names in brackets, e.g. <code>[a]</code>, with their values.</td>
</tr> </tr>
</table> </table>
## Workflow ## Workflow
An example workflow is embedded in the image below and can be opened in ComfyUI. An example workflow is embedded in the image below and can be opened in ComfyUI.
![workflow](https://github.com/Zuellni/ComfyUI-ExLlama-Nodes/assets/123005779/bf688acb-6f7a-4410-98ff-cf22b6937ae7)