Update README.md
This commit is contained in:
@@ -1,4 +1,4 @@
|
||||
# ComfyUI ExLlamaV2 Nodes
|
||||
# ComfyUI ExLlama Nodes
|
||||
A simple local text generator for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) using [ExLlamaV2](https://github.com/turboderp/exllamav2).
|
||||
|
||||
## Installation
|
||||
@@ -15,9 +15,9 @@ pip install flash_attn-X.X.X+cuXXX.torch2.X.X-cp3XX-cp3XX-win_amd64.whl
|
||||
```
|
||||
|
||||
## Usage
|
||||
Only EXL2, 4-bit GPTQ and unquantized models are supported. You can find them on [Hugging Face](https://huggingface.co).
|
||||
Only EXL2, 4-bit GPTQ and FP16 models are supported. You can find them on [Hugging Face](https://huggingface.co).
|
||||
|
||||
To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in `models/llm`.
|
||||
To use a model with the nodes, you should clone its repository with `git` or manually download all the files and place them in a folder in `models/llm`.
|
||||
For example, if you want to download the 6-bit [Llama-3-8B-Instruct](https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2), use the following command:
|
||||
```
|
||||
git install lfs
|
||||
@@ -29,6 +29,9 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo
|
||||
|
||||
## Nodes
|
||||
<table>
|
||||
<tr>
|
||||
<td colspan="3" align="center"><b>ExLlama Nodes</b></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Loader</b></td>
|
||||
<td colspan="2">Loads models from the <code>llm</code> directory.</td>
|
||||
@@ -45,17 +48,44 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo
|
||||
</tr>
|
||||
<tr>
|
||||
<td></td>
|
||||
<td><i>flash_attention</i></td>
|
||||
<td>Enabling reduces VRAM usage, not supported on cards with compute capability below <code>8.0</code>.</td>
|
||||
<td ><i>flash_attention</i></td>
|
||||
<td>Enabling reduces VRAM usage, not supported on cards with compute capability lower than <code>8.0</code>.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td></td>
|
||||
<td><i>max_seq_len</i></td>
|
||||
<td>Max context, higher value equals higher VRAM usage. <code>0</code> will default to model config.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Formatter</b></td>
|
||||
<td colspan="2">Formats messages using the model's chat template.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td></td>
|
||||
<td><i>add_assistant_role</i></td>
|
||||
<td>Appends an assistant role to the formatted output.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Tokenizer</b></td>
|
||||
<td colspan="2">Tokenizes input text using the model's tokenizer.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td></td>
|
||||
<td><i>add_bos_token</i></td>
|
||||
<td>Prepends the input with a <code>bos</code> token if enabled.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td></td>
|
||||
<td><i>encode_special_tokens</i></td>
|
||||
<td>Encodes special tokens such as <code>bos</code> and <code>eos</code> if enabled, otherwise treats them as normal strings.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Settings</b></td>
|
||||
<td colspan="2">Optional sampler settings node. Refer to <a href="https://docs.sillytavern.app/usage/common-settings/#sampler-parameters">SillyTavern</a> for parameters.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Generator</b></td>
|
||||
<td colspan="2">Generates text based on the given prompt. Refer to <a href="https://docs.sillytavern.app/usage/common-settings/#sampler-parameters">SillyTavern</a> for sampler parameters.</td>
|
||||
<td colspan="2">Generates text based on the given input.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td></td>
|
||||
@@ -65,24 +95,29 @@ git clone https://huggingface.co/turboderp/Llama-3-8B-Instruct-exl2 -b 6.0bpw mo
|
||||
<tr>
|
||||
<td></td>
|
||||
<td><i>stop_conditions</i></td>
|
||||
<td>List of strings to stop generation on, e.g. <code>["\n"]</code> to stop on newline. Leave empty to only stop on <code>eos</code> token.</td>
|
||||
<td>A list of strings to stop generation on, e.g. <code>"\n"</code> to stop on newline. Leave empty to only stop on <code>eos</code>.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td></td>
|
||||
<td><i>max_tokens</i></td>
|
||||
<td>Max new tokens, <code>0</code> will use available context.</td>
|
||||
<td>Max new tokens to generate. <code>0</code> will use available context.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Previewer</b></td>
|
||||
<td colspan="3" align="center"><b>Text Nodes</b></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Message</b></td>
|
||||
<td colspan="2">A message for the <code>Formatter</code> node. Can be chained to create a conversation.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Preview</b></td>
|
||||
<td colspan="2">Displays generated text in the UI.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>Replacer</b></td>
|
||||
<td><b>Replace</b></td>
|
||||
<td colspan="2">Replaces variable names in brackets, e.g. <code>[a]</code>, with their values.</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
## Workflow
|
||||
An example workflow is embedded in the image below and can be opened in ComfyUI.
|
||||
|
||||

|
||||
|
||||
Reference in New Issue
Block a user