Files
Zuellni-ComfyUI-ExLlama-Nodes/README.md
T
2023-11-25 14:55:51 +01:00

78 lines
3.4 KiB
Markdown

# ComfyUI ExLlama Nodes
A simple text generator for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) utilizing [ExLlamaV2](https://github.com/turboderp/exllamav2).
## Installation
Navigate to the root ComfyUI directory, clone the repository to `custom_nodes` and install dependencies:
```
git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes custom_nodes/ComfyUI-ExLlama-Nodes
pip install -r custom_nodes/ComfyUI-ExLlama-Nodes/requirements.txt
```
Optionally, you can install [flash-attention](https://github.com/Dao-AILab/flash-attention) by uncommenting the relevant lines in the requirements file. It should lower VRAM usage but your mileage may vary.
> [!IMPORTANT]
> The wheels included in the requirements file should match the latest portable ComfyUI build. If you see any ExLlama-related errors while loading the nodes, try to install it manually following the [official instructions](https://github.com/turboderp/exllamav2#installation).
## Usage
Only EXL2 and 4-bit GPTQ models are supported. You can find a lot of them on [Hugging Face](https://huggingface.co/TheBloke). Refer to the model card in each repository for details about quant differences and instruction formats.
To use a model with the nodes, you should clone its repository with git or manually download all the files and place them in `models/llm`. For example, if you'd like to download the 4-bit 32g version of [Zephyr 7B Beta](https://huggingface.co/TheBloke/zephyr-7B-beta-GPTQ), use the following command:
```
git clone https://huggingface.co/TheBloke/zephyr-7B-beta-GPTQ -b gptq-4bit-32g-actorder_True models/llm/zephyr-7b-gptq-32g
```
> [!TIP]
> You can add your own `llm` path to the [extra_model_paths.yaml](https://github.com/comfyanonymous/ComfyUI/blob/master/extra_model_paths.yaml.example) file and place the models there instead.
## Nodes
<table>
<tr>
<td><b>Loader</b></td>
<td colspan="2">Loads models from the <code>llm</code> directory.</td>
</tr>
<tr>
<td></td>
<td><i>gpu_split</i></td>
<td>Comma-separated VRAM in GB per GPU, eg <code>6.9, 8</code>.</td>
</tr>
<tr>
<td></td>
<td><i>cache_8bit</i></td>
<td>Lower VRAM usage but also lower speed.</td>
</tr>
<tr>
<td></td>
<td><i>max_seq_len</i></td>
<td>Max context, higher number equals higher VRAM usage. <code>0</code> will default to config.</td>
</tr>
<tr>
<td><b>Generator</b></td>
<td colspan="2">Generates text based on the given prompt. Refer to <a href="https://github.com/oobabooga/text-generation-webui/wiki/03-%E2%80%90-Parameters-Tab#parameters-description">text-generation-webui</a> for parameters.</td>
</tr>
<tr>
<td></td>
<td><i>unload</i></td>
<td>Unloads the model after each generation.</td>
</tr>
<tr>
<td></td>
<td><i>single_line</i></td>
<td>Stops the generation on newline.</td>
</tr>
<tr>
<td></td>
<td><i>max_tokens</i></td>
<td>Max new tokens, <code>0</code> will use available context.</td>
</tr>
<tr>
<td><b>Preview</b></td>
<td colspan="2">Displays generated text in the UI.</td>
</tr>
<tr>
<td><b>Replace</b></td>
<td colspan="2">Replaces variable names enclosed in brackets, eg <code>[a]</code>, with their values.</td>
</tr>
</table>
## Workflow
The example workflow is embedded in the image below and can be opened in ComfyUI.
![workflow](https://github.com/Zuellni/ComfyUI-ExLlama-Nodes/assets/123005779/cb20b040-9856-4dab-aed0-4b318bc2d805)