diff --git a/README.md b/README.md index e7e5f35..85c36f9 100644 --- a/README.md +++ b/README.md @@ -1,17 +1,17 @@ -# ExLlama nodes for ComfyUI +# ComfyUI ExLlama Nodes A simple prompt generator for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) utilizing [ExLlama](https://github.com/turboderp/exllama). ## Installation -Clone the repository to `custom_nodes` in your ComfyUI directory and install dependencies: +Clone the repository to `custom_nodes` in your ComfyUI directory and install the dependencies: ``` git clone https://github.com/Zuellni/ComfyUI-ExLlama-Nodes python -m pip install -r requirements.txt ``` -Next, install the latest pre-built ExLlama wheel from [here](https://github.com/jllllll/exllama/releases/latest). +Next, install the latest pre-built ExLlama wheel from https://github.com/jllllll/exllama/releases/latest. Choose the version matching your platform, Python, and PyTorch CUDA/ROCm. -Example for Windows with Python 3.10 and CUDA 11.8 which should match the portable ComfyUI build: +Example for Windows with Python 3.10 and CUDA 11.8, which should match the portable ComfyUI build: ``` python -m pip install https://github.com/jllllll/exllama/releases/download/0.0.17/exllama-0.0.17+cu118-cp310-cp310-win_amd64.whl ``` @@ -19,11 +19,11 @@ python -m pip install https://github.com/jllllll/exllama/releases/download/0.0.1 ## Nodes Name | Description :--- | :--- -Loader | Loads 4-bit GPTQ Llama/2 models. You can find a lot of them on [Hugging Face](https://huggingface.co/TheBloke).
Clone the model repository or download all the files in it to an empty directory, then point to it in `model_dir`. The `model.safetensors` file won't work on its own.

ExLlama allocates memory based on `max_seq_len`. Lowering it is a good way to save on VRAM. It's currently not possible to [offload](https://github.com/turboderp/exllama/issues/177) the model to RAM. -Generator | Returns a `string` based on the given `prompt` for use with other nodes. Default values correspond to the `simple-1` preset from [text-generation-webui](https://github.com/oobabooga/text-generation-webui). ExLlama isn't [deterministic](https://github.com/turboderp/exllama/issues/201), so the outputs may differ even with the same seed.

To load a LoRA specify its directory in `lora_dir`. It should contain `adapter_model.bin` and `adapter_config.json`. +Loader | Loads 4-bit GPTQ Llama/2 models. You can find a lot of them on [Hugging Face](https://huggingface.co/TheBloke).
Clone the model repository or download all the files in it and place them in an empty directory, then specify the path in `model_dir`. The `model.safetensors` file won't work on its own.

ExLlama allocates memory based on `max_seq_len`. Lowering it is a good way to save on VRAM. It's currently not possible to [offload](https://github.com/turboderp/exllama/issues/177) the model to RAM. +Generator | Returns a `string` based on the given `prompt` for use with other nodes. Default values correspond to the `simple-1` preset from [text-generation-webui](https://github.com/oobabooga/text-generation-webui). ExLlama isn't [deterministic](https://github.com/turboderp/exllama/issues/201), so the outputs may differ slightly even with the same seed.

To load a LoRA specify the path to its directory in `lora_dir`. It should contain `adapter_model.bin` and `adapter_config.json`. Previewer | Displays generated outputs in the UI. ## Workflow -Can be loaded directly in ComfyUI. Model used: [MythoLogic-Mini-7B](https://huggingface.co/TheBloke/MythoLogic-Mini-7B-GPTQ). +The workflow below can be loaded directly in ComfyUI. Model used: [MythoLogic-Mini-7B](https://huggingface.co/TheBloke/MythoLogic-Mini-7B-GPTQ). ![workflow](https://github.com/Zuellni/ComfyUI-ExLlama-Nodes/assets/123005779/c6821ba6-3a7a-4dd2-9852-372f79f63569)