From e0201ed007d98ce14c98c56f7fd3553fcb6390e6 Mon Sep 17 00:00:00 2001 From: Zuellni <123005779+Zuellni@users.noreply.github.com> Date: Tue, 19 Sep 2023 10:05:02 +0200 Subject: [PATCH] Add a note on model downloading --- README.md | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 4a0228f..01e98ee 100644 --- a/README.md +++ b/README.md @@ -20,9 +20,12 @@ pip install https://github.com/jllllll/exllama/releases/download/0.0.17/exllama- Comes with the following nodes: ### Loader -Used to load 4-bit GPTQ Llama/2 models. You can find a lot of them over at [Hugging Face](https://huggingface.co/TheBloke). +Used to load 4-bit GPTQ Llama/2 models. You can find a lot of them over at [Hugging Face](https://huggingface.co/TheBloke). + +You should either clone the model repository or download all the files in it manually, then point to the directory in `model_dir`. The `model.safetensors` file on its own is not enough to work. + ExLlama allocates [memory](https://github.com/turboderp/exllama/issues/259) according to `max_seq_len`. Lowering it is a good way to save on GPU RAM. -It's currently not possible to [offload](https://github.com/turboderp/exllama/issues/177) the model to CPU RAM. +It's currently not possible to [offload](https://github.com/turboderp/exllama/issues/177) the models to CPU RAM. ### Generator Generates a `string` based on the given `prompt` for use with other nodes.