From 98ec700a5c2a75ff9173c8da4169b14c63b57d1d Mon Sep 17 00:00:00 2001
From: Zuellni <123005779+Zuellni@users.noreply.github.com>
Date: Sat, 25 Nov 2023 13:21:19 +0100
Subject: [PATCH] Remove the tensor numel check, should be redundant
---
README.md | 2 +-
exllama.py | 14 ++++++--------
2 files changed, 7 insertions(+), 9 deletions(-)
diff --git a/README.md b/README.md
index 3db5bc6..bd5a46f 100644
--- a/README.md
+++ b/README.md
@@ -24,7 +24,7 @@ git clone https://huggingface.co/TheBloke/zephyr-7B-beta-GPTQ -b gptq-4bit-32g-a
## Nodes
Name | Description
:--- | :---
-Loader | Loads models from the `llm` directory.
`gpu_split` - comma-separated VRAM in GB per GPU, if using more than one.
`cache_8bit` - lower VRAM usage but also lower speed if set to `True`.
`max_seq_len` - max context length, higher number equals higher VRAM usage. Setting it to `0` will make the model use default context length from its config file.
+Loader | Loads models from the `llm` directory.
`gpu_split` - comma-separated VRAM in GB per GPU, eg `6.9, 8`, if using more than one.
`cache_8bit` - lower VRAM usage but also lower speed if set to `True`.
`max_seq_len` - max context length, higher number equals higher VRAM usage. Setting it to `0` will make the model use the default context length from its config file.
Generator | Generates text based on the given prompt. Refer to [text-generation-webui](https://github.com/oobabooga/text-generation-webui/wiki/03-%E2%80%90-Parameters-Tab#parameters-description) for parameter explanations.
`unload` - unloads the model after each generation if set to `True`, freeing all the VRAM used.
`single_line` - stops generation on new line.
`max_tokens` - max new tokens to generate, setting it to `0` will make the model use all available context.
Preview | Displays generated text in the UI.
Replace | Replaces variable names enclosed in brackets, such as `[a]`, with their values.
diff --git a/exllama.py b/exllama.py
index 67b7720..ec74a50 100644
--- a/exllama.py
+++ b/exllama.py
@@ -174,18 +174,16 @@ class Generator:
model.generator.begin_stream(input, settings, token_healing=True)
progress = ProgressBar(max_tokens)
eos = False
- chunks = ""
+ output = ""
tokens = 0
while not eos and tokens < max_tokens:
- chunk, eos, tensor = model.generator.stream()
+ chunk, eos, _ = model.generator.stream()
+ progress.update(1)
+ output += chunk
+ tokens += 1
- if token := tensor.numel():
- progress.update(token)
- chunks += chunk
- tokens += token
-
- output = chunks.strip()
+ output = output.strip()
total = round(time() - start, 2)
speed = round(tokens / total, 2)