diff --git a/README.md b/README.md index 3db5bc6..bd5a46f 100644 --- a/README.md +++ b/README.md @@ -24,7 +24,7 @@ git clone https://huggingface.co/TheBloke/zephyr-7B-beta-GPTQ -b gptq-4bit-32g-a ## Nodes Name | Description :--- | :--- -Loader | Loads models from the `llm` directory.
`gpu_split` - comma-separated VRAM in GB per GPU, if using more than one.
`cache_8bit` - lower VRAM usage but also lower speed if set to `True`.
`max_seq_len` - max context length, higher number equals higher VRAM usage. Setting it to `0` will make the model use default context length from its config file. +Loader | Loads models from the `llm` directory.
`gpu_split` - comma-separated VRAM in GB per GPU, eg `6.9, 8`, if using more than one.
`cache_8bit` - lower VRAM usage but also lower speed if set to `True`.
`max_seq_len` - max context length, higher number equals higher VRAM usage. Setting it to `0` will make the model use the default context length from its config file. Generator | Generates text based on the given prompt. Refer to [text-generation-webui](https://github.com/oobabooga/text-generation-webui/wiki/03-%E2%80%90-Parameters-Tab#parameters-description) for parameter explanations.
`unload` - unloads the model after each generation if set to `True`, freeing all the VRAM used.
`single_line` - stops generation on new line.
`max_tokens` - max new tokens to generate, setting it to `0` will make the model use all available context. Preview | Displays generated text in the UI. Replace | Replaces variable names enclosed in brackets, such as `[a]`, with their values. diff --git a/exllama.py b/exllama.py index 67b7720..ec74a50 100644 --- a/exllama.py +++ b/exllama.py @@ -174,18 +174,16 @@ class Generator: model.generator.begin_stream(input, settings, token_healing=True) progress = ProgressBar(max_tokens) eos = False - chunks = "" + output = "" tokens = 0 while not eos and tokens < max_tokens: - chunk, eos, tensor = model.generator.stream() + chunk, eos, _ = model.generator.stream() + progress.update(1) + output += chunk + tokens += 1 - if token := tensor.numel(): - progress.update(token) - chunks += chunk - tokens += token - - output = chunks.strip() + output = output.strip() total = round(time() - start, 2) speed = round(tokens / total, 2)