diff --git a/README.md b/README.md index 657eeae..3b2b60c 100644 --- a/README.md +++ b/README.md @@ -19,6 +19,12 @@ Hope you will enjoy your enhanced prompts and inference capabilities of these mo Added support for quantized models. They are performing exceptionally well. Check metrics below. +**Important** +To be able to use these models you will need to install AutoGPTQ library. +`pip install auto-gptq` +If you are on Windows you will need to install this from source to enable CUDA extensions. +[AutoGPTQ](https://github.com/AutoGPTQ/AutoGPTQ/) + Model `alexwww94/glm-4v-9b-gptq-4bit` is significatly more lightweight than the original and will hold ~8.5 GB of hdd space. Model `alexwww94/glm-4v-9b-gptq-3bit` is even more lightweight and will hold ~7.6 GB of hdd space.