updated README.md with information about AutoGPTQ

This commit is contained in:
Nojahhh
2024-12-04 21:27:48 +01:00
parent 2f0fd328e0
commit d73d3415c7
+6
View File
@@ -19,6 +19,12 @@ Hope you will enjoy your enhanced prompts and inference capabilities of these mo
Added support for quantized models. They are performing exceptionally well. Check metrics below.
**Important**
To be able to use these models you will need to install AutoGPTQ library.
`pip install auto-gptq`
If you are on Windows you will need to install this from source to enable CUDA extensions.
[AutoGPTQ](https://github.com/AutoGPTQ/AutoGPTQ/)
Model `alexwww94/glm-4v-9b-gptq-4bit` is significatly more lightweight than the original and will hold ~8.5 GB of hdd space.
Model `alexwww94/glm-4v-9b-gptq-3bit` is even more lightweight and will hold ~7.6 GB of hdd space.