updated README.md with information about AutoGPTQ
This commit is contained in:
@@ -19,6 +19,12 @@ Hope you will enjoy your enhanced prompts and inference capabilities of these mo
|
||||
|
||||
Added support for quantized models. They are performing exceptionally well. Check metrics below.
|
||||
|
||||
**Important**
|
||||
To be able to use these models you will need to install AutoGPTQ library.
|
||||
`pip install auto-gptq`
|
||||
If you are on Windows you will need to install this from source to enable CUDA extensions.
|
||||
[AutoGPTQ](https://github.com/AutoGPTQ/AutoGPTQ/)
|
||||
|
||||
Model `alexwww94/glm-4v-9b-gptq-4bit` is significatly more lightweight than the original and will hold ~8.5 GB of hdd space.
|
||||
|
||||
Model `alexwww94/glm-4v-9b-gptq-3bit` is even more lightweight and will hold ~7.6 GB of hdd space.
|
||||
|
||||
Reference in New Issue
Block a user