clarified quant option when using GPTQ models of vision model

This commit is contained in:
Nojahhh
2024-10-20 23:42:07 +02:00
parent f97ddf16d1
commit e2eef3bbb8
+1 -1
View File
@@ -64,7 +64,7 @@ The `GLM4ModelLoader` class is responsible for loading GLM-4 models. It supports
`THUDM/glm-4v-9b` requires `bf16` and is set to run in 4-bit by default based on it's size.
`alexwww94/glm-4v-9b-gptq-4bit` requires `bf16` and is set to run in 4-bit by default.
`alexwww94/glm-4v-9b-gptq-3bit` requires `bf16` and is set to run in 3-bit by default.
- **quantization**: Set the number of bits for quantization (`4`, `8`, `16`). Default value of `4`.
- **quantization**: Set the number of bits for quantization (`4`, `8`, `16`). Default value of `4`. (This option is bypassed when using the GPTQ-models).
#### Output