Commit Graph
155 Commits
Author SHA1 Message Date
Zuellni 49ace07a0e Update README.md 2024-06-14 12:51:36 +02:00
Zuellni e5e13c40ed Update README.md 2024-06-05 09:53:12 +02:00
Zuellni d7413b2ebc Fix vram usage with higher context 2024-05-17 22:55:50 +02:00
Zuellni c005e3c637 Update config init 2024-05-04 09:13:01 +02:00
Zuellni 4422f0b0d0 Update readme 2024-04-22 18:30:33 +02:00
Zuellni 4633de563a Update workflow with llama-3 format 2024-04-22 16:18:18 +02:00
Zuellni 87810044ea Change default cache bits to 16, reorganize imports, update readme 2024-04-22 15:50:50 +02:00
Zuellni 8c2f17829a Clarify instructions 2024-04-08 03:03:53 +02:00
Zuellni c5765ffb8c Update with a cascade workflow 2024-04-07 20:19:55 +02:00
Zuellni fb7b7ff683 Update readme 2024-04-07 20:00:44 +02:00
Zuellni fb5b6d95ef Update link 2024-04-07 19:55:34 +02:00
Zuellni d2543cbd59 Simplify instructions 2024-04-07 19:52:53 +02:00
Zuellni baed69aba5 Bump minimum exllamav2 version, add autosplit and q4 cache, some reformatting 2024-04-07 19:35:32 +02:00
Zuellni d2b4d7a796 More sanity checks for model unloading, bump exllama 2024-02-26 13:24:03 +01:00
Zuellni f7a108315e Update readme 2024-02-21 23:06:35 +01:00
Zuellni 217304e0ef Only unload other models if unload is checked 2024-02-17 16:53:00 +01:00
Zuellni 8c4b9d058f Update workflow 2024-02-17 15:14:46 +01:00
Zuellni 861aa6fc51 Update readme 2024-02-17 15:13:53 +01:00
Zuellni 8b7b957b3e Add gc again 2024-02-17 14:53:54 +01:00
Zuellni e1c8e61291 Unload other models first before loading llms, rename some params, update requirements 2024-02-17 14:40:46 +01:00
Zuellni ca41704f38 Update readme 2024-02-14 20:26:46 +01:00
Zuellni bf64bcc0e4 Split requirements to file, neutralize samplers by default 2024-02-14 20:23:27 +01:00
Zuellni 9ffbacf080 Update requirements 2024-02-12 16:38:52 +01:00
Zuellni 6cc728d7cb Update readme 2024-02-09 16:13:30 +01:00
Zuellni ec6443816e Update readme 2024-02-06 14:51:02 +01:00
Zuellni 6831b951c5 Update readme 2024-02-03 15:53:52 +01:00
Zuellni 54bedbed53 Update workflow 2024-02-03 15:52:58 +01:00
Zuellni 88a826fa30 Add top_a sampling 2024-02-03 15:14:38 +01:00
Zuellni 18c7854d76 Add missing import, look for paths recursively 2024-02-03 15:13:30 +01:00
Zuellni 261a73a119 Merge pull request #14 from ScottNealon/patch-1
Parse all possible folders mapped to "llm"
2024-02-03 15:11:39 +01:00
Scott Nealon 348aeea258 Parse all possible folders mapped to "llm" 2024-02-02 22:12:50 -05:00
Zuellni 73e6a0b0e4 Bump exllama 2024-01-26 15:39:58 +01:00
Zuellni 331b594f00 Bump flash-attn 2023-12-27 13:13:16 +01:00
Zuellni de401d308d Bump exllama 2023-12-21 19:51:05 +01:00
Zuellni 04f0c87e31 Fix error when model isn't fully unloaded and bump requirements 2023-12-02 23:27:41 +01:00
Zuellni eca815e45a Fix linux requirements 2023-11-25 20:16:56 +01:00
Zuellni 15a224ccbb Even better table 2023-11-25 14:55:51 +01:00
Zuellni 485e2c4c51 Prettify table 2023-11-25 13:48:56 +01:00
Zuellni 98ec700a5c Remove the tensor numel check, should be redundant 2023-11-25 13:21:19 +01:00
Zuellni 69b70b1b41 Clarify gpu split should be in GB, not MB 2023-11-25 13:09:40 +01:00
Zuellni 1b9803ae34 Update wheels 2023-11-25 12:15:44 +01:00
Zuellni 677bc6c388 Update readme 2023-11-25 12:10:53 +01:00
Zuellni 1bfa440ef1 Add progress bar to model loader 2023-11-24 18:49:50 +01:00
Zuellni 09de8351b9 Update flash attention wheel 2023-11-24 18:09:04 +01:00
Zuellni 3e2c6d6759 Update link 2023-11-24 17:42:03 +01:00
Zuellni 125cdf51cb Prettify readme 2023-11-24 14:58:06 +01:00
Zuellni 8ca924691f Update readme 2023-11-23 22:47:49 +01:00
Zuellni 68f563eb4b Add field descriptions 2023-11-23 22:44:15 +01:00
Zuellni d881503d4e Add some info about model downloading 2023-11-23 22:15:10 +01:00
Zuellni 9cbbcb1e72 Update readme 2023-11-23 20:22:45 +01:00