Commit Graph
100 Commits
Author SHA1 Message Date
Zuellni 5c63d0609f Remove the hidden inputs since they don't seem to be working anyway (no prompt in workflow), touch-up the convert node a bit 2024-07-29 19:22:39 +02:00
Zuellni 1b02390317 Update workflow 2024-07-29 16:53:02 +02:00
Zuellni 42d621d4fd Merge pull request #25 from Zuellni/dev
Overhaul
2024-07-29 16:52:19 +02:00
Zuellni 66a538e316 Update README.md 2024-07-29 15:51:34 +02:00
Zuellni 2db2d93844 Added a convert node, a string node and fixed some js (probably) 2024-07-29 15:41:49 +02:00
Zuellni 2acc730ad3 Update README.md 2024-07-28 21:31:21 +02:00
Zuellni 7ae7ef9a24 Split some nodes, added a format node that uses tokenizer config template 2024-07-28 17:33:07 +02:00
Zuellni 8311351697 Cache should be a multiple of 256 for paged attention 2024-06-26 16:01:08 +02:00
Zuellni d5feb5a293 Clean up cache 2024-06-26 13:07:37 +02:00
Zuellni 0bd03f4dae Clean up imports 2024-06-24 12:33:30 +02:00
Zuellni cc6c50e2ff Merge pull request #23 from Zuellni/dev
Switch to dynamic generator, bump minimum exllamav2 version
2024-06-14 18:52:24 +02:00
Zuellni 5a90464b6c Update README.md 2024-06-14 18:48:05 +02:00
Zuellni 51728c0102 Update requirements.txt 2024-06-14 13:28:48 +02:00
Zuellni cac9d1fdf4 Update README.md 2024-06-14 13:24:38 +02:00
Zuellni f54093078e Clean up README 2024-06-14 13:21:42 +02:00
Zuellni add411a7c3 Merge remote-tracking branch 'origin/main' into dev 2024-06-14 13:15:47 +02:00
Zuellni 49ace07a0e Update README.md 2024-06-14 12:51:36 +02:00
Zuellni 4754f78c70 Add option to enable/disable flash attention and fast tensors 2024-06-14 11:50:22 +02:00
Zuellni f988703614 Add custom stop conditions 2024-06-14 11:33:10 +02:00
Zuellni e91a1ce2c6 Switch to dynamic generator 2024-06-14 10:57:28 +02:00
Zuellni e5e13c40ed Update README.md 2024-06-05 09:53:12 +02:00
Zuellni d7413b2ebc Fix vram usage with higher context 2024-05-17 22:55:50 +02:00
Zuellni c005e3c637 Update config init 2024-05-04 09:13:01 +02:00
Zuellni 4422f0b0d0 Update readme 2024-04-22 18:30:33 +02:00
Zuellni 4633de563a Update workflow with llama-3 format 2024-04-22 16:18:18 +02:00
Zuellni 87810044ea Change default cache bits to 16, reorganize imports, update readme 2024-04-22 15:50:50 +02:00
Zuellni 8c2f17829a Clarify instructions 2024-04-08 03:03:53 +02:00
Zuellni c5765ffb8c Update with a cascade workflow 2024-04-07 20:19:55 +02:00
Zuellni fb7b7ff683 Update readme 2024-04-07 20:00:44 +02:00
Zuellni fb5b6d95ef Update link 2024-04-07 19:55:34 +02:00
Zuellni d2543cbd59 Simplify instructions 2024-04-07 19:52:53 +02:00
Zuellni baed69aba5 Bump minimum exllamav2 version, add autosplit and q4 cache, some reformatting 2024-04-07 19:35:32 +02:00
Zuellni d2b4d7a796 More sanity checks for model unloading, bump exllama 2024-02-26 13:24:03 +01:00
Zuellni f7a108315e Update readme 2024-02-21 23:06:35 +01:00
Zuellni 217304e0ef Only unload other models if unload is checked 2024-02-17 16:53:00 +01:00
Zuellni 8c4b9d058f Update workflow 2024-02-17 15:14:46 +01:00
Zuellni 861aa6fc51 Update readme 2024-02-17 15:13:53 +01:00
Zuellni 8b7b957b3e Add gc again 2024-02-17 14:53:54 +01:00
Zuellni e1c8e61291 Unload other models first before loading llms, rename some params, update requirements 2024-02-17 14:40:46 +01:00
Zuellni ca41704f38 Update readme 2024-02-14 20:26:46 +01:00
Zuellni bf64bcc0e4 Split requirements to file, neutralize samplers by default 2024-02-14 20:23:27 +01:00
Zuellni 9ffbacf080 Update requirements 2024-02-12 16:38:52 +01:00
Zuellni 6cc728d7cb Update readme 2024-02-09 16:13:30 +01:00
Zuellni ec6443816e Update readme 2024-02-06 14:51:02 +01:00
Zuellni 6831b951c5 Update readme 2024-02-03 15:53:52 +01:00
Zuellni 54bedbed53 Update workflow 2024-02-03 15:52:58 +01:00
Zuellni 88a826fa30 Add top_a sampling 2024-02-03 15:14:38 +01:00
Zuellni 18c7854d76 Add missing import, look for paths recursively 2024-02-03 15:13:30 +01:00
Zuellni 261a73a119 Merge pull request #14 from ScottNealon/patch-1
Parse all possible folders mapped to "llm"
2024-02-03 15:11:39 +01:00
Zuellni 73e6a0b0e4 Bump exllama 2024-01-26 15:39:58 +01:00
Zuellni 331b594f00 Bump flash-attn 2023-12-27 13:13:16 +01:00
Zuellni de401d308d Bump exllama 2023-12-21 19:51:05 +01:00
Zuellni 04f0c87e31 Fix error when model isn't fully unloaded and bump requirements 2023-12-02 23:27:41 +01:00
Zuellni eca815e45a Fix linux requirements 2023-11-25 20:16:56 +01:00
Zuellni 15a224ccbb Even better table 2023-11-25 14:55:51 +01:00
Zuellni 485e2c4c51 Prettify table 2023-11-25 13:48:56 +01:00
Zuellni 98ec700a5c Remove the tensor numel check, should be redundant 2023-11-25 13:21:19 +01:00
Zuellni 69b70b1b41 Clarify gpu split should be in GB, not MB 2023-11-25 13:09:40 +01:00
Zuellni 1b9803ae34 Update wheels 2023-11-25 12:15:44 +01:00
Zuellni 677bc6c388 Update readme 2023-11-25 12:10:53 +01:00
Zuellni 1bfa440ef1 Add progress bar to model loader 2023-11-24 18:49:50 +01:00
Zuellni 09de8351b9 Update flash attention wheel 2023-11-24 18:09:04 +01:00
Zuellni 3e2c6d6759 Update link 2023-11-24 17:42:03 +01:00
Zuellni 125cdf51cb Prettify readme 2023-11-24 14:58:06 +01:00
Zuellni 8ca924691f Update readme 2023-11-23 22:47:49 +01:00
Zuellni 68f563eb4b Add field descriptions 2023-11-23 22:44:15 +01:00
Zuellni d881503d4e Add some info about model downloading 2023-11-23 22:15:10 +01:00
Zuellni 9cbbcb1e72 Update readme 2023-11-23 20:22:45 +01:00
Zuellni 4d65f3c4db Update readme 2023-11-23 20:12:02 +01:00
Zuellni 2f1affecb4 Simplify model loader
Load from custom 'llm' dir specified in 'extra_model_paths.yaml' or from 'models/llm'
2023-11-23 20:04:14 +01:00
Zuellni 8b75af13b3 Update workflow 2023-11-22 16:15:23 +01:00
Zuellni 72a44cbcca Fix preview output 2023-11-22 16:14:28 +01:00
Zuellni 51bb1bcf84 Revert autosplit 2023-11-22 16:04:01 +01:00
Zuellni 996fdd012d Update readme 2023-11-22 14:28:23 +01:00
Zuellni 51ff962f14 Add autosplit, temperature_last 2023-11-22 14:15:00 +01:00
Zuellni 957fe3e6a5 Bump exllamav2 2023-11-22 12:38:08 +01:00
Zuellni 05a30fce3f Remove unnecessary line in preview 2023-11-21 01:56:57 +01:00
Zuellni b6cd5f0de8 Update workflow 2023-11-20 23:37:17 +01:00
Zuellni a33a1ac395 Add gpu split, add 8bit cache toggle,
add min_p, encode specal tokens,
update requirements versions
2023-11-20 21:44:52 +01:00
Zuellni 9d20724d02 Update readme 2023-11-07 14:04:16 +01:00
Zuellni 47e38dabbc Update workflow 2023-11-07 13:04:37 +01:00
Zuellni cb7f195068 Merge pull request #9 from Zuellni/0.0.7
0.0.7
2023-11-07 13:03:56 +01:00
Zuellni 7e97b26af9 Fix preview size 2023-11-07 12:55:39 +01:00
Zuellni f48ebd1c68 Some minor changes 2023-11-07 12:46:50 +01:00
Zuellni 55871388b6 Move for clarity 2023-10-27 15:54:24 +02:00
Zuellni dcaade09c2 Remove loras for now, there seems to be a memory leak and idk how to fix it
Remove allowed strings, they don't seem very useful
2023-10-27 15:51:15 +02:00
Zuellni 2b3ddde76b Update README.md 2023-10-25 19:59:58 +02:00
Zuellni 0b94a076ab Update README.md 2023-10-25 19:57:16 +02:00
Zuellni d2c554b69d Add loras, 8bit cache, fix random seed, unloading 2023-10-25 19:53:44 +02:00
Zuellni 420b1fbd2d Update workflow 2023-10-10 22:21:58 +02:00
Zuellni ebb32253a9 Allow only some tokens in output, add condition node
save prompt without having to preview, auto set seq len and tokens if 0
and some other stuff I already forgot, it should all work, but likely won't
2023-10-10 22:20:39 +02:00
Zuellni f16a2ab515 Update readme link 2023-10-08 20:09:23 +02:00
Zuellni ebf8ccab60 Garbage collect model before reloading, should fix vram issues 2023-10-07 22:59:45 +02:00
Zuellni aa44ddb7b5 Change stop_on_newline to boolean 2023-10-07 22:39:47 +02:00
Zuellni 09bc77ba94 Check if the prompt's not empty before loading anything 2023-10-07 11:02:33 +02:00
Zuellni 6e78068c1a Split nodes into files, allow loading each separately without cloning the repo, some other minor stuff, install exllamav2 pip package by default 2023-10-07 10:56:18 +02:00
Zuellni 796a6b8f25 Bump exllamav2 2023-10-05 18:16:29 +02:00
Zuellni ba03e64315 Update workflow 2023-09-29 18:33:22 +02:00
Zuellni 164b9e95e0 Rename formatter to replacer 2023-09-29 18:17:36 +02:00
Zuellni 8f39f98a7d Update readme 2023-09-29 14:46:45 +02:00