Update README.md
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# WORK IN PROGRESS
|
||||
|
||||
## Currently requires `flash_attn` !
|
||||
## Note: Sampling is slow without `flash_attn` !
|
||||
|
||||
For Linux users this doesn't mean anything but `pip install flash_attn`.
|
||||
|
||||
@@ -8,6 +8,8 @@ However doing same on Windows currently will most likely fail if you do not have
|
||||
Alternative for Windows can be pre-built wheels from here, has to match your python environment:
|
||||
https://github.com/bdashore3/flash-attention/releases
|
||||
|
||||
If flash_attn is not installed, attention code will fallback to torch SDP attention, which is at least twice as slow and memory hungry.
|
||||
|
||||
## Text encoder setup
|
||||
|
||||
Lumina-next uses Google's Gemma-2b -LLM: https://huggingface.co/google/gemma-2b
|
||||
|
||||
Reference in New Issue
Block a user