diff --git a/README.md b/README.md index 90f40f2..ec853bf 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ # WORK IN PROGRESS -## Currently requires `flash_attn` ! +## Note: Sampling is slow without `flash_attn` ! For Linux users this doesn't mean anything but `pip install flash_attn`. @@ -8,6 +8,8 @@ However doing same on Windows currently will most likely fail if you do not have Alternative for Windows can be pre-built wheels from here, has to match your python environment: https://github.com/bdashore3/flash-attention/releases +If flash_attn is not installed, attention code will fallback to torch SDP attention, which is at least twice as slow and memory hungry. + ## Text encoder setup Lumina-next uses Google's Gemma-2b -LLM: https://huggingface.co/google/gemma-2b