From ab3e63ca1276761b123ddaf6e7ac49d273e48a99 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Jukka=20Sepp=C3=A4nen?= <40791699+kijai@users.noreply.github.com> Date: Mon, 17 Jun 2024 19:50:23 +0300 Subject: [PATCH] Update README.md --- README.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/README.md b/README.md index 90f40f2..ec853bf 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ # WORK IN PROGRESS -## Currently requires `flash_attn` ! +## Note: Sampling is slow without `flash_attn` ! For Linux users this doesn't mean anything but `pip install flash_attn`. @@ -8,6 +8,8 @@ However doing same on Windows currently will most likely fail if you do not have Alternative for Windows can be pre-built wheels from here, has to match your python environment: https://github.com/bdashore3/flash-attention/releases +If flash_attn is not installed, attention code will fallback to torch SDP attention, which is at least twice as slow and memory hungry. + ## Text encoder setup Lumina-next uses Google's Gemma-2b -LLM: https://huggingface.co/google/gemma-2b