add llama-cpp-python with GPU Support (CUDA) and files to run on windows.linux/macOs . Add dropdown list to load gguf for CPU/GPU
This commit is contained in:
@@ -2,18 +2,17 @@
|
||||
|
||||
A custom node pack for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) that allows you to run Large Language Models (LLMs) locally and use them for prompt generation and other text tasks directly within your ComfyUI workflows.
|
||||
|
||||
This pack provides nodes to connect to and utilize local LLMs in Hugging Face (PyTorch) or GGUF format, eliminating the need for external API calls. It's designed to integrate seamlessly with prompt generation workflows, such as those involving image description nodes like Florence-2.
|
||||
This pack provides nodes to connect to and utilize local LLMs (like Llama, Phi, Gemma, Hermes in Hugging Face PyTorch format, or GGUF models) without needing external API calls. It's designed to integrate seamlessly with prompt generation workflows, such as those involving image description nodes like Florence-2, and to simplify the creation of complex prompts for models like Flux Kontext Dev.
|
||||
|
||||
## Features
|
||||
|
||||
* **Local LLM Execution (Hugging Face & GGUF):** Run powerful LLMs directly on your machine (CPU or GPU) using either standard Hugging Face models or efficient GGUF models.
|
||||
* **Set Local LLM Service Connector Node (Hugging Face):** Select and configure your local Hugging Face format LLM model (models must be placed in `ComfyUI/models/LLM/`).
|
||||
* **Set Local GGUF LLM Service Connector Node:** Select and configure your local GGUF format LLM model file (`.gguf` files must be placed in `ComfyUI/models/LLM/`).
|
||||
* **Local Kontext Prompt Generator Node:** Generate detailed image prompts by combining descriptions and edit instructions, leveraging your connected local LLM.
|
||||
* **Set Local LLM Service Connector Node (HuggingFace):** Select and configure your local Hugging Face format LLM model (models must be placed in `ComfyUI/models/LLM/`).
|
||||
* **Set Local GGUF LLM Service Connector Node:** Select and configure your local GGUF format LLM model file (`.gguf` files must be placed in `ComfyUI/models/LLM/`). **Includes dropdown for device selection (CPU/GPU) and `n_gpu_layers` slider for fine-grained control.**
|
||||
* **Local Kontext Prompt Generator Node:** **(Key Feature)** Generates detailed image prompts by intelligently combining image descriptions (e.g., from Florence-2) with simple user instructions. Designed to work with local LLM connectors to produce high-quality prompts for advanced models like Flux Kontext Dev, simplifying the user's task.
|
||||
* **User Preset Management:** Add and remove custom prompt generation presets using dedicated nodes.
|
||||
* **Compatibility:** Designed to work with the `LLMServiceConnector` type identifier, ensuring compatibility with standard MieNodes prompt generators (e.g., `KontextPromptGenerator`) if needed.
|
||||
* **VRAM Optimization (Hugging Face):** Includes commented code examples for integrating Hugging Face model quantization (4-bit/8-bit) using `bitsandbytes` to reduce memory footprint.
|
||||
* **Efficient GGUF Models:** GGUF models are inherently quantized, offering lower memory usage and often good performance, especially on CPU.
|
||||
* **VRAM Optimization Ready:** Includes commented code examples for integrating quantization (4-bit/8-bit using `bitsandbytes` for Hugging Face models, or controlling `n_gpu_layers` for GGUF) to reduce memory footprint for running alongside large image models like Flux.
|
||||
* **Simplified User Experience:** Allows users to provide simple, natural language requests (e.g., "Make it look like it's being used in a luxury spa") and translates them into complex, Flux-ready prompts using the connected local LLM.
|
||||
|
||||
## Installation
|
||||
|
||||
@@ -24,24 +23,78 @@ This pack provides nodes to connect to and utilize local LLMs in Hugging Face (P
|
||||
git clone https://github.com/your_username/ComfyUI_LocalLLMNodes.git
|
||||
# Or download the zip and extract it into a folder named ComfyUI_LocalLLMNodes
|
||||
```
|
||||
4. **Install Dependencies:** Navigate into the `ComfyUI_LocalLLMNodes` directory and install the required Python packages.
|
||||
4. **Install Dependencies:** You can install dependencies using `pip` and the provided scripts or `requirements.txt`.
|
||||
* **Option 1: Using Installation Scripts (Recommended for GPU Support):**
|
||||
* Navigate to the `ComfyUI_LocalLLMNodes` directory:
|
||||
```bash
|
||||
cd ComfyUI_LocalLLMNodes
|
||||
```
|
||||
* **Linux/macOS:**
|
||||
```bash
|
||||
# Make the script executable
|
||||
chmod +x install_deps.sh
|
||||
# Run the script
|
||||
./install_deps.sh
|
||||
```
|
||||
* **Windows (Command Prompt):**
|
||||
```cmd
|
||||
install_deps.bat
|
||||
```
|
||||
* **Windows (PowerShell):**
|
||||
```powershell
|
||||
# You might need to adjust execution policy first:
|
||||
# Set-ExecutionPolicy RemoteSigned -Scope CurrentUser
|
||||
.\install_deps.ps1
|
||||
```
|
||||
* These scripts will install core dependencies and specifically `llama-cpp-python` with CUDA support (essential for GPU acceleration with GGUF models).
|
||||
* **Option 2: Standard `pip install`:**
|
||||
```bash
|
||||
cd ComfyUI_LocalLLMNodes
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
*Note: Ensure you are installing these packages in the same Python environment that you use to run ComfyUI.*
|
||||
*Note on `llama-cpp-python`: The standard `pip install llama-cpp-python` often installs a CPU-only version. For GPU acceleration (highly recommended for GGUF models), use the installation scripts above or follow the manual steps below.*
|
||||
|
||||
### Installing `llama-cpp-python` with GPU Support (CUDA) - Important for GGUF Nodes
|
||||
|
||||
To leverage your GPU for running GGUF models via the `Set Local GGUF LLM Service Connector 🐑` node, you need to install `llama-cpp-python` with CUDA support compiled in. The standard `pip install llama-cpp-python` often installs a CPU-only version.
|
||||
|
||||
**Manual Installation (if not using scripts):**
|
||||
|
||||
1. **Ensure CUDA Toolkit is Installed:** You need the NVIDIA CUDA toolkit installed on your system, matching the version compatible with your GPU drivers. Check NVIDIA's website for instructions.
|
||||
2. **Set Environment Variables and Install:**
|
||||
* **Linux/macOS:**
|
||||
```bash
|
||||
# Replace cu118/cu121/cu124 with the CUDA version you have installed (e.g., cu118 for CUDA 11.8, cu121 for CUDA 12.1)
|
||||
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --force-reinstall --no-cache-dir
|
||||
```
|
||||
* **Windows (Command Prompt):**
|
||||
```cmd
|
||||
set CMAKE_ARGS=-DGGML_CUDA=on
|
||||
pip install llama-cpp-python --force-reinstall --no-cache-dir
|
||||
```
|
||||
* **Windows (PowerShell):**
|
||||
```powershell
|
||||
$env:CMAKE_ARGS = "-DGGML_CUDA=on"
|
||||
pip install llama-cpp-python --force-reinstall --no-cache-dir
|
||||
```
|
||||
3. **Verify Installation:** After installation, you can check if CUDA support is enabled by running Python and trying to import:
|
||||
```bash
|
||||
cd ComfyUI_LocalLLMNodes
|
||||
pip install -r requirements.txt
|
||||
python -c "import llama_cpp; print('llama_cpp imported successfully')"
|
||||
# A successful import without errors related to CUDA libraries usually indicates it's compiled correctly.
|
||||
# Detailed logs during model loading (like `load_tensors: layer X assigned to device CUDA0`) will confirm GPU usage.
|
||||
```
|
||||
*Note: Ensure you are installing these packages in the same Python environment that you use to run ComfyUI.*
|
||||
*Note on `llama-cpp-python`: Installing with GPU support (CUDA) requires specific environment variables during installation (e.g., `CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install llama-cpp-python --force-reinstall --no-cache-dir`). See the `llama-cpp-python` documentation for details. The CPU-only installation (`pip install llama-cpp-python`) is simpler and sufficient for CPU inference.
|
||||
|
||||
## Usage
|
||||
|
||||
1. **Download a Local LLM:**
|
||||
* **Hugging Face Format:**
|
||||
* Obtain a Hugging Face format LLM (e.g., `TinyLlama/TinyLlama-1.1B-Chat-v1.0`, `microsoft/Phi-3-mini-4k-instruct`, `NousResearch/Hermes-2-Pro-Llama-3-8B`).
|
||||
* Obtain a Hugging Face format LLM (e.g., `TinyLlama/TinyLlama-1.1B-Chat-v1.0`, `microsoft/Phi-3-mini-4k-instruct`, `NousResearch/Hermes-2-Pro-Llama-3-8B`, `mistralai/Mistral-Nemo-Instruct-2407`).
|
||||
* Download the model files into a subdirectory within your `ComfyUI/models/LLM/` folder.
|
||||
* Example: `ComfyUI/models/LLM/Phi-3-mini-4k-instruct/` should contain `config.json`, `pytorch_model.bin` (or `.safetensors`), `tokenizer_config.json`, etc.
|
||||
* **GGUF Format:**
|
||||
* Obtain a GGUF format LLM file (e.g., `mistral-7b-instruct-v0.3.Q8_0.gguf`).
|
||||
* Place the `.gguf` file directly in your `ComfyUI/models/LLM/` folder, or within a subfolder (e.g., `ComfyUI/models/LLM/Mistral-7B-Instruct/` containing `mistral-7b-instruct-v0.3.Q8_0.gguf`).
|
||||
* Obtain a GGUF format LLM file (e.g., `mistral-nemo-instruct-2407.Q8_0.gguf`, `Hermes-2-Pro-Llama-3-8B-Q8_0.gguf`).
|
||||
* Place the `.gguf` file directly in your `ComfyUI/models/LLM/` folder, or within a subfolder (e.g., `ComfyUI/models/LLM/Mistral-Nemo-Instruct-2407/` containing `mistral-nemo-instruct-2407.Q8_0.gguf`).
|
||||
2. **Restart ComfyUI** to load the new nodes.
|
||||
3. **Find the Nodes:** Look for the new nodes in the ComfyUI node library under the category:
|
||||
* `Local LLM Nodes/LLM Connectors`
|
||||
@@ -51,31 +104,40 @@ This pack provides nodes to connect to and utilize local LLMs in Hugging Face (P
|
||||
* Select your downloaded Hugging Face model directory from the dropdown menu.
|
||||
* **For GGUF Models:**
|
||||
* Add the **"Set Local GGUF LLM Service Connector 🐑"** node to your graph.
|
||||
* Select your downloaded GGUF model file (or its containing directory) from the dropdown menu.
|
||||
* **Common Steps:**
|
||||
* **Select your downloaded GGUF model file (or its containing directory) from the dropdown menu.**
|
||||
* **Use the `device` dropdown:** Choose `GPU` to attempt GPU acceleration (requires `llama-cpp-python` installed with CUDA support as described above). Choose `CPU` to force CPU execution.
|
||||
* **Adjust `n_gpu_layers` slider:**
|
||||
* If `device` is `GPU`, the slider defaults to `-1`, meaning "offload as many layers as possible to the GPU".
|
||||
* If `device` is `CPU`, the slider defaults to `-1`, but the node internally sets `n_gpu_layers=0`.
|
||||
* You can manually adjust the slider to offload a specific number of layers (e.g., `30` out of `32`) if desired, overriding the default behavior.
|
||||
* **Common Steps for Prompt Generation:**
|
||||
* Add the **"Local Kontext Prompt Generator 🐑"** node.
|
||||
* Connect the output of the chosen "Set Local ... LLM Service Connector 🐑" node to the `llm_service_connector` input of the "Local Kontext Prompt Generator 🐑" node.
|
||||
* Provide inputs like `image1_description` (e.g., from Florence-2), `edit_instruction`, and select a `preset`.
|
||||
* Connect the `kontext_prompt` output to your desired node (e.g., an image generator).
|
||||
* **Provide Inputs:**
|
||||
* Connect the output of an image description node (like Florence-2) to the `image1_description` input.
|
||||
* Provide a simple, natural language `edit_instruction` in the node's text field (e.g., "Make it look like it's being used in a luxury spa", "Change the background to a beach").
|
||||
* Select a suitable `preset` from the dropdown (e.g., "User Intent -> Flux Prompt", "Product - 产品摄影").
|
||||
* Connect the `kontext_prompt` output to your desired node (e.g., an image generator like Flux Kontext Dev).
|
||||
|
||||
## Memory Optimization
|
||||
|
||||
Running large LLMs alongside large image models (like SDXL or Flux) can strain system resources (RAM/VRAM).
|
||||
|
||||
* **Hugging Face Models:**
|
||||
* **Quantization:** The `local_llm_connector.py` file includes commented code examples showing how to implement 4-bit or 8-bit quantization using the `bitsandbytes` library. This can significantly reduce the LLM's VRAM usage.
|
||||
* **Quantization:** The `local_llm_connector.py` file includes commented code examples for integrating Hugging Face model quantization (4-bit/8-bit) using the `bitsandbytes` library. This can significantly reduce the LLM's VRAM usage.
|
||||
* To use quantization:
|
||||
1. Ensure `bitsandbytes` is installed (`pip install bitsandbytes`).
|
||||
2. Uncomment and adjust the quantization configuration section in the `_load_model` method within `local_llm_connector.py`.
|
||||
3. Restart ComfyUI.
|
||||
* **GGUF Models:**
|
||||
* GGUF models are pre-quantized (e.g., Q4, Q5, Q8). Choosing a more quantized version (like Q8_0 vs. f16) inherently uses less memory. For GPU acceleration with GGUF models, configure the `n_gpu_layers` parameter during loading (if supported by your `llama-cpp-python` build).
|
||||
* GGUF models are inherently quantized (e.g., Q4_K_M, Q5_K, Q8_0). Choosing a more quantized version (like Q8_0 vs. f16) inherently uses less memory.
|
||||
* For GPU acceleration with GGUF models, configure the `n_gpu_layers` parameter during loading (if supported by your `llama-cpp-python` build and setup). The `Set Local GGUF LLM Service Connector 🐑` node provides explicit controls for this.
|
||||
|
||||
## Nodes Included
|
||||
|
||||
* `SetLocalLLMServiceConnector`: Selects and prepares a connection to a local Hugging Face format LLM model.
|
||||
* `SetLocalGGUFLLMServiceConnector`: Selects and prepares a connection to a local GGUF format LLM model file.
|
||||
* `LocalKontextPromptGenerator`: Generates prompts using a connected local LLM based on descriptions and instructions.
|
||||
* `SetLocalGGUFLLMServiceConnector`: **(Updated)** Selects and prepares a connection to a local GGUF format LLM model file. Includes `device` dropdown (`CPU`/`GPU`) and `n_gpu_layers` slider for controlling offloading.
|
||||
* `LocalKontextPromptGenerator`: **(Key Node)** Generates prompts using a connected local LLM by combining image descriptions and simple user instructions, optimized for advanced image models like Flux Kontext Dev.
|
||||
* `AddUserLocalKontextPreset`: Adds a custom preset for prompt generation.
|
||||
* `RemoveUserLocalKontextPreset`: Removes a custom preset.
|
||||
|
||||
@@ -85,10 +147,10 @@ Running large LLMs alongside large image models (like SDXL or Flux) can strain s
|
||||
* Python Libraries (see `requirements.txt` for versions):
|
||||
* `transformers`
|
||||
* `torch`
|
||||
* `llama-cpp-python` (Optional, for GGUF model support - **install with GPU flags as described for GPU acceleration**)
|
||||
* `bitsandbytes` (Optional, for Hugging Face model quantization)
|
||||
* `llama-cpp-python` (Optional, for GGUF model support)
|
||||
* Other dependencies as listed in `requirements.txt`
|
||||
|
||||
## Acknowledgements
|
||||
|
||||
This node pack builds upon concepts and structures found in the excellent [ComfyUI-MieNodes](https://github.com/MieMieeeee/ComfyUI-MieNodes) extension, particularly the `KontextPromptGenerator` and LLM service connector patterns.
|
||||
This node pack builds upon concepts and structures found in the excellent [ComfyUI-MieNodes](https://github.com/MieMieeeee/ComfyUI-MieNodes) extension, particularly the `KontextPromptGenerator` and LLM service connector patterns.
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
@echo off
|
||||
REM install_deps.bat - Installs dependencies for ComfyUI_LocalLLMNodes (Command Prompt)
|
||||
|
||||
REM Exit immediately if a command exits with a non-zero status.
|
||||
if "%~1"=="-CI" set CI=1
|
||||
setlocal enabledelayedexpansion
|
||||
|
||||
echo ===== Installing ComfyUI_LocalLLMNodes Dependencies =====
|
||||
|
||||
REM --- Install core dependencies from requirements.txt ---
|
||||
REM This installs transformers, torch, bitsandbytes (if listed), and llama-cpp-python (CPU version potentially)
|
||||
echo 1/3: Installing core dependencies from requirements.txt...
|
||||
pip install -r requirements.txt
|
||||
if errorlevel 1 (
|
||||
echo Error occurred during pip install -r requirements.txt
|
||||
exit /b 1
|
||||
)
|
||||
|
||||
REM --- Install llama-cpp-python with CUDA support ---
|
||||
REM This overwrites/ensures the GPU-accelerated version is installed
|
||||
echo 2/3: Installing llama-cpp-python with CUDA support...
|
||||
echo (This might take a while and download/build components...)
|
||||
REM Set environment variable and install llama-cpp-python
|
||||
REM Adjust cu118/cu121/cu124 based on your CUDA version
|
||||
set CMAKE_ARGS=-DGGML_CUDA=on
|
||||
pip install llama-cpp-python --force-reinstall --no-cache-dir
|
||||
if errorlevel 1 (
|
||||
echo Error occurred during pip install llama-cpp-python
|
||||
exit /b 1
|
||||
)
|
||||
|
||||
REM --- Completion ---
|
||||
echo 3/3: Installation process completed.
|
||||
echo ===== Installation Finished =====
|
||||
echo.
|
||||
echo Please ensure you have the NVIDIA CUDA toolkit installed
|
||||
echo and that your environment is set up correctly for ComfyUI.
|
||||
echo You can now start ComfyUI.
|
||||
|
||||
endlocal
|
||||
@@ -0,0 +1,33 @@
|
||||
# install_deps.ps1 - Installs dependencies for ComfyUI_LocalLLMNodes (PowerShell)
|
||||
|
||||
Write-Host "===== Installing ComfyUI_LocalLLMNodes Dependencies ====="
|
||||
|
||||
# --- Install core dependencies from requirements.txt ---
|
||||
# This installs transformers, torch, bitsandbytes (if listed), and llama-cpp-python (CPU version potentially)
|
||||
Write-Host "1/3: Installing core dependencies from requirements.txt..."
|
||||
pip install -r requirements.txt
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
Write-Error "Error occurred during pip install -r requirements.txt"
|
||||
exit 1
|
||||
}
|
||||
|
||||
# --- Install llama-cpp-python with CUDA support ---
|
||||
# This overwrites/ensures the GPU-accelerated version is installed
|
||||
Write-Host "2/3: Installing llama-cpp-python with CUDA support..."
|
||||
Write-Host " (This might take a while and download/build components...)"
|
||||
# Set environment variable and install llama-cpp-python
|
||||
# Adjust cu118/cu121/cu124 based on your CUDA version
|
||||
$env:CMAKE_ARGS = "-DGGML_CUDA=on"
|
||||
pip install llama-cpp-python --force-reinstall --no-cache-dir
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
Write-Error "Error occurred during pip install llama-cpp-python"
|
||||
exit 1
|
||||
}
|
||||
|
||||
# --- Completion ---
|
||||
Write-Host "3/3: Installation process completed."
|
||||
Write-Host "===== Installation Finished ====="
|
||||
Write-Host ""
|
||||
Write-Host "Please ensure you have the NVIDIA CUDA toolkit installed"
|
||||
Write-Host "and that your environment is set up correctly for ComfyUI."
|
||||
Write-Host "You can now start ComfyUI."
|
||||
@@ -0,0 +1,29 @@
|
||||
#!/bin/bash
|
||||
|
||||
# install_deps.sh - Installs dependencies for ComfyUI_LocalLLMNodes
|
||||
|
||||
# Exit immediately if a command exits with a non-zero status.
|
||||
set -e
|
||||
|
||||
echo "===== Installing ComfyUI_LocalLLMNodes Dependencies ====="
|
||||
|
||||
# --- Install core dependencies from requirements.txt ---
|
||||
# This installs transformers, torch, bitsandbytes (if listed), and llama-cpp-python (CPU version potentially)
|
||||
echo "1/3: Installing core dependencies from requirements.txt..."
|
||||
pip install -r requirements.txt
|
||||
|
||||
# --- Install llama-cpp-python with CUDA support ---
|
||||
# This overwrites/ensures the GPU-accelerated version is installed
|
||||
echo "2/3: Installing llama-cpp-python with CUDA support..."
|
||||
echo " (This might take a while and download/build components...)"
|
||||
# Set environment variable and install llama-cpp-python
|
||||
# Adjust cu118/cu121/cu124 based on your CUDA version (check `nvcc --version`)
|
||||
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --force-reinstall --no-cache-dir
|
||||
|
||||
# --- Completion ---
|
||||
echo "3/3: Installation process completed."
|
||||
echo "===== Installation Finished ====="
|
||||
echo ""
|
||||
echo "Please ensure you have the NVIDIA CUDA toolkit installed"
|
||||
echo "and that your environment is set up correctly for ComfyUI."
|
||||
echo "You can now start ComfyUI."
|
||||
+77
-71
@@ -59,18 +59,29 @@ class SetLocalGGUFLLMServiceConnector:
|
||||
return {
|
||||
"required": {
|
||||
"local_gguf_model_name": (model_names, {"default": model_names[0] if model_names else "No_Local_GGUF_Models_Found"}),
|
||||
# Optional parameters for llama-cpp-python loading
|
||||
# "n_ctx": ("INT", {"default": 4096, "min": 1, "max": 100000}),
|
||||
# "n_gpu_layers": ("INT", {"default": 0, "min": -1, "max": 100}), # -1 = all
|
||||
# --- Add device selection dropdown ---
|
||||
"device": (["GPU", "CPU"], {"default": "GPU"}), # Default to GPU
|
||||
# --- n_gpu_layers slider ---
|
||||
"n_gpu_layers": ("INT", {
|
||||
"default": -1, # This default will be overridden based on 'device'
|
||||
"min": -1, # -1 for 'all possible'
|
||||
"max": 100, # Adjust based on typical model layer counts if needed
|
||||
"step": 1,
|
||||
"display": "slider"
|
||||
}),
|
||||
# --- Optional: Expose other common parameters ---
|
||||
# "n_threads": ("INT", {"default": 8, "min": 1, "max": 64}),
|
||||
# "n_ctx": ("INT", {"default": 4096, "min": 1, "max": 32768}), # Max depends on model
|
||||
},
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("LLMServiceConnector",) # Use the same type for compatibility
|
||||
FUNCTION = "get_connector"
|
||||
CATEGORY = LOCAL_LLM_CATEGORY
|
||||
CATEGORY = LOCAL_LLM_CATEGORY # Should be defined earlier, e.g., "Local LLM Nodes/LLM Connectors"
|
||||
OUTPUT_NODE = False
|
||||
|
||||
def get_connector(self, local_gguf_model_name): # , n_ctx=4096, n_gpu_layers=0):
|
||||
# --- Update get_connector to accept 'device' and set n_gpu_layers default dynamically ---
|
||||
def get_connector(self, local_gguf_model_name, device, n_gpu_layers): # , n_threads=8, n_ctx=4096): # Add other params if exposed
|
||||
"""Returns a connector object for the selected GGUF model."""
|
||||
if not LLAMA_CPP_AVAILABLE:
|
||||
raise Exception("The 'llama-cpp-python' library is required for the Local GGUF LLM node but is not installed.")
|
||||
@@ -78,30 +89,49 @@ class SetLocalGGUFLLMServiceConnector:
|
||||
if local_gguf_model_name == "No_Local_GGUF_Models_Found":
|
||||
raise Exception("No local GGUF models found in models/LLM directory. Please place your .gguf files there.")
|
||||
|
||||
# Determine the base path selected by the user
|
||||
# Determine the full path to the .gguf file
|
||||
base_models_dir = os.path.join(folder_paths.models_dir, "LLM")
|
||||
model_path = os.path.join(base_models_dir, local_gguf_model_name)
|
||||
|
||||
gguf_file_path = None
|
||||
# Case 1: User selected a name derived from a file directly in models/LLM (e.g., 'model' for 'models/LLM/model.gguf')
|
||||
# model_path would be 'models/LLM/model' (not a .gguf file itself)
|
||||
if os.path.isfile(model_path) and model_path.endswith('.gguf'):
|
||||
gguf_file_path = model_path
|
||||
# Case 2: User selected a name derived from a directory (e.g., 'ModelDir' for 'models/LLM/ModelDir/')
|
||||
# model_path would be 'models/LLM/ModelDir' (a directory)
|
||||
elif os.path.isdir(model_path):
|
||||
# Search for the .gguf file inside the directory
|
||||
# Search for .gguf file inside the directory
|
||||
for file in os.listdir(model_path):
|
||||
if file.endswith('.gguf'):
|
||||
gguf_file_path = os.path.join(model_path, file)
|
||||
break # Use the first .gguf file found
|
||||
break
|
||||
|
||||
# Final check: Ensure we found a valid .gguf file path
|
||||
if not gguf_file_path or not os.path.exists(gguf_file_path):
|
||||
raise FileNotFoundError(f"[LocalGGUFLLMConnector] GGUF model file not found for selection: '{local_gguf_model_name}'. Expected file: '{gguf_file_path}'")
|
||||
raise FileNotFoundError(f"[LocalGGUFLLMConnector] GGUF model file not found for selection: {local_gguf_model_name}")
|
||||
|
||||
# Create and return the GGUF connector instance, passing the resolved .gguf file path
|
||||
connector = LocalGGUFLLMServiceConnector(gguf_file_path) # Pass kwargs like n_ctx, n_gpu_layers if added
|
||||
# --- Key Change: Set n_gpu_layers default based on device selection ---
|
||||
# Determine the final n_gpu_layers value to pass to the connector
|
||||
final_n_gpu_layers = n_gpu_layers
|
||||
|
||||
# Apply default logic based on device and slider state
|
||||
# If device is CPU and n_gpu_layers is the slider's default (-1), assume user wants CPU mode
|
||||
if device == "CPU" and n_gpu_layers == -1:
|
||||
final_n_gpu_layers = 0
|
||||
log(f"[LocalGGUFLLMConnector] Device set to CPU, overriding n_gpu_layers to 0.")
|
||||
# If device is GPU and n_gpu_layers is the slider's default (-1), keep -1 for max offload
|
||||
elif device == "GPU" and n_gpu_layers == -1:
|
||||
final_n_gpu_layers = -1
|
||||
log(f"[LocalGGUFLLMConnector] Device set to GPU, keeping n_gpu_layers=-1 (max offload).")
|
||||
else:
|
||||
# If user explicitly set n_gpu_layers (slider moved), respect that value regardless of device dropdown
|
||||
log(f"[LocalGGUFLLMConnector] Using user-provided n_gpu_layers={n_gpu_layers} (device={device}).")
|
||||
# --- End of Key Change ---
|
||||
|
||||
# --- Pass the potentially adjusted n_gpu_layers (and others if exposed) to the connector ---
|
||||
connector = LocalGGUFLLMServiceConnector(
|
||||
gguf_file_path,
|
||||
n_gpu_layers=final_n_gpu_layers
|
||||
# n_threads=n_threads, # Pass if exposed
|
||||
# n_ctx=n_ctx # Pass if exposed
|
||||
)
|
||||
# --- End of Passing Parameters ---
|
||||
return (connector,)
|
||||
|
||||
|
||||
@@ -110,12 +140,15 @@ class LocalGGUFLLMServiceConnector:
|
||||
Represents the connection to a specific local GGUF LLM using llama-cpp-python.
|
||||
The `gguf_file_path` should point directly to the .gguf file.
|
||||
"""
|
||||
def __init__(self, gguf_file_path): # , n_ctx=4096, n_gpu_layers=0):
|
||||
# --- Update __init__ to accept and store parameters ---
|
||||
def __init__(self, gguf_file_path, n_gpu_layers=-1): # , n_ctx=4096, n_threads=8):
|
||||
self.gguf_file_path = gguf_file_path
|
||||
self.n_gpu_layers = n_gpu_layers # Store n_gpu_layers
|
||||
# self.n_ctx = n_ctx # Store n_ctx if exposed
|
||||
# self.n_threads = n_threads # Store n_threads if exposed
|
||||
self.model = None
|
||||
self.is_loaded = False
|
||||
# self.n_ctx = n_ctx
|
||||
# self.n_gpu_layers = n_gpu_layers
|
||||
# --- End of Update ---
|
||||
|
||||
def _load_model(self):
|
||||
"""Loads the GGUF model using llama-cpp-python."""
|
||||
@@ -125,36 +158,35 @@ class LocalGGUFLLMServiceConnector:
|
||||
try:
|
||||
log(f"[LocalGGUFLLMConnector] Loading GGUF model from: {self.gguf_file_path}")
|
||||
# --- Model Loading Configuration for llama-cpp-python ---
|
||||
# Adjust these parameters based on your needs and system capabilities.
|
||||
# Use the parameters passed from the node
|
||||
model_kwargs = {
|
||||
"n_ctx": 4096, # Context window size (adjust if needed, larger uses more memory)
|
||||
"n_threads": 8, # Number of CPU threads to use
|
||||
# "n_threads_batch": 8, # For batch processing (if applicable)
|
||||
# --- GPU Acceleration (if llama-cpp-python was built with CUDA support) ---
|
||||
# "n_gpu_layers": self.n_gpu_layers, # Number of layers to offload to GPU (e.g., 33 for 7B models)
|
||||
# --- Verbosity ---
|
||||
# "verbose": False, # Set to False to reduce llama.cpp logging
|
||||
"n_ctx": getattr(self, 'n_ctx', 4096), # Use default if not passed/store, or just hardcode if not exposed
|
||||
"n_threads": getattr(self, 'n_threads', 8), # Use default if not passed/store, or just hardcode if not exposed
|
||||
# --- Key Change: Add n_gpu_layers ---
|
||||
"n_gpu_layers": self.n_gpu_layers, # Use the value set by the user/node (default adjusted by SetLocalGGUFLLMServiceConnector)
|
||||
# --- Optional: Adjust verbosity ---
|
||||
# "verbose": False, # Set to False to reduce llama.cpp logging if needed
|
||||
}
|
||||
# --- End of Key Change ---
|
||||
|
||||
# --- Load Model ---
|
||||
# Crucially, pass the full path to the .gguf file
|
||||
# Crucially, pass the full path to the .gguf file and the kwargs
|
||||
self.model = llama_cpp.Llama(model_path=self.gguf_file_path, **model_kwargs)
|
||||
|
||||
self.is_loaded = True
|
||||
log(f"[LocalGGUFLLMConnector] GGUF model loaded successfully from: {self.gguf_file_path}")
|
||||
log(f"[LocalGGUFLLMConnector] GGUF model loaded successfully.")
|
||||
except Exception as e:
|
||||
error_msg = f"[LocalGGUFLLMConnector] Failed to load GGUF model from '{self.gguf_file_path}': {e}"
|
||||
error_msg = f"[LocalGGUFLLMConnector] Failed to load GGUF model: {e}"
|
||||
log(error_msg)
|
||||
# Re-raise the exception to halt the node execution
|
||||
raise Exception(error_msg) from e # Chain the exception
|
||||
|
||||
# ... (rest of the class: invoke method remains largely the same) ...
|
||||
def invoke(self, messages, **generation_kwargs):
|
||||
"""
|
||||
Generates text using the local GGUF LLM based on the messages list.
|
||||
:param messages: List of message dictionaries (like OpenAI format).
|
||||
Example: [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Hello!"}]
|
||||
:param generation_kwargs: Additional arguments for text generation (e.g., max_tokens, temperature).
|
||||
These will be filtered for compatibility with llama-cpp-python.
|
||||
:return: The generated text string.
|
||||
"""
|
||||
try:
|
||||
@@ -162,49 +194,23 @@ class LocalGGUFLLMServiceConnector:
|
||||
self._load_model() # Load model on first invocation
|
||||
|
||||
if not self.model:
|
||||
raise Exception("[LocalGGUFLLMConnector] Local GGUF LLM model failed to load or is not initialized.")
|
||||
raise Exception("[LocalGGUFLLMConnector] Local GGUF LLM model failed to load.")
|
||||
|
||||
# --- Prepare generation parameters for llama-cpp-python ---
|
||||
# Filter the provided kwargs to only include ones accepted by create_chat_completion
|
||||
# Commonly accepted kwargs for llama-cpp-python (check its documentation for the latest)
|
||||
accepted_kwargs = [
|
||||
'temperature', 'top_p', 'top_k', 'max_tokens', 'presence_penalty',
|
||||
'frequency_penalty', 'repeat_penalty', 'seed', 'stop', 'stream',
|
||||
'mirostat_mode', 'mirostat_tau', 'mirostat_eta'
|
||||
# Add others as needed/allowed by llama-cpp-python's API
|
||||
]
|
||||
filtered_kwargs = {k: v for k, v in generation_kwargs.items() if k in accepted_kwargs}
|
||||
# Filter kwargs for llama-cpp-python chat completion
|
||||
chat_kwargs = {
|
||||
k: v for k, v in generation_kwargs.items()
|
||||
if k in ['temperature', 'top_p', 'top_k', 'max_tokens', 'presence_penalty', 'frequency_penalty', 'repeat_penalty', 'seed']
|
||||
}
|
||||
if 'max_new_tokens' in generation_kwargs and 'max_tokens' not in chat_kwargs:
|
||||
chat_kwargs['max_tokens'] = generation_kwargs['max_new_tokens']
|
||||
|
||||
# Handle 'max_new_tokens' if passed (common in Transformers, convert to 'max_tokens' for llama-cpp)
|
||||
# Note: The logic here prioritizes 'max_tokens' if both are somehow passed.
|
||||
if 'max_new_tokens' in generation_kwargs and 'max_tokens' not in filtered_kwargs:
|
||||
filtered_kwargs['max_tokens'] = generation_kwargs['max_new_tokens']
|
||||
# Example of handling 'seed' if passed directly (llama-cpp might use it differently internally)
|
||||
# The `seed` is often handled by setting the global random state before generation
|
||||
# if 'seed' in generation_kwargs and generation_kwargs['seed'] is not None:
|
||||
# # llama_cpp.Llama.sample_seed might be relevant, or rely on global state
|
||||
# pass # llama-cpp-python often handles seed within the generation call if passed
|
||||
|
||||
# --- Call the LLM ---
|
||||
# Use the chat completion interface which is generally preferred and handles templates
|
||||
response = self.model.create_chat_completion(
|
||||
messages=messages,
|
||||
**filtered_kwargs # Pass the filtered and potentially adjusted kwargs
|
||||
)
|
||||
# --- Extract the generated text ---
|
||||
# llama-cpp-python's create_chat_completion returns a dict
|
||||
# response = {'id': '...', 'object': 'chat.completion', 'created': ..., 'model': '...', 'choices': [{'index': 0, 'message': {'role': 'assistant', 'content': '...'}, 'finish_reason': 'stop'}], 'usage': {...}}
|
||||
response = self.model.create_chat_completion(messages=messages, **chat_kwargs)
|
||||
generated_text = response['choices'][0]['message']['content']
|
||||
if generated_text is None:
|
||||
generated_text = "" # Handle potential None if model produces no content
|
||||
|
||||
# Ensure the output is a clean string
|
||||
final_text = generated_text.strip()
|
||||
return final_text # <-- Return ONLY the generated string
|
||||
return generated_text.strip() if generated_text else ""
|
||||
|
||||
except Exception as e:
|
||||
# Catch any error that occurred within the try block and log it
|
||||
error_msg = f"[LocalGGUFLLMConnector] Error in invoke method: {str(e)}"
|
||||
log(error_msg)
|
||||
# Re-raise the exception so the calling node (e.g., LocalKontextPromptGenerator) knows it failed
|
||||
raise e # Or raise Exception(error_msg) from e
|
||||
raise e # Re-raise
|
||||
|
||||
# ... (rest of the file) ...
|
||||
Reference in New Issue
Block a user