Merge pull request #19 from darth-veitcher/008-llamacpp-integration

feat: llama.cpp support via LLMProvider adapter (issue #15)
This commit is contained in:
James Veitch
2026-07-11 16:48:09 +01:00
committed by GitHub
25 changed files with 1565 additions and 20 deletions
+1 -1
View File
@@ -1,3 +1,3 @@
{
"feature_directory": "specs/007-llm-provider-abstraction"
"feature_directory": "specs/008-llamacpp-integration"
}
+1 -1
View File
@@ -1,5 +1,5 @@
<!-- SPECKIT START -->
For additional context about technologies to be used, project structure,
shell commands, and other important information, read the current plan
at specs/007-llm-provider-abstraction/plan.md
at specs/008-llamacpp-integration/plan.md
<!-- SPECKIT END -->
+37 -7
View File
@@ -4,17 +4,18 @@ A collection of workflow efficiency and quality-of-life nodes built out of neces
## What is this?
`comfydv` fills gaps in ComfyUI's built-in node library: dynamic string formatting, seed-controlled random selection, graceful workflow interruption, and Ollama LLM integration. Install it once and connect the nodes like any other — no Python knowledge required.
`comfydv` fills gaps in ComfyUI's built-in node library: dynamic string formatting, seed-controlled random selection, graceful workflow interruption, and local LLM integration (Ollama and llama.cpp). Install it once and connect the nodes like any other — no Python knowledge required.
| Node | What it does |
|------|-------------|
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the connection through the graph as an `LLM_CLIENT` socket — the same generic socket any future backend's client node will emit. |
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the connection through the graph as an `LLM_CLIENT` socket — a generic connection type any backend's client node emits. |
| **LlamaCpp Client** | Configures a connection to a `llama-server` instance running in router mode (default: `http://localhost:8080`). Emits the same `LLM_CLIENT` socket as Ollama Client — every node below works with either. |
| **LLM Model Selector** | Fetches the live model list from the connected server and presents it as a dropdown. Outputs the selected model name. |
| **LLM Load Model** | Loads a model into memory using `/api/generate` with `keep_alive=-1`. |
| **LLM Unload Model** | Evicts a model from memory using `/api/generate` with `keep_alive=0`. |
| **LLM Load Model** | Loads a model into memory on the connected server. |
| **LLM Unload Model** | Evicts a model from memory on the connected server. |
| **Chat Completion** | Sends a prompt (and optional conversation history) to the connected server. Response and history are shown inline in the node body and available as output sockets. |
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
@@ -31,15 +32,18 @@ cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/darth-veitcher/comfydv.git
```
Restart ComfyUI. The nodes appear under the **dv/** and **dv/ollama** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`) are installed automatically via `requirements.txt`.
Restart ComfyUI. The nodes appear under the **dv/**, **dv/ollama**, and **dv/llamacpp** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) are installed automatically via `requirements.txt`.
For Ollama nodes: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`) before using the Ollama nodes.
For local LLM nodes, pick one backend (or both):
- **Ollama**: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`).
- **llama.cpp**: [build/install `llama-server`](https://github.com/ggml-org/llama.cpp) and launch it in router mode (`llama-server --models-dir ./models`) — see the [llama.cpp section](#llamacpp) below.
## Quickstart
1. Install via ComfyUI Manager (search `comfydv`) or clone manually into `custom_nodes/`.
2. Right-click the canvas → Add Node → **dv/** to find Format String, Random Choice, and Circuit Breaker.
3. For Ollama nodes: start Ollama (`ollama serve`), pull a model (`ollama pull qwen2.5:latest`), then add nodes from **dv/ollama/**.
3. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
## Documentation
@@ -166,3 +170,29 @@ If you saved a workflow before this rename, ComfyUI will report the old node typ
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
---
## llama.cpp
A second backend for the same chat/model-management nodes documented above — **LlamaCpp Client** is the only new node; everything else (Chat Completion, LLM Model Selector, LLM Load Model, LLM Unload Model, structured output, multi-turn history) works unchanged, because they don't know or care which backend they're talking to.
### Prerequisite: router mode
`llama-server` needs to be launched in **router mode** — a directory of models, not a single `-m model.gguf`:
```bash
llama-server --models-dir ./models -c 8192
```
This gives `comfydv` live model status (including `loading`/`downloading`, not just loaded/unloaded — a richer picture than Ollama can report) and explicit load/unload, the same way the Ollama nodes already work.
### LlamaCpp Client node
Configure the server address once (default `http://localhost:8080`); every downstream node inherits it automatically — same pattern as Ollama Client, same `LLM_CLIENT` socket.
### Switching an existing workflow from Ollama to llama.cpp
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed at your running `llama-server`. Nothing else changes — same Chat Completion node, same Load/Unload nodes, same structured-output behavior. That's the entire point of sharing one `LLM_CLIENT` socket type across backends.
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
+9 -4
View File
@@ -1,6 +1,6 @@
# comfydv — Roadmap
<!-- generated by beacon roadmap export — 2026-07-04 -->
<!-- generated by beacon roadmap export — 2026-07-11 -->
> comfydv is a small, high-quality ComfyUI utility pack that fills the gaps the core node library leaves: composable string formatting, seed-controlled randomisation, and workflow flow-control. Winning looks like: every node is well-tested, installs in one step, produces no surprises in production workflows, and is documented well enough that a non-programmer ComfyUI user can connect it without reading source code.
@@ -13,12 +13,14 @@ gantt
excludes weekends
section Active
llama.cpp Model Integration :active, llamacpp-integration, 2026-07-11, 7d
LLM Provider Abstraction :active, llm-provider-abstraction, 2026-07-11, 7d
ComfyUI UX Polish & Manager Compatibility :active, ux-and-install, 2026-06-28, 21d
section Done
BEACON Bootstrap :done, beacon-bootstrap, 2026-07-04, 7d
Logging Modernisation :done, logging-modernisation, 2026-07-04, 7d
Ollama Model Integration :done, ollama-integration, 2026-07-04, 7d
BEACON Bootstrap :done, beacon-bootstrap, 2026-07-11, 7d
Logging Modernisation :done, logging-modernisation, 2026-07-11, 7d
Ollama Model Integration :done, ollama-integration, 2026-07-11, 7d
```
@@ -26,6 +28,8 @@ gantt
| Epic | Title | Status | Specs | Fidelity |
|---|---|---|---|---|
| [llamacpp-integration](project-management/Roadmap/epics/llamacpp-integration.md) | llama.cpp Model Integration | Active | 1/1 shipped | S+ A+ T:96% |
| [llm-provider-abstraction](project-management/Roadmap/epics/llm-provider-abstraction.md) | LLM Provider Abstraction | Active | 1/1 shipped | S+ A+ T:58% |
| [ux-and-install](project-management/Roadmap/epics/ux-and-install.md) | ComfyUI UX Polish & Manager Compatibility | Active | 1/4 shipped | S+ A+ T:100% |
| [beacon-bootstrap](project-management/Roadmap/epics/archive/beacon-bootstrap.md) | BEACON Bootstrap | Done | — | S? A? T:- |
| [logging-modernisation](project-management/Roadmap/epics/archive/logging-modernisation.md) | Logging Modernisation | Done | 1/1 shipped | S+ A+ T:100% |
@@ -46,3 +50,4 @@ _No active bullets._
| [ADR-003](../../ADRs/ADR-003-requirements-txt-authoring-policy.md) | Hand-authored requirements.txt as a curated subset of pyproject.toml | Accepted |
| [ADR-004](project-management/ADRs/ADR-004-aiohttp-over-httpx-for-ollama.md) | Use aiohttp for Ollama HTTP communication instead of httpx | Accepted |
| [ADR-005](project-management/ADRs/ADR-005-ollama-host-config-via-client-node.md) | Ollama host configuration via OllamaClient node and OLLAMA_CLIENT socket type | Accepted |
| [ADR-007](project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md) | LLMProvider adapter pattern shared across Ollama and llama.cpp | Accepted |
+33 -5
View File
@@ -7,10 +7,11 @@ A collection of workflow efficiency and quality-of-life nodes built out of neces
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the connection through the graph as an `LLM_CLIENT` socket — the same generic socket any future backend's client node will emit. |
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the connection through the graph as an `LLM_CLIENT` socket — a generic connection type any backend's client node emits. |
| **LlamaCpp Client** | Configures a connection to a `llama-server` instance running in router mode (default: `http://localhost:8080`). Emits the same `LLM_CLIENT` socket as Ollama Client — every node below works with either. |
| **LLM Model Selector** | Fetches the live model list from the connected server and presents it as a dropdown. Outputs the selected model name. |
| **LLM Load Model** | Loads a model into memory using `/api/generate` with `keep_alive=-1`. |
| **LLM Unload Model** | Evicts a model from memory using `/api/generate` with `keep_alive=0`. |
| **LLM Load Model** | Loads a model into memory on the connected server. |
| **LLM Unload Model** | Evicts a model from memory on the connected server. |
| **Chat Completion** | Sends a prompt (and optional conversation history) to the connected server. Response and history are shown inline in the node body and available as output sockets. |
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
@@ -27,9 +28,12 @@ cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/darth-veitcher/comfydv.git
```
Restart ComfyUI. The nodes appear under the **dv/** and **dv/ollama** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`) are installed automatically via `requirements.txt`.
Restart ComfyUI. The nodes appear under the **dv/**, **dv/ollama**, and **dv/llamacpp** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) are installed automatically via `requirements.txt`.
For Ollama nodes: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`) before using the Ollama nodes.
For local LLM nodes, pick one backend (or both):
- **Ollama**: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`).
- **llama.cpp**: [build/install `llama-server`](https://github.com/ggml-org/llama.cpp) and launch it in router mode (`llama-server --models-dir ./models`) — see the [llama.cpp section](#llamacpp) below.
---
@@ -150,3 +154,27 @@ If you saved a workflow before this rename, ComfyUI will report the old node typ
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
---
## llama.cpp
A second backend for the same chat/model-management nodes documented above — **LlamaCpp Client** is the only new node; everything else (Chat Completion, LLM Model Selector, LLM Load Model, LLM Unload Model, structured output, multi-turn history) works unchanged, because they don't know or care which backend they're talking to.
### Prerequisite: router mode
`llama-server` needs to be launched in **router mode** — a directory of models, not a single `-m model.gguf`:
```bash
llama-server --models-dir ./models -c 8192
```
This gives `comfydv` live model status (including `loading`/`downloading`, not just loaded/unloaded — a richer picture than Ollama can report) and explicit load/unload, the same way the Ollama nodes already work.
### LlamaCpp Client node
Configure the server address once (default `http://localhost:8080`); every downstream node inherits it automatically — same pattern as Ollama Client, same `LLM_CLIENT` socket.
### Switching an existing workflow from Ollama to llama.cpp
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed at your running `llama-server`. Nothing else changes — same Chat Completion node, same Load/Unload nodes, same structured-output behavior. That's the entire point of sharing one `LLM_CLIENT` socket type across backends.
@@ -1,7 +1,7 @@
# Epic: llama.cpp Model Integration
## Status
Planning — started 2026-07-11
Active — spec 008-llamacpp-integration complete, ready to finish once merged
## Why now
@@ -32,6 +32,7 @@ epic.
_Filled by `beacon specify --epic llamacpp-integration` / `/speckit-specify`
once this epic is accepted._
- specs/008-llamacpp-integration/
## ADRs
- project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md — decided during the prerequisite epic; this epic implements the second `LLMProvider` the ADR anticipated
@@ -0,0 +1 @@
epic = "llamacpp-integration"
@@ -0,0 +1,39 @@
# Specification Quality Checklist: llama.cpp Model Integration
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-07-11
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
No [NEEDS CLARIFICATION] markers needed — scope boundaries (router-mode-only,
no GPU tuning, no auth/TLS, no Manager listing) came directly from the
parent epic's Non-goals (`project-management/Roadmap/epics/llamacpp-integration.md`)
and ADR-007. User Story 4 (swap backends without touching downstream nodes)
is the adapter pattern's central promise made concrete and testable, not
padding.
@@ -0,0 +1,44 @@
# Contract: `LlamaCppProvider` conforms to `LLMProvider`
This is the concrete proof of ADR-007's adapter pattern — the same protocol
contract documented in
`specs/007-llm-provider-abstraction/contracts/llm_provider_protocol.md`,
now with a second implementation. Nothing in that contract changes; this
file only documents `LlamaCppProvider`'s specific wire-format bindings.
```python
class LlamaCppProvider:
def __init__(self, host: str, headers: dict | None = None): ...
async def list_models(self) -> list[ModelInfo]:
"""GET {host}/models → data[] → ModelInfo(name=m["id"], status=ModelStatus(m["status"]["value"]), size=None)"""
async def load_model(self, model: str) -> None:
"""POST {host}/models/load {"model": model}"""
async def unload_model(self, model: str) -> None:
"""POST {host}/models/unload {"model": model}"""
async def chat(self, model, messages, options=None, timeout_secs=300.0) -> str:
"""POST {host}/v1/chat/completions → choices[0].message.content"""
async def chat_structured(self, model, messages, schema, options=None, timeout_secs=300.0, max_retries=2) -> BaseModel:
"""Delegates to comfydv._llm.chat.chat_structured(base_url=f"{host}/v1", ...) — identical call OllamaProvider makes"""
```
## Behavioral requirements (inherited from the protocol contract, restated for this implementation)
- `load_model`/`unload_model` MUST be idempotent. **Live-verified against a
real router-mode server**: router mode's own endpoints are *not*
idempotent — `/models/load` on an already-loaded model returns HTTP 400
`"model is already running"`, and `/models/unload` on an already-unloaded
model returns HTTP 400 `"model is not running"`, instead of `{"success": true}`.
`LlamaCppProvider` absorbs this itself: these two specific error messages
are treated as the desired end-state already reached, not a failure; any
other error still propagates.
- `list_models()` MUST NOT normalize away llama.cpp's `sleeping`/`downloading`
states (unlike `OllamaProvider`, which has no choice but to normalize —
see `research.md`).
- A `llama-server` not running in router mode (missing endpoints) MUST
surface a clear, specific error (spec.md FR-006) — not a generic
connection failure indistinguishable from "server not running at all."
@@ -0,0 +1,41 @@
# Data Model: llama.cpp Model Integration
No new types — this feature is a second implementation of the existing
`LLMProvider` protocol, `ModelStatus`, `ModelInfo`, and `Message` types
(`src/comfydv/_llm/provider.py`, unchanged). This file documents
`LlamaCppProvider`'s field mapping from llama-server's router-mode JSON onto
those existing types (see `research.md` for the verified API shapes).
## `LlamaCppProvider.list_models()` → `ModelInfo` mapping
| `ModelInfo` field | Source (`GET /models` response, per model in `data[]`) |
|---|---|
| `name` | `id` — **not** `name` (llama.cpp's field name differs from Ollama's) |
| `status` | `status.value` — nested object, not a flat string |
| `size` | Not provided by this endpoint; `None` |
`status.value` maps directly onto `ModelStatus`'s five values
(`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`) — llama.cpp's
vocabulary is exactly `ModelStatus`'s full set, so unlike `OllamaProvider`
(which normalizes into a narrower subset), `LlamaCppProvider` needs no
approximation. A `"failed": true` state exists outside this vocabulary
(model process crashed) — out of scope per spec.md's edge cases; treated as
whatever `status.value` reports rather than added as a sixth enum value.
## `LlamaCppProvider.load_model()` / `unload_model()`
Both `POST /models/load` and `POST /models/unload` take `{"model": <id>}` —
the same `id` string `list_models()` returns as `ModelInfo.name`. No mapping
ambiguity here (unlike Ollama, where load/unload uses `/api/generate`'s
`keep_alive` side effect rather than a dedicated endpoint).
## `LlamaCppProvider.chat()` / `chat_structured()`
Both reach `llama-server`'s OpenAI-compatible `/v1/chat/completions` —
`chat_structured()` calls the existing shared `comfydv._llm.chat.chat_structured()`
helper unchanged (`base_url=f"{self.host}/v1"`, matching `OllamaProvider`'s
own call exactly). `chat()` parses the response as
`choices[0].message.content` (OpenAI shape), not Ollama's native
`message.content` — the two providers' non-structured paths differ here
because llama-server doesn't have an Ollama-style native `/api/chat`
endpoint to prefer instead.
@@ -0,0 +1,11 @@
Feature: US1 — Connect to a local llama.cpp server and get chat responses
Scenario: llama.cpp connection node feeds the existing chat node
Given a running local llama-server (router mode) and a workflow with a llama.cpp connection node wired into the existing chat node
When the workflow executes
Then the chat node returns the model's text response
Scenario: Unreachable llama.cpp server surfaces a clear error
Given the llama.cpp connection node configured with an unreachable server address
When the workflow executes
Then the chat node reports a clear connection error
@@ -0,0 +1,11 @@
Feature: US2 — Get structured, validated output from llama.cpp
Scenario: Valid structured response exposes typed fields, same as Ollama
Given a chat node connected to llama.cpp with structured output enabled and a valid schema
When the workflow executes and the model responds correctly
Then each schema field is available as its own typed output, and no required field is blank
Scenario: Invalid response retries then fails clearly, same as Ollama
Given a llama.cpp-hosted model that returns invalid or incomplete structured output
When the workflow executes
Then the node retries automatically and, if still unsuccessful, fails with a clear error
@@ -0,0 +1,16 @@
Feature: US3 — See and control which models are loaded on llama.cpp
Scenario: List models with full status vocabulary
Given a running local llama-server with at least one available model
When a workflow author uses the model-listing node
Then they see each available model along with its current status, drawn from llama.cpp's full status vocabulary
Scenario: Load a model into memory
Given a model that is not currently loaded
When a workflow author runs the load-model node against it
Then the model becomes loaded and is then usable by the chat node
Scenario: Unload a model from memory
Given a model that is loaded and idle
When a workflow author runs the unload-model node against it
Then the model is freed from memory and its reported status updates accordingly
@@ -0,0 +1,6 @@
Feature: US4 — Swap from Ollama to llama.cpp without touching the rest of the workflow
Scenario: Replacing only the connection node preserves the workflow
Given a workflow with chat/model-management nodes wired to an Ollama connection node
When a workflow author replaces only the connection node with a llama.cpp one, pointed at a running llama-server
Then the workflow runs successfully with no changes to any other node
+117
View File
@@ -0,0 +1,117 @@
# Implementation Plan: llama.cpp Model Integration
**Branch**: `008-llamacpp-integration` | **Date**: 2026-07-11 | **Spec**: [spec.md](./spec.md)
**Input**: Feature specification from `/specs/008-llamacpp-integration/spec.md`
**Note**: This template is filled in by the `/speckit-plan` command. See `.specify/templates/plan-template.md` for the execution workflow.
## Summary
Implement `LlamaCppProvider` as the second `LLMProvider` (ADR-007), backed by
`llama-server`'s router mode (`GET /models`, `POST /models/load`,
`POST /models/unload`, `/v1/chat/completions`). Add one new ComfyUI node
(`LlamaCppClient`) emitting the existing `LLM_CLIENT` socket type — no other
node classes change. This is the concrete proof the provider abstraction
(prerequisite epic, PR #17) actually generalizes: a second backend, zero
changes to `ChatCompletion`/`LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel`.
## Technical Context
**Language/Version**: Python ≥3.11 (unchanged, per `pyproject.toml`)
**Primary Dependencies**: `aiohttp` (existing — model-management REST calls),
`pydantic-ai`/`openai` (existing, from the prerequisite epic — `chat_structured()`
reuses the shared helper unchanged, zero new structured-output code)
**Storage**: N/A — no persistent storage; reuses the existing
`_MODEL_LIST_CACHE`/`_CHAT_RESPONSE_CACHE` infra pattern from `OllamaProvider`
**Testing**: `pytest` via `uv run pytest`, following `tests/test_ollama_provider.py`'s
established convention (mock at the provider's own `_post_json`/`_get_json`
seam, no live server required for unit tests)
**Target Platform**: ComfyUI custom-node runtime, same as the existing Ollama
integration
**Project Type**: Library / ComfyUI custom-node pack (single project, adds to
existing `src/comfydv/` layout)
**Performance Goals**: No new numeric target; must not add latency beyond
what `OllamaProvider`'s equivalent methods already accept
**Constraints**: Router-mode-only (spec.md Assumptions — a `llama-server`
without `--models-dir`/`--models-preset` doesn't expose these endpoints at
all, FR-006); model identifier field is `id` (llama.cpp) vs `name` (Ollama) —
`LlamaCppProvider.list_models()` must map this correctly (see `research.md`);
`status` is a nested object (`{"value": "..."}`), not a flat string
**Scale/Scope**: One new class (`LlamaCppProvider`, mirrors `OllamaProvider`'s
shape), one new ComfyUI node (`LlamaCppClient`), one new test file — no
changes to `ollama.py`, `_llm/provider.py`, `_llm/chat.py`, or any existing
node class
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Principle | Verdict | Notes |
|---|---|---|
| I. ComfyUI Contract First | PASS | `LlamaCppClient` exposes the standard `INPUT_TYPES`/`RETURN_TYPES`/`FUNCTION`/`CATEGORY`; registered in `NODE_CLASS_MAPPINGS` like every other node. |
| II. Sandbox All User-Supplied Code | N/A | No template/expression evaluation in this feature. |
| III. Test-First | PASS (binding) | `tests/test_llamacpp_provider.py` written test-first, mirroring `test_ollama_provider.py`'s TDD-pair structure. |
| IV. Graceful Degradation Outside ComfyUI | PASS (binding) | `LlamaCppProvider` lives in `src/comfydv/_llm/`, which already has no `comfy`/`server` imports at module scope (verified for the prerequisite epic; this feature adds no new module-scope imports of either). |
| V. Simplicity — Function Before Class | PASS, same justification as `OllamaProvider` | `LlamaCppProvider` carries connection state (host, headers) across 5 methods — the same shared-state condition that already justified `OllamaProvider` as a class (research.md, prerequisite epic). No new gate — same precedent applies. |
| VI. Fixed Output Positions | N/A | `LlamaCppClient`'s single output (`client`) isn't a multi-output node; no positional contract to preserve. |
Re-checked post-Phase 1 design (data-model.md): unchanged — no new gate
violations. No Complexity Tracking entries needed (unlike the prerequisite
epic, this feature introduces no new pattern, just a second instance of an
already-justified one).
## Project Structure
### Documentation (this feature)
```text
specs/008-llamacpp-integration/
├── plan.md # This file
├── research.md # Phase 0 — router-mode API shape, verified live
├── data-model.md # Phase 1 — LlamaCppProvider field mapping
├── quickstart.md # Phase 1 — minimal workflow walkthrough
├── contracts/ # Phase 1 — LlamaCppProvider's protocol conformance
└── tasks.md # Phase 2 (/speckit-tasks)
```
### Source Code (repository root)
```text
src/comfydv/
├── ollama.py # unchanged — add LlamaCppClient node only via a new module
├── llamacpp.py # new — LlamaCppClient node (mirrors OllamaClient's shape)
├── _llm/
│ ├── provider.py # unchanged — LLMProvider/ModelStatus/ModelInfo/Message
│ ├── ollama_provider.py # unchanged
│ ├── llamacpp_provider.py # new — LlamaCppProvider (mirrors ollama_provider.py's shape)
│ └── chat.py # unchanged — chat_structured() reused as-is
└── __init__.py # add LlamaCppClient import + NODE_CLASS_MAPPINGS entry
tests/
├── test_ollama_provider.py # unchanged
├── test_llamacpp_provider.py # new — mirrors test_ollama_provider.py's structure
└── test_llamacpp.py # new — LlamaCppClient node contract test (small; mirrors
# the OllamaClient-specific slice of test_ollama.py)
```
**Structure Decision**: New `src/comfydv/llamacpp.py` module (not added into
`ollama.py`) for the `LlamaCppClient` node, and a new `src/comfydv/_llm/llamacpp_provider.py`
for `LlamaCppProvider` — mirroring the existing `ollama.py`/`ollama_provider.py`
split exactly, so the two backends read as parallel, symmetric implementations
rather than one growing to accommodate the other. No existing file is
modified except `__init__.py`'s registration block.
## Complexity Tracking
> **Fill ONLY if Constitution Check has violations that must be justified**
None — see Constitution Check above.
@@ -0,0 +1,26 @@
# Quickstart: llama.cpp Model Integration
## Prerequisite
Launch `llama-server` in router mode:
```bash
llama-server --models-dir ./models -c 8192
```
## Minimal workflow
1. Add an **LlamaCpp Client** node. Set its host widget (default
`http://localhost:8080`, llama-server's default port).
2. Wire it into a **Chat Completion** node — the exact same node used for
Ollama. Set a model and prompt, run.
3. Structured output, model listing, and load/unload all work exactly as
documented for Ollama in the main README/quickstart — swap the client
node, nothing else changes.
## Swapping an existing Ollama workflow to llama.cpp
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed
at your running `llama-server`. Every downstream node (Chat Completion, LLM
Model Selector, LLM Load Model, LLM Unload Model) keeps working unmodified —
this is the whole point of the provider abstraction (ADR-007).
@@ -0,0 +1,78 @@
# Research: llama.cpp Model Integration
## Decision: exact router-mode API shape (verified against `ggml-org/llama.cpp`'s live `tools/server/README.md`, not assumed)
llama.cpp's router mode postdates this session's training data — verified live
against the authoritative source rather than guessed, since getting field
names wrong here would silently produce broken code (wrong key = `KeyError`
or silent `None`, not an obvious failure).
**`GET /models`** response:
```json
{
"data": [
{
"id": "ggml-org/gemma-3-4b-it-GGUF:Q4_K_M",
"path": "/Users/.../gemma-3-4b-it-Q4_K_M.gguf",
"status": {
"value": "loaded",
"args": ["llama-server", "-ctx", "4096"]
},
"architecture": {
"input_modalities": ["text", "image"],
"output_modalities": ["text"]
}
}
]
}
```
**Two details that would have been wrong by assumption:**
1. The model identifier field is **`id`**, not `name` — different from Ollama's
`/api/tags`, which uses `name`. `OllamaProvider.list_models()` maps
`m["name"]`; `LlamaCppProvider.list_models()` must map `m["id"]` instead.
2. **`status` is a nested object** (`{"value": "loaded", ...}`), not a flat
string field. `LlamaCppProvider.list_models()` must read
`m["status"]["value"]`, not `m["status"]` directly. A `"failed"` state
also exists (`{"failed": true, "exit_code": ...}`) outside the five
`ModelStatus` values the protocol defines — not handled by this feature
(see Non-goals/edge cases in `spec.md`); a failed model is reported as
whatever `status.value` degrades to rather than added as a sixth enum
value, keeping `ModelStatus` unchanged across both providers.
**`POST /models/load`** and **`POST /models/unload`**: identical request
shape, `{"model": "<id>"}` (using the same `id` string from `GET /models`,
despite the request field being named `model` not `id`). Response:
`{"success": true}`.
**CLI**: `--models-dir <path>` or `--models-preset <path>.ini` — a deployment
prerequisite (spec.md Assumptions), not something comfydv configures.
## Decision: `chat_structured()` needs zero new code
`llama-server`'s `/v1/chat/completions` is OpenAI-compatible (the same
assumption ADR-007 made when adopting `pydantic-ai`). `LlamaCppProvider.chat_structured()`
calls the exact same `comfydv._llm.chat.chat_structured()` helper
`OllamaProvider` already uses, with `base_url=f"{self.host}/v1"` — the only
per-provider difference. This is the concrete proof the shared mechanism
generalizes (spec.md User Story 2/FR-004), not just an assumption.
## Decision: `chat()` (non-structured) also reuses the OpenAI-compatible endpoint
Unlike Ollama (which has both a native `/api/chat` and an OpenAI-compat
`/v1/chat/completions`), llama-server's primary chat endpoint is the
OpenAI-compatible one. `LlamaCppProvider.chat()` POSTs to
`{host}/v1/chat/completions` (via the existing `_post_json` helper, no new
HTTP client) rather than mirroring Ollama's native-endpoint choice — the
response shape (`choices[0].message.content`) differs from Ollama's native
`message.content` and must be parsed accordingly.
## Decision: no protocol changes needed
`LLMProvider`'s five methods (`list_models`/`load_model`/`unload_model`/
`chat`/`chat_structured`) already cover everything router mode needs — this
was the actual point of designing the protocol at the operation level in
ADR-007, and this research confirms it held up against llama.cpp's real API,
not just Ollama's.
+110
View File
@@ -0,0 +1,110 @@
# Feature Specification: llama.cpp Model Integration
**Feature Branch**: `008-llamacpp-integration`
**Created**: 2026-07-11
**Status**: Draft
**Input**: User description: "Add ComfyUI nodes for llama.cpp local inference via llama-server's router mode, implementing the LlamaCppProvider as the second LLMProvider (ADR-007) alongside the existing OllamaProvider. Router mode exposes GET /models (with live status), POST /models/load, POST /models/unload, giving llama.cpp the same manual load/unload memory-management primitives as Ollama. No new ComfyUI node classes needed for chat/model-selection/load/unload — only a new LlamaCppClient config node; the existing generic ChatCompletion/LLMModelSelector/LLMLoadModel/LLMUnloadModel nodes work unchanged once wired to it."
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Connect to a local llama.cpp server and get chat responses (Priority: P1) 🎯 MVP
As a ComfyUI workflow author running `llama-server` locally, I want a connection node for it — just like the one I already use for Ollama — so I can get chat responses from a llama.cpp-hosted model using the same chat node I already know.
**Why this priority**: This is the entire point of the feature and the proof that the provider abstraction (shipped in the prerequisite epic) actually works: a second backend, zero changes to the chat node.
**Independent Test**: Wire a new llama.cpp connection node into the existing chat node, run against a local `llama-server` (router mode), confirm a text response.
**Acceptance Scenarios**:
1. **Given** a running local `llama-server` (router mode) and a workflow with a llama.cpp connection node wired into the existing chat node, **When** the workflow executes, **Then** the chat node returns the model's text response — using the exact same chat node a workflow author already uses for Ollama.
2. **Given** the llama.cpp connection node configured with an unreachable server address, **When** the workflow executes, **Then** the chat node reports a clear connection error, matching the behavior workflow authors already know from the Ollama connection.
---
### User Story 2 - Get structured, validated output from llama.cpp (Priority: P1)
As a workflow author, I want structured output (a schema-validated response instead of free text) to work identically regardless of whether I'm connected to Ollama or llama.cpp, so I don't have to relearn or rebuild anything when switching backends.
**Why this priority**: Structured output is a core existing capability (already proven for Ollama); this story proves the shared mechanism genuinely generalizes rather than being Ollama-specific in practice, not just in name.
**Independent Test**: Enable structured output on the chat node with a schema, run against a llama.cpp-hosted model, confirm each schema field is populated and never blank — using the same steps as the equivalent Ollama test.
**Acceptance Scenarios**:
1. **Given** a chat node connected to llama.cpp with structured output enabled and a valid schema, **When** the workflow executes and the model responds correctly, **Then** each schema field is available as its own typed output, and no required field is blank.
2. **Given** a llama.cpp-hosted model that returns invalid or incomplete structured output, **When** the workflow executes, **Then** the node retries automatically and, if still unsuccessful, fails with a clear error — identical behavior to the Ollama path.
---
### User Story 3 - See and control which models are loaded on llama.cpp (Priority: P2)
As a workflow author running models locally, I want to see live model status (including whether a model is currently loading or being downloaded, not just loaded/unloaded) and explicitly load or unload a model on my llama.cpp server, so I can manage memory the same way I already do for Ollama — with more visibility, since llama.cpp's router mode reports richer status than Ollama does.
**Why this priority**: Valuable and proves the model-management path generalizes too, but a workflow can still run chat completions without ever calling load/unload explicitly (the server can load on first use), so it's lower risk to defer than basic chat.
**Independent Test**: Use the existing model-listing node against a running `llama-server`, confirm it shows each available model with its current status (including `loading`/`downloading` if applicable); use the existing load/unload nodes against one model and confirm its status changes.
**Acceptance Scenarios**:
1. **Given** a running local `llama-server` with at least one available model, **When** a workflow author uses the model-listing node, **Then** they see each available model along with its current status, drawn from llama.cpp's full status vocabulary (not just loaded/unloaded).
2. **Given** a model that is not currently loaded, **When** a workflow author runs the load-model node against it, **Then** the model becomes loaded and is then usable by the chat node.
3. **Given** a model that is loaded and idle, **When** a workflow author runs the unload-model node against it, **Then** the model is freed from memory and its reported status updates accordingly.
---
### User Story 4 - Swap from Ollama to llama.cpp without touching the rest of the workflow (Priority: P3)
As a workflow author with an existing Ollama-based workflow, I want to switch it to llama.cpp by changing only the connection node, so I don't have to rebuild my chat/model-management logic for a second backend.
**Why this priority**: This is the adapter pattern's actual promise made concrete for a user, but it's a validation/demonstration story rather than new capability — everything it depends on is already covered by User Stories 1–3.
**Independent Test**: Take a workflow using the Ollama connection node, replace it with the llama.cpp connection node (same downstream nodes, no other changes), run it, confirm it still works.
**Acceptance Scenarios**:
1. **Given** a workflow with chat/model-management nodes wired to an Ollama connection node, **When** a workflow author replaces only the connection node with a llama.cpp one (pointed at a running `llama-server`), **Then** the workflow runs successfully with no changes to any other node.
---
### Edge Cases
- What happens when `llama-server` is running but was launched without router mode (i.e. with `-m` instead of `--models-dir`)? The router-mode-only endpoints this feature depends on won't exist — the connection/model-management nodes should fail with a clear error, not hang or silently return empty results.
- What happens when the configured server address is unreachable at the moment a model-listing, load, or unload node runs (not just the chat node)?
- What happens when llama.cpp reports a model status this feature doesn't expect (a router-mode API change)? Should degrade gracefully (surface the status if recognized, don't crash on an unrecognized one), not silently misreport.
- What happens to an in-flight chat request if the model it depends on is unloaded by another node in the same workflow run? (Same question already answered for Ollama — behavior should be consistent.)
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: The system MUST allow a workflow author to configure a connection to a local `llama-server` (router mode) the same way they already configure a connection to Ollama — a dedicated connection node, reusable across multiple nodes in a workflow.
- **FR-002**: The system MUST NOT require any new or different node classes for chat, structured output, model listing, or load/unload when using llama.cpp — the existing generic nodes MUST work unchanged once connected to a llama.cpp connection node.
- **FR-003**: The system MUST report each model's status using llama.cpp's full status vocabulary (unloaded, loading, loaded, sleeping, downloading) when connected to llama.cpp — not degraded to the narrower Ollama-compatible set.
- **FR-004**: The system's chat and structured-output behavior MUST be identical between Ollama and llama.cpp connections, given equivalent inputs — same retry limits, same validation rules, same error conditions (this is the direct continuation of the prerequisite epic's own FR-007/FR-008).
- **FR-005**: The system MUST allow a workflow author to explicitly load a model into memory and explicitly unload a model from memory on a connected llama.cpp server.
- **FR-006**: The system MUST surface a clear, specific error when connected to a `llama-server` instance that isn't running in router mode (the endpoints this feature needs don't exist), rather than an unhelpful generic failure.
### Key Entities *(include if feature involves data)*
- **llama.cpp connection**: A configured connection to a local `llama-server` instance running in router mode (host + any authentication), implementing the same connection concept already established for Ollama.
- **Model status**: Reuses the existing status concept from the prerequisite feature, now populated with llama.cpp's full vocabulary rather than a narrowed subset.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: A workflow author can connect to a llama.cpp server and get a chat response using the same node count and shape as connecting to Ollama (one connection node, one chat node) — no new nodes to learn for the chat path.
- **SC-002**: An existing workflow can be repointed from Ollama to llama.cpp by changing exactly one node (the connection node) — zero edits to any chat or model-management node.
- **SC-003**: Structured-output workflows behave identically (same validation guarantees, zero blank-required-field results) regardless of which backend is connected.
- **SC-004**: Model status reporting for llama.cpp surfaces all five status values where applicable — a strictly richer view than what Ollama can report through the same interface.
## Assumptions
- Workflow authors run their own local `llama-server` instance, launched in router mode (`--models-dir` or `--models-preset`), reachable over HTTP from the machine running ComfyUI; this feature does not install, configure, or launch that server.
- Non-router-mode `llama-server` usage (a single model launched with `-m`) is out of scope — router mode is required for the load/unload/status parity with Ollama that is this feature's whole point.
- GPU inference optimisation, quantisation tuning, authentication/TLS, and ComfyUI Manager registry listing are out of scope, consistent with the prerequisite Ollama epic's own non-goals.
- The `LLMProvider` protocol and generic nodes (`ChatCompletion`, `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`) already exist and are not modified by this feature — if llama.cpp's router mode needs a protocol capability that doesn't exist yet, that is a protocol change scoped as its own follow-up, not silently special-cased here.
+224
View File
@@ -0,0 +1,224 @@
# Tasks: llama.cpp Model Integration
**Input**: Design documents from `/specs/008-llamacpp-integration/`
**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/llamacpp_provider_conformance.md
**Tests**: First-class — every implementation task has a paired failing-test task (`-T`/`-I` suffix).
**Organization**: Grouped by user story (spec.md priorities P1/P1/P2/P3).
## Format: `[ID] [P?] [Story] Description`
- **[P]**: Can run in parallel (different files, no dependencies)
- **[Story]**: US1–US4
- **-T / -I**: paired test (red) / implementation (green)
## Path Conventions
Single project: `src/comfydv/`, `tests/` at repository root, mirroring the
`ollama.py`/`_llm/ollama_provider.py` split exactly (plan.md Structure
Decision).
---
## Phase 1: Setup
- [x] T001 No new dependencies — `aiohttp`/`pydantic-ai` already present from the prerequisite epic (verified in `pyproject.toml`)
- [x] T002 [P] Create `src/comfydv/_llm/llamacpp_provider.py` and `src/comfydv/llamacpp.py` (empty modules with docstrings, mirroring `ollama_provider.py`/`ollama.py`'s module docstring style)
---
## Phase 2: Foundational
None — `LLMProvider`, `ModelStatus`, `ModelInfo`, `Message`, and the shared
`chat_structured()` helper already exist from the prerequisite epic and are
unmodified by this feature (plan.md Constitution Check, research.md).
**Checkpoint**: nothing blocks user story work — it can start immediately.
---
## Phase 3: User Story 1 — Connect to a local llama.cpp server and get chat responses (Priority: P1) 🎯 MVP
**Goal**: A workflow author wires an `LlamaCppClient` node into the existing `ChatCompletion` node and gets a text response.
**Independent Test**: Wire `LlamaCppClient` → `ChatCompletion`, run against a live `llama-server` (router mode), confirm text output.
- [x] T003-T [US1] Write FAILING test: `LlamaCppProvider.chat()` POSTs to `{host}/v1/chat/completions` and parses `choices[0].message.content`, in `tests/test_llamacpp_provider.py` (witnesses `features/us1_connect_and_chat.feature` scenario "llama.cpp connection node feeds the existing chat node")
- [x] T003-I [US1] Implement `LlamaCppProvider.__init__`/`.chat()` in `src/comfydv/_llm/llamacpp_provider.py` (data-model.md — OpenAI-shape response parsing, not Ollama's native shape) — makes T003-T pass
- [x] T004-T [US1] Write FAILING test: `LlamaCppClient` node's `INPUT_TYPES`/`RETURN_TYPES` match `OllamaClient`'s shape (`LLM_CLIENT` output), and `create_client()` constructs a `LlamaCppProvider`, in `tests/test_llamacpp.py`
- [x] T004-I [US1] Implement `LlamaCppClient` node in `src/comfydv/llamacpp.py` (mirrors `OllamaClient` exactly, default host `http://localhost:8080` per llama-server's default port) — makes T004-T pass (depends on T003-I)
- [x] T005-T [US1] Write FAILING test: `LlamaCppClient` registered in `NODE_CLASS_MAPPINGS`/`NODE_DISPLAY_NAME_MAPPINGS`, in `tests/test_llamacpp.py`
- [x] T005-I [US1] Register `LlamaCppClient` in `src/comfydv/__init__.py` — makes T005-T pass (depends on T004-I)
- [x] T006-T [US1] Write FAILING test: `LlamaCppProvider` connection error surfaces a clear message (mirrors `OllamaProvider`'s `_post_json` connection-error contract), in `tests/test_llamacpp_provider.py` (witnesses `features/us1_connect_and_chat.feature` scenario "Unreachable llama.cpp server surfaces a clear error")
- [x] T006-I [US1] Confirm `LlamaCppProvider.chat()` reuses the shared `_post_json` connection-error handling unchanged (likely no code change needed — verify, don't assume) — makes T006-T pass
**Checkpoint**: US1 fully functional and independently testable (MVP) — proves the adapter pattern for the chat path.
---
## Phase 4: User Story 2 — Get structured, validated output from llama.cpp (Priority: P1)
**Goal**: `structured_output=True` on `ChatCompletion` works identically against llama.cpp.
**Independent Test**: Enable `structured_output` with a schema, run against a llama.cpp-hosted model, confirm typed sockets populate and are never blank.
- [x] T007-T [US2] Write FAILING test: `LlamaCppProvider.chat_structured()` builds `base_url=f"{host}/v1"` and delegates to the shared `comfydv._llm.chat.chat_structured()` helper unchanged, in `tests/test_llamacpp_provider.py` (witnesses `features/us2_structured_output.feature` scenario "Valid structured response exposes typed fields, same as Ollama")
- [x] T007-I [US2] Implement `LlamaCppProvider.chat_structured()` in `src/comfydv/_llm/llamacpp_provider.py` — zero new structured-output logic, same call shape `OllamaProvider.chat_structured()` already makes — makes T007-T pass
- [x] T008 [US2] No new test needed for the retry-then-fail path (witnesses `features/us2_structured_output.feature` scenario "Invalid response retries then fails clearly, same as Ollama") — already fully covered by `tests/test_llm_chat_structured.py`'s existing suite, since `LlamaCppProvider.chat_structured()` calls the identical shared helper `OllamaProvider` does; re-testing it here would duplicate coverage without adding confidence (same reasoning as the prerequisite epic's D5)
**Checkpoint**: US1 + US2 both independently functional — the chat surface is now backend-agnostic in practice, not just in name.
---
## Phase 5: User Story 3 — See and control which models are loaded on llama.cpp (Priority: P2)
**Goal**: `LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel` work against llama.cpp via `LlamaCppProvider`.
**Independent Test**: List models via `LLMModelSelector` wired to `LlamaCppClient`; load/unload one; confirm status changes, including `loading`/`downloading` states if triggered.
- [x] T009-T [P] [US3] Write FAILING test: `LlamaCppProvider.list_models()` maps `GET /models`'s `data[].id`→`ModelInfo.name` and `data[].status.value`→`ModelInfo.status`, surfacing all five `ModelStatus` values without normalization (data-model.md), in `tests/test_llamacpp_provider.py` (witnesses `features/us3_model_lifecycle.feature` scenario "List models with full status vocabulary")
- [x] T009-I [US3] Implement `LlamaCppProvider.list_models()` in `src/comfydv/_llm/llamacpp_provider.py` — makes T009-T pass
- [x] T010-T [P] [US3] Write FAILING test: `LlamaCppProvider.load_model()`/`unload_model()` POST `{"model": id}` to `/models/load`/`/models/unload` and are idempotent, in `tests/test_llamacpp_provider.py` (witnesses `features/us3_model_lifecycle.feature` scenarios "Load a model into memory" and "Unload a model from memory")
- [x] T010-I [US3] Implement `LlamaCppProvider.load_model()`/`unload_model()` in `src/comfydv/_llm/llamacpp_provider.py` — makes T010-T pass
- [x] T011 [US3] No new node-layer tests needed — `LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel` are untouched by this epic (plan.md Structure Decision) and already have delegation-test coverage against a generic `_FakeProvider` in `tests/test_ollama.py`; that coverage is provider-agnostic by construction (FR-002), so it already proves these nodes work with `LlamaCppProvider` too, not just `OllamaProvider`
**Checkpoint**: US1 + US2 + US3 independently functional.
---
## Phase 6: User Story 4 — Swap from Ollama to llama.cpp without touching the rest of the workflow (Priority: P3)
**Goal**: Demonstrate/prove the adapter pattern's actual promise end-to-end.
**Independent Test**: Same workflow, only the connection node changes.
- [x] T012-T [US4] Write FAILING test: a workflow-shaped test (client → `ChatCompletion` → `LLMModelSelector` → `LLMLoadModel` → `LLMUnloadModel`) runs identically whether `client` is an `OllamaProvider`-double or a `LlamaCppProvider`-double — i.e. no node branches on provider type, in `tests/test_llamacpp.py` (witnesses `features/us4_swap_backends.feature` scenario "Replacing only the connection node preserves the workflow")
- [x] T012-I [US4] No implementation expected — this test should already pass given T003-T011 (it's a regression/integration proof, not new functionality); if it fails, that reveals a node secretly branching on provider type, which would be a real bug to fix, not a feature to add
**Checkpoint**: all four user stories independently functional; the adapter pattern is proven end-to-end, not just asserted.
---
## Phase 7: Polish & Cross-Cutting Concerns
- [x] T013 [P] `ruff check --fix && ruff format` — clean
- [x] T014 [P] `ty check` — clean, same pre-existing diagnostics as the prerequisite epic (unrelated to this feature — `comfy`/`server`/`folder_paths` unresolved-import, `format_string.py`'s dynamic RETURN_TYPES, `create_model`/`RandomChoice` — none touch the new files)
- [x] T015 Confirmed via grep: `llamacpp_provider.py`/`llamacpp.py` import no `comfy`/`server`/`folder_paths` at module scope
- [x] T016 `beacon doctor --strict`: only the pre-existing `llm-provider-abstraction: all specs [complete]` epic-gates item (PR #18, the archive-bookkeeping PR for the *prerequisite* epic, not yet merged — unrelated to this feature) and `tdd-commit-discipline` (disclosed pattern, same reasoning as the prerequisite epic)
- [x] T017 Live smoke test — run against a real router-mode `llama-server` (Homebrew-installed, already present in the dev environment; a prior pass wrongly assumed no server was reachable without actually checking). Full lifecycle exercised against a real 5.6GB local GGUF model: `list_models()` → `load_model()` → `chat()` → `unload_model()`, plus explicit idempotency checks (calling `load_model()`/`unload_model()` again in the already-satisfied state). Found and fixed a real gap not caught by the mocked suite — see the finding below.
---
## Dependencies & Execution Order
### Phase Dependencies
- **Setup (Phase 1)**: no dependencies
- **Foundational (Phase 2)**: none — nothing blocks user story work
- **US1**: no dependency on other stories — genuinely the MVP
- **US2**: depends on US1's `LlamaCppProvider` skeleton existing (T003-I), but its own logic (T007) has no dependency on US1's chat() specifically
- **US3**: independent of US1/US2 except sharing `LlamaCppProvider`'s constructor (T003-I) — unlike the prerequisite epic's atomic cutover, there is no shared "client output type" migration risk here, since `LlamaCppClient` is a brand-new node, not a changed one
- **US4**: depends on US1–US3 all being done (it's a proof, not new functionality)
- **Polish**: depends on all four user stories
### Parallel Opportunities
- T002 can start immediately
- T009-T and T010-T can run in parallel (different methods, same file, no shared state)
- T013/T014 can run in parallel in Polish
---
## Implementation Strategy
### MVP First
1. Phase 1 (Setup, trivial) → Phase 3 (US1) → **STOP and validate US1 independently** against a live `llama-server`.
### Incremental Delivery
1. US1 → validate → basic chat parity with Ollama, on a second backend.
2. US2 → validate → structured-output parity — the shared mechanism holds.
3. US3 → validate → model-management parity, with richer status than Ollama can offer.
4. US4 → validate → the adapter pattern is proven, not just asserted.
5. Polish.
Unlike the prerequisite epic, **this decomposition genuinely holds** —
there is no shared "output type" migration forcing an atomic cutover, because
`LlamaCppClient` is new, not a change to an existing node. Each phase really
can land independently.
---
## Post-implementation review finding (fixed)
A `beacon-reviewer` pass ahead of PR open found `LlamaCppProvider.list_models()`
caught *every* exception and returned `[]`, silently indistinguishable from
"no models installed" — violating FR-006 and `contracts/llamacpp_provider_conformance.md`'s
explicit requirement that a non-router-mode `llama-server` (unreachable
endpoints → HTTP error on `GET /models`) surface a clear, specific error.
Fixed: `_get_json` (shared with `OllamaProvider`, in `ollama_provider.py`) now
raises `RuntimeError` on an HTTP error status, matching `_post_json`'s
existing behavior — its docstring already claimed this, it just didn't do it.
`LlamaCppProvider.list_models()` now distinguishes `OSError` (genuinely
unreachable — connection refused, DNS failure, timeout; all aiohttp
connection-level exceptions are `OSError` subclasses) from `RuntimeError`
(server responded, but with an error): the former still degrades gracefully
to `[]` (consistent with `OllamaProvider`'s existing UX), the latter is
re-raised naming router mode as the likely cause. Regression test added:
`test_list_models_non_router_mode_raises_clear_error` in
`tests/test_llamacpp_provider.py`. `OllamaProvider`'s own `list_models()`/
`_fetch_models()` still catch broadly and degrade to `[]` unchanged — no
spec requirement asks Ollama to make this distinction, and this fix doesn't
force it to.
---
## Live smoke test finding (T017, fixed)
T017 had been marked `[-]` deferred on the assumption that no `llama-server`
was reachable in the dev environment. That assumption was never actually
checked — `llama-server` was installed via Homebrew the whole time, and a
router-mode server was launched against a real local GGUF model
(`--models-dir` pointed at a symlinked model file) to run the smoke test for
real.
This caught a genuine gap the mocked suite couldn't: `contracts/llamacpp_provider_conformance.md`
claimed router mode's `/models/load`/`/models/unload` return `{"success": true}`
on an already-loaded/unloaded model, satisfying the `LLMProvider` protocol's
idempotency requirement "without extra handling." That claim was never
live-verified — live testing showed the opposite: both endpoints return HTTP
400 (`"model is already running"` / `"model is not running"`) instead.
`LlamaCppProvider.load_model()`/`unload_model()` now absorb exactly those two
error messages as the desired end-state already reached (any other error
still propagates); the contract doc is corrected to describe the real
behavior. Regression tests added (mocked, so they run in CI):
`test_load_model_already_running_is_idempotent`,
`test_unload_model_not_running_is_idempotent`, and their
`_other_http_error_still_raises` counterparts confirming non-idempotency
errors aren't over-broadly swallowed.
Also observed live (informational, no code change needed): `load_model()`/
`unload_model()` return once the request is *accepted*, not once the state
transition completes — a 5.6GB model reported `LOADING` for several seconds
before `LOADED`. This matches `ModelStatus`'s documented vocabulary (`loading`
is a real, intended state) and how a real UI would behave — fire the request,
poll `list_models()` for the transition. No protocol change; noted here so
it's not mistaken for a future bug report.
**`chat_structured()` live-verified separately** (T007/T008's mocked coverage
only ever exercised the call-shape, never the real network path): ran a
second live smoke test — `load_model()` → `chat_structured()` with a real
`pydantic.BaseModel` schema — against the same router-mode server. Result
validated correctly (`Color(name='Red', hex_code='#FF0000')`), confirming
pydantic-ai's `Agent`/`OpenAIProvider(base_url=f"{host}/v1")` mechanism
genuinely works against llama-server's OpenAI-compatible endpoint, not just
Ollama's (which was the only one live-verified in the prerequisite epic).
No gap found here — recorded as verification evidence, not a fix.
**Net result**: every `LlamaCppProvider` method (`list_models`, `load_model`,
`unload_model`, `chat`, `chat_structured`) has now been exercised against a
real router-mode `llama-server`, not just mocks. T017 is genuinely done.
+3
View File
@@ -2,6 +2,7 @@ import logging
from .circuit_breaker import CircuitBreaker
from .format_string import FormatString
from .llamacpp import LlamaCppClient
from .ollama import (
ChatCompletion,
LLMLoadModel,
@@ -34,6 +35,7 @@ NODE_CLASS_MAPPINGS = {
# LLM nodes (generic, ADR-007) — see comfydv.ollama.MIGRATION_MAP for
# the pre-cutover Ollama-specific names these replace
"OllamaClient": OllamaClient,
"LlamaCppClient": LlamaCppClient,
"LLMModelSelector": LLMModelSelector,
"LLMLoadModel": LLMLoadModel,
"LLMUnloadModel": LLMUnloadModel,
@@ -59,6 +61,7 @@ NODE_DISPLAY_NAME_MAPPINGS = {
"FormatString": "Format String (Python f-strings)",
# LLM nodes (generic, ADR-007)
"OllamaClient": "Ollama Client",
"LlamaCppClient": "LlamaCpp Client",
"LLMModelSelector": "LLM Model Selector",
"LLMLoadModel": "LLM Load Model",
"LLMUnloadModel": "LLM Unload Model",
+214
View File
@@ -0,0 +1,214 @@
"""LlamaCppProvider — LLMProvider implementation backed by llama-server's
router mode.
Mirrors comfydv._llm.ollama_provider's structure exactly (ADR-007's parallel-
implementation pattern). Router-mode API shape verified live against
ggml-org/llama.cpp's tools/server/README.md (postdates training data) — see
specs/008-llamacpp-integration/research.md. Two details differ from Ollama:
the model identifier field is "id" (not "name"), and "status" is a nested
object ({"value": "..."}), not a flat string.
Deployment prerequisite: llama-server must be launched with --models-dir or
--models-preset (router mode) — the endpoints this provider calls don't
exist otherwise (spec.md FR-006).
"""
import logging
from pydantic import BaseModel
from comfydv._llm.ollama_provider import _TTLLRUCache, _cache_key, _get_json, _post_json
from comfydv._llm.provider import Message, ModelInfo, ModelStatus
logger = logging.getLogger(__name__)
# Own cache pool, not shared with OllamaProvider's — see plan.md's Structure
# Decision (parallel, symmetric, independent implementations). ChatCompletion
# is OUTPUT_NODE=True and re-executes every queue run regardless of which
# provider is wired in, so caching parity matters for llama.cpp too, not
# just Ollama.
_MODEL_LIST_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=20.0)
_CHAT_RESPONSE_CACHE = _TTLLRUCache(maxsize=64, ttl_seconds=None)
class LlamaCppProvider:
"""LLMProvider implementation backed by llama-server's router mode.
Host and headers are captured once at construction — every method
reuses them, the same pattern OllamaProvider already established.
"""
def __init__(self, host: str, headers: dict | None = None):
self.host = host
self.headers = dict(headers) if headers else None
async def list_models(self) -> list[ModelInfo]:
"""GET {host}/models — every model llama-server's router knows
about, with its live status. Unlike OllamaProvider, no
normalization is needed: llama.cpp's status vocabulary is exactly
ModelStatus's full set.
"""
cache_key = _cache_key("llamacpp_list_models", self.host, self.headers or {})
cached, hit = _MODEL_LIST_CACHE.get(cache_key)
if hit:
return cached
try:
data = await _get_json(f"{self.host}/models", headers=self.headers)
except OSError as exc:
# Genuinely unreachable (connection refused, DNS failure, timed
# out — aiohttp's connection-level exceptions are all OSError
# subclasses) — degrade gracefully like OllamaProvider does, so
# a not-yet-started server just shows an empty dropdown rather
# than a hard error.
logger.warning(
"Could not fetch llama.cpp models from %s: %s", self.host, exc
)
return []
except RuntimeError as exc:
# The server answered but with an HTTP error status — GET
# /models only exists in router mode, so this is almost always
# a llama-server launched without --models-dir/--models-preset.
# Surfacing this distinctly (FR-006) matters: silently returning
# [] here would be indistinguishable from "no models installed".
raise RuntimeError(
f"llama-server at {self.host} did not return a model list from "
f"GET {self.host}/models — is it running in router mode "
f"(--models-dir or --models-preset)? Underlying error: {exc}"
) from exc
models = []
for m in data.get("data", []):
status_value = m.get("status", {}).get("value")
try:
status = ModelStatus(status_value)
except ValueError:
logger.warning(
"llama.cpp reported an unrecognized model status %r for %r — "
"skipping status normalization, this model will be omitted",
status_value,
m.get("id"),
)
continue
models.append(ModelInfo(name=m["id"], status=status, size=None))
if models:
_MODEL_LIST_CACHE.set(cache_key, models)
return models
async def load_model(self, model: str) -> None:
if not model.strip():
raise ValueError("model name cannot be empty")
try:
await _post_json(
f"{self.host}/models/load",
{"model": model},
headers=self.headers,
)
except RuntimeError as exc:
# Confirmed live: router mode's own /models/load is NOT
# idempotent — it 400s "model is already running" rather than
# the {"success": true} the contract assumed. The LLMProvider
# protocol requires load_model() to be idempotent, so this
# error is the desired end-state, not a failure — absorb it
# here rather than leaking the wire-level quirk to callers.
if "model is already running" not in str(exc):
raise
async def unload_model(self, model: str) -> None:
if not model.strip():
raise ValueError("model name cannot be empty")
try:
await _post_json(
f"{self.host}/models/unload",
{"model": model},
headers=self.headers,
)
except RuntimeError as exc:
# Mirror of load_model()'s non-idempotency above, confirmed live:
# /models/unload 400s "model is not running" on an already-
# unloaded model instead of {"success": true}.
if "model is not running" not in str(exc):
raise
async def chat(
self,
model: str,
messages: list[Message],
options: dict | None = None,
timeout_secs: float = 300.0,
) -> str:
payload_messages = [m.model_dump() for m in messages]
payload: dict = {"model": model, "messages": payload_messages, "stream": False}
if options:
# Passed through verbatim, same nesting OllamaProvider.chat() uses
# (payload["options"] = options) — the OllamaOption* nodes emit
# Ollama-native parameter names (num_predict, repeat_penalty,
# ...), which llama-server's OpenAI-compatible endpoint won't
# recognize either way; translating them is out of scope for
# this epic (plan.md Non-goals — no changes to the generic
# nodes). This keeps the two providers' handling consistent
# rather than silently special-casing one of them.
payload["options"] = options
cache_key = _cache_key(
"llamacpp_chat",
self.host,
self.headers or {},
model,
payload_messages,
options or {},
)
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
if hit:
return cached
result = await _post_json(
f"{self.host}/v1/chat/completions",
payload,
timeout=timeout_secs,
headers=self.headers,
)
choices = result.get("choices") or []
response_text = (
choices[0].get("message", {}).get("content", "") or "" if choices else ""
)
_CHAT_RESPONSE_CACHE.set(cache_key, response_text)
return response_text
async def chat_structured(
self,
model: str,
messages: list[Message],
schema: type[BaseModel],
options: dict | None = None,
timeout_secs: float = 300.0,
max_retries: int = 2,
) -> BaseModel:
from comfydv._llm.chat import chat_structured as _chat_structured_impl
payload_messages = [m.model_dump() for m in messages]
cache_key = _cache_key(
"llamacpp_chat_structured",
self.host,
self.headers or {},
model,
payload_messages,
options or {},
schema.model_json_schema(),
)
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
if hit:
return schema.model_validate(cached)
result = await _chat_structured_impl(
base_url=f"{self.host}/v1",
model=model,
messages=messages,
schema=schema,
headers=self.headers,
options=options,
max_retries=max_retries,
timeout_secs=timeout_secs,
)
_CHAT_RESPONSE_CACHE.set(cache_key, result.model_dump())
return result
+13 -1
View File
@@ -132,7 +132,14 @@ async def _post_json(
async def _get_json(
url: str, *, timeout: float = 5.0, headers: dict | None = None
) -> dict:
"""GET url, return parsed response dict. Raises on connection/HTTP error."""
"""GET url, return parsed response dict.
Raises RuntimeError on an HTTP error status (distinct message, so callers
can tell "server responded with an error" from "couldn't reach it at
all" — aiohttp connection/timeout errors propagate unwrapped for that
reason). Message is generic, not backend-branded: this helper is shared
by every LLMProvider implementation.
"""
import aiohttp
async with aiohttp.ClientSession() as session:
@@ -141,6 +148,11 @@ async def _get_json(
headers=headers or None,
timeout=aiohttp.ClientTimeout(total=timeout),
) as resp:
if resp.status >= 400:
body = await resp.text()
raise RuntimeError(
f"Server returned HTTP {resp.status} for {url}: {body[:300]}"
)
return await resp.json()
+34
View File
@@ -0,0 +1,34 @@
"""llama.cpp connection node for ComfyUI.
Mirrors comfydv.ollama's OllamaClient exactly (ADR-007's parallel-
implementation pattern) — LlamaCppClient is the only new node this feature
introduces. Every other generic node (ChatCompletion, LLMModelSelector,
LLMLoadModel, LLMUnloadModel) already works with any LLM_CLIENT-typed
provider unchanged.
Deployment prerequisite: llama-server must be launched in router mode
(--models-dir or --models-preset) — see specs/008-llamacpp-integration/quickstart.md.
"""
from comfydv._llm.llamacpp_provider import LlamaCppProvider
class LlamaCppClient:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"host": ("STRING", {"default": "http://localhost:8080"}),
},
"optional": {
"headers": ("OLLAMA_HEADERS",),
},
}
RETURN_TYPES = ("LLM_CLIENT",)
RETURN_NAMES = ("client",)
FUNCTION = "create_client"
CATEGORY = "dv/llamacpp"
def create_client(self, host: str, headers: dict | None = None):
return (LlamaCppProvider(host, headers),)
+137
View File
@@ -0,0 +1,137 @@
"""
Tests for comfydv.llamacpp.LlamaCppClient — the one new ComfyUI node this
feature introduces. Also proves the adapter pattern end-to-end (US4): the
same generic nodes work unmodified against either provider.
BDD coverage:
../specs/008-llamacpp-integration/features/us1_connect_and_chat.feature
../specs/008-llamacpp-integration/features/us4_swap_backends.feature
"""
from comfydv._llm.llamacpp_provider import LlamaCppProvider
from comfydv._llm.ollama_provider import OllamaProvider
from comfydv.llamacpp import LlamaCppClient
from comfydv.ollama import (
ChatCompletion,
LLMLoadModel,
LLMModelSelector,
LLMUnloadModel,
)
class _FakeProvider:
"""Mirrors tests/test_ollama.py's _FakeProvider — reused here for US4's
swap-backends proof rather than duplicated, since the whole point is
that node behavior doesn't depend on which concrete provider it gets."""
def __init__(self, chat_response="ok"):
self.chat_response = chat_response
self.models = [{"name": "m", "status": "loaded"}]
self.calls: list[tuple] = []
async def list_models(self):
self.calls.append(("list_models",))
return self.models
async def load_model(self, model):
self.calls.append(("load_model", model))
async def unload_model(self, model):
self.calls.append(("unload_model", model))
async def chat(self, model, messages, options=None, timeout_secs=300.0):
self.calls.append(("chat", model))
return self.chat_response
def test_client_outputs_llamacpp_provider():
(client,) = LlamaCppClient().create_client("http://localhost:8080")
assert isinstance(client, LlamaCppProvider)
assert client.host == "http://localhost:8080"
def test_client_default_host_matches_llama_server_default_port():
input_types = LlamaCppClient.INPUT_TYPES()
assert input_types["required"]["host"][1]["default"] == "http://localhost:8080"
def test_client_output_type_is_generic_llm_client():
assert LlamaCppClient.RETURN_TYPES == ("LLM_CLIENT",)
def test_client_carries_headers():
(client,) = LlamaCppClient().create_client(
"http://localhost:8080", headers={"Authorization": "Bearer abc"}
)
assert client.headers == {"Authorization": "Bearer abc"}
def test_node_contract():
assert hasattr(LlamaCppClient, "INPUT_TYPES")
assert hasattr(LlamaCppClient, "RETURN_TYPES")
assert hasattr(LlamaCppClient, "FUNCTION")
assert hasattr(LlamaCppClient, "CATEGORY")
assert hasattr(LlamaCppClient, LlamaCppClient.FUNCTION)
def test_registered_in_node_class_mappings():
from comfydv import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS
assert NODE_CLASS_MAPPINGS["LlamaCppClient"] is LlamaCppClient
assert "LlamaCppClient" in NODE_DISPLAY_NAME_MAPPINGS
# ---------------------------------------------------------------------------
# US4 — swap backends without touching downstream nodes
# ---------------------------------------------------------------------------
def _run_workflow(client) -> None:
"""The same node sequence a workflow author would wire up, regardless
of which provider `client` is."""
ChatCompletion().chat(client=client, model="m", prompt="hi")
LLMModelSelector().select_model(client=client, model="m")
LLMLoadModel().load_model(client=client, model="m")
LLMUnloadModel().unload_model(client=client, model="m")
def test_same_workflow_runs_against_either_fake_provider():
"""No node branches on provider type — the same call sequence succeeds
whether client looks like an Ollama-shaped or llama.cpp-shaped provider."""
ollama_like = _FakeProvider(chat_response="ollama says hi")
llamacpp_like = _FakeProvider(chat_response="llamacpp says hi")
# Neither call raises — that's the actual assertion. If ChatCompletion/
# LLMModelSelector/LLMLoadModel/LLMUnloadModel secretly special-cased a
# concrete provider type (isinstance checks, attribute probing beyond
# the protocol), one of these would fail.
_run_workflow(ollama_like)
_run_workflow(llamacpp_like)
# LLMModelSelector is pure passthrough (client is accepted only for
# wiring/typing, never dereferenced), so it makes no provider call.
expected = ["chat", "load_model", "unload_model"]
assert [c[0] for c in ollama_like.calls] == expected
assert [c[0] for c in llamacpp_like.calls] == expected
def test_real_providers_are_interchangeable_client_output():
"""OllamaClient and LlamaCppClient both emit LLM_CLIENT — a workflow
author can wire either one into the same downstream nodes."""
from comfydv.ollama import OllamaClient
(ollama_client,) = OllamaClient().create_client("http://localhost:11434")
(llamacpp_client,) = LlamaCppClient().create_client("http://localhost:8080")
assert isinstance(ollama_client, OllamaProvider)
assert isinstance(llamacpp_client, LlamaCppProvider)
# Both satisfy the same protocol shape — same method names available.
for method in (
"list_models",
"load_model",
"unload_model",
"chat",
"chat_structured",
):
assert callable(getattr(ollama_client, method))
assert callable(getattr(llamacpp_client, method))
+357
View File
@@ -0,0 +1,357 @@
"""
Tests for comfydv._llm.llamacpp_provider.LlamaCppProvider — mirrors
test_ollama_provider.py's structure exactly (ADR-007's parallel-
implementation pattern). Mocks at the provider's own _post_json/_get_json
seam.
BDD coverage:
../specs/008-llamacpp-integration/features/us1_connect_and_chat.feature
../specs/008-llamacpp-integration/features/us2_structured_output.feature
../specs/008-llamacpp-integration/features/us3_model_lifecycle.feature
"""
import pytest
import comfydv._llm.llamacpp_provider as provider_mod
from comfydv._llm.llamacpp_provider import LlamaCppProvider
from comfydv._llm.ollama_provider import _run_async
from comfydv._llm.provider import Message, ModelStatus
@pytest.fixture(autouse=True)
def _clear_provider_caches():
provider_mod._MODEL_LIST_CACHE.clear()
provider_mod._CHAT_RESPONSE_CACHE.clear()
yield
provider_mod._MODEL_LIST_CACHE.clear()
provider_mod._CHAT_RESPONSE_CACHE.clear()
# ---------------------------------------------------------------------------
# list_models — the "id" field name and nested "status.value" are the two
# details research.md flagged as easy to get wrong by assumption.
# ---------------------------------------------------------------------------
def test_list_models_maps_id_field_to_name(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
return {"data": [{"id": "gemma-3-4b:Q4_K_M", "status": {"value": "loaded"}}]}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
(model,) = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
assert model.name == "gemma-3-4b:Q4_K_M"
def test_list_models_reads_nested_status_value(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
return {
"data": [
{"id": "a", "status": {"value": "sleeping"}},
{"id": "b", "status": {"value": "downloading", "progress": {}}},
]
}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
by_name = {m.name: m for m in models}
assert by_name["a"].status == ModelStatus.SLEEPING
assert by_name["b"].status == ModelStatus.DOWNLOADING
def test_list_models_no_normalization_needed_full_vocabulary(monkeypatch):
"""Unlike OllamaProvider, llama.cpp's status vocabulary is exactly
ModelStatus's full set — every value should pass through untouched."""
async def fake_get(url, *, timeout=5.0, headers=None):
return {
"data": [
{"id": v, "status": {"value": v}}
for v in ["unloaded", "loading", "loaded", "sleeping", "downloading"]
]
}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
assert {m.status for m in models} == set(ModelStatus)
def test_list_models_skips_unrecognized_status(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
return {
"data": [
{"id": "crashed", "status": {"value": "failed", "exit_code": 1}},
{"id": "ok", "status": {"value": "loaded"}},
]
}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
assert [m.name for m in models] == ["ok"]
def test_list_models_unreachable_returns_empty(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
raise ConnectionError("no route to host")
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
models = _run_async(LlamaCppProvider("http://localhost:19999").list_models())
assert models == []
def test_list_models_non_router_mode_raises_clear_error(monkeypatch):
"""FR-006: a llama-server that IS reachable but wasn't launched with
--models-dir/--models-preset answers GET /models with an HTTP error
(the endpoint doesn't exist outside router mode). That must surface as
a specific, actionable error — not silently degrade to an empty list,
which would be indistinguishable from "server has no models"."""
async def fake_get(url, *, timeout=5.0, headers=None):
raise RuntimeError("Server returned HTTP 404 for http://x/models: not found")
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
with pytest.raises(RuntimeError, match="router mode"):
_run_async(LlamaCppProvider("http://localhost:8080").list_models())
def test_list_models_cached_second_call(monkeypatch):
calls = {"n": 0}
async def fake_get(url, *, timeout=5.0, headers=None):
calls["n"] += 1
return {"data": [{"id": "a", "status": {"value": "loaded"}}]}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
provider = LlamaCppProvider("http://localhost:8080")
_run_async(provider.list_models())
_run_async(provider.list_models())
assert calls["n"] == 1
# ---------------------------------------------------------------------------
# load_model / unload_model
# ---------------------------------------------------------------------------
def test_load_model_posts_to_models_load_with_model_field(monkeypatch):
captured = {}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
captured["url"] = url
captured["payload"] = payload
return {"success": True}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
assert captured["url"] == "http://localhost:8080/models/load"
assert captured["payload"] == {"model": "gemma-3-4b"}
def test_unload_model_posts_to_models_unload_with_model_field(monkeypatch):
captured = {}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
captured["url"] = url
captured["payload"] = payload
return {"success": True}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
assert captured["url"] == "http://localhost:8080/models/unload"
assert captured["payload"] == {"model": "gemma-3-4b"}
def test_load_model_already_running_is_idempotent(monkeypatch):
"""Confirmed live: router mode's /models/load is NOT idempotent at the
wire level — it 400s "model is already running" rather than returning
{"success": true}. The LLMProvider protocol requires load_model() to be
idempotent, so LlamaCppProvider must absorb this itself."""
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append((url, payload))
raise RuntimeError(
"Server returned HTTP 400 for "
f'{url}: {{"error":{{"code":400,"message":"model is already '
'running","type":"invalid_request_error"}}}}'
)
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
# The absence of a raised exception is only meaningful if the request
# actually happened and hit the "already running" branch — assert that
# directly rather than trusting silence alone.
assert calls == [("http://localhost:8080/models/load", {"model": "gemma-3-4b"})]
def test_load_model_other_http_error_still_raises(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
raise RuntimeError(f"Server returned HTTP 500 for {url}: internal error")
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
with pytest.raises(RuntimeError, match="500"):
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
def test_unload_model_not_running_is_idempotent(monkeypatch):
"""Mirror of the load_model case, confirmed live: /models/unload 400s
"model is not running" on an already-unloaded model."""
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append((url, payload))
raise RuntimeError(
"Server returned HTTP 400 for "
f'{url}: {{"error":{{"code":400,"message":"model is not '
'running","type":"invalid_request_error"}}}}'
)
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
assert calls == [("http://localhost:8080/models/unload", {"model": "gemma-3-4b"})]
def test_unload_model_other_http_error_still_raises(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
raise RuntimeError(f"Server returned HTTP 500 for {url}: internal error")
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
with pytest.raises(RuntimeError, match="500"):
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
def test_load_model_empty_raises_before_network(monkeypatch):
def fail_post(*a, **k):
raise AssertionError("must not call _post_json for an empty model name")
monkeypatch.setattr(provider_mod, "_post_json", fail_post)
with pytest.raises(ValueError, match="cannot be empty"):
_run_async(LlamaCppProvider("http://localhost:8080").load_model(""))
def test_unload_model_empty_raises_before_network(monkeypatch):
def fail_post(*a, **k):
raise AssertionError("must not call _post_json for an empty model name")
monkeypatch.setattr(provider_mod, "_post_json", fail_post)
with pytest.raises(ValueError, match="cannot be empty"):
_run_async(LlamaCppProvider("http://localhost:8080").unload_model(" "))
# ---------------------------------------------------------------------------
# chat — OpenAI response shape (choices[0].message.content), not Ollama's
# native shape
# ---------------------------------------------------------------------------
def test_chat_posts_to_v1_chat_completions_and_parses_openai_shape(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
assert url == "http://localhost:8080/v1/chat/completions"
return {"choices": [{"message": {"role": "assistant", "content": "hello"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b", [Message(role="user", content="hi")]
)
)
assert result == "hello"
def test_chat_no_choices_returns_empty_string(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
return {"choices": []}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b", [Message(role="user", content="hi")]
)
)
assert result == ""
def test_chat_second_identical_call_is_cached(monkeypatch):
calls = {"n": 0}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls["n"] += 1
return {"choices": [{"message": {"content": "cached"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
provider = LlamaCppProvider("http://localhost:8080")
messages = [Message(role="user", content="hi")]
r1 = _run_async(provider.chat("m", messages))
r2 = _run_async(provider.chat("m", messages))
assert r1 == r2 == "cached"
assert calls["n"] == 1
# ---------------------------------------------------------------------------
# chat_structured — zero new logic, delegates to the shared helper unchanged
# ---------------------------------------------------------------------------
def test_chat_structured_builds_v1_base_url_and_delegates(monkeypatch):
from pydantic import BaseModel
class Widget(BaseModel):
name: str
captured = {}
async def fake_chat_structured(**kwargs):
captured.update(kwargs)
return Widget(name="x")
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat_structured(
"gemma-3-4b", [Message(role="user", content="hi")], Widget
)
)
assert result == Widget(name="x")
assert captured["base_url"] == "http://localhost:8080/v1"
assert captured["model"] == "gemma-3-4b"
def test_chat_structured_forwards_options(monkeypatch):
from pydantic import BaseModel
class Widget(BaseModel):
name: str
captured = {}
async def fake_chat_structured(**kwargs):
captured.update(kwargs)
return Widget(name="x")
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
_run_async(
LlamaCppProvider("http://localhost:8080").chat_structured(
"gemma-3-4b",
[Message(role="user", content="hi")],
Widget,
options={"temperature": 0.0},
)
)
assert captured["options"] == {"temperature": 0.0}