docs(beacon): DESIGN artifacts for llama.cpp Model Integration (spec 008)

Full spec/plan/tasks/BDD scaffold for the llamacpp-integration epic, now
that its dependency (llm-provider-abstraction, PR #17) is satisfied.
Researched llama-server's router-mode API shape live (postdates training
data) rather than assuming it — two details that would have been wrong by
assumption: the model identifier field is "id" not "name" (differs from
Ollama's /api/tags), and "status" is a nested object ({"value": "..."}) not
a flat string.

Unlike the prerequisite epic, this decomposition genuinely holds as 4
independent user stories — LlamaCppClient is a new node, not a changed one,
so there's no shared "output type" migration forcing an atomic cutover.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
This commit is contained in:
James Veitch
2026-07-11 15:17:54 +01:00
co-authored by Claude Sonnet 5
parent 9ca65bc29a
commit d3ee35bfda
16 changed files with 649 additions and 2 deletions
+1 -1
View File
@@ -1,3 +1,3 @@
{
"feature_directory": "specs/007-llm-provider-abstraction"
"feature_directory": "specs/008-llamacpp-integration"
}
+1 -1
View File
@@ -1,5 +1,5 @@
<!-- SPECKIT START -->
For additional context about technologies to be used, project structure,
shell commands, and other important information, read the current plan
at specs/007-llm-provider-abstraction/plan.md
at specs/008-llamacpp-integration/plan.md
<!-- SPECKIT END -->
@@ -32,6 +32,7 @@ epic.
_Filled by `beacon specify --epic llamacpp-integration` / `/speckit-specify`
once this epic is accepted._
- specs/008-llamacpp-integration/
## ADRs
- project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md — decided during the prerequisite epic; this epic implements the second `LLMProvider` the ADR anticipated
@@ -0,0 +1 @@
epic = "llamacpp-integration"
@@ -0,0 +1,39 @@
# Specification Quality Checklist: llama.cpp Model Integration
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-07-11
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
No [NEEDS CLARIFICATION] markers needed — scope boundaries (router-mode-only,
no GPU tuning, no auth/TLS, no Manager listing) came directly from the
parent epic's Non-goals (`project-management/Roadmap/epics/llamacpp-integration.md`)
and ADR-007. User Story 4 (swap backends without touching downstream nodes)
is the adapter pattern's central promise made concrete and testable, not
padding.
@@ -0,0 +1,39 @@
# Contract: `LlamaCppProvider` conforms to `LLMProvider`
This is the concrete proof of ADR-007's adapter pattern — the same protocol
contract documented in
`specs/007-llm-provider-abstraction/contracts/llm_provider_protocol.md`,
now with a second implementation. Nothing in that contract changes; this
file only documents `LlamaCppProvider`'s specific wire-format bindings.
```python
class LlamaCppProvider:
def __init__(self, host: str, headers: dict | None = None): ...
async def list_models(self) -> list[ModelInfo]:
"""GET {host}/models → data[] → ModelInfo(name=m["id"], status=ModelStatus(m["status"]["value"]), size=None)"""
async def load_model(self, model: str) -> None:
"""POST {host}/models/load {"model": model}"""
async def unload_model(self, model: str) -> None:
"""POST {host}/models/unload {"model": model}"""
async def chat(self, model, messages, options=None, timeout_secs=300.0) -> str:
"""POST {host}/v1/chat/completions → choices[0].message.content"""
async def chat_structured(self, model, messages, schema, options=None, timeout_secs=300.0, max_retries=2) -> BaseModel:
"""Delegates to comfydv._llm.chat.chat_structured(base_url=f"{host}/v1", ...) — identical call OllamaProvider makes"""
```
## Behavioral requirements (inherited from the protocol contract, restated for this implementation)
- `load_model`/`unload_model` MUST be idempotent — router mode's own
`{"success": true}` response on an already-loaded/unloaded model satisfies
this without extra handling.
- `list_models()` MUST NOT normalize away llama.cpp's `sleeping`/`downloading`
states (unlike `OllamaProvider`, which has no choice but to normalize —
see `research.md`).
- A `llama-server` not running in router mode (missing endpoints) MUST
surface a clear, specific error (spec.md FR-006) — not a generic
connection failure indistinguishable from "server not running at all."
@@ -0,0 +1,41 @@
# Data Model: llama.cpp Model Integration
No new types — this feature is a second implementation of the existing
`LLMProvider` protocol, `ModelStatus`, `ModelInfo`, and `Message` types
(`src/comfydv/_llm/provider.py`, unchanged). This file documents
`LlamaCppProvider`'s field mapping from llama-server's router-mode JSON onto
those existing types (see `research.md` for the verified API shapes).
## `LlamaCppProvider.list_models()` → `ModelInfo` mapping
| `ModelInfo` field | Source (`GET /models` response, per model in `data[]`) |
|---|---|
| `name` | `id` — **not** `name` (llama.cpp's field name differs from Ollama's) |
| `status` | `status.value` — nested object, not a flat string |
| `size` | Not provided by this endpoint; `None` |
`status.value` maps directly onto `ModelStatus`'s five values
(`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`) — llama.cpp's
vocabulary is exactly `ModelStatus`'s full set, so unlike `OllamaProvider`
(which normalizes into a narrower subset), `LlamaCppProvider` needs no
approximation. A `"failed": true` state exists outside this vocabulary
(model process crashed) — out of scope per spec.md's edge cases; treated as
whatever `status.value` reports rather than added as a sixth enum value.
## `LlamaCppProvider.load_model()` / `unload_model()`
Both `POST /models/load` and `POST /models/unload` take `{"model": <id>}` —
the same `id` string `list_models()` returns as `ModelInfo.name`. No mapping
ambiguity here (unlike Ollama, where load/unload uses `/api/generate`'s
`keep_alive` side effect rather than a dedicated endpoint).
## `LlamaCppProvider.chat()` / `chat_structured()`
Both reach `llama-server`'s OpenAI-compatible `/v1/chat/completions` —
`chat_structured()` calls the existing shared `comfydv._llm.chat.chat_structured()`
helper unchanged (`base_url=f"{self.host}/v1"`, matching `OllamaProvider`'s
own call exactly). `chat()` parses the response as
`choices[0].message.content` (OpenAI shape), not Ollama's native
`message.content` — the two providers' non-structured paths differ here
because llama-server doesn't have an Ollama-style native `/api/chat`
endpoint to prefer instead.
@@ -0,0 +1,11 @@
Feature: US1 — Connect to a local llama.cpp server and get chat responses
Scenario: llama.cpp connection node feeds the existing chat node
Given a running local llama-server (router mode) and a workflow with a llama.cpp connection node wired into the existing chat node
When the workflow executes
Then the chat node returns the model's text response
Scenario: Unreachable llama.cpp server surfaces a clear error
Given the llama.cpp connection node configured with an unreachable server address
When the workflow executes
Then the chat node reports a clear connection error
@@ -0,0 +1,11 @@
Feature: US2 — Get structured, validated output from llama.cpp
Scenario: Valid structured response exposes typed fields, same as Ollama
Given a chat node connected to llama.cpp with structured output enabled and a valid schema
When the workflow executes and the model responds correctly
Then each schema field is available as its own typed output, and no required field is blank
Scenario: Invalid response retries then fails clearly, same as Ollama
Given a llama.cpp-hosted model that returns invalid or incomplete structured output
When the workflow executes
Then the node retries automatically and, if still unsuccessful, fails with a clear error
@@ -0,0 +1,16 @@
Feature: US3 — See and control which models are loaded on llama.cpp
Scenario: List models with full status vocabulary
Given a running local llama-server with at least one available model
When a workflow author uses the model-listing node
Then they see each available model along with its current status, drawn from llama.cpp's full status vocabulary
Scenario: Load a model into memory
Given a model that is not currently loaded
When a workflow author runs the load-model node against it
Then the model becomes loaded and is then usable by the chat node
Scenario: Unload a model from memory
Given a model that is loaded and idle
When a workflow author runs the unload-model node against it
Then the model is freed from memory and its reported status updates accordingly
@@ -0,0 +1,6 @@
Feature: US4 — Swap from Ollama to llama.cpp without touching the rest of the workflow
Scenario: Replacing only the connection node preserves the workflow
Given a workflow with chat/model-management nodes wired to an Ollama connection node
When a workflow author replaces only the connection node with a llama.cpp one, pointed at a running llama-server
Then the workflow runs successfully with no changes to any other node
+117
View File
@@ -0,0 +1,117 @@
# Implementation Plan: llama.cpp Model Integration
**Branch**: `008-llamacpp-integration` | **Date**: 2026-07-11 | **Spec**: [spec.md](./spec.md)
**Input**: Feature specification from `/specs/008-llamacpp-integration/spec.md`
**Note**: This template is filled in by the `/speckit-plan` command. See `.specify/templates/plan-template.md` for the execution workflow.
## Summary
Implement `LlamaCppProvider` as the second `LLMProvider` (ADR-007), backed by
`llama-server`'s router mode (`GET /models`, `POST /models/load`,
`POST /models/unload`, `/v1/chat/completions`). Add one new ComfyUI node
(`LlamaCppClient`) emitting the existing `LLM_CLIENT` socket type — no other
node classes change. This is the concrete proof the provider abstraction
(prerequisite epic, PR #17) actually generalizes: a second backend, zero
changes to `ChatCompletion`/`LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel`.
## Technical Context
**Language/Version**: Python ≥3.11 (unchanged, per `pyproject.toml`)
**Primary Dependencies**: `aiohttp` (existing — model-management REST calls),
`pydantic-ai`/`openai` (existing, from the prerequisite epic — `chat_structured()`
reuses the shared helper unchanged, zero new structured-output code)
**Storage**: N/A — no persistent storage; reuses the existing
`_MODEL_LIST_CACHE`/`_CHAT_RESPONSE_CACHE` infra pattern from `OllamaProvider`
**Testing**: `pytest` via `uv run pytest`, following `tests/test_ollama_provider.py`'s
established convention (mock at the provider's own `_post_json`/`_get_json`
seam, no live server required for unit tests)
**Target Platform**: ComfyUI custom-node runtime, same as the existing Ollama
integration
**Project Type**: Library / ComfyUI custom-node pack (single project, adds to
existing `src/comfydv/` layout)
**Performance Goals**: No new numeric target; must not add latency beyond
what `OllamaProvider`'s equivalent methods already accept
**Constraints**: Router-mode-only (spec.md Assumptions — a `llama-server`
without `--models-dir`/`--models-preset` doesn't expose these endpoints at
all, FR-006); model identifier field is `id` (llama.cpp) vs `name` (Ollama) —
`LlamaCppProvider.list_models()` must map this correctly (see `research.md`);
`status` is a nested object (`{"value": "..."}`), not a flat string
**Scale/Scope**: One new class (`LlamaCppProvider`, mirrors `OllamaProvider`'s
shape), one new ComfyUI node (`LlamaCppClient`), one new test file — no
changes to `ollama.py`, `_llm/provider.py`, `_llm/chat.py`, or any existing
node class
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Principle | Verdict | Notes |
|---|---|---|
| I. ComfyUI Contract First | PASS | `LlamaCppClient` exposes the standard `INPUT_TYPES`/`RETURN_TYPES`/`FUNCTION`/`CATEGORY`; registered in `NODE_CLASS_MAPPINGS` like every other node. |
| II. Sandbox All User-Supplied Code | N/A | No template/expression evaluation in this feature. |
| III. Test-First | PASS (binding) | `tests/test_llamacpp_provider.py` written test-first, mirroring `test_ollama_provider.py`'s TDD-pair structure. |
| IV. Graceful Degradation Outside ComfyUI | PASS (binding) | `LlamaCppProvider` lives in `src/comfydv/_llm/`, which already has no `comfy`/`server` imports at module scope (verified for the prerequisite epic; this feature adds no new module-scope imports of either). |
| V. Simplicity — Function Before Class | PASS, same justification as `OllamaProvider` | `LlamaCppProvider` carries connection state (host, headers) across 5 methods — the same shared-state condition that already justified `OllamaProvider` as a class (research.md, prerequisite epic). No new gate — same precedent applies. |
| VI. Fixed Output Positions | N/A | `LlamaCppClient`'s single output (`client`) isn't a multi-output node; no positional contract to preserve. |
Re-checked post-Phase 1 design (data-model.md): unchanged — no new gate
violations. No Complexity Tracking entries needed (unlike the prerequisite
epic, this feature introduces no new pattern, just a second instance of an
already-justified one).
## Project Structure
### Documentation (this feature)
```text
specs/008-llamacpp-integration/
├── plan.md # This file
├── research.md # Phase 0 — router-mode API shape, verified live
├── data-model.md # Phase 1 — LlamaCppProvider field mapping
├── quickstart.md # Phase 1 — minimal workflow walkthrough
├── contracts/ # Phase 1 — LlamaCppProvider's protocol conformance
└── tasks.md # Phase 2 (/speckit-tasks)
```
### Source Code (repository root)
```text
src/comfydv/
├── ollama.py # unchanged — add LlamaCppClient node only via a new module
├── llamacpp.py # new — LlamaCppClient node (mirrors OllamaClient's shape)
├── _llm/
│ ├── provider.py # unchanged — LLMProvider/ModelStatus/ModelInfo/Message
│ ├── ollama_provider.py # unchanged
│ ├── llamacpp_provider.py # new — LlamaCppProvider (mirrors ollama_provider.py's shape)
│ └── chat.py # unchanged — chat_structured() reused as-is
└── __init__.py # add LlamaCppClient import + NODE_CLASS_MAPPINGS entry
tests/
├── test_ollama_provider.py # unchanged
├── test_llamacpp_provider.py # new — mirrors test_ollama_provider.py's structure
└── test_llamacpp.py # new — LlamaCppClient node contract test (small; mirrors
# the OllamaClient-specific slice of test_ollama.py)
```
**Structure Decision**: New `src/comfydv/llamacpp.py` module (not added into
`ollama.py`) for the `LlamaCppClient` node, and a new `src/comfydv/_llm/llamacpp_provider.py`
for `LlamaCppProvider` — mirroring the existing `ollama.py`/`ollama_provider.py`
split exactly, so the two backends read as parallel, symmetric implementations
rather than one growing to accommodate the other. No existing file is
modified except `__init__.py`'s registration block.
## Complexity Tracking
> **Fill ONLY if Constitution Check has violations that must be justified**
None — see Constitution Check above.
@@ -0,0 +1,26 @@
# Quickstart: llama.cpp Model Integration
## Prerequisite
Launch `llama-server` in router mode:
```bash
llama-server --models-dir ./models -c 8192
```
## Minimal workflow
1. Add an **LlamaCpp Client** node. Set its host widget (default
`http://localhost:8080`, llama-server's default port).
2. Wire it into a **Chat Completion** node — the exact same node used for
Ollama. Set a model and prompt, run.
3. Structured output, model listing, and load/unload all work exactly as
documented for Ollama in the main README/quickstart — swap the client
node, nothing else changes.
## Swapping an existing Ollama workflow to llama.cpp
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed
at your running `llama-server`. Every downstream node (Chat Completion, LLM
Model Selector, LLM Load Model, LLM Unload Model) keeps working unmodified —
this is the whole point of the provider abstraction (ADR-007).
@@ -0,0 +1,78 @@
# Research: llama.cpp Model Integration
## Decision: exact router-mode API shape (verified against `ggml-org/llama.cpp`'s live `tools/server/README.md`, not assumed)
llama.cpp's router mode postdates this session's training data — verified live
against the authoritative source rather than guessed, since getting field
names wrong here would silently produce broken code (wrong key = `KeyError`
or silent `None`, not an obvious failure).
**`GET /models`** response:
```json
{
"data": [
{
"id": "ggml-org/gemma-3-4b-it-GGUF:Q4_K_M",
"path": "/Users/.../gemma-3-4b-it-Q4_K_M.gguf",
"status": {
"value": "loaded",
"args": ["llama-server", "-ctx", "4096"]
},
"architecture": {
"input_modalities": ["text", "image"],
"output_modalities": ["text"]
}
}
]
}
```
**Two details that would have been wrong by assumption:**
1. The model identifier field is **`id`**, not `name` — different from Ollama's
`/api/tags`, which uses `name`. `OllamaProvider.list_models()` maps
`m["name"]`; `LlamaCppProvider.list_models()` must map `m["id"]` instead.
2. **`status` is a nested object** (`{"value": "loaded", ...}`), not a flat
string field. `LlamaCppProvider.list_models()` must read
`m["status"]["value"]`, not `m["status"]` directly. A `"failed"` state
also exists (`{"failed": true, "exit_code": ...}`) outside the five
`ModelStatus` values the protocol defines — not handled by this feature
(see Non-goals/edge cases in `spec.md`); a failed model is reported as
whatever `status.value` degrades to rather than added as a sixth enum
value, keeping `ModelStatus` unchanged across both providers.
**`POST /models/load`** and **`POST /models/unload`**: identical request
shape, `{"model": "<id>"}` (using the same `id` string from `GET /models`,
despite the request field being named `model` not `id`). Response:
`{"success": true}`.
**CLI**: `--models-dir <path>` or `--models-preset <path>.ini` — a deployment
prerequisite (spec.md Assumptions), not something comfydv configures.
## Decision: `chat_structured()` needs zero new code
`llama-server`'s `/v1/chat/completions` is OpenAI-compatible (the same
assumption ADR-007 made when adopting `pydantic-ai`). `LlamaCppProvider.chat_structured()`
calls the exact same `comfydv._llm.chat.chat_structured()` helper
`OllamaProvider` already uses, with `base_url=f"{self.host}/v1"` — the only
per-provider difference. This is the concrete proof the shared mechanism
generalizes (spec.md User Story 2/FR-004), not just an assumption.
## Decision: `chat()` (non-structured) also reuses the OpenAI-compatible endpoint
Unlike Ollama (which has both a native `/api/chat` and an OpenAI-compat
`/v1/chat/completions`), llama-server's primary chat endpoint is the
OpenAI-compatible one. `LlamaCppProvider.chat()` POSTs to
`{host}/v1/chat/completions` (via the existing `_post_json` helper, no new
HTTP client) rather than mirroring Ollama's native-endpoint choice — the
response shape (`choices[0].message.content`) differs from Ollama's native
`message.content` and must be parsed accordingly.
## Decision: no protocol changes needed
`LLMProvider`'s five methods (`list_models`/`load_model`/`unload_model`/
`chat`/`chat_structured`) already cover everything router mode needs — this
was the actual point of designing the protocol at the operation level in
ADR-007, and this research confirms it held up against llama.cpp's real API,
not just Ollama's.
+110
View File
@@ -0,0 +1,110 @@
# Feature Specification: llama.cpp Model Integration
**Feature Branch**: `008-llamacpp-integration`
**Created**: 2026-07-11
**Status**: Draft
**Input**: User description: "Add ComfyUI nodes for llama.cpp local inference via llama-server's router mode, implementing the LlamaCppProvider as the second LLMProvider (ADR-007) alongside the existing OllamaProvider. Router mode exposes GET /models (with live status), POST /models/load, POST /models/unload, giving llama.cpp the same manual load/unload memory-management primitives as Ollama. No new ComfyUI node classes needed for chat/model-selection/load/unload — only a new LlamaCppClient config node; the existing generic ChatCompletion/LLMModelSelector/LLMLoadModel/LLMUnloadModel nodes work unchanged once wired to it."
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Connect to a local llama.cpp server and get chat responses (Priority: P1) 🎯 MVP
As a ComfyUI workflow author running `llama-server` locally, I want a connection node for it — just like the one I already use for Ollama — so I can get chat responses from a llama.cpp-hosted model using the same chat node I already know.
**Why this priority**: This is the entire point of the feature and the proof that the provider abstraction (shipped in the prerequisite epic) actually works: a second backend, zero changes to the chat node.
**Independent Test**: Wire a new llama.cpp connection node into the existing chat node, run against a local `llama-server` (router mode), confirm a text response.
**Acceptance Scenarios**:
1. **Given** a running local `llama-server` (router mode) and a workflow with a llama.cpp connection node wired into the existing chat node, **When** the workflow executes, **Then** the chat node returns the model's text response — using the exact same chat node a workflow author already uses for Ollama.
2. **Given** the llama.cpp connection node configured with an unreachable server address, **When** the workflow executes, **Then** the chat node reports a clear connection error, matching the behavior workflow authors already know from the Ollama connection.
---
### User Story 2 - Get structured, validated output from llama.cpp (Priority: P1)
As a workflow author, I want structured output (a schema-validated response instead of free text) to work identically regardless of whether I'm connected to Ollama or llama.cpp, so I don't have to relearn or rebuild anything when switching backends.
**Why this priority**: Structured output is a core existing capability (already proven for Ollama); this story proves the shared mechanism genuinely generalizes rather than being Ollama-specific in practice, not just in name.
**Independent Test**: Enable structured output on the chat node with a schema, run against a llama.cpp-hosted model, confirm each schema field is populated and never blank — using the same steps as the equivalent Ollama test.
**Acceptance Scenarios**:
1. **Given** a chat node connected to llama.cpp with structured output enabled and a valid schema, **When** the workflow executes and the model responds correctly, **Then** each schema field is available as its own typed output, and no required field is blank.
2. **Given** a llama.cpp-hosted model that returns invalid or incomplete structured output, **When** the workflow executes, **Then** the node retries automatically and, if still unsuccessful, fails with a clear error — identical behavior to the Ollama path.
---
### User Story 3 - See and control which models are loaded on llama.cpp (Priority: P2)
As a workflow author running models locally, I want to see live model status (including whether a model is currently loading or being downloaded, not just loaded/unloaded) and explicitly load or unload a model on my llama.cpp server, so I can manage memory the same way I already do for Ollama — with more visibility, since llama.cpp's router mode reports richer status than Ollama does.
**Why this priority**: Valuable and proves the model-management path generalizes too, but a workflow can still run chat completions without ever calling load/unload explicitly (the server can load on first use), so it's lower risk to defer than basic chat.
**Independent Test**: Use the existing model-listing node against a running `llama-server`, confirm it shows each available model with its current status (including `loading`/`downloading` if applicable); use the existing load/unload nodes against one model and confirm its status changes.
**Acceptance Scenarios**:
1. **Given** a running local `llama-server` with at least one available model, **When** a workflow author uses the model-listing node, **Then** they see each available model along with its current status, drawn from llama.cpp's full status vocabulary (not just loaded/unloaded).
2. **Given** a model that is not currently loaded, **When** a workflow author runs the load-model node against it, **Then** the model becomes loaded and is then usable by the chat node.
3. **Given** a model that is loaded and idle, **When** a workflow author runs the unload-model node against it, **Then** the model is freed from memory and its reported status updates accordingly.
---
### User Story 4 - Swap from Ollama to llama.cpp without touching the rest of the workflow (Priority: P3)
As a workflow author with an existing Ollama-based workflow, I want to switch it to llama.cpp by changing only the connection node, so I don't have to rebuild my chat/model-management logic for a second backend.
**Why this priority**: This is the adapter pattern's actual promise made concrete for a user, but it's a validation/demonstration story rather than new capability — everything it depends on is already covered by User Stories 1–3.
**Independent Test**: Take a workflow using the Ollama connection node, replace it with the llama.cpp connection node (same downstream nodes, no other changes), run it, confirm it still works.
**Acceptance Scenarios**:
1. **Given** a workflow with chat/model-management nodes wired to an Ollama connection node, **When** a workflow author replaces only the connection node with a llama.cpp one (pointed at a running `llama-server`), **Then** the workflow runs successfully with no changes to any other node.
---
### Edge Cases
- What happens when `llama-server` is running but was launched without router mode (i.e. with `-m` instead of `--models-dir`)? The router-mode-only endpoints this feature depends on won't exist — the connection/model-management nodes should fail with a clear error, not hang or silently return empty results.
- What happens when the configured server address is unreachable at the moment a model-listing, load, or unload node runs (not just the chat node)?
- What happens when llama.cpp reports a model status this feature doesn't expect (a router-mode API change)? Should degrade gracefully (surface the status if recognized, don't crash on an unrecognized one), not silently misreport.
- What happens to an in-flight chat request if the model it depends on is unloaded by another node in the same workflow run? (Same question already answered for Ollama — behavior should be consistent.)
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: The system MUST allow a workflow author to configure a connection to a local `llama-server` (router mode) the same way they already configure a connection to Ollama — a dedicated connection node, reusable across multiple nodes in a workflow.
- **FR-002**: The system MUST NOT require any new or different node classes for chat, structured output, model listing, or load/unload when using llama.cpp — the existing generic nodes MUST work unchanged once connected to a llama.cpp connection node.
- **FR-003**: The system MUST report each model's status using llama.cpp's full status vocabulary (unloaded, loading, loaded, sleeping, downloading) when connected to llama.cpp — not degraded to the narrower Ollama-compatible set.
- **FR-004**: The system's chat and structured-output behavior MUST be identical between Ollama and llama.cpp connections, given equivalent inputs — same retry limits, same validation rules, same error conditions (this is the direct continuation of the prerequisite epic's own FR-007/FR-008).
- **FR-005**: The system MUST allow a workflow author to explicitly load a model into memory and explicitly unload a model from memory on a connected llama.cpp server.
- **FR-006**: The system MUST surface a clear, specific error when connected to a `llama-server` instance that isn't running in router mode (the endpoints this feature needs don't exist), rather than an unhelpful generic failure.
### Key Entities *(include if feature involves data)*
- **llama.cpp connection**: A configured connection to a local `llama-server` instance running in router mode (host + any authentication), implementing the same connection concept already established for Ollama.
- **Model status**: Reuses the existing status concept from the prerequisite feature, now populated with llama.cpp's full vocabulary rather than a narrowed subset.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: A workflow author can connect to a llama.cpp server and get a chat response using the same node count and shape as connecting to Ollama (one connection node, one chat node) — no new nodes to learn for the chat path.
- **SC-002**: An existing workflow can be repointed from Ollama to llama.cpp by changing exactly one node (the connection node) — zero edits to any chat or model-management node.
- **SC-003**: Structured-output workflows behave identically (same validation guarantees, zero blank-required-field results) regardless of which backend is connected.
- **SC-004**: Model status reporting for llama.cpp surfaces all five status values where applicable — a strictly richer view than what Ollama can report through the same interface.
## Assumptions
- Workflow authors run their own local `llama-server` instance, launched in router mode (`--models-dir` or `--models-preset`), reachable over HTTP from the machine running ComfyUI; this feature does not install, configure, or launch that server.
- Non-router-mode `llama-server` usage (a single model launched with `-m`) is out of scope — router mode is required for the load/unload/status parity with Ollama that is this feature's whole point.
- GPU inference optimisation, quantisation tuning, authentication/TLS, and ComfyUI Manager registry listing are out of scope, consistent with the prerequisite Ollama epic's own non-goals.
- The `LLMProvider` protocol and generic nodes (`ChatCompletion`, `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`) already exist and are not modified by this feature — if llama.cpp's router mode needs a protocol capability that doesn't exist yet, that is a protocol change scoped as its own follow-up, not silently special-cased here.
+151
View File
@@ -0,0 +1,151 @@
# Tasks: llama.cpp Model Integration
**Input**: Design documents from `/specs/008-llamacpp-integration/`
**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/llamacpp_provider_conformance.md
**Tests**: First-class — every implementation task has a paired failing-test task (`-T`/`-I` suffix).
**Organization**: Grouped by user story (spec.md priorities P1/P1/P2/P3).
## Format: `[ID] [P?] [Story] Description`
- **[P]**: Can run in parallel (different files, no dependencies)
- **[Story]**: US1–US4
- **-T / -I**: paired test (red) / implementation (green)
## Path Conventions
Single project: `src/comfydv/`, `tests/` at repository root, mirroring the
`ollama.py`/`_llm/ollama_provider.py` split exactly (plan.md Structure
Decision).
---
## Phase 1: Setup
- [x] T001 No new dependencies — `aiohttp`/`pydantic-ai` already present from the prerequisite epic (verified in `pyproject.toml`)
- [ ] T002 [P] Create `src/comfydv/_llm/llamacpp_provider.py` and `src/comfydv/llamacpp.py` (empty modules with docstrings, mirroring `ollama_provider.py`/`ollama.py`'s module docstring style)
---
## Phase 2: Foundational
None — `LLMProvider`, `ModelStatus`, `ModelInfo`, `Message`, and the shared
`chat_structured()` helper already exist from the prerequisite epic and are
unmodified by this feature (plan.md Constitution Check, research.md).
**Checkpoint**: nothing blocks user story work — it can start immediately.
---
## Phase 3: User Story 1 — Connect to a local llama.cpp server and get chat responses (Priority: P1) 🎯 MVP
**Goal**: A workflow author wires an `LlamaCppClient` node into the existing `ChatCompletion` node and gets a text response.
**Independent Test**: Wire `LlamaCppClient` → `ChatCompletion`, run against a live `llama-server` (router mode), confirm text output.
- [ ] T003-T [US1] Write FAILING test: `LlamaCppProvider.chat()` POSTs to `{host}/v1/chat/completions` and parses `choices[0].message.content`, in `tests/test_llamacpp_provider.py` (witnesses `features/us1_connect_and_chat.feature` scenario "llama.cpp connection node feeds the existing chat node")
- [ ] T003-I [US1] Implement `LlamaCppProvider.__init__`/`.chat()` in `src/comfydv/_llm/llamacpp_provider.py` (data-model.md — OpenAI-shape response parsing, not Ollama's native shape) — makes T003-T pass
- [ ] T004-T [US1] Write FAILING test: `LlamaCppClient` node's `INPUT_TYPES`/`RETURN_TYPES` match `OllamaClient`'s shape (`LLM_CLIENT` output), and `create_client()` constructs a `LlamaCppProvider`, in `tests/test_llamacpp.py`
- [ ] T004-I [US1] Implement `LlamaCppClient` node in `src/comfydv/llamacpp.py` (mirrors `OllamaClient` exactly, default host `http://localhost:8080` per llama-server's default port) — makes T004-T pass (depends on T003-I)
- [ ] T005-T [US1] Write FAILING test: `LlamaCppClient` registered in `NODE_CLASS_MAPPINGS`/`NODE_DISPLAY_NAME_MAPPINGS`, in `tests/test_llamacpp.py`
- [ ] T005-I [US1] Register `LlamaCppClient` in `src/comfydv/__init__.py` — makes T005-T pass (depends on T004-I)
- [ ] T006-T [US1] Write FAILING test: `LlamaCppProvider` connection error surfaces a clear message (mirrors `OllamaProvider`'s `_post_json` connection-error contract), in `tests/test_llamacpp_provider.py` (witnesses `features/us1_connect_and_chat.feature` scenario "Unreachable llama.cpp server surfaces a clear error")
- [ ] T006-I [US1] Confirm `LlamaCppProvider.chat()` reuses the shared `_post_json` connection-error handling unchanged (likely no code change needed — verify, don't assume) — makes T006-T pass
**Checkpoint**: US1 fully functional and independently testable (MVP) — proves the adapter pattern for the chat path.
---
## Phase 4: User Story 2 — Get structured, validated output from llama.cpp (Priority: P1)
**Goal**: `structured_output=True` on `ChatCompletion` works identically against llama.cpp.
**Independent Test**: Enable `structured_output` with a schema, run against a llama.cpp-hosted model, confirm typed sockets populate and are never blank.
- [ ] T007-T [US2] Write FAILING test: `LlamaCppProvider.chat_structured()` builds `base_url=f"{host}/v1"` and delegates to the shared `comfydv._llm.chat.chat_structured()` helper unchanged, in `tests/test_llamacpp_provider.py` (witnesses `features/us2_structured_output.feature` scenario "Valid structured response exposes typed fields, same as Ollama")
- [ ] T007-I [US2] Implement `LlamaCppProvider.chat_structured()` in `src/comfydv/_llm/llamacpp_provider.py` — zero new structured-output logic, same call shape `OllamaProvider.chat_structured()` already makes — makes T007-T pass
- [ ] T008 [US2] No new test needed for the retry-then-fail path (witnesses `features/us2_structured_output.feature` scenario "Invalid response retries then fails clearly, same as Ollama") — already fully covered by `tests/test_llm_chat_structured.py`'s existing suite, since `LlamaCppProvider.chat_structured()` calls the identical shared helper `OllamaProvider` does; re-testing it here would duplicate coverage without adding confidence (same reasoning as the prerequisite epic's D5)
**Checkpoint**: US1 + US2 both independently functional — the chat surface is now backend-agnostic in practice, not just in name.
---
## Phase 5: User Story 3 — See and control which models are loaded on llama.cpp (Priority: P2)
**Goal**: `LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel` work against llama.cpp via `LlamaCppProvider`.
**Independent Test**: List models via `LLMModelSelector` wired to `LlamaCppClient`; load/unload one; confirm status changes, including `loading`/`downloading` states if triggered.
- [ ] T009-T [P] [US3] Write FAILING test: `LlamaCppProvider.list_models()` maps `GET /models`'s `data[].id`→`ModelInfo.name` and `data[].status.value`→`ModelInfo.status`, surfacing all five `ModelStatus` values without normalization (data-model.md), in `tests/test_llamacpp_provider.py` (witnesses `features/us3_model_lifecycle.feature` scenario "List models with full status vocabulary")
- [ ] T009-I [US3] Implement `LlamaCppProvider.list_models()` in `src/comfydv/_llm/llamacpp_provider.py` — makes T009-T pass
- [ ] T010-T [P] [US3] Write FAILING test: `LlamaCppProvider.load_model()`/`unload_model()` POST `{"model": id}` to `/models/load`/`/models/unload` and are idempotent, in `tests/test_llamacpp_provider.py` (witnesses `features/us3_model_lifecycle.feature` scenarios "Load a model into memory" and "Unload a model from memory")
- [ ] T010-I [US3] Implement `LlamaCppProvider.load_model()`/`unload_model()` in `src/comfydv/_llm/llamacpp_provider.py` — makes T010-T pass
- [ ] T011 [US3] No new node-layer tests needed — `LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel` are untouched by this epic (plan.md Structure Decision) and already have delegation-test coverage against a generic `_FakeProvider` in `tests/test_ollama.py`; that coverage is provider-agnostic by construction (FR-002), so it already proves these nodes work with `LlamaCppProvider` too, not just `OllamaProvider`
**Checkpoint**: US1 + US2 + US3 independently functional.
---
## Phase 6: User Story 4 — Swap from Ollama to llama.cpp without touching the rest of the workflow (Priority: P3)
**Goal**: Demonstrate/prove the adapter pattern's actual promise end-to-end.
**Independent Test**: Same workflow, only the connection node changes.
- [ ] T012-T [US4] Write FAILING test: a workflow-shaped test (client → `ChatCompletion` → `LLMModelSelector` → `LLMLoadModel` → `LLMUnloadModel`) runs identically whether `client` is an `OllamaProvider`-double or a `LlamaCppProvider`-double — i.e. no node branches on provider type, in `tests/test_llamacpp.py` (witnesses `features/us4_swap_backends.feature` scenario "Replacing only the connection node preserves the workflow")
- [ ] T012-I [US4] No implementation expected — this test should already pass given T003-T011 (it's a regression/integration proof, not new functionality); if it fails, that reveals a node secretly branching on provider type, which would be a real bug to fix, not a feature to add
**Checkpoint**: all four user stories independently functional; the adapter pattern is proven end-to-end, not just asserted.
---
## Phase 7: Polish & Cross-Cutting Concerns
- [ ] T013 [P] `uv run ruff check --fix && uv run ruff format` across `src/comfydv/_llm/llamacpp_provider.py`, `src/comfydv/llamacpp.py`, `src/comfydv/__init__.py`, and the new test files
- [ ] T014 [P] `uv run ty check` — resolve any new typing errors
- [ ] T015 Confirm Constitution Principle IV: `llamacpp_provider.py`/`llamacpp.py` import no `comfy`/`server` at module scope outside the existing guarded pattern
- [ ] T016 `beacon doctor --strict` — resolve any new findings (pre-existing/disclosed items from the prerequisite epic are not this feature's concern)
- [ ] T017 Manual/live smoke test against a real `llama-server` (router mode) if reachable in the environment — mirrors T-CUT-12's approach from the prerequisite epic (run live if possible, degrade to a documented walkthrough if not)
---
## Dependencies & Execution Order
### Phase Dependencies
- **Setup (Phase 1)**: no dependencies
- **Foundational (Phase 2)**: none — nothing blocks user story work
- **US1**: no dependency on other stories — genuinely the MVP
- **US2**: depends on US1's `LlamaCppProvider` skeleton existing (T003-I), but its own logic (T007) has no dependency on US1's chat() specifically
- **US3**: independent of US1/US2 except sharing `LlamaCppProvider`'s constructor (T003-I) — unlike the prerequisite epic's atomic cutover, there is no shared "client output type" migration risk here, since `LlamaCppClient` is a brand-new node, not a changed one
- **US4**: depends on US1–US3 all being done (it's a proof, not new functionality)
- **Polish**: depends on all four user stories
### Parallel Opportunities
- T002 can start immediately
- T009-T and T010-T can run in parallel (different methods, same file, no shared state)
- T013/T014 can run in parallel in Polish
---
## Implementation Strategy
### MVP First
1. Phase 1 (Setup, trivial) → Phase 3 (US1) → **STOP and validate US1 independently** against a live `llama-server`.
### Incremental Delivery
1. US1 → validate → basic chat parity with Ollama, on a second backend.
2. US2 → validate → structured-output parity — the shared mechanism holds.
3. US3 → validate → model-management parity, with richer status than Ollama can offer.
4. US4 → validate → the adapter pattern is proven, not just asserted.
5. Polish.
Unlike the prerequisite epic, **this decomposition genuinely holds** —
there is no shared "output type" migration forcing an atomic cutover, because
`LlamaCppClient` is new, not a change to an existing node. Each phase really
can land independently.