Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f16da579bc | ||
|
|
7a40d22230 | ||
|
|
42c614d5f8 | ||
|
|
a0f391f212 | ||
|
|
3c5136af70 | ||
|
|
97249ee281 | ||
|
|
1f138e5991 | ||
|
|
2724f535f5 | ||
|
|
a613006f5a | ||
|
|
56c7d4ca56 | ||
|
|
cf7cfeb6aa | ||
|
|
c9aefef915 | ||
|
|
f060698e7f | ||
|
|
0f4b2234db | ||
|
|
ecba52999d | ||
|
|
24119e3ab6 | ||
|
|
1f30c2ff2f | ||
|
|
ecfc7eca59 | ||
|
|
e534fac121 | ||
|
|
6c3c341913 | ||
|
|
8cb9a74462 | ||
|
|
5e678c2700 | ||
|
|
3854b27cba | ||
|
|
e432054b47 | ||
|
|
09fe013ff9 | ||
|
|
0f73f3c2cc | ||
|
|
24917be802 | ||
|
|
46c14a872d | ||
|
|
0cfe0d22ff | ||
|
|
bc7b33922b | ||
|
|
1028315d37 | ||
|
|
a2f2464fdb | ||
|
|
4a631cc8dc | ||
|
|
fd9054c03f | ||
|
|
23d96f1559 | ||
|
|
9ece5c0efa | ||
|
|
e225dcaaa3 | ||
|
|
1c72e0b2ea | ||
|
|
7e56f2e5bf | ||
|
|
f4a3390f13 | ||
|
|
a40c1ba373 | ||
|
|
eab86b0324 | ||
|
|
45e49a98eb | ||
|
|
87ec79be1b | ||
|
|
1dcef77a2b | ||
|
|
09a173e552 | ||
|
|
92da4fc117 | ||
|
|
1fe3bef663 | ||
|
|
3bd857f0ab | ||
|
|
834ae3e9e3 | ||
|
|
2df9e87838 | ||
|
|
27b26ec7a4 | ||
|
|
16b7e944f0 | ||
|
|
36e2cd123c | ||
|
|
1e9a8a3b6a | ||
|
|
d59a85dba6 | ||
|
|
15bb25efa8 | ||
|
|
282198a7df | ||
|
|
ac248e1e99 | ||
|
|
ea66ceee9d | ||
|
|
cf3f7fc174 | ||
|
|
5af502d07e | ||
|
|
d3ee35bfda | ||
|
|
27916f4fb0 | ||
|
|
689aa2a391 | ||
|
|
9ca65bc29a | ||
|
|
c32b6a7e42 | ||
|
|
3df38485d2 | ||
|
|
8e227790a9 | ||
|
|
6c2c906c4a | ||
|
|
1879ca2265 | ||
|
|
d90ae98fed | ||
|
|
fb111c0142 | ||
|
|
22dbddc8eb | ||
|
|
ef2464aebb | ||
|
|
736ffd1e77 | ||
|
|
1b1c7830d5 | ||
|
|
9d90f39854 | ||
|
|
fb0556fff3 | ||
|
|
b3374c88c4 | ||
|
|
50076d7c6b | ||
|
|
aa7ff8dae3 | ||
|
|
253569d483 | ||
|
|
40e851588d | ||
|
|
23f8f55493 | ||
|
|
d38cd5e794 | ||
|
|
e42913d587 | ||
|
|
335ff5e31a |
@@ -0,0 +1,19 @@
|
||||
/beacon:continue --full-auto keep delivering and - if none exist to complete - use the product and engineering triumvirate and identify the next epics and specs required. Use github issues to communicate with me and to track issues. Don't manufacture marginal work.
|
||||
|
||||
Use `gh issue create` to raise topics/questions/blockers async (and `gh issue comment` to post progress/track them); reserve `AskUserQuestion` for genuinely loop-blocking decisions only. Check for new comments, issues, and state changes as part of the loop assessment.
|
||||
|
||||
<separation_of_concerns>
|
||||
Use specialist and targeted subagents for deliver whilst you own the orchestration, planning, and review of their outputs.
|
||||
</separation_of_concerns>
|
||||
|
||||
<the_token_trap>
|
||||
Avoid the urge to produce content simply because it is rewarded. Perfection is when there is nothing left to remove. Keep in mind that les - invariably - is more.
|
||||
|
||||
Challenge yourself to deliver outcomes without flamboyance, excessive verbosity, or unnecessary codebase bloat.
|
||||
</the_token_trap>
|
||||
|
||||
**Note:** At the end of each iteration critically review the state of our README and documentation and then spawn a dedicated subagent to keep things fresh. Documentation and the README should never feel like an incremental read (referring to older versions) and should always read like a fresh "this is the the thing and how it works", not "the thing was this, we've done X, and now it's Y". Write this from a Product and UX perspective so that it is engaging and remember that humans are visual creatures. Screenshots are important both to explain and to engage.
|
||||
|
||||
DO NOT MAKE DESIGN DESIGNS. If the plan is underdeliverable raise an issue and abort. If you deviate but can justify and demonstrate it delivers the scope to the specification please document this clearly with rationale and update the necessary designs, decisions, and user facing documentation where required.
|
||||
|
||||
THREE (!) SENIOR DEVELOPERS WITH MORE THAN 30 YEARS EXPERIENCE WILL REVIEW YOUR PR. They will not accept shortcuts, monkeypatched tests, fake tests, hacky approaches or deviations from the plan. Produce Senior dev grade production code aligned to the plan at all times.
|
||||
@@ -29,5 +29,5 @@ jobs:
|
||||
|
||||
- name: Deploy docs with mike
|
||||
run: |
|
||||
VERSION=$(uv run python -c "import tomllib; d=tomllib.load(open('pyproject.toml','rb')); print(d['project']['version'])")
|
||||
VERSION=$(uv run python -c "from importlib.metadata import version; print(version('comfydv'))")
|
||||
uv run mike deploy --push --update-aliases "$VERSION" stable
|
||||
|
||||
@@ -1,3 +1,3 @@
|
||||
{
|
||||
"feature_directory": "specs/006-ollama-model-integration"
|
||||
"feature_directory": "specs/009-vlm-image-input"
|
||||
}
|
||||
|
||||
@@ -10,6 +10,17 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
||||
- BEACON framework bootstrap: problem statement, constitution, roadmap, architecture document
|
||||
- `CHANGELOG.md` (this file)
|
||||
- README: What-is-this, Install, and Quickstart sections
|
||||
- `LLMProvider` protocol (`comfydv._llm`) — a shared adapter boundary so ComfyUI LLM nodes work with any backend that implements it, starting with `OllamaProvider`. Structured output now goes through `pydantic-ai` (ADR-007), superseding the hand-rolled Ollama tool-calling approach.
|
||||
- **Chat Completion** now accepts an optional `image` input for vision-capable models (VLMs): wire a ComfyUI `IMAGE` and the connected model can describe or reason about it. Works identically on both backends (Ollama multimodal models; llama.cpp launched with `--mmproj`), and composes with structured output and multi-turn history. Images are carried on `Message.images` and translated to each backend's native shape (Ollama's flat `images` array, llama.cpp's OpenAI `image_url` parts, pydantic-ai `BinaryContent` on the structured path) — ADR-008, extending ADR-007's adapter pattern to a second input modality. Text-only workflows are unchanged when no image is wired.
|
||||
- **Ollama Option — Disable Thinking** node: turn off (or explicitly re-enable) a "thinking"-capable model's chain-of-thought reasoning. Chains into the same composable `OLLAMA_OPTIONS` socket every other `OllamaOption*` node uses, but works for both backends — each `LLMProvider` implementation pops the `think` key back out and translates it to its own wire shape (Ollama: a top-level `think` field; llama.cpp: `chat_template_kwargs`/`reasoning_effort` request-body fields, not live-verified — see ADR-010).
|
||||
|
||||
### Changed
|
||||
- **Breaking:** `OllamaChatCompletion` → `ChatCompletion`, `OllamaModelSelector` → `LLMModelSelector`, `OllamaLoadModel` → `LLMLoadModel`, `OllamaUnloadModel` → `LLMUnloadModel`, and the `OLLAMA_CLIENT` socket type → `LLM_CLIENT` — these nodes are now backend-generic. `OllamaClient` is unchanged by name but now outputs an `OllamaProvider` rather than a plain string; existing saved workflows using the old node/socket names need reconnecting (see `comfydv.ollama.MIGRATION_MAP` for the full old→new mapping).
|
||||
- `ChatCompletion`'s `structured_output=True` path now routes Ollama through Ollama's native `/api/chat` + `"format"` instead of the shared `pydantic-ai` OpenAI-compat path — Ollama's OpenAI-compatible endpoint was found to silently reload the model at its default context size on every call, discarding any `options` (e.g. `num_ctx`) override. `LlamaCppProvider` is unaffected and keeps the shared path, switched to `pydantic-ai`'s `NativeOutput` mode (ADR-009).
|
||||
|
||||
### Fixed
|
||||
- `structured_output=True` requests could fail validation ("token limit exceeded before any response was generated") against "thinking"-capable models, which spent their whole token budget on chain-of-thought reasoning before ever producing the structured response (ADR-009).
|
||||
- A non-required structured-output schema field rejected an explicit `null` value from the model (only an *omitted* field was tolerated), even though models routinely emit explicit `null` for absent optional fields.
|
||||
|
||||
## [0.1.0] — 2026-06-01
|
||||
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
<!-- SPECKIT START -->
|
||||
For additional context about technologies to be used, project structure,
|
||||
shell commands, and other important information, read the current plan
|
||||
at specs/006-ollama-model-integration/plan.md
|
||||
at specs/009-vlm-image-input/plan.md
|
||||
<!-- SPECKIT END -->
|
||||
|
||||
@@ -1,28 +1,37 @@
|
||||
# comfydv
|
||||
|
||||
A collection of workflow efficiency and quality-of-life nodes built out of necessity for personal ComfyUI use.
|
||||
**Quality-of-life nodes for ComfyUI, built to disappear into your workflow.**
|
||||
|
||||
## What is this?
|
||||
`comfydv` fills the gaps ComfyUI's built-in library leaves on the table: string templates that build their own sockets as you type, seed-controlled randomisation, graceful mid-queue interruption, and a local-LLM integration that doesn't care whether you're running Ollama or llama.cpp. No Python required — install it, drop the nodes on your canvas, wire them up.
|
||||
|
||||
`comfydv` fills gaps in ComfyUI's built-in node library: dynamic string formatting, seed-controlled random selection, graceful workflow interruption, and Ollama LLM integration. Install it once and connect the nodes like any other — no Python knowledge required.
|
||||

|
||||
|
||||
## What is comfydv?
|
||||
|
||||
A small, focused ComfyUI utility pack. It exists because:
|
||||
|
||||
- **It reads your intent, not just your syntax.** Format String detects `{variables}` in a template and adds/removes input sockets live, as you type — no manual socket wrangling.
|
||||
- **One LLM integration, any local backend.** Wire a Chat Completion node once; swap between Ollama and llama.cpp by changing a single upstream client node. Structured output, multi-turn history, and model load/unload work identically on both.
|
||||
- **It fails politely.** Circuit Breaker halts a queue run cleanly instead of throwing a stack trace at you; a disconnected LLM server gets a specific, actionable error instead of a silent empty dropdown.
|
||||
- **Small, tested, boring in the best way.** Every node is unit-tested and the local-LLM nodes are verified against real running servers, not just mocks.
|
||||
|
||||
## What's inside
|
||||
|
||||
| Node | What it does |
|
||||
|------|-------------|
|
||||
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
|
||||
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
|
||||
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
|
||||
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the host URL through the graph as an `OLLAMA_CLIENT` socket. |
|
||||
| **Ollama Model Selector** | Fetches the live model list from Ollama and presents it as a dropdown. Outputs the selected model name. |
|
||||
| **Ollama Load Model** | Loads a model into Ollama's memory using `/api/generate` with `keep_alive=-1`. |
|
||||
| **Ollama Unload Model** | Evicts a model from Ollama's memory using `/api/generate` with `keep_alive=0`. |
|
||||
| **Ollama Chat Completion** | Sends a prompt (and optional conversation history) to Ollama `/api/chat`. Response and history are shown inline in the node body and available as output sockets. |
|
||||
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
|
||||
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
|
||||
| **Ollama History Length** | Returns the number of messages in an `OLLAMA_HISTORY` list as an integer. |
|
||||
|------|---------------|
|
||||
| **Format String** | Renders a Python f-string or Jinja2 template. Sockets appear and disappear automatically as you type variables. |
|
||||
| **Random Choice** | Accepts any number of typed inputs and returns one at random, with a seed for reproducibility. |
|
||||
| **Circuit Breaker** | Halts the current queue run gracefully — no crash, just a clean stop — when a condition isn't met. |
|
||||
| **Ollama Client** / **LlamaCpp Client** | Configure a connection to a local Ollama or llama.cpp server. Both emit the same `LLM_CLIENT` socket — every node below works with either. |
|
||||
| **LLM Model Selector** | Live dropdown of models available on the connected server. |
|
||||
| **LLM Load Model** / **LLM Unload Model** | Explicit VRAM management — pin a model in memory before inference, evict it after. |
|
||||
| **Chat Completion** | Send a prompt (optionally with history) to the connected server; response shown inline and as an output socket. |
|
||||
| **Ollama Option — \*** | Seven composable parameter nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into Chat Completion's `options` input. |
|
||||
| **Ollama Debug History** / **Ollama History Length** | Inspect an `OLLAMA_HISTORY` conversation list — pretty-print it or count its messages. |
|
||||
|
||||
## Install
|
||||
|
||||
**Via ComfyUI Manager** (recommended): search for `comfydv` and click Install.
|
||||
**Via ComfyUI Manager** (recommended): search for `comfydv`, click Install.
|
||||
|
||||
**Manual:**
|
||||
|
||||
@@ -31,124 +40,147 @@ cd /path/to/ComfyUI/custom_nodes
|
||||
git clone https://github.com/darth-veitcher/comfydv.git
|
||||
```
|
||||
|
||||
Restart ComfyUI. The nodes appear under the **dv/** and **dv/ollama** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`) are installed automatically via `requirements.txt`.
|
||||
Restart ComfyUI. Nodes appear under **dv/**, **dv/ollama**, and **dv/llamacpp** in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) install automatically via `requirements.txt`.
|
||||
|
||||
For Ollama nodes: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`) before using the Ollama nodes.
|
||||
For local-LLM nodes, bring your own backend:
|
||||
|
||||
- **Ollama** — [install Ollama](https://ollama.com/download), pull a model (`ollama pull qwen2.5:latest`).
|
||||
- **llama.cpp** — [build/install `llama-server`](https://github.com/ggml-org/llama.cpp), launch it in [router mode](#llamacpp).
|
||||
|
||||
## Quickstart
|
||||
|
||||
1. Install via ComfyUI Manager (search `comfydv`) or clone manually into `custom_nodes/`.
|
||||
2. Right-click the canvas → Add Node → **dv/** to find Format String, Random Choice, and Circuit Breaker.
|
||||
3. For Ollama nodes: start Ollama (`ollama serve`), pull a model (`ollama pull qwen2.5:latest`), then add nodes from **dv/ollama/**.
|
||||
1. Right-click the canvas → Add Node → **dv/** for Format String, Random Choice, and Circuit Breaker.
|
||||
2. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
|
||||
|
||||
## Documentation
|
||||
|
||||
Full documentation: [darth-veitcher.github.io/comfydv](https://darth-veitcher.github.io/comfydv/stable/)
|
||||
Full documentation, including every node's inputs/outputs: **[darth-veitcher.github.io/comfydv](https://darth-veitcher.github.io/comfydv/stable/)**
|
||||
|
||||
---
|
||||
|
||||
## Format String
|
||||
|
||||
Formats text from a Python f-string or Jinja2 template. As you type the template, input sockets appear and disappear automatically — one per variable detected.
|
||||
|
||||
### Python f-strings
|
||||
|
||||
Type `{variable_name}` and a socket appears. Wire it to any string output in your workflow.
|
||||
Type a template, get sockets. `{variable_name}` in f-string mode, `{{ variable_name }}` in Jinja2 mode — either way, comfydv watches what you type and keeps the node's inputs in sync automatically.
|
||||
|
||||

|
||||
|
||||
Outputs are always in a stable order:
|
||||
|
||||
| Output | Content |
|
||||
|--------|---------|
|
||||
| `formatted_string` | The rendered result |
|
||||
| `saved_file_path` | Path written to disk (if `save_path` is set) |
|
||||
| `<var>` … | Pass-through of each input value, for easy chaining |
|
||||
| `saved_file_path` | Where it was written, if `save_path` is set |
|
||||
| `<var>` … | Each input passed through unchanged, for easy chaining |
|
||||
|
||||
### Jinja2 templates
|
||||
|
||||
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`, …), conditionals (`{% if %}…{% endif %}`), and loops.
|
||||
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`), conditionals, and loops:
|
||||
|
||||

|
||||
|
||||
Variables detected in `{{ }}` expressions become input sockets exactly as in Simple mode. See the [Jinja2 documentation](https://jinja.palletsprojects.com/en/latest/) for the full filter/test reference.
|
||||
|
||||
---
|
||||
|
||||
## Random Choice
|
||||
|
||||
Connect any number of inputs of the same type. Each run picks one at random. Set `seed` for reproducibility.
|
||||
Wire in any number of same-typed inputs — images, strings, conditioning, anything ComfyUI can carry over a socket — and get one back at random.
|
||||
|
||||

|
||||
|
||||
- Accepts any ComfyUI type (STRING, IMAGE, CONDITIONING, …)
|
||||
- Add as many inputs as you like; unused slots are removed automatically when disconnected
|
||||
- `seed = 0` randomises on every run; any other value locks the selection
|
||||
`seed = 0` randomises every run; any other value locks the selection. Unused input slots vanish automatically when you disconnect them.
|
||||
|
||||
---
|
||||
|
||||
## Circuit Breaker
|
||||
|
||||
Stops the queue gracefully when a condition isn't met — no crash, no error, just a clean halt.
|
||||
Stop a queue run cleanly when a condition isn't met, instead of letting a downstream node crash on bad input.
|
||||
|
||||

|
||||
|
||||
Wire an image (or any trigger) into `trigger` and a boolean into `status`. When `status` is **false** the node raises `InterruptProcessingException`, which tells ComfyUI to stop the current run cleanly. When `status` is **true** the image passes through unchanged.
|
||||
Wire a trigger (an image, or anything) into `trigger` and a boolean into `status`. `status = false` raises `InterruptProcessingException` — ComfyUI stops the run without an error dialog. `status = true` passes the trigger straight through.
|
||||
|
||||
Typical use: skip an expensive upscale step when a quality-check node says the draft is already good enough.
|
||||
Typical use: skip an expensive upscale pass when an upstream quality-check node says the draft's already good enough.
|
||||
|
||||
---
|
||||
|
||||
## Ollama
|
||||
## Local LLMs
|
||||
|
||||
14 nodes for integrating a local Ollama LLM into your ComfyUI workflow. The host URL is configured once in **Ollama Client** and threaded through the graph — all downstream nodes receive it via the `OLLAMA_CLIENT` socket.
|
||||
One set of nodes, two interchangeable backends. Configure a connection once with **Ollama Client** or **LlamaCpp Client** — both output the same `LLM_CLIENT` socket — and every downstream node (model selection, load/unload, chat, structured output, multi-turn history) works exactly the same way regardless of which one you picked. Swapping backends means rewiring one node, not rebuilding your graph.
|
||||
|
||||
### Ollama Client node
|
||||
|
||||
Configure the server address once; all downstream Ollama nodes inherit it automatically.
|
||||
### Connect and chat
|
||||
|
||||

|
||||
|
||||
### Model lifecycle (load and unload)
|
||||
1. **Ollama Client** (default `http://localhost:11434`) or **LlamaCpp Client** (default `http://localhost:8080`) — set the host.
|
||||
2. **LLM Model Selector** — pick a model from the live dropdown, or wire a model name straight into Chat Completion.
|
||||
3. **Chat Completion** — wire in client, model, and prompt. The response renders inline in the node and is also available as an output socket.
|
||||
|
||||
On memory-constrained machines and single-GPU setups, explicitly loading and unloading the model before and after inference is critical. **Ollama Load Model** pins the model into VRAM (`keep_alive=-1`); **Ollama Unload Model** evicts it immediately (`keep_alive=0`), freeing memory for image generation or other models.
|
||||

|
||||
|
||||

|
||||
A complete graph looks like this:
|
||||
|
||||
The correct chain is **Load → Chat → Unload**, enforced through data dependencies:
|
||||

|
||||
|
||||
1. Wire `OllamaLoadModel.model_name` → `OllamaChatCompletion.model`. This creates the data dependency that guarantees Load runs before Chat and passes the model name into the Chat node's plain-string `model` input.
|
||||
2. Wire `OllamaChatCompletion.model_name` → `OllamaUnloadModel.model`. This guarantees Unload runs after Chat completes.
|
||||
3. Optionally wire `OllamaChatCompletion.response` → `OllamaUnloadModel.passthrough` — Unload returns the response unchanged so the rest of your workflow can still consume it.
|
||||
### Describing images (vision)
|
||||
|
||||
### Minimal chat workflow
|
||||
**Chat Completion** has an optional **image** input. Wire any `IMAGE` into it and, with a vision-capable model loaded, the model can describe or reason about the picture — captioning, visual Q&A, reading text in an image, whatever the model supports.
|
||||
|
||||
1. **Ollama Client** → set host (default `http://localhost:11434`)
|
||||
2. **Ollama Model Selector** → pick a model from the live dropdown (or type/wire a model name directly into Chat Completion's `model` input)
|
||||
3. **Ollama Chat Completion** → wire client + model + prompt; the response appears inline in the node body and is also available as an output socket
|
||||
- **Ollama** — use a multimodal model (e.g. a llava-class model).
|
||||
- **llama.cpp** — launch `llama-server` with a multimodal projector: `--mmproj <projector.gguf>` alongside the model.
|
||||
|
||||

|
||||
Image input works the same on both backends — same node, same wiring — and composes with everything else: structured output (schema-validated fields pulled straight from the image), multi-turn history, and the option nodes. A batch of images is sent as multiple images on the turn. Leave the input unwired and Chat Completion behaves exactly as before, text only.
|
||||
|
||||
Wire multiple nodes together for a complete end-to-end workflow:
|
||||
### Manual memory management
|
||||
|
||||

|
||||
Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline.
|
||||
|
||||
### Option nodes
|
||||

|
||||
|
||||
Chain any combination of **Ollama Option —** nodes before Chat Completion to override inference parameters:
|
||||
The **Load → Chat → Unload** chain is enforced by data dependencies, not by convention:
|
||||
|
||||
| Option node | Ollama param |
|
||||
|-------------|-------------|
|
||||
1. `LLMLoadModel.model_name` → `ChatCompletion.model` — guarantees Load runs before Chat, and feeds the model name straight in.
|
||||
2. `ChatCompletion.model_name` → `LLMUnloadModel.model` — guarantees Unload runs after Chat completes.
|
||||
3. *(Optional)* `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged, so the rest of your workflow can still consume it.
|
||||
|
||||
### Tuning generation
|
||||
|
||||
Chain any combination of **Ollama Option —** nodes ahead of Chat Completion to override inference parameters:
|
||||
|
||||

|
||||
|
||||
| Option node | Parameter |
|
||||
|-------------|-----------|
|
||||
| Temperature | `temperature` |
|
||||
| Seed | `seed` |
|
||||
| Max Tokens | `num_predict` |
|
||||
| Top P | `top_p` |
|
||||
| Top K | `top_k` |
|
||||
| Repeat Penalty | `repeat_penalty` |
|
||||
| Extra Body | arbitrary JSON merged into options |
|
||||
|
||||

|
||||
| Extra Body | arbitrary JSON, merged into `options` |
|
||||
|
||||
### Multi-turn conversations
|
||||
|
||||
`OLLAMA_HISTORY` flows out of Chat Completion as a list of `{"role", "content"}` dicts. Wire it back into the next Chat Completion for multi-turn conversations, or inspect it with **Ollama Debug History** / **Ollama History Length**.
|
||||
`OLLAMA_HISTORY` flows out of Chat Completion as a `{"role", "content"}` list. Feed it back into the next Chat Completion call for multi-turn context, or inspect it with **Ollama Debug History** / **Ollama History Length**.
|
||||
|
||||
### llama.cpp
|
||||
|
||||
Everything above works unchanged against llama.cpp — swap in a **LlamaCpp Client** and the rest of the graph doesn't know the difference. The one thing llama.cpp needs that Ollama doesn't: **router mode**, a directory of models rather than a single `-m model.gguf`:
|
||||
|
||||
```bash
|
||||
llama-server --models-dir ./models -c 8192
|
||||
```
|
||||
|
||||
In exchange, router mode gives comfydv a richer live status than Ollama can report — `loading` and `downloading`, not just loaded/unloaded — plus the same explicit load/unload primitives Ollama's nodes already use.
|
||||
|
||||
### Upgrading a workflow saved before this rename
|
||||
|
||||
Nodes were renamed once, to make them backend-generic (`OllamaChatCompletion` → `ChatCompletion`, etc.). If ComfyUI reports old node types as missing when you reopen a saved workflow, reconnect using this table — behavior is unchanged, only the names are:
|
||||
|
||||
| Old | New |
|
||||
|-----|-----|
|
||||
| `OllamaChatCompletion` | `ChatCompletion` |
|
||||
| `OllamaModelSelector` | `LLMModelSelector` |
|
||||
| `OllamaLoadModel` | `LLMLoadModel` |
|
||||
| `OllamaUnloadModel` | `LLMUnloadModel` |
|
||||
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
|
||||
|
||||
`OllamaClient` kept its name — delete and re-add any node showing as missing, then rewire it to the same `OllamaClient` node.
|
||||
|
||||
---
|
||||
|
||||
## License
|
||||
|
||||
[AGPL-3.0](LICENSE)
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# comfydv — Roadmap
|
||||
|
||||
<!-- generated by beacon roadmap export — 2026-06-28 -->
|
||||
<!-- generated by beacon roadmap export — 2026-07-11 -->
|
||||
|
||||
> comfydv is a small, high-quality ComfyUI utility pack that fills the gaps the core node library leaves: composable string formatting, seed-controlled randomisation, and workflow flow-control. Winning looks like: every node is well-tested, installs in one step, produces no surprises in production workflows, and is documented well enough that a non-programmer ComfyUI user can connect it without reading source code.
|
||||
|
||||
@@ -13,14 +13,14 @@ gantt
|
||||
excludes weekends
|
||||
|
||||
section Active
|
||||
llama.cpp Model Integration :active, llamacpp-integration, 2026-07-11, 7d
|
||||
ComfyUI UX Polish & Manager Compatibility :active, ux-and-install, 2026-06-28, 21d
|
||||
|
||||
section Planned
|
||||
Ollama Model Integration :ollama-integration, 2026-06-28, 7d
|
||||
|
||||
section Done
|
||||
BEACON Bootstrap :done, beacon-bootstrap, 2026-06-28, 7d
|
||||
Logging Modernisation :done, logging-modernisation, 2026-06-28, 7d
|
||||
BEACON Bootstrap :done, beacon-bootstrap, 2026-07-11, 7d
|
||||
LLM Provider Abstraction :done, llm-provider-abstraction, 2026-07-11, 7d
|
||||
Logging Modernisation :done, logging-modernisation, 2026-07-11, 7d
|
||||
Ollama Model Integration :done, ollama-integration, 2026-07-11, 7d
|
||||
|
||||
```
|
||||
|
||||
@@ -28,10 +28,12 @@ gantt
|
||||
|
||||
| Epic | Title | Status | Specs | Fidelity |
|
||||
|---|---|---|---|---|
|
||||
| [ollama-integration](project-management/Roadmap/epics/ollama-integration.md) | Ollama Model Integration | Planning | — | S? A? T:- |
|
||||
| [llamacpp-integration](project-management/Roadmap/epics/llamacpp-integration.md) | llama.cpp Model Integration | Active | 1/1 shipped | S+ A+ T:96% |
|
||||
| [ux-and-install](project-management/Roadmap/epics/ux-and-install.md) | ComfyUI UX Polish & Manager Compatibility | Active | 1/4 shipped | S+ A+ T:100% |
|
||||
| [beacon-bootstrap](project-management/Roadmap/epics/archive/beacon-bootstrap.md) | BEACON Bootstrap | Done | — | S? A? T:- |
|
||||
| [llm-provider-abstraction](project-management/Roadmap/epics/archive/llm-provider-abstraction.md) | LLM Provider Abstraction | Done | 1/1 shipped | S+ A+ T:58% |
|
||||
| [logging-modernisation](project-management/Roadmap/epics/archive/logging-modernisation.md) | Logging Modernisation | Done | 1/1 shipped | S+ A+ T:100% |
|
||||
| [ollama-integration](project-management/Roadmap/epics/archive/ollama-integration.md) | Ollama Model Integration | Done | 1/1 shipped | S+ A+ T:100% |
|
||||
|
||||
_Fidelity: `S+/S?` = has specs / none · `A+/A?` = has ADRs / none · `T:N%` = task completion_
|
||||
|
||||
@@ -46,3 +48,6 @@ _No active bullets._
|
||||
| [ADR-001](../../../ADRs/ADR-001-stdlib-logging-over-console-libraries.md) | Use stdlib logging instead of colored console output libraries | Accepted |
|
||||
| [ADR-002](../../../ADRs/ADR-002-nullhandler-pattern-for-library-loggers.md) | NullHandler pattern for the comfydv package root logger | Accepted |
|
||||
| [ADR-003](../../ADRs/ADR-003-requirements-txt-authoring-policy.md) | Hand-authored requirements.txt as a curated subset of pyproject.toml | Accepted |
|
||||
| [ADR-004](project-management/ADRs/ADR-004-aiohttp-over-httpx-for-ollama.md) | Use aiohttp for Ollama HTTP communication instead of httpx | Accepted |
|
||||
| [ADR-005](project-management/ADRs/ADR-005-ollama-host-config-via-client-node.md) | Ollama host configuration via OllamaClient node and OLLAMA_CLIENT socket type | Accepted |
|
||||
| [ADR-007](project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md) | LLMProvider adapter pattern shared across Ollama and llama.cpp | Accepted |
|
||||
|
||||
@@ -10,10 +10,10 @@
|
||||
"Random Choice",
|
||||
"Circuit Breaker",
|
||||
"Ollama Client",
|
||||
"Ollama Model Selector",
|
||||
"Ollama Load Model",
|
||||
"Ollama Unload Model",
|
||||
"Ollama Chat Completion",
|
||||
"LLM Model Selector",
|
||||
"LLM Load Model",
|
||||
"LLM Unload Model",
|
||||
"Chat Completion",
|
||||
"Ollama Option — Temperature",
|
||||
"Ollama Option — Seed",
|
||||
"Ollama Option — Max Tokens",
|
||||
|
||||
@@ -17,5 +17,7 @@ services:
|
||||
- sh
|
||||
- -c
|
||||
- |
|
||||
git clone https://github.com/Comfy-Org/ComfyUI-Manager.git /app/ComfyUI/custom_nodes/comfyui-manager
|
||||
pip install -q -r /app/ComfyUI/custom_nodes/comfyui-manager/requirements.txt
|
||||
pip install -q -r /app/ComfyUI/custom_nodes/comfydv/requirements.txt
|
||||
exec python main.py --cpu --listen 0.0.0.0 --port 8188
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
FROM python:3.11-slim
|
||||
FROM python:3.13-slim
|
||||
|
||||
WORKDIR /app
|
||||
|
||||
@@ -10,7 +10,7 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# Clone ComfyUI (pinned tag for reproducibility; bump manually)
|
||||
ARG COMFYUI_VERSION=v0.3.44
|
||||
ARG COMFYUI_VERSION=v0.27.0
|
||||
RUN git clone --depth 1 --branch ${COMFYUI_VERSION} \
|
||||
https://github.com/comfyanonymous/ComfyUI.git /app/ComfyUI
|
||||
|
||||
|
||||
|
Before Width: | Height: | Size: 12 KiB After Width: | Height: | Size: 12 KiB |
|
Before Width: | Height: | Size: 32 KiB After Width: | Height: | Size: 33 KiB |
|
Before Width: | Height: | Size: 40 KiB After Width: | Height: | Size: 41 KiB |
|
After Width: | Height: | Size: 13 KiB |
|
After Width: | Height: | Size: 37 KiB |
|
Before Width: | Height: | Size: 38 KiB After Width: | Height: | Size: 55 KiB |
|
Before Width: | Height: | Size: 11 KiB After Width: | Height: | Size: 12 KiB |
|
Before Width: | Height: | Size: 50 KiB After Width: | Height: | Size: 55 KiB |
|
Before Width: | Height: | Size: 40 KiB After Width: | Height: | Size: 41 KiB |
|
Before Width: | Height: | Size: 58 KiB After Width: | Height: | Size: 64 KiB |
|
Before Width: | Height: | Size: 19 KiB After Width: | Height: | Size: 18 KiB |
|
After Width: | Height: | Size: 68 KiB |
@@ -1,24 +1,37 @@
|
||||
# comfydv
|
||||
|
||||
A collection of workflow efficiency and quality-of-life nodes built out of necessity for personal ComfyUI use.
|
||||
**Quality-of-life nodes for ComfyUI, built to disappear into your workflow.**
|
||||
|
||||
`comfydv` fills the gaps ComfyUI's built-in library leaves on the table: string templates that build their own sockets as you type, seed-controlled randomisation, graceful mid-queue interruption, and a local-LLM integration that doesn't care whether you're running Ollama or llama.cpp. No Python required — install it, drop the nodes on your canvas, wire them up.
|
||||
|
||||

|
||||
|
||||
## What is comfydv?
|
||||
|
||||
A small, focused ComfyUI utility pack. It exists because:
|
||||
|
||||
- **It reads your intent, not just your syntax.** Format String detects `{variables}` in a template and adds/removes input sockets live, as you type — no manual socket wrangling.
|
||||
- **One LLM integration, any local backend.** Wire a Chat Completion node once; swap between Ollama and llama.cpp by changing a single upstream client node. Structured output, multi-turn history, and model load/unload work identically on both.
|
||||
- **It fails politely.** Circuit Breaker halts a queue run cleanly instead of throwing a stack trace at you; a disconnected LLM server gets a specific, actionable error instead of a silent empty dropdown.
|
||||
- **Small, tested, boring in the best way.** Every node is unit-tested and the local-LLM nodes are verified against real running servers, not just mocks.
|
||||
|
||||
## What's inside
|
||||
|
||||
| Node | What it does |
|
||||
|------|-------------|
|
||||
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
|
||||
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
|
||||
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
|
||||
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the host URL through the graph as an `OLLAMA_CLIENT` socket. |
|
||||
| **Ollama Model Selector** | Fetches the live model list from Ollama and presents it as a dropdown. Outputs the selected model name. |
|
||||
| **Ollama Load Model** | Loads a model into Ollama's memory using `/api/generate` with `keep_alive=-1`. |
|
||||
| **Ollama Unload Model** | Evicts a model from Ollama's memory using `/api/generate` with `keep_alive=0`. |
|
||||
| **Ollama Chat Completion** | Sends a prompt (and optional conversation history) to Ollama `/api/chat`. Response and history are shown inline in the node body and available as output sockets. |
|
||||
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
|
||||
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
|
||||
| **Ollama History Length** | Returns the number of messages in an `OLLAMA_HISTORY` list as an integer. |
|
||||
|------|---------------|
|
||||
| **Format String** | Renders a Python f-string or Jinja2 template. Sockets appear and disappear automatically as you type variables. |
|
||||
| **Random Choice** | Accepts any number of typed inputs and returns one at random, with a seed for reproducibility. |
|
||||
| **Circuit Breaker** | Halts the current queue run gracefully — no crash, just a clean stop — when a condition isn't met. |
|
||||
| **Ollama Client** / **LlamaCpp Client** | Configure a connection to a local Ollama or llama.cpp server. Both emit the same `LLM_CLIENT` socket — every node below works with either. |
|
||||
| **LLM Model Selector** | Live dropdown of models available on the connected server. |
|
||||
| **LLM Load Model** / **LLM Unload Model** | Explicit VRAM management — pin a model in memory before inference, evict it after. |
|
||||
| **Chat Completion** | Send a prompt (optionally with history) to the connected server; response shown inline and as an output socket. |
|
||||
| **Ollama Option — \*** | Seven composable parameter nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into Chat Completion's `options` input. |
|
||||
| **Ollama Debug History** / **Ollama History Length** | Inspect an `OLLAMA_HISTORY` conversation list — pretty-print it or count its messages. |
|
||||
|
||||
## Install
|
||||
|
||||
**Via ComfyUI Manager** (recommended): search for `comfydv` and click Install.
|
||||
**Via ComfyUI Manager** (recommended): search for `comfydv`, click Install.
|
||||
|
||||
**Manual:**
|
||||
|
||||
@@ -27,112 +40,130 @@ cd /path/to/ComfyUI/custom_nodes
|
||||
git clone https://github.com/darth-veitcher/comfydv.git
|
||||
```
|
||||
|
||||
Restart ComfyUI. The nodes appear under the **dv/** and **dv/ollama** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`) are installed automatically via `requirements.txt`.
|
||||
Restart ComfyUI. Nodes appear under **dv/**, **dv/ollama**, and **dv/llamacpp** in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) install automatically via `requirements.txt`.
|
||||
|
||||
For Ollama nodes: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`) before using the Ollama nodes.
|
||||
For local-LLM nodes, bring your own backend:
|
||||
|
||||
- **Ollama** — [install Ollama](https://ollama.com/download), pull a model (`ollama pull qwen2.5:latest`).
|
||||
- **llama.cpp** — [build/install `llama-server`](https://github.com/ggml-org/llama.cpp), launch it in [router mode](#llamacpp).
|
||||
|
||||
## Quickstart
|
||||
|
||||
1. Right-click the canvas → Add Node → **dv/** for Format String, Random Choice, and Circuit Breaker.
|
||||
2. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
|
||||
|
||||
---
|
||||
|
||||
## Format String
|
||||
|
||||
Formats text from a Python f-string or Jinja2 template. As you type the template, input sockets appear and disappear automatically — one per variable detected.
|
||||
|
||||
### Python f-strings
|
||||
|
||||
Type `{variable_name}` and a socket appears. Wire it to any string output in your workflow.
|
||||
Type a template, get sockets. `{variable_name}` in f-string mode, `{{ variable_name }}` in Jinja2 mode — either way, comfydv watches what you type and keeps the node's inputs in sync automatically.
|
||||
|
||||

|
||||
|
||||
| Output | Content |
|
||||
|--------|---------|
|
||||
| `formatted_string` | The rendered result |
|
||||
| `saved_file_path` | Path written to disk (if `save_path` is set) |
|
||||
| `<var>` … | Pass-through of each input value, for easy chaining |
|
||||
| `saved_file_path` | Where it was written, if `save_path` is set |
|
||||
| `<var>` … | Each input passed through unchanged, for easy chaining |
|
||||
|
||||
### Jinja2 templates
|
||||
|
||||
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`, …), conditionals (`{% if %}…{% endif %}`), and loops.
|
||||
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`), conditionals, and loops:
|
||||
|
||||

|
||||
|
||||
Variables detected in `{{ }}` expressions become input sockets exactly as in Simple mode. See the [Jinja2 documentation](https://jinja.palletsprojects.com/en/latest/) for the full filter/test reference.
|
||||
|
||||
---
|
||||
|
||||
## Random Choice
|
||||
|
||||
Connect any number of inputs of the same type. Each run picks one at random. Set `seed` for reproducibility.
|
||||
Wire in any number of same-typed inputs — images, strings, conditioning, anything ComfyUI can carry over a socket — and get one back at random.
|
||||
|
||||

|
||||
|
||||
- Accepts any ComfyUI type (STRING, IMAGE, CONDITIONING, …)
|
||||
- Add as many inputs as you like; unused slots are removed automatically when disconnected
|
||||
- `seed = 0` randomises on every run; any other value locks the selection
|
||||
`seed = 0` randomises every run; any other value locks the selection. Unused input slots vanish automatically when you disconnect them.
|
||||
|
||||
---
|
||||
|
||||
## Circuit Breaker
|
||||
|
||||
Stops the queue gracefully when a condition isn't met — no crash, no error, just a clean halt.
|
||||
Stop a queue run cleanly when a condition isn't met, instead of letting a downstream node crash on bad input.
|
||||
|
||||

|
||||
|
||||
Wire an image (or any trigger) into `trigger` and a boolean into `status`. When `status` is **false** the node raises `InterruptProcessingException`, which tells ComfyUI to stop the current run cleanly. When `status` is **true** the image passes through unchanged.
|
||||
Wire a trigger (an image, or anything) into `trigger` and a boolean into `status`. `status = false` raises `InterruptProcessingException` — ComfyUI stops the run without an error dialog. `status = true` passes the trigger straight through.
|
||||
|
||||
Typical use: skip an expensive upscale step when a quality-check node says the draft is already good enough.
|
||||
Typical use: skip an expensive upscale pass when an upstream quality-check node says the draft's already good enough.
|
||||
|
||||
---
|
||||
|
||||
## Ollama
|
||||
## Local LLMs
|
||||
|
||||
14 nodes for integrating a local Ollama LLM into your ComfyUI workflow. The host URL is configured once in **Ollama Client** and threaded through the graph — all downstream nodes receive it via the `OLLAMA_CLIENT` socket.
|
||||
One set of nodes, two interchangeable backends. Configure a connection once with **Ollama Client** or **LlamaCpp Client** — both output the same `LLM_CLIENT` socket — and every downstream node (model selection, load/unload, chat, structured output, multi-turn history) works exactly the same way regardless of which one you picked. Swapping backends means rewiring one node, not rebuilding your graph.
|
||||
|
||||
### Ollama Client node
|
||||
|
||||
Configure the server address once; all downstream Ollama nodes inherit it automatically.
|
||||
### Connect and chat
|
||||
|
||||

|
||||
|
||||
### Model lifecycle (load and unload)
|
||||
1. **Ollama Client** (default `http://localhost:11434`) or **LlamaCpp Client** (default `http://localhost:8080`) — set the host.
|
||||
2. **LLM Model Selector** — pick a model from the live dropdown, or wire a model name straight into Chat Completion.
|
||||
3. **Chat Completion** — wire in client, model, and prompt. The response renders inline in the node and is also available as an output socket.
|
||||
|
||||
On memory-constrained machines and single-GPU setups, explicitly loading and unloading the model before and after inference is critical. **Ollama Load Model** pins the model into VRAM (`keep_alive=-1`); **Ollama Unload Model** evicts it immediately (`keep_alive=0`), freeing memory for image generation or other models.
|
||||

|
||||
|
||||

|
||||
A complete graph looks like this:
|
||||
|
||||
The correct chain is **Load → Chat → Unload**, enforced through data dependencies:
|
||||

|
||||
|
||||
1. Wire `OllamaLoadModel.model_name` → `OllamaChatCompletion.model`. This creates the data dependency that guarantees Load runs before Chat and passes the model name into the Chat node's plain-string `model` input.
|
||||
2. Wire `OllamaChatCompletion.model_name` → `OllamaUnloadModel.model`. This guarantees Unload runs after Chat completes.
|
||||
3. Optionally wire `OllamaChatCompletion.response` → `OllamaUnloadModel.passthrough` — Unload returns the response unchanged so the rest of your workflow can still consume it.
|
||||
### Manual memory management
|
||||
|
||||
### Minimal chat workflow
|
||||
Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline.
|
||||
|
||||
1. **Ollama Client** → set host (default `http://localhost:11434`)
|
||||
2. **Ollama Model Selector** → pick a model from the live dropdown (or type/wire a model name directly into Chat Completion's `model` input)
|
||||
3. **Ollama Chat Completion** → wire client + model + prompt; the response appears inline in the node body and is also available as an output socket
|
||||

|
||||
|
||||

|
||||
The **Load → Chat → Unload** chain is enforced by data dependencies, not by convention:
|
||||
|
||||
Wire multiple nodes together for a complete end-to-end workflow:
|
||||
1. `LLMLoadModel.model_name` → `ChatCompletion.model` — guarantees Load runs before Chat, and feeds the model name straight in.
|
||||
2. `ChatCompletion.model_name` → `LLMUnloadModel.model` — guarantees Unload runs after Chat completes.
|
||||
3. *(Optional)* `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged, so the rest of your workflow can still consume it.
|
||||
|
||||

|
||||
### Tuning generation
|
||||
|
||||
### Option nodes
|
||||
Chain any combination of **Ollama Option —** nodes ahead of Chat Completion to override inference parameters:
|
||||
|
||||
Chain any combination of **Ollama Option —** nodes before Chat Completion to override inference parameters:
|
||||

|
||||
|
||||
| Option node | Ollama param |
|
||||
|-------------|-------------|
|
||||
| Option node | Parameter |
|
||||
|-------------|-----------|
|
||||
| Temperature | `temperature` |
|
||||
| Seed | `seed` |
|
||||
| Max Tokens | `num_predict` |
|
||||
| Top P | `top_p` |
|
||||
| Top K | `top_k` |
|
||||
| Repeat Penalty | `repeat_penalty` |
|
||||
| Extra Body | arbitrary JSON merged into options |
|
||||
|
||||

|
||||
| Extra Body | arbitrary JSON, merged into `options` |
|
||||
|
||||
### Multi-turn conversations
|
||||
|
||||
`OLLAMA_HISTORY` flows out of Chat Completion as a list of `{"role", "content"}` dicts. Wire it back into the next Chat Completion for multi-turn conversations, or inspect it with **Ollama Debug History** / **Ollama History Length**.
|
||||
`OLLAMA_HISTORY` flows out of Chat Completion as a `{"role", "content"}` list. Feed it back into the next Chat Completion call for multi-turn context, or inspect it with **Ollama Debug History** / **Ollama History Length**.
|
||||
|
||||
### llama.cpp
|
||||
|
||||
Everything above works unchanged against llama.cpp — swap in a **LlamaCpp Client** and the rest of the graph doesn't know the difference. The one thing llama.cpp needs that Ollama doesn't: **router mode**, a directory of models rather than a single `-m model.gguf`:
|
||||
|
||||
```bash
|
||||
llama-server --models-dir ./models -c 8192
|
||||
```
|
||||
|
||||
In exchange, router mode gives comfydv a richer live status than Ollama can report — `loading` and `downloading`, not just loaded/unloaded — plus the same explicit load/unload primitives Ollama's nodes already use.
|
||||
|
||||
### Upgrading a workflow saved before this rename
|
||||
|
||||
Nodes were renamed once, to make them backend-generic (`OllamaChatCompletion` → `ChatCompletion`, etc.). If ComfyUI reports old node types as missing when you reopen a saved workflow, reconnect using this table — behavior is unchanged, only the names are:
|
||||
|
||||
| Old | New |
|
||||
|-----|-----|
|
||||
| `OllamaChatCompletion` | `ChatCompletion` |
|
||||
| `OllamaModelSelector` | `LLMModelSelector` |
|
||||
| `OllamaLoadModel` | `LLMLoadModel` |
|
||||
| `OllamaUnloadModel` | `LLMUnloadModel` |
|
||||
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
|
||||
|
||||
`OllamaClient` kept its name — delete and re-add any node showing as missing, then rewire it to the same `OllamaClient` node.
|
||||
|
||||
@@ -5,3 +5,9 @@
|
||||
# the bullet survives a fresh clone (`beacon doctor`'s active-bullet check then
|
||||
# works in CI) and no per-branch file is left behind on the trunk after a merge.
|
||||
# `git log -p` on this file is the audit trail of who started which bullet when.
|
||||
|
||||
[bullets."claude/chatcompletion-image-input-7mm6q5"]
|
||||
title = "VLM image input for ChatCompletion"
|
||||
owner = "noreply@anthropic.com"
|
||||
started = "2026-07-22T19:11:08+00:00"
|
||||
epic = "vlm-image-input"
|
||||
|
||||
@@ -0,0 +1,172 @@
|
||||
# ADR-006: Structured Ollama output via OpenAI-compatible tool-calling + dynamic pydantic validation, not pydantic-ai
|
||||
|
||||
## Status
|
||||
|
||||
> Superseded by [ADR-007](ADR-007-llm-provider-adapter-pattern.md)
|
||||
|
||||
_Date:_ 2026-07-09
|
||||
_Deciders:_ darth-veitcher
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
`OllamaChatCompletion` sends free-text prompts to Ollama's `/api/chat` and
|
||||
returns `result["message"]["content"]` as-is, with no constraint on what the
|
||||
model may emit. In practice this is unreliable in three concrete ways:
|
||||
models prepend commentary ("Here you are:"), wrap responses in ` ``` ` code
|
||||
fences, or occasionally return blank content — all things prompt wording
|
||||
alone (e.g. a stricter `system` message) cannot reliably prevent, since it
|
||||
only *asks* the model to behave, it doesn't constrain what tokens the
|
||||
decoder is able to produce.
|
||||
|
||||
The obvious library to reach for is `pydantic-ai`, which offers structured,
|
||||
validated LLM output as a first-class feature. Its Ollama support, however,
|
||||
goes through an OpenAI-compatible client — which pulls in the `openai` SDK,
|
||||
which depends on `httpx`. This directly reverses
|
||||
[ADR-004](ADR-004-aiohttp-over-httpx-for-ollama.md), which explicitly
|
||||
rejected `httpx` as "a new runtime dep for functionality aiohttp already
|
||||
provides" (ComfyUI's own server is aiohttp-based, so aiohttp is a guaranteed
|
||||
transitive dependency; httpx is not).
|
||||
|
||||
Two Ollama-native mechanisms can force structured output without a new HTTP
|
||||
client, since both are reachable over plain JSON POST via the existing
|
||||
`aiohttp`-based `_post_json`:
|
||||
|
||||
1. **Native `/api/chat` `"format"` field** — a JSON Schema that constrains
|
||||
**decoding itself** (grammar-constrained sampling): the model's sampler
|
||||
is restricted to only emit tokens matching the schema.
|
||||
2. **OpenAI-compatible `/v1/chat/completions` tool-calling** — a `tools`
|
||||
array plus `tool_choice` forcing a single named function call, relying on
|
||||
the model's own trained function-calling behavior rather than a
|
||||
grammar-to-token mapping.
|
||||
|
||||
Both were tried against a real local model
|
||||
(`lukey03/qwen3.5-9b-abliterated-vision`) during implementation. The native
|
||||
`format` field was silently ignored — the model returned plain unstructured
|
||||
text (`"pong"`) despite the schema constraint, reproduced twice. Inspecting
|
||||
`/api/show` revealed this model's `TEMPLATE` is a degenerate `{{ .Prompt }}`
|
||||
with no role/message structure — consistent with a community "abliteration"
|
||||
process having modified the tokenizer/vocab in a way that breaks Ollama's
|
||||
grammar-to-token mapping, causing it to silently fall back to unconstrained
|
||||
generation instead of erroring. Tool-calling doesn't depend on that mapping;
|
||||
it succeeded on its first test against the same model. (A follow-up
|
||||
tool-calling call with `options` included did also fail — this specific
|
||||
model appears broadly unreliable, consistent with a degraded fine-tune, so
|
||||
this evidence is suggestive rather than conclusive. No second generative
|
||||
model was available locally to get a cleaner signal.)
|
||||
|
||||
## Decision
|
||||
|
||||
Use Ollama's OpenAI-compatible tool-calling (`/v1/chat/completions`,
|
||||
`tools`/`tool_choice` forcing a single call) for `structured_output=True`
|
||||
requests, sent through the existing `_post_json` helper — no new HTTP
|
||||
client, no `pydantic-ai`, no `openai` SDK. Non-structured requests are
|
||||
completely unaffected and keep using native `/api/chat`.
|
||||
|
||||
Use plain `pydantic` (not `pydantic-ai`) purely as a validation layer:
|
||||
given the user-supplied JSON Schema (`output_schema` input on
|
||||
`OllamaChatCompletion`), dynamically build a `pydantic.BaseModel` via
|
||||
`pydantic.create_model(...)` and validate/parse the tool call's `arguments`
|
||||
JSON against it. Required *string* fields get `min_length=1` — JSON
|
||||
Schema's `"required"` only checks presence, so a model could satisfy it
|
||||
with `""`, silently reintroducing the "blank output" problem. On validation
|
||||
failure (invalid JSON, missing/empty required field, or the model not
|
||||
calling the tool at all — observed to happen even with `tool_choice`
|
||||
forcing it), retry with fresh network calls (bounded by a `max_retries`
|
||||
input, clamped to 0–5); if every attempt fails, raise a clear
|
||||
`RuntimeError` naming the model, the attempt count, and a truncated
|
||||
snippet of the last invalid response — never silently degrade to
|
||||
unvalidated content.
|
||||
|
||||
This is opt-in: a new `structured_output: BOOLEAN` input on the existing
|
||||
`OllamaChatCompletion` node, default `False`. When off, behavior is
|
||||
unchanged — no `tools`/`tool_choice` sent, native `/api/chat` used, no
|
||||
dynamic outputs, `RETURN_TYPES` stays the original fixed 3-tuple. When on,
|
||||
one additional ComfyUI output socket is exposed per schema property
|
||||
(mirroring `FormatString`'s existing dynamic-output-socket pattern via
|
||||
`unique_id`/`RETURN_TYPES` mutation), so downstream nodes can consume
|
||||
individually typed fields instead of parsing JSON themselves.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Easier:**
|
||||
- Fixes the three concrete unreliability problems without depending on a
|
||||
model/tokenizer-sensitive grammar-constraint mechanism that was observed
|
||||
to fail silently on at least one real model.
|
||||
- No new HTTP stack: `pydantic` is a validation-only dependency, not a
|
||||
client library. ADR-004's aiohttp-only stance is preserved.
|
||||
- Fully backward compatible — `structured_output` defaults off, and
|
||||
non-structured requests still use native `/api/chat` exactly as before.
|
||||
|
||||
**Harder / constrained:**
|
||||
- Tool-calling depends on the model having usable trained function-calling
|
||||
behavior. Models with no tool-calling training may perform worse here
|
||||
than they would under grammar-constrained `format` decoding — this
|
||||
repo's only local test model was itself too unreliable to fully confirm
|
||||
either mechanism's ceiling. If well-behaved-model testing later shows
|
||||
native `format` is meaningfully more reliable in the common case, this
|
||||
decision should be revisited rather than treated as permanent.
|
||||
- Only a flat `properties: {name: {type: ...}}` shape is interpreted into
|
||||
typed ComfyUI sockets. Complex JSON Schema constructs (`$ref`,
|
||||
`oneOf`/`anyOf`/`allOf`, `enum`, nested `object`/`array` item schemas) are
|
||||
still forwarded to Ollama verbatim as the tool's `parameters`, but
|
||||
comfydv's own type mapping falls back to `STRING` for anything it doesn't
|
||||
recognize — no nested typed sockets.
|
||||
- `RETURN_TYPES`/`RETURN_NAMES` are class-level state, shared across every
|
||||
`OllamaChatCompletion` instance in a graph (same accepted limitation
|
||||
`FormatString.update_widget` already ships with) — the first execution
|
||||
after toggling `structured_output` or editing `output_schema` may show
|
||||
stale downstream socket typing until it runs once.
|
||||
- Neither mechanism is guaranteed 100% across all versions/models — hence
|
||||
the retry-then-raise defense-in-depth, rather than trusting either
|
||||
constraint blindly.
|
||||
|
||||
**Debt introduced:**
|
||||
- None. `pydantic` is a widely-used, low-conflict-risk dependency; ComfyUI
|
||||
itself is expected to already bundle it for its own API layer, though
|
||||
this repo lists it explicitly in both `pyproject.toml` and
|
||||
`requirements.txt` per ADR-003 rather than assume so.
|
||||
|
||||
## Considered Alternatives
|
||||
|
||||
### Alternative A: `pydantic-ai`
|
||||
|
||||
**Why rejected:** Its Ollama support goes through an OpenAI-compatible
|
||||
client, reintroducing `httpx` + the `openai` SDK — the exact dependency
|
||||
ADR-004 evaluated and rejected. The reliability benefit it offers is the
|
||||
same tool-calling/validation mechanism this ADR adopts directly over plain
|
||||
`aiohttp`, without the added dependency weight.
|
||||
|
||||
### Alternative B: Native `/api/chat` `"format"` field (JSON-Schema-constrained decoding)
|
||||
|
||||
**Why rejected as primary:** Theoretically the stronger guarantee — a
|
||||
sampler-level constraint rather than learned behavior — and remains a
|
||||
reasonable mechanism for well-behaved models. Rejected here because it
|
||||
failed outright (silently ignored, not even erroring) against the one real
|
||||
model available for testing, traced to that model's modified tokenizer
|
||||
breaking Ollama's grammar-to-token mapping. Tool-calling succeeded where it
|
||||
failed. See "Harder / constrained" above — this may be revisited if
|
||||
broader testing shows native `format` is more reliable in the common case.
|
||||
|
||||
### Alternative C: Prompt-only enforcement (stricter `system` message)
|
||||
|
||||
**Why rejected:** `system` already exists as an input and users can already
|
||||
try this — it's what led to the reported problem in the first place.
|
||||
Wording can reduce commentary/fences/blank output but cannot guarantee
|
||||
their absence, since nothing constrains the actual token stream.
|
||||
|
||||
### Alternative D: Response-side post-processing (regex-strip fences/preamble)
|
||||
|
||||
**Why rejected as the primary fix:** Cheap, but fundamentally reactive —
|
||||
it can strip a fence wrapper after the fact but can't recover genuinely
|
||||
blank output, and heuristics for "commentary" are unreliable across models.
|
||||
Not pursued as a fallback either, to keep this change minimal and avoid two
|
||||
competing "make output clean" mechanisms with unclear precedence.
|
||||
|
||||
---
|
||||
|
||||
## Links
|
||||
|
||||
- Related ADRs: [ADR-003](ADR-003-requirements-txt-authoring-policy.md), [ADR-004](ADR-004-aiohttp-over-httpx-for-ollama.md), [ADR-005](ADR-005-ollama-host-config-via-client-node.md)
|
||||
- Originating epic (archived, scope predates this decision): `project-management/Roadmap/epics/archive/ollama-integration.md`
|
||||
@@ -0,0 +1,198 @@
|
||||
# ADR-007: LLMProvider adapter pattern shared across Ollama and llama.cpp
|
||||
|
||||
## Status
|
||||
|
||||
> Accepted
|
||||
|
||||
_Date:_ 2026-07-11
|
||||
_Deciders:_ darth-veitcher
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
GitHub issue #15 asks comfydv to add ComfyUI nodes for llama.cpp, mirroring
|
||||
the existing Ollama integration
|
||||
(`project-management/Roadmap/epics/archive/ollama-integration.md`), and
|
||||
explicitly poses the design question: separate nodes per backend, or an
|
||||
adapter pattern that shares code? It's motivated by llama.cpp's new "router
|
||||
mode" (`llama-server --models-dir <dir>`, via
|
||||
[llama.cpp PR #18228](https://github.com/ggml-org/llama.cpp/pull/18228),
|
||||
merged 2025-12-21), which exposes `GET /models`, `POST /models/load`,
|
||||
`POST /models/unload`, and `--sleep-idle-seconds` auto-unload — giving
|
||||
llama.cpp the same manual load/unload memory-management primitives comfydv
|
||||
already relies on for Ollama.
|
||||
|
||||
`src/comfydv/ollama.py` (1056 lines, 17 node classes) has no abstraction
|
||||
layer today: two module-level free functions (`_post_json`, `_fetch_models`)
|
||||
are called directly by every node, with Ollama endpoint paths hardcoded
|
||||
inline; socket types (`OLLAMA_CLIENT`, `OLLAMA_OPTIONS`, `OLLAMA_HISTORY`)
|
||||
and `PromptServer` routes (`/dv/ollama/...`) are Ollama-named throughout.
|
||||
|
||||
**Two things needed resolving to answer issue #15's question honestly:**
|
||||
|
||||
1. **Does model lifecycle management (list/load/unload) actually converge
|
||||
between backends, or not?** At the wire-protocol level, no: Ollama uses
|
||||
`/api/tags` + a per-request `keep_alive` TTL on `/api/generate`;
|
||||
llama.cpp router mode uses `/models` + explicit `/models/load` /
|
||||
`/models/unload` + named status states
|
||||
(`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`). Judged at that
|
||||
level, a shared interface looks forced. But at the *conceptual* level,
|
||||
both APIs support exactly the same four operations — list models with
|
||||
status, load a model, unload a model, and generate/chat against a loaded
|
||||
model — just with different mechanics. Ollama's `keep_alive`-based
|
||||
load/unload already *is* `load_model()`/`unload_model()`, mechanically
|
||||
implemented as a side effect of a `/api/generate` call rather than a
|
||||
dedicated endpoint. A `Protocol` boundary at the operation level, not the
|
||||
wire-format level, fits both backends without forcing anything.
|
||||
|
||||
2. **Does `pydantic-ai` make sense now that a second backend exists?**
|
||||
[ADR-006](ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md)
|
||||
(2026-07-09) rejected `pydantic-ai` for a single backend because its
|
||||
Ollama support pulls in `httpx` + the `openai` SDK, reversing
|
||||
[ADR-004](ADR-004-aiohttp-over-httpx-for-ollama.md)'s aiohttp-only
|
||||
stance. A research pass against current `pydantic-ai` docs/source (this
|
||||
moves fast and postdates training data, so verified live rather than
|
||||
assumed) found: `httpx` is a **base dependency of `pydantic-ai-slim`
|
||||
itself**, not merely pulled in by an OpenAI-specific extra; `openai` SDK
|
||||
+ `tiktoken` are additionally required to reach any OpenAI-compatible
|
||||
backend; there is no aiohttp transport option anywhere in pydantic-ai.
|
||||
This is a **fixed, one-time dependency tax**, not one that grows per
|
||||
backend. Separately, `pydantic.create_model()`-built `BaseModel`
|
||||
subclasses (comfydv's existing pattern for validating against a
|
||||
user-supplied JSON-Schema string at workflow-execution time) work as
|
||||
pydantic-ai's `output_type` with no special-casing — the dynamic-schema
|
||||
requirement is not a blocker. And `OpenAIProvider(base_url=...)` is the
|
||||
exact generic mechanism pydantic-ai's own `OllamaProvider` is built on
|
||||
internally, so llama.cpp's `/v1/chat/completions` reaches an identical
|
||||
code path with a different `base_url` — genuinely shared implementation,
|
||||
not just a shared shape.
|
||||
|
||||
Both findings point the same direction: a real `Protocol`-based adapter,
|
||||
with `pydantic-ai` as the mechanism behind its structured-output method.
|
||||
|
||||
## Decision
|
||||
|
||||
Define a `LLMProvider` `Protocol` (new internal module, e.g.
|
||||
`src/comfydv/_llm/provider.py`) with the common surface:
|
||||
|
||||
```python
|
||||
class LLMProvider(Protocol):
|
||||
async def list_models(self) -> list[ModelInfo]: ...
|
||||
async def load_model(self, model: str) -> None: ...
|
||||
async def unload_model(self, model: str) -> None: ...
|
||||
async def chat(self, model: str, messages: ..., options: ...) -> str: ...
|
||||
async def chat_structured(self, model: str, messages: ..., schema: type[BaseModel], options: ...) -> BaseModel: ...
|
||||
```
|
||||
|
||||
`OllamaProvider` and `LlamaCppProvider` each implement it, absorbing their
|
||||
own REST mechanics internally (Ollama: `/api/tags`, `/api/generate` with
|
||||
`keep_alive`; llama.cpp: `/models`, `/models/load`, `/models/unload`) via
|
||||
`aiohttp`, unchanged from ADR-004's stance for non-chat calls. Both
|
||||
implement `chat_structured()` via the same `pydantic-ai` `Agent`/
|
||||
`output_type` call through `OpenAIProvider(base_url=...)` — one shared
|
||||
implementation, differing only in `base_url` and model name.
|
||||
|
||||
ComfyUI nodes become **generic, not per-backend**: `OllamaClient` and
|
||||
`LlamaCppClient` both output the same `LLM_CLIENT` socket type (each
|
||||
internally constructs the matching provider); a single `LLMModelSelector`,
|
||||
`LLMLoadModel`, `LLMUnloadModel`, and `ChatCompletion` node operate against
|
||||
`LLM_CLIENT` generically. Swapping providers on the canvas means rewiring
|
||||
which client node feeds the chat/management nodes, not swapping node
|
||||
classes — this is the direct answer to issue #15's question: **adapter
|
||||
pattern**, implemented as a protocol boundary at the operation level.
|
||||
|
||||
This **supersedes ADR-006**: `OllamaChatCompletion`'s `structured_output=True`
|
||||
path moves from hand-rolled tool-calling to `pydantic-ai` via the protocol.
|
||||
|
||||
This **narrows ADR-004's scope**: aiohttp remains the transport for every
|
||||
non-chat REST call inside each provider; `httpx`/`openai` enter the
|
||||
dependency tree scoped specifically to `chat_structured()`, via
|
||||
`pydantic-ai`.
|
||||
|
||||
**Documented approximation:** `ModelStatus` includes `sleeping` and
|
||||
`downloading`, states that exist in llama.cpp router mode but not in
|
||||
Ollama's API. `OllamaProvider.list_models()` normalizes into the same enum
|
||||
rather than inventing Ollama-specific states — a model that's resident and
|
||||
idle maps to `loaded` (Ollama has no distinct "kept warm but not serving"
|
||||
signal via this API), and `downloading` is simply never emitted by
|
||||
`OllamaProvider` (Ollama's pull/download flow is out of scope per the
|
||||
original Ollama epic's non-goals). This is an accepted, explicit
|
||||
approximation, not a silent gap.
|
||||
|
||||
**Confirmed:** adopting generic node/socket names (`LLM_CLIENT`,
|
||||
`ChatCompletion`, etc.) means renaming away from `OLLAMA_CLIENT`,
|
||||
`OllamaChatCompletion`, and similar — a breaking change for any saved
|
||||
workflow using the current names. The Ollama integration shipped
|
||||
2026-07-04, so the blast radius is small. Confirmed 2026-07-11: rename in
|
||||
place now rather than carry Ollama-prefixed names forward or maintain
|
||||
deprecated aliases indefinitely.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Easier:**
|
||||
- Issue #15's question gets a real answer: one generic node set works with
|
||||
any backend that implements `LLMProvider`, including future ones (a third
|
||||
local server, or a hosted OpenAI/Anthropic provider) without new node
|
||||
classes.
|
||||
- One implementation of tool-calling/structured-output logic instead of
|
||||
duplicating it per backend; `pydantic-ai` brings built-in retry/validation
|
||||
machinery, replacing ADR-006's hand-rolled retry loop.
|
||||
- The protocol boundary keeps each backend's REST quirks contained inside
|
||||
its provider — the graph never has to know Ollama uses `keep_alive` while
|
||||
llama.cpp uses explicit load/unload endpoints.
|
||||
|
||||
**Harder / constrained:**
|
||||
- New dependencies (`pydantic-ai`, `openai`, `tiktoken`, `httpx`) land in a
|
||||
project that was previously aiohttp-only.
|
||||
- This is a nontrivial migration of tested, shipped Ollama code (the
|
||||
`ollama-integration` epic is Done) — not purely additive work. Must be
|
||||
proven regression-safe before it's trusted as the foundation for
|
||||
llama.cpp.
|
||||
- `ModelStatus` is not perfectly symmetric across backends — the
|
||||
`sleeping`/`downloading` states are llama.cpp-only in practice; documented
|
||||
above, but still a leak of llama.cpp's richer vocabulary into a
|
||||
nominally-generic type.
|
||||
- The node/socket rename is a breaking change for existing saved workflows —
|
||||
confirmed acceptable given the small blast radius (see above).
|
||||
|
||||
**Debt introduced:**
|
||||
- None deliberately, contingent on the migration preserving existing
|
||||
Ollama behavior exactly (verified against `tests/test_ollama.py`).
|
||||
|
||||
## Considered Alternatives
|
||||
|
||||
### Alternative A: Shared `pydantic-ai` chat layer only; separate per-backend management nodes
|
||||
|
||||
**Why rejected:** This was the first-pass design — judged convergence at
|
||||
the REST wire-protocol level (Ollama's `/api/tags`+`keep_alive` vs.
|
||||
llama.cpp's `/models`+`/models/load`+`/models/unload` don't look alike) and
|
||||
concluded a shared interface would be forced. That framing was wrong: the
|
||||
right level to judge convergence is the *operation* (list/load/unload/chat),
|
||||
not the wire format. Both backends genuinely support the same four
|
||||
operations; only their REST mechanics differ, and those differences belong
|
||||
inside each provider implementation, not on the graph.
|
||||
|
||||
### Alternative B: Hand-roll llama.cpp's structured output too (duplicate ADR-006's approach)
|
||||
|
||||
**Why rejected:** Two independent implementations of the same
|
||||
OpenAI-compatible tool-calling mechanism is the DRY violation issue #15
|
||||
raises in the first place, with no offsetting benefit now that a second
|
||||
backend exists to justify a shared layer.
|
||||
|
||||
### Alternative C: Keep `pydantic-ai` rejected; extract a shared aiohttp-based internal helper instead
|
||||
|
||||
**Why rejected:** Avoids new dependencies entirely, but forces re-deriving
|
||||
`pydantic-ai`'s retry/validation machinery by hand for no benefit beyond
|
||||
dependency-avoidance — and doesn't change the model-management convergence
|
||||
question at all (that's orthogonal to which HTTP client the chat path
|
||||
uses). The one-time dependency tax is judged worth paying for the fuller
|
||||
abstraction, now that two backends exist to amortize it against.
|
||||
|
||||
---
|
||||
|
||||
## Links
|
||||
|
||||
- Related epics: `project-management/Roadmap/epics/llm-provider-abstraction.md`, `project-management/Roadmap/epics/llamacpp-integration.md`
|
||||
- Related ADRs: [ADR-004](ADR-004-aiohttp-over-httpx-for-ollama.md) (narrowed), [ADR-005](ADR-005-ollama-host-config-via-client-node.md) (client-node pattern generalized to `LLM_CLIENT`), [ADR-006](ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md) (superseded)
|
||||
- External reference: [llama.cpp PR #18228](https://github.com/ggml-org/llama.cpp/pull/18228) (router mode, merged 2025-12-21), GitHub issue #15
|
||||
@@ -0,0 +1,158 @@
|
||||
# ADR-008: Multimodal image input carried on the Message across the LLMProvider boundary
|
||||
|
||||
## Status
|
||||
|
||||
> Proposed
|
||||
|
||||
_Date:_ 2026-07-22
|
||||
_Deciders:_ darth-veitcher
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
The `ChatCompletion` node and the `LLMProvider` protocol (ADR-007) are
|
||||
text-only today. `Message` (`src/comfydv/_llm/provider.py`) carries a single
|
||||
`content: str`; `ChatCompletion`'s `INPUT_TYPES` (`src/comfydv/ollama.py`)
|
||||
exposes no `IMAGE` socket. Users want to couple the existing chat node with a
|
||||
vision-capable model (a VLM) to describe or reason about an image produced
|
||||
elsewhere in a ComfyUI workflow.
|
||||
|
||||
ADR-007 deliberately scoped this out: it defined `Message` as text-only and
|
||||
recorded that "if llama.cpp's router mode needs a protocol capability that
|
||||
doesn't exist yet, that is a protocol change scoped as its own follow-up, not
|
||||
silently special-cased." This ADR is that follow-up — it extends the same
|
||||
adapter pattern to a second input modality.
|
||||
|
||||
The two shipped backends carry images very differently on the wire, and the
|
||||
node has two distinct code paths (free-text vs structured), so a decision is
|
||||
needed about **where** an image lives as it crosses the provider boundary and
|
||||
**who** translates it into each backend's native shape:
|
||||
|
||||
- **Ollama free-text** — `OllamaProvider.chat()` posts `[m.model_dump() for m
|
||||
in messages]` to the native `/api/chat`, which accepts a per-message
|
||||
`images` field: an array of base64-encoded image data alongside the text
|
||||
`content`.
|
||||
- **llama.cpp free-text** — `LlamaCppProvider.chat()` posts to the
|
||||
OpenAI-compatible `/v1/chat/completions`, where a message's `content` is a
|
||||
list of typed parts (`{"type": "text", ...}`,
|
||||
`{"type": "image_url", "image_url": {"url": "data:image/...;base64,..."}}`)
|
||||
— a flat sibling `images` field is not understood.
|
||||
- **Structured output (both backends)** — routed through the shared
|
||||
`chat_structured()` helper (`src/comfydv/_llm/chat.py`) over pydantic-ai,
|
||||
which represents images as typed multimodal content
|
||||
(`BinaryContent` / `ImageUrl`) inside the user prompt, not as a raw request
|
||||
field.
|
||||
|
||||
The competing concern is DRY vs. leakage: a single carrier keeps the graph and
|
||||
the node backend-agnostic (ADR-007's whole point), but the per-backend wire
|
||||
shapes are irreducibly different and must be translated somewhere.
|
||||
|
||||
## Decision
|
||||
|
||||
**Carry images as an optional field on `Message`, and make each provider
|
||||
responsible for translating that field into its own native wire shape** — the
|
||||
exact same division of responsibility ADR-007 established for text and model
|
||||
management (operation-level protocol, wire-format quirks contained inside each
|
||||
provider).
|
||||
|
||||
1. **Protocol** — extend `Message` with an optional
|
||||
`images: list[str] | None = None`, where each entry is a base64-encoded
|
||||
image. `content` stays required; a text-only message sets `images=None` and
|
||||
is byte-for-byte unchanged from today (`model_dump()` omits it or emits
|
||||
`null`), so all existing Ollama/llama.cpp behavior is preserved.
|
||||
|
||||
2. **Node** — `ChatCompletion` gains one **optional** `image: ("IMAGE",)`
|
||||
input. When wired, the node encodes the ComfyUI `IMAGE` tensor to base64
|
||||
and attaches it to the user `Message` it already constructs. When not
|
||||
wired, the node builds exactly the message it builds today. The node never
|
||||
branches on which concrete provider it holds — consistent with ADR-007.
|
||||
|
||||
3. **Per-provider translation** (the leakage lives here, deliberately):
|
||||
- `OllamaProvider.chat()` — the flat `images` field on the dumped message
|
||||
already matches Ollama's native `/api/chat` schema; it flows through with
|
||||
no transform.
|
||||
- `LlamaCppProvider.chat()` — maps a message's `images` into OpenAI-style
|
||||
`image_url` content parts before POSTing to `/v1/chat/completions`.
|
||||
- `chat_structured()` (shared) — maps the last user message's `images` into
|
||||
pydantic-ai multimodal content on the `user_prompt`; both backends inherit
|
||||
this single implementation, mirroring how they already share the
|
||||
structured text path.
|
||||
|
||||
4. **No new node classes and no new socket types** — image support is an
|
||||
additional optional input on the *existing* generic node, so a workflow
|
||||
author gains vision by wiring one socket, not by learning a new node. This
|
||||
is the direct extension of ADR-007's "generic, not per-backend" node stance.
|
||||
|
||||
The base64 string is the neutral interchange form at the boundary because it
|
||||
is the one representation every target consumes (Ollama's `images` array,
|
||||
OpenAI's `data:` URI, and pydantic-ai's `BinaryContent` all accept it),
|
||||
keeping the `Message` carrier itself provider-agnostic.
|
||||
|
||||
_Wire specifics (exact Ollama `/api/chat` image field, llama.cpp multimodal
|
||||
readiness via `mmproj`, and pydantic-ai's multimodal content type) are
|
||||
verified live in this feature's `research.md`/`plan.md` per project
|
||||
convention, not assumed from training data._
|
||||
|
||||
## Consequences
|
||||
|
||||
**Easier:**
|
||||
- One carrier (`Message.images`) and one node change unlock vision on both
|
||||
backends at once; a future third provider implements image translation in
|
||||
its own `chat()` exactly as it implements text, with no protocol churn.
|
||||
- The graph and the node stay backend-agnostic — swapping providers still
|
||||
means rewiring one client node, now including the image path.
|
||||
- Text-only workflows are entirely unaffected (additive optional field +
|
||||
optional socket).
|
||||
|
||||
**Harder / constrained:**
|
||||
- `Message` is no longer a trivially-uniform text struct; each provider's
|
||||
`chat()` (and the shared structured helper) must handle the `images` field,
|
||||
even if only to pass it through. This is accepted leakage, localized to the
|
||||
provider layer — the same tradeoff ADR-007 already made for `keep_alive` vs
|
||||
explicit load/unload.
|
||||
- Vision requires a model actually loaded with multimodal weights (Ollama
|
||||
multimodal models; llama.cpp launched with an `mmproj` projector). A
|
||||
text-only model receiving images degrades to a backend error, not a node
|
||||
crash — surfacing that clearly is a spec requirement, not something this
|
||||
boundary can prevent.
|
||||
|
||||
**Debt introduced:**
|
||||
- None deliberately, contingent on text-only requests remaining byte-identical
|
||||
to today (guarded by the existing Ollama/llama.cpp provider tests, which must
|
||||
stay green).
|
||||
|
||||
## Considered Alternatives
|
||||
|
||||
### Alternative A: A separate `images` parameter threaded through `chat()`/`chat_structured()` signatures
|
||||
|
||||
**Why rejected:** Widens every provider method signature and the protocol for
|
||||
a value that is conceptually part of a message turn. Images belong to a
|
||||
specific message (which turn the picture accompanies), and multi-turn vision
|
||||
histories need per-message association — a single side-channel parameter can't
|
||||
express that. Putting it on `Message` keeps turn/image association intact and
|
||||
leaves method signatures unchanged.
|
||||
|
||||
### Alternative B: A dedicated multimodal node / socket type separate from `ChatCompletion`
|
||||
|
||||
**Why rejected:** Reintroduces exactly the per-capability node proliferation
|
||||
ADR-007 eliminated. A workflow author would maintain two chat nodes and
|
||||
relearn one for vision. An optional input on the existing node is strictly
|
||||
simpler and keeps the "one generic node set" promise.
|
||||
|
||||
### Alternative C: Normalize images to OpenAI content-parts at the boundary; make Ollama un-translate
|
||||
|
||||
**Why rejected:** Picks OpenAI's shape as the canonical form and forces the
|
||||
Ollama provider — whose native API wants the simpler flat `images` array — to
|
||||
convert *away* from it. That inverts the "each provider owns its own wire
|
||||
format" principle and does more work on the currently-simpler path. A neutral
|
||||
base64 carrier that every backend adapts *from* is the orthogonal choice.
|
||||
|
||||
---
|
||||
|
||||
## Links
|
||||
|
||||
- Related epic: `project-management/Roadmap/epics/vlm-image-input.md`
|
||||
- Related spec: `specs/009-vlm-image-input/`
|
||||
- Related ADRs: [ADR-007](ADR-007-llm-provider-adapter-pattern.md) (extended — same adapter pattern, second input modality), [ADR-005](ADR-005-ollama-host-config-via-client-node.md) (client-node pattern, unchanged)
|
||||
- External reference: GitHub issue #15 (llama.cpp parity), Ollama multimodal `/api/chat` `images`, OpenAI vision `image_url` content parts
|
||||
@@ -0,0 +1,82 @@
|
||||
# ADR-009: Provider-specific structured output — NativeOutput for llama.cpp, hand-rolled native `/api/chat` for Ollama
|
||||
|
||||
## Status
|
||||
|
||||
> Accepted
|
||||
|
||||
_Date:_ 2026-07-25
|
||||
_Deciders:_ darth-veitcher
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
`chat_structured()` (`src/comfydv/_llm/chat.py`, ADR-007) builds its `pydantic-ai` `Agent` with a bare `output_type=schema`. Passing a raw `pydantic.BaseModel` subclass this way makes `pydantic-ai` default to **tool-calling** (a synthetic forced function call) for structured output — inherited silently from ADR-007's move to `pydantic-ai`, never a deliberate re-decision. ADR-007 doesn't discuss output-mode choice at all.
|
||||
|
||||
Testing a real multi-agent ComfyUI workflow (`workflows/ltx-i2v-pipeline.json`) against a live local Ollama server surfaced this as a real reliability problem: against a "thinking"-capable model (`qwen3.5:9b`), tool-calling failed consistently — the model spent its entire token budget on internal chain-of-thought reasoning and never emitted the tool call, so every attempt failed pydantic validation ("token limit exceeded before any response was generated"), each attempt taking 5-7+ minutes before giving up.
|
||||
|
||||
### First fix attempt: `NativeOutput` + a priming call (superseded within this same ADR)
|
||||
|
||||
`pydantic-ai` 2.9.0 exposes `NativeOutput`, which makes the `Agent` use `response_format: {"type": "json_schema", ...}` over the OpenAI-compatible endpoint instead of tool-calling. Live-tested directly against Ollama via `curl` before touching code: `/v1/chat/completions` with `response_format: json_schema` returned clean, schema-valid JSON, with the model's reasoning in a separate `message.reasoning` field — fast, and reasoning no longer competed with structured output for token budget. This part of the fix is real and is kept — see Decision §1.
|
||||
|
||||
A second, separate problem was also found: Ollama's OpenAI-compatible endpoint doesn't honor a per-request `options` override (e.g. `num_ctx`) — sending the same request to native `/api/chat` reloaded the model at the requested context size; `/v1/chat/completions` silently kept whatever was already loaded. The first fix attempt worked around this with a priming call: hit native `/api/generate` with the desired `options` immediately before the real `/v1/chat/completions` request, on the theory that Ollama would keep the just-loaded context for the next call.
|
||||
|
||||
**This did not work, and the failure mode looked exactly like the original bug** — confirmed while re-testing the actual workflow end-to-end (`workflows/ltx-i2v-pipeline.json`, Agent 2 "Scene Grounder": long system prompt + 9-property schema + image), which kept failing with the identical "token limit exceeded" error even after the priming fix landed, tests passed, and `max_tokens`/`num_ctx`/`timeout_secs` were all raised generously. Isolated the exact mechanism with a direct, non-ComfyUI-mediated `curl` sequence:
|
||||
|
||||
1. `POST /api/generate` with `options: {num_ctx: 20480}`, `keep_alive: -1` → confirmed via `GET /api/ps`: `context_length: 20480`, loaded "forever".
|
||||
2. Immediately `POST /v1/chat/completions` for the same model — **even with the identical `options: {num_ctx: 20480}` included in that request's body** → `GET /api/ps` immediately after: `context_length: 4096` (back to default), `expires_at` reset to a normal ~5-minute keep-alive.
|
||||
|
||||
So `/v1/chat/completions` doesn't merely *ignore* `options.num_ctx` — every call to it silently **reloads the model at the default context size**, discarding whatever was primed, regardless of what that same call's own `options` field says. A priming call immediately before the real request is structurally incapable of working, because the real request itself is what undoes the priming.
|
||||
|
||||
Re-ran the same sequence against native `/api/chat` instead of `/v1/chat/completions`: the primed `context_length: 20480` was preserved through and after the call. Native `/api/chat` also accepts `"format": <json schema>` directly, giving grammar-constrained structured output in the same request that correctly honors `options` — no separate priming call needed at all.
|
||||
|
||||
### Re-litigating ADR-006's model concern
|
||||
|
||||
This also revisits [ADR-006](ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md), which rejected native `format`-based output — but its rejection was scoped to one specific model, `lukey03/qwen3.5-9b-abliterated-vision`, whose degenerate chat template silently ignored the constraint. ADR-006 explicitly flagged this as revisitable: *"If well-behaved-model testing later shows native `format` is meaningfully more reliable in the common case, this decision should be revisited rather than treated as permanent."* Re-tested that exact model against native structured output: it no longer silently ignores the constraint (ADR-006's specific failure mode) — it returns schema-valid JSON, but the *content* is still garbled (`"ponáp∵49\n"` instead of the requested `"pong"`), consistent with ADR-006's "degenerate tokenizer" diagnosis. That model was also already failing under the tool-calling path (hanging without completing, observed live during this same investigation). So neither part of this decision regresses that model — it was already unusable for structured output either way. `structured_output` has no production users yet, so there is no back-compat concern in making this change.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. `LlamaCppProvider` keeps the shared `pydantic-ai` path, switched to `NativeOutput`.** `chat.py`'s `_build_agent()` builds `Agent(chat_model, output_type=NativeOutput(schema), retries=0)` instead of a bare `output_type=schema`. llama-server's OpenAI-compatible endpoint is its genuine native structured-output surface (no equivalent context-reload bug found or expected — llama-server's context is fixed at process launch via `--ctx-size`, not a per-request concern, so there's nothing for a request to silently reset), so the shared-implementation architecture from ADR-007 stays intact for this provider.
|
||||
|
||||
**2. `OllamaProvider.chat_structured()` no longer uses `chat.py` at all.** It hand-rolls its own call to Ollama's **native** `/api/chat` with a `"format"` JSON schema, mirroring the request-building and retry/validation contract `chat.py` established (bounded retries 0–5, `RuntimeError` naming the model/attempt-count/truncated-response on exhaustion) but without pydantic-ai in the loop for this provider — there is no bare-metal native-JSON-schema mode in pydantic-ai's OpenAI-compatible model class to point at Ollama's native (non-OpenAI-shaped) endpoint, so this is a direct `_post_json` call, parsed with `schema.model_validate_json(...)`, retried on `pydantic.ValidationError` (which pydantic v2 also raises for malformed JSON, not just schema mismatches). `options` (from `OllamaOption*` nodes) is included directly in this same request's `"options"` field and is correctly honored, since it's the native endpoint — no separate priming call, because none is needed: structured output and context sizing now apply atomically in one request.
|
||||
|
||||
The retry/validation contract itself (bounded retries, `RuntimeError` naming the model/attempt-count/truncated-response on exhaustion) is unchanged and now implemented twice — once in `chat.py` for llama.cpp, once directly in `ollama_provider.py` for Ollama — rather than shared, which is the real cost of this decision (see Consequences).
|
||||
|
||||
## Consequences
|
||||
|
||||
**Easier:**
|
||||
- Structured output is now reliable against "thinking"-capable models on both providers — reasoning and structured content are separate response fields (`message.reasoning`/`message.thinking` vs `message.content`) under both `NativeOutput` and Ollama's native `format`, rather than competing for the same token stream under tool-calling.
|
||||
- `num_ctx` and other Ollama-native options now actually apply to structured-output requests — genuinely fixed this time, confirmed by re-running the actual failing workflow agent, not just by a passing test suite (the first fix attempt passed every test and still didn't work end-to-end).
|
||||
- Meaningfully faster in the success case than tool-calling against a thinking model.
|
||||
- Closes ADR-006's own explicit "revisit later" flag with concrete evidence rather than leaving it open indefinitely.
|
||||
|
||||
**Harder / constrained:**
|
||||
- `OllamaProvider` and `LlamaCppProvider` now have two independent structured-output implementations instead of one shared one — ADR-007's "share one implementation" goal no longer holds for this piece. A future structured-output feature (e.g. plumbing reasoning content back to the caller) needs to land in both places.
|
||||
- Structured output guarantees schema-*shape* validity, not semantic correctness, on both providers now — a genuinely broken model (degenerate tokenizer, as with the abliterated test model) can still return valid-JSON garbage instead of raising a clear error. Downstream consumers should not treat "returned without error" as "returned correct content" for low-quality/unreliable models.
|
||||
- Ollama's native `/api/chat` endpoint's `"format"` field is only checked against top-level `type`/`properties`/`required` the same way the OpenAI-compat `response_format` was — no change to `_build_structured_model`'s shallow-schema behavior in `ollama.py`.
|
||||
|
||||
**Debt introduced:**
|
||||
- Two structured-output code paths (per provider) instead of one shared one, as noted above — accepted because the two providers' actual constraints (Ollama's context-reset-per-OpenAI-compat-call bug vs. llama-server's fixed-at-launch context) are genuinely different, not incidentally different.
|
||||
- Not addressed here (flagged for a future ADR if pursued): a model's reasoning/thinking content is available (`message.reasoning` natively for Ollama, parsed into pydantic-ai's `ThinkingPart` for llama.cpp) but discarded by both `chat_structured()` implementations, which only return the validated schema instance. Plumbing this back to `ChatCompletion` as a node output would need a `LLMProvider.chat_structured()` return-type change — a `Protocol`-level change affecting both providers, out of scope here.
|
||||
- Not addressed here (pre-existing, unrelated): `LlamaCppProvider.chat()`'s own code comments already note that `OllamaOption*` nodes emit Ollama-native option names llama-server's OpenAI-compatible endpoint doesn't recognize — an accepted gap from the llama.cpp integration epic, unrelated to this decision.
|
||||
|
||||
## Considered Alternatives
|
||||
|
||||
### Alternative A: `NativeOutput` + priming call for both providers (the first fix attempt)
|
||||
|
||||
**Why rejected:** This is what ADR-009 originally shipped as. It passed every test (including a new one added specifically for the priming call) and one successful live single-agent ComfyUI run, but failed to actually fix the real workflow — confirmed by re-running the full pipeline and hitting the identical original failure on a later, more complex agent. Root-caused only after that: `/v1/chat/completions` unconditionally reloads the model at default context on *every* call, so priming immediately before the real call is undone by the real call itself. No amount of retrying, raising `max_tokens`, or raising `timeout_secs` fixes a context-size problem that the request itself keeps resetting.
|
||||
|
||||
### Alternative B: Fully switch both providers off `pydantic-ai`, hand-roll native structured output everywhere
|
||||
|
||||
**Why rejected:** Unnecessary for `LlamaCppProvider` — no evidence llama-server's OpenAI-compatible endpoint has Ollama's context-reset behavior (its context is fixed at process launch regardless of request), so `NativeOutput` over the existing shared path is strictly simpler there and keeps ADR-007's sharing goal intact for at least one provider.
|
||||
|
||||
### Alternative C: Do nothing, document Ollama's context-reset behavior as a known limitation
|
||||
|
||||
**Why rejected:** The underlying failure mode (indefinite-looking hangs, or outright failures, against any Ollama model needing more than the default 4096-token context while using structured output) is common enough — any long system prompt plus a non-trivial schema hits it — that documenting around it would leave `structured_output=True` effectively broken for Ollama in exactly the cases where structured output is most useful (complex, multi-field extraction tasks).
|
||||
|
||||
---
|
||||
|
||||
## Links
|
||||
|
||||
- Related ADRs: [ADR-006](ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md) (superseded rationale, not superseded status — ADR-006's tool-calling-vs-native evidence and reasoning stand as historical record; this ADR only revisits its "revisit later" flag), [ADR-007](ADR-007-llm-provider-adapter-pattern.md) (provider abstraction this decision partially steps outside of, for Ollama only)
|
||||
- Discovered while building/testing: `workflows/ltx-i2v-pipeline.json`
|
||||
@@ -0,0 +1,64 @@
|
||||
# ADR-010: `"think"` as an options-carried, per-provider-translated toggle
|
||||
|
||||
## Status
|
||||
|
||||
> Accepted
|
||||
|
||||
_Date:_ 2026-07-25
|
||||
_Deciders:_ darth-veitcher
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
ADR-009's investigation into structured-output reliability surfaced, as a side effect, how expensive a "thinking"-capable model's chain-of-thought reasoning is: on a real workflow, a single agent call could spend 5-7+ minutes and its entire token budget on reasoning before ever producing the requested response. Both Ollama and llama-server can turn this off, but neither exposes it through the generic `options` dict `ChatCompletion` already forwards — each has a completely different, incompatible wire shape:
|
||||
|
||||
- **Ollama** — live-tested directly: `"think": false` must be a **top-level** field on `/api/chat` (and `/v1/chat/completions`). Nested inside `options` (`{"options": {"think": false}}`) it's silently ignored — confirmed live (`eval_count: 223` reasoning tokens burned vs. `eval_count: 2` with it top-level). One existing test (`test_multi_turn_receives_context`) already carried `options={"think": False}` in its docstring's stated intent; it was a no-op the whole time.
|
||||
- **llama.cpp** — not live-tested (no router-mode `llama-server` instance available; user explicitly chose doc-based research over waiting for one). Per `tools/server/README.md`: `chat_template_kwargs: {"enable_thinking": false}` (Qwen3-style HF chat-template convention) and/or `reasoning_effort: "none"` (a more model-agnostic OpenAI-style convention llama-server also accepts) — both as request-body fields on `/v1/chat/completions`, not nested in `options` either.
|
||||
|
||||
Two shapes were considered for exposing this from comfydv:
|
||||
|
||||
1. **A first-class `ChatCompletion` input + `LLMProvider` protocol parameter** — mirroring how `Message.images` crossed the provider boundary (ADR-008). Initially implemented this way.
|
||||
2. **A composable `OllamaOption*`-style node merging a `"think"` key into the existing `OLLAMA_OPTIONS` chain**, with each provider popping that one key back out and translating it before building its own request — proposed as a simplification once (1) was drafted, since every other tunable knob already flows through this exact composition pattern and a new top-level node parameter would be the only one that doesn't.
|
||||
|
||||
## Decision
|
||||
|
||||
Went with option 2. `OllamaOptionDisableThinking` (`src/comfydv/ollama.py`) is a new node, identical in shape to `OllamaOptionTemperature`/`OllamaOptionSeed`/etc.: `disable_thinking: BOOLEAN` (default `True`), merges `{"think": not disable_thinking}` into whatever `OLLAMA_OPTIONS` chain it's wired into. No `ChatCompletion` or `LLMProvider` protocol signature change.
|
||||
|
||||
Both providers now start `chat()`/`chat_structured()` by popping `"think"` out of the incoming `options` dict (`_pop_think()`, `ollama_provider.py`, shared by both — a pure function, doesn't mutate the caller's dict) and translate it into their own shape before building the request:
|
||||
|
||||
- `OllamaProvider`: sets `payload["think"]` at the top level (both `chat()`'s native `/api/chat` call and `chat_structured()`'s, per ADR-009's native-endpoint rewrite).
|
||||
- `LlamaCppProvider`: sets `chat_template_kwargs`/`reasoning_effort` — directly in its own hand-rolled `chat()` payload, and via `chat.py`'s `extra_body` (the same mechanism `options` itself uses) for `chat_structured()`, which still shares the pydantic-ai path per ADR-009.
|
||||
|
||||
Despite living in `ollama.py` and following the `OllamaOption*` naming convention (matching every other option node in that module, all genuinely Ollama-native and untranslated for llama.cpp — see `LlamaCppProvider.chat()`'s own comment), `"think"` is the one key from that chain **both** providers recognize and translate; it isn't itself Ollama's native wire format, it's a comfydv-level convention that happens to reuse Ollama's own field name since Ollama's is the more literal of the two backends' conventions.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Easier:**
|
||||
- One node works for both backends, reusing the exact composition pattern (`OLLAMA_OPTIONS` chaining into `ChatCompletion`'s `options` input) every other tunable parameter already uses — no new socket type, no `ChatCompletion.INPUT_TYPES` change, no `LLMProvider` protocol change.
|
||||
- Fixes an existing test's stated-but-unfulfilled intent for free: `test_multi_turn_receives_context` and `test_structured_output_against_unreliable_model_stays_schema_valid` already passed `options={"think": False}` and now it actually works.
|
||||
- Meaningfully faster for any thinking-capable model, and directly reduces the token-budget pressure ADR-009 had to fix around.
|
||||
|
||||
**Harder / constrained:**
|
||||
- The llama.cpp translation is not live-verified — sourced from the server's documented request-body fields, not confirmed against a running `llama-server`. Verify against your own deployment before relying on it; a follow-up should close this gap once an instance is available (the user explicitly chose this tradeoff over waiting).
|
||||
- `"think"` living among genuinely-Ollama-native `OllamaOption*` nodes (which llama.cpp does *not* translate — see that class's own code comment) is a small naming/mental-model inconsistency: one key out of that whole chain is special-cased by both providers. Documented here and in `_pop_think()`'s own docstring so it doesn't read as an oversight later.
|
||||
|
||||
**Debt introduced:**
|
||||
- None. No new dependency, no new socket type.
|
||||
|
||||
## Considered Alternatives
|
||||
|
||||
### Alternative A: First-class `ChatCompletion` input + protocol parameter (mirroring `Message.images`, ADR-008)
|
||||
|
||||
**Why rejected:** Correct in principle (this is a cross-provider concern needing real translation, exactly like images), but heavier than necessary — a new node input plus a `LLMProvider.chat()`/`chat_structured()` signature change plus threading a new parameter through every call site, when the existing `options` dict composition already has a clean seam (`_pop_think`) for a value that needs per-provider translation before hitting the wire. Started implementing this way; reverted once the composable-option alternative was raised.
|
||||
|
||||
### Alternative B: Separate provider-specific nodes (`OllamaOptionDisableThinking` / a llama.cpp-only equivalent)
|
||||
|
||||
**Why rejected:** Splits one concept into two nodes for no real benefit — both backends' translation lives in code either way, so there's no cost to having one node recognize the same key on both.
|
||||
|
||||
---
|
||||
|
||||
## Links
|
||||
|
||||
- Related ADRs: [ADR-007](ADR-007-llm-provider-adapter-pattern.md) (the `LLMProvider` boundary this operates within), [ADR-008](ADR-008-multimodal-image-input-across-llmprovider-boundary.md) (the pattern this ADR considered and didn't need — cross-provider concerns don't always require a protocol change), [ADR-009](ADR-009-native-structured-output-mode.md) (the investigation that surfaced how expensive unmanaged thinking is)
|
||||
- llama.cpp server docs (request-body fields, not live-verified): `tools/server/README.md` in `ggml-org/llama.cpp`
|
||||
@@ -31,3 +31,4 @@ Superseded ADRs keep their file; update their status to `Superseded by ADR-###`.
|
||||
| ADR | Title | Status | Date |
|
||||
|-----|-------|--------|------|
|
||||
| [ADR-000](ADR-000-template.md) | Template | — | — |
|
||||
| [ADR-008](ADR-008-multimodal-image-input-across-llmprovider-boundary.md) | Multimodal image input across the LLMProvider boundary | Proposed | 2026-07-22 |
|
||||
|
||||
@@ -26,6 +26,9 @@
|
||||
- **BEACON bootstrap** — `epics/archive/beacon-bootstrap.md` — ✅ DONE — BEACON framework wired up; problem statement, constitution, roadmap, and ADR template populated; quality gates clean
|
||||
- **Logging modernisation** — `epics/logging-modernisation.md` — ✅ DONE — stdlib logging, NullHandler, silent-by-default; colorama/rich/termcolor removed; 11 tests
|
||||
- **ComfyUI UX Polish & Manager Compatibility** — `epics/ux-and-install.md` — 🔄 ACTIVE — Fix installation, core UX bugs (debounce, connection drops, alert dialogs), correctness bugs (class-level mutation, IS_CHANGED, seed=0), and metadata drift
|
||||
- **LLM Provider Abstraction** — `epics/archive/llm-provider-abstraction.md` — ✅ DONE — shared `LLMProvider` protocol (list/load/unload/chat/structured-output) and generic ComfyUI nodes; Ollama integration migrated onto it (ADR-007, supersedes ADR-006); merged via PR #17
|
||||
- **llama.cpp Model Integration** — `epics/llamacpp-integration.md` — 🔄 ACTIVE — Add a `LlamaCppProvider` implementing the shared protocol via llama-server's router mode (GitHub issue #15); dependency on LLM Provider Abstraction now satisfied
|
||||
- **VLM Image Input for ChatCompletion** — `epics/vlm-image-input.md` — 📋 PLANNING — Wire a ComfyUI IMAGE into the existing generic ChatCompletion node so a vision-capable model can describe/understand images; images carried on the `Message` and translated per-provider (ADR-008 extends ADR-007)
|
||||
|
||||
For the live rollup (specs per epic, % tasks complete, last-commit age):
|
||||
|
||||
@@ -40,6 +43,7 @@ beacon epic list --detailed
|
||||
- **BEACON bootstrap** is a prerequisite for all other epics (quality gates need to pass before new work merges)
|
||||
- **Test hardening** is independent of documentation and can run in parallel
|
||||
- **Documentation** depends on the final node API (output order, input names) being stable — start after test hardening locks the contracts
|
||||
- **llama.cpp Model Integration** depends on **LLM Provider Abstraction** landing first — its `LlamaCppProvider` implements the protocol that epic defines, and reuses its generic nodes as-is
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
# Epic: LLM Provider Abstraction
|
||||
|
||||
## Status
|
||||
Done — completed 2026-07-11
|
||||
|
||||
## Why now
|
||||
|
||||
GitHub issue #15 asks for llama.cpp support "similar to Ollama," and
|
||||
explicitly raises the question of separate nodes vs. an adapter pattern.
|
||||
[ADR-007](../../ADRs/ADR-007-llm-provider-adapter-pattern.md) answers that
|
||||
with a real adapter: a `LLMProvider` protocol (`list_models`/`load_model`/
|
||||
`unload_model`/`chat`/`chat_structured`) that any backend implements, backing
|
||||
a set of generic ComfyUI nodes (`ChatCompletion`, `LLMModelSelector`,
|
||||
`LLMLoadModel`, `LLMUnloadModel`) that work with whichever provider is wired
|
||||
in. For that to be real — not just aspirational — the existing Ollama
|
||||
integration has to migrate onto the protocol first, including moving its
|
||||
structured-output mechanism from ADR-006's hand-rolled tool-calling onto
|
||||
`pydantic-ai`. This epic is that migration; it's a prerequisite for
|
||||
`llamacpp-integration`, not additive scope on top of it — building
|
||||
llama.cpp against generic nodes that don't exist yet isn't possible.
|
||||
|
||||
## Dependencies
|
||||
|
||||
_None to start — this epic can begin immediately._ The
|
||||
`llamacpp-integration` epic depends on this one landing first: its
|
||||
`LlamaCppProvider` implements the protocol this epic defines, and its
|
||||
`LlamaCppClient` node emits the same `LLM_CLIENT` socket type this epic
|
||||
introduces.
|
||||
|
||||
## Specs
|
||||
|
||||
_Filled by `beacon specify --epic llm-provider-abstraction` /
|
||||
`/speckit-specify` once this epic is accepted._
|
||||
|
||||
- specs/007-llm-provider-abstraction/
|
||||
## ADRs
|
||||
|
||||
- project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md — defines the `LLMProvider` protocol as the adapter boundary; supersedes ADR-006; narrows ADR-004's scope to non-chat REST calls per-provider
|
||||
|
||||
## Success criteria
|
||||
|
||||
- `LLMProvider` protocol defined (new internal module, e.g. `src/comfydv/_llm/provider.py`): `list_models()`, `load_model(name)`, `unload_model(name)`, `chat(...)`, `chat_structured(..., schema)`
|
||||
- `ModelStatus` enum defined (`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`) per ADR-007's documented approximation
|
||||
- `OllamaProvider` implements the protocol, wrapping all existing Ollama REST logic (`/api/tags`, `/api/generate` with `keep_alive`) over `aiohttp` — behavior-preserving port of the current `_post_json`/`_fetch_models` logic, not a rewrite of the underlying calls
|
||||
- `chat_structured()` implemented via `pydantic-ai`'s `Agent`/`output_type`, called through `OpenAIProvider(base_url=<host>/v1)`, using the same dynamic `pydantic.create_model()`-from-JSON-Schema pattern as ADR-006 — same dynamic-socket UX, same retry/validation contract (bounded `max_retries`, required-string-non-empty check, clear `RuntimeError` on exhaustion)
|
||||
- `pydantic-ai` and `openai` added to `pyproject.toml`, curated into `requirements.txt` per [ADR-003](../../ADRs/ADR-003-requirements-txt-authoring-policy.md)
|
||||
- Generic ComfyUI nodes replace the current Ollama-specific ones: `OllamaClient` now outputs `LLM_CLIENT` (constructing an `OllamaProvider` internally); `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`, `ChatCompletion` operate against `LLM_CLIENT` generically
|
||||
- The non-structured-output chat path (native `/api/chat`) is behavior-unchanged
|
||||
- All existing `tests/test_ollama.py` coverage passes against the migrated implementation (adjusted for renamed node/socket types, unchanged in behavior otherwise)
|
||||
- Model-management calls remain entirely on `aiohttp`, inside `OllamaProvider` — untouched transport-wise
|
||||
- CI smoke test passes
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No `LlamaCppProvider` or llama.cpp nodes in this epic — that's `llamacpp-integration`
|
||||
- No behavior change to non-structured-output chat
|
||||
- No tracing/observability integration (e.g. Logfire), even though `pydantic-ai` supports it
|
||||
- No multi-turn agentic tool use beyond the existing single structured-output call
|
||||
- No backward-compat aliases for the old `Ollama`-prefixed node/socket names — confirmed 2026-07-11 to rename in place (see ADR-007)
|
||||
|
||||
## Notes
|
||||
|
||||
This is the riskiest part of the whole llama.cpp proposal: it changes the
|
||||
implementation of tested, shipped code from a Done epic
|
||||
(`archive/ollama-integration.md`), not just adding new code, and it's a
|
||||
breaking rename (`OLLAMA_CLIENT`→`LLM_CLIENT`,
|
||||
`OllamaChatCompletion`→`ChatCompletion`, etc.) for anyone with saved
|
||||
workflows using the current node/socket names. Confirmed 2026-07-11 (per
|
||||
ADR-007): rename in place now — the Ollama integration only shipped
|
||||
2026-07-04, so the blast radius is small — rather than carrying
|
||||
`Ollama`-prefixed generic nodes forward or maintaining deprecated aliases
|
||||
indefinitely.
|
||||
|
||||
Recommend an adversarial pass (`/beacon:review` or `/beacon:engineering`)
|
||||
before merging, specifically checking that the retry/validation contract
|
||||
from ADR-006's `## Decision` section is preserved exactly by the
|
||||
`pydantic-ai` reimplementation, and that `OllamaProvider`'s REST calls are a
|
||||
faithful port of the current `_post_json`/`_fetch_models` logic.
|
||||
|
||||
**2026-07-11 — mid-build correction, tracked in [issue #16](https://github.com/darth-veitcher/comfydv/issues/16):**
|
||||
the Foundational layer (`LLMProvider` protocol + `OllamaProvider` skeleton)
|
||||
shipped safely, but the planned per-user-story incremental cutover doesn't
|
||||
hold — `OllamaClient` is a single shared producer for every downstream
|
||||
Ollama node, so the node-layer rename/cutover (`tasks.md`'s US1 + US3) must
|
||||
land as one atomic change, not four independent ones. Confirmed by
|
||||
independent product + engineering review. Re-scoped as its own dedicated
|
||||
follow-up BUILD session — see `specs/007-llm-provider-abstraction/tasks.md`'s
|
||||
correction note for full detail. Open question for the next session: does
|
||||
this take priority over `ux-and-install` (active, 1/4 specs shipped), since
|
||||
llama.cpp (issue #15) has no deadline.
|
||||
|
||||
**2026-07-11 — properly specced:** the deferred cutover is now fully
|
||||
inventoried and planned in
|
||||
`specs/007-llm-provider-abstraction/atomic-cutover-plan.md` — a full
|
||||
line-by-line read of the ~125 affected references (not an estimate), the
|
||||
design decisions it surfaced (cache-singleton duplication, `client ==
|
||||
"<string>"` equality breaking, bare-string-client backward compat removal,
|
||||
and a test-layer split so the 35 `_post_json` monkeypatches land at the
|
||||
right architectural seam), and a 12-step sequenced task list (T-CUT-01 …
|
||||
T-CUT-12). This is now the authoritative implementation plan for the
|
||||
cutover — the next BUILD session executes it directly rather than
|
||||
re-deriving the approach.
|
||||
|
||||
`pydantic-ai`'s `StructuredDict` (raw-JSON-Schema output, no Python class)
|
||||
was considered as a lighter-weight alternative to `create_model()` during
|
||||
research and rejected: it performs no pydantic validation at all, which
|
||||
would silently drop the "reject blank required strings" safeguard ADR-006
|
||||
introduced. Stick with `create_model()`-built `BaseModel` subclasses.
|
||||
|
||||
**2026-07-11 — cutover executed, PR open:** T-CUT-01 through T-CUT-12
|
||||
complete (`specs/007-llm-provider-abstraction/tasks.md` — every task `[x]`
|
||||
or explicitly `[-]` superseded/deferred with a reason). `beacon epic
|
||||
refresh` reports 1/1 owned specs complete. An independent `beacon-reviewer`
|
||||
pass caught one real regression (`options` silently dropped in
|
||||
structured-output mode) before merge — fixed and re-verified clear. README
|
||||
and docs/index.md updated for the rename, including the FR-009 migration
|
||||
table. **PR: [#17](https://github.com/darth-veitcher/comfydv/pull/17)** —
|
||||
not yet merged; `beacon epic finish` waits for that, per
|
||||
`beacon epic refresh`'s own guidance.
|
||||
@@ -1,7 +1,7 @@
|
||||
# Epic: Ollama Model Integration
|
||||
|
||||
## Status
|
||||
Planning — started 2026-06-28
|
||||
Done — completed 2026-07-04
|
||||
|
||||
## Why now
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
# Epic: llama.cpp Model Integration
|
||||
|
||||
## Status
|
||||
Active — spec 008-llamacpp-integration complete, ready to finish once merged
|
||||
|
||||
## Why now
|
||||
|
||||
GitHub issue #15 requests llama.cpp support "similar to Ollama." This is now
|
||||
practical because llama.cpp's `llama-server` gained a "router mode" via
|
||||
[llama.cpp PR #18228](https://github.com/ggml-org/llama.cpp/pull/18228)
|
||||
(merged 2025-12-21): launched with `--models-dir <dir>` (or
|
||||
`--models-preset <file>.ini`) instead of `-m`, it exposes `GET /models`
|
||||
(with live status: `unloaded`/`loading`/`loaded`/`sleeping`/`downloading`),
|
||||
`POST /models/load`, `POST /models/unload`, and `--sleep-idle-seconds`
|
||||
auto-unload — giving llama.cpp the same manual load/unload
|
||||
memory-management primitives comfydv already relies on for Ollama. With the
|
||||
`llm-provider-abstraction` epic in place, adding llama.cpp is now a matter
|
||||
of implementing one more `LLMProvider`, not building a parallel set of
|
||||
ComfyUI nodes.
|
||||
|
||||
## Dependencies
|
||||
|
||||
Depended on the `llm-provider-abstraction` epic landing first — **satisfied**
|
||||
2026-07-11, merged via [PR #17](https://github.com/darth-veitcher/comfydv/pull/17)
|
||||
(`project-management/Roadmap/epics/archive/llm-provider-abstraction.md`).
|
||||
This epic's `LlamaCppProvider` implements the `LLMProvider` protocol that
|
||||
epic defined, and its `LlamaCppClient` node emits the same `LLM_CLIENT`
|
||||
socket type the generic `ChatCompletion`/`LLMModelSelector`/`LLMLoadModel`/
|
||||
`LLMUnloadModel` nodes already consume — none of those node classes are
|
||||
touched by this epic.
|
||||
|
||||
## Specs
|
||||
|
||||
_Filled by `beacon specify --epic llamacpp-integration` / `/speckit-specify`
|
||||
once this epic is accepted._
|
||||
|
||||
- specs/008-llamacpp-integration/
|
||||
## ADRs
|
||||
|
||||
- project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md — decided during the prerequisite epic; this epic implements the second `LLMProvider` the ADR anticipated
|
||||
|
||||
## Success criteria
|
||||
|
||||
- `LlamaCppProvider` implements the `LLMProvider` protocol from the prerequisite epic:
|
||||
- `list_models()` via `GET /models`, surfacing native status (`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`) directly — no normalization needed, since llama.cpp's vocabulary is the `ModelStatus` enum's superset
|
||||
- `load_model()` / `unload_model()` via `POST /models/load` / `POST /models/unload`
|
||||
- `chat_structured()` via the same shared `pydantic-ai` mechanism as `OllamaProvider`, `OpenAIProvider(base_url=<llama-server host>/v1)` — no new structured-output code, just a different `base_url`
|
||||
- `LlamaCppClient` config node (reuses the [ADR-005](../../ADRs/ADR-005-ollama-host-config-via-client-node.md) config-node pattern), outputs the same `LLM_CLIENT` socket type `OllamaClient` does
|
||||
- No new node classes for model selection, load/unload, or chat — the generic `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`, and `ChatCompletion` nodes from the prerequisite epic work unchanged once a `LlamaCppClient` is wired in
|
||||
- `LlamaCppClient` registered in `NODE_CLASS_MAPPINGS` / `NODE_DISPLAY_NAME_MAPPINGS`
|
||||
- Test coverage for `LlamaCppProvider` mirrors the `OllamaProvider` test conventions established in the prerequisite epic
|
||||
- No new runtime dependencies beyond what the prerequisite epic already introduced (`aiohttp` for model management, `pydantic-ai`/`openai` for chat)
|
||||
- CI smoke test passes
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No support for llama-server's non-router single-model launch mode (`-m`) — router mode only, since that's what gives load/unload parity with Ollama
|
||||
- No GPU inference optimisation or quantisation tuning — CPU-first dev harness, consistent with the Ollama epic's own non-goal
|
||||
- No auth/TLS/remote-serving hardening — localhost/configurable host via client node only, consistent with the Ollama epic
|
||||
- No ComfyUI Manager registry listing in this epic
|
||||
- No changes to the generic nodes or `LLMProvider` protocol themselves — if llama.cpp's router mode needs something the protocol doesn't support, that's a protocol change scoped back into the prerequisite epic's follow-up, not silently special-cased here
|
||||
|
||||
## Notes
|
||||
|
||||
Router mode is a deployment prerequisite, not something comfydv configures:
|
||||
the user must launch `llama-server` with `--models-dir`/`--models-preset`
|
||||
themselves. Document this clearly in the eventual spec/node tooltips.
|
||||
|
||||
Reference: [llama.cpp PR #18228](https://github.com/ggml-org/llama.cpp/pull/18228), GitHub issue #15.
|
||||
@@ -0,0 +1,61 @@
|
||||
# Epic: VLM Image Input for ChatCompletion
|
||||
|
||||
## Status
|
||||
Planning — started 2026-07-22
|
||||
|
||||
## Why now
|
||||
|
||||
The generic `ChatCompletion` node and the `LLMProvider` protocol landed
|
||||
text-only (ADR-007), which explicitly deferred any protocol change for a new
|
||||
capability as "its own follow-up." Both shipped backends can already serve
|
||||
vision models — Ollama multimodal models via `/api/chat`'s per-message
|
||||
`images`, and llama.cpp via a multimodal projector (`mmproj`) on the same
|
||||
OpenAI-compatible `/v1/chat/completions` the provider already calls — so the
|
||||
gap is entirely on comfydv's client side, not the servers'. Coupling the chat
|
||||
node with a VLM to describe or reason about images produced elsewhere in a
|
||||
workflow is a frequently-wanted next step, and the adapter pattern makes it a
|
||||
small, symmetric addition rather than a new node family.
|
||||
|
||||
## Specs
|
||||
_SpecKit specs that contribute to this epic._
|
||||
|
||||
- specs/009-vlm-image-input/ — wire a ComfyUI IMAGE into the existing ChatCompletion node; images carried on the Message and translated per-provider
|
||||
|
||||
## ADRs
|
||||
_Cross-cutting decisions this epic required._
|
||||
|
||||
- project-management/ADRs/ADR-008-multimodal-image-input-across-llmprovider-boundary.md — carry images as an optional `Message.images` field; each provider translates to its own wire shape (extends ADR-007's adapter pattern to a second input modality)
|
||||
- project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md — the adapter pattern this epic extends; the generic node/protocol it adds an image path to
|
||||
|
||||
## Success criteria
|
||||
|
||||
- `Message` carries an optional `images` field; text-only requests remain byte-for-byte unchanged (existing Ollama + llama.cpp provider tests stay green)
|
||||
- `ChatCompletion` gains one **optional** `IMAGE` input — no new node classes, no new socket types; a workflow author gains vision by wiring one socket
|
||||
- A wired image reaches a vision model and produces a description/answer on **both** backends:
|
||||
- Ollama: flat per-message `images` passes through `/api/chat` untransformed
|
||||
- llama.cpp: mapped to OpenAI `image_url` content parts on `/v1/chat/completions`
|
||||
- Structured output with an image works via the shared `chat_structured()` (pydantic-ai multimodal content) — one implementation, both backends
|
||||
- A text-only model that receives an image degrades to a clear backend error surfaced by the node, not a crash
|
||||
- Test coverage mirrors the `OllamaProvider`/`LlamaCppProvider` conventions (mock at the provider's own transport seam); CI smoke test passes
|
||||
- No new runtime dependencies beyond what ADR-007 already introduced
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No image **output** or image generation — input-to-VLM only
|
||||
- No new node classes or socket types — additive optional input on the existing generic node
|
||||
- No changes to the client/config nodes (`OllamaClient`, `LlamaCppClient`) or model-management nodes
|
||||
- No auto-provisioning of vision models — the user must have a multimodal model loaded (Ollama multimodal model; llama.cpp launched with an `mmproj` projector); this epic does not install or configure it
|
||||
- No video, audio, or document/PDF modalities — still images only
|
||||
- No image preprocessing beyond what's needed to hand a ComfyUI IMAGE tensor to a backend (no resizing policy, tiling, or OCR of our own)
|
||||
- No `OllamaOption*` parameter translation work — inherited unchanged from ADR-007's scope
|
||||
|
||||
## Notes
|
||||
|
||||
Multimodal readiness is a deployment prerequisite, not something comfydv
|
||||
configures: document in node tooltips that the wired model must be
|
||||
vision-capable, and that llama.cpp needs `--mmproj`. The exact wire shapes
|
||||
(Ollama `/api/chat` `images`, OpenAI `image_url`, pydantic-ai `BinaryContent`)
|
||||
are verified live in the spec's `research.md`/`plan.md`, consistent with how
|
||||
the llama.cpp epic verified router-mode endpoints.
|
||||
|
||||
Reference: ADR-008, ADR-007, GitHub issue #15.
|
||||
@@ -1,6 +1,6 @@
|
||||
[project]
|
||||
name = "comfydv"
|
||||
version = "0.1.0"
|
||||
dynamic = ["version"]
|
||||
requires-python = ">=3.11"
|
||||
description = "Quality of life ComfyUI nodes: dynamic string formatting, seed-controlled random selection, and conditional queue interruption."
|
||||
readme = "README.md"
|
||||
@@ -10,6 +10,8 @@ authors = [
|
||||
dependencies = [
|
||||
"aiohttp>=3.9.0",
|
||||
"jinja2>=3.1.6",
|
||||
"pydantic>=2.0",
|
||||
"pydantic-ai-slim[openai]>=2.9.0",
|
||||
]
|
||||
|
||||
[project.urls]
|
||||
@@ -17,11 +19,15 @@ Repository = "https://github.com/darth-veitcher/comfydv"
|
||||
Documentation = "https://darth-veitcher.github.io/comfydv/stable/"
|
||||
|
||||
[build-system]
|
||||
requires = ["hatchling"]
|
||||
requires = ["hatchling", "hatch-vcs"]
|
||||
build-backend = "hatchling.build"
|
||||
|
||||
[tool.hatch.version]
|
||||
source = "vcs"
|
||||
|
||||
[dependency-groups]
|
||||
dev = [
|
||||
"pillow>=10.0.0",
|
||||
"playwright>=1.60.0",
|
||||
"pytest>=8.4.2",
|
||||
"pytest-cov>=6.0.0",
|
||||
|
||||
@@ -2,3 +2,5 @@
|
||||
# Lists only deps NOT already provided by ComfyUI's own environment.
|
||||
# See ADR-003: never auto-generate this file with `uv export`.
|
||||
jinja2>=3.1.6
|
||||
pydantic>=2.0
|
||||
pydantic-ai-slim[openai]>=2.9.0
|
||||
|
||||
@@ -117,7 +117,11 @@ async def _capture(
|
||||
pad: int = 60,
|
||||
) -> None:
|
||||
"""Screenshot the canvas, optionally cropped tightly around the node."""
|
||||
canvas = page.locator("canvas#graph-canvas, canvas").first
|
||||
# ComfyUI's frontend added a minimap canvas since this script was last
|
||||
# verified — it appears before #graph-canvas in DOM order, so a bare
|
||||
# union selector's `.first` silently grabbed the 250x200 minimap
|
||||
# instead of the real graph. Target #graph-canvas explicitly.
|
||||
canvas = page.locator("canvas#graph-canvas")
|
||||
box = await canvas.bounding_box()
|
||||
if box is None:
|
||||
await page.screenshot(path=str(out))
|
||||
@@ -307,13 +311,13 @@ async def scene_ollama_client(page: Page, out: Path) -> None:
|
||||
|
||||
|
||||
async def scene_ollama_chat(page: Page, out: Path) -> None:
|
||||
"""OllamaChatCompletion — showing the live model dropdown and prompt widget."""
|
||||
"""ChatCompletion — showing the live model dropdown and prompt widget."""
|
||||
await _clear(page)
|
||||
|
||||
info = await page.evaluate(
|
||||
"""
|
||||
async () => {
|
||||
const node = LiteGraph.createNode("OllamaChatCompletion");
|
||||
const node = LiteGraph.createNode("ChatCompletion");
|
||||
node.pos = [60, 60];
|
||||
window.app.graph.add(node);
|
||||
|
||||
@@ -327,7 +331,7 @@ async def scene_ollama_chat(page: Page, out: Path) -> None:
|
||||
} else {
|
||||
// Manually fetch and populate the COMBO
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
|
||||
if (resp.ok) {
|
||||
const data = await resp.json();
|
||||
const models = data.models || [];
|
||||
@@ -357,8 +361,142 @@ async def scene_ollama_chat(page: Page, out: Path) -> None:
|
||||
await _capture(page, out, info["pos"], info["size"])
|
||||
|
||||
|
||||
async def scene_structured_output(page: Page, out: Path) -> None:
|
||||
"""ChatCompletion with structured_output enabled — the live dynamic
|
||||
output sockets (summary/sentiment/confidence) that appear as soon as
|
||||
output_schema is edited, no graph run required. Exercises the real
|
||||
js/ollama.js widget callbacks (structuredWidget.callback /
|
||||
schemaWidget.callback), the same code path a live user's checkbox click
|
||||
and schema edit trigger — not a hand-simulated approximation."""
|
||||
await _clear(page)
|
||||
|
||||
schema = (
|
||||
'{"type": "object", "properties": '
|
||||
'{"summary": {"type": "string"}, '
|
||||
'"sentiment": {"type": "string"}, '
|
||||
'"confidence": {"type": "number"}}, '
|
||||
'"required": ["summary", "sentiment", "confidence"]}'
|
||||
)
|
||||
|
||||
info = await page.evaluate(
|
||||
f"""
|
||||
async () => {{
|
||||
const node = LiteGraph.createNode("ChatCompletion");
|
||||
node.pos = [60, 60];
|
||||
window.app.graph.add(node);
|
||||
|
||||
const promptWidget = node.widgets.find(w => w.name === "prompt");
|
||||
if (promptWidget) promptWidget.value =
|
||||
"The new render pipeline cut our export time in half and the team is thrilled.";
|
||||
|
||||
const structuredWidget = node.widgets.find(w => w.name === "structured_output");
|
||||
const schemaWidget = node.widgets.find(w => w.name === "output_schema");
|
||||
if (structuredWidget) structuredWidget.value = true;
|
||||
if (schemaWidget) schemaWidget.value = {json.dumps(schema)};
|
||||
|
||||
// Fire the same callbacks js/ollama.js attaches on node creation —
|
||||
// real widget-edit code path, not a re-implementation.
|
||||
if (structuredWidget?.callback) await structuredWidget.callback(true);
|
||||
if (schemaWidget?.callback) await schemaWidget.callback({json.dumps(schema)});
|
||||
|
||||
await new Promise(r => setTimeout(r, 600));
|
||||
window.app.canvas.setDirty(true, true);
|
||||
window.app.canvas.draw(true, true);
|
||||
return {{ pos: [node.pos[0], node.pos[1]], size: [node.size[0], node.size[1]] }};
|
||||
}}
|
||||
"""
|
||||
)
|
||||
|
||||
await asyncio.sleep(0.8)
|
||||
await _redraw(page)
|
||||
await _frame_node(page, info["pos"], info["size"])
|
||||
await _capture(page, out, info["pos"], info["size"])
|
||||
|
||||
|
||||
async def scene_llamacpp_client(page: Page, out: Path) -> None:
|
||||
"""LlamaCppClient — single node showing the router-mode host URL widget."""
|
||||
await _clear(page)
|
||||
|
||||
info = await page.evaluate(
|
||||
"""
|
||||
() => {
|
||||
const node = LiteGraph.createNode("LlamaCppClient");
|
||||
node.pos = [60, 60];
|
||||
window.app.graph.add(node);
|
||||
const hostWidget = node.widgets && node.widgets.find(w => w.name === "host");
|
||||
if (hostWidget) hostWidget.value = "http://localhost:8080";
|
||||
window.app.canvas.setDirty(true, true);
|
||||
window.app.canvas.draw(true, true);
|
||||
return { pos: [node.pos[0], node.pos[1]], size: [node.size[0], node.size[1]] };
|
||||
}
|
||||
"""
|
||||
)
|
||||
|
||||
await _frame_node(page, info["pos"], info["size"])
|
||||
await _capture(page, out, info["pos"], info["size"])
|
||||
|
||||
|
||||
async def scene_llamacpp_workflow(page: Page, out: Path) -> None:
|
||||
"""LlamaCppClient → the same ChatCompletion node the Ollama workflow
|
||||
uses, unmodified — the actual point of the adapter pattern. Attempts a
|
||||
live model-list refresh via backend=llamacpp against
|
||||
host.docker.internal:8080; degrades gracefully (same as a real user's
|
||||
"no server running yet" state) if nothing is listening there."""
|
||||
await _clear(page)
|
||||
|
||||
info = await page.evaluate(
|
||||
"""
|
||||
async () => {
|
||||
const graph = window.app.graph;
|
||||
|
||||
const client = LiteGraph.createNode("LlamaCppClient");
|
||||
client.pos = [40, 60];
|
||||
graph.add(client);
|
||||
const hostWidget = client.widgets.find(w => w.name === "host");
|
||||
if (hostWidget) hostWidget.value = "http://host.docker.internal:8080";
|
||||
|
||||
const chat = LiteGraph.createNode("ChatCompletion");
|
||||
chat.pos = [380, 40];
|
||||
graph.add(chat);
|
||||
const promptWidget = chat.widgets.find(w => w.name === "prompt");
|
||||
if (promptWidget) promptWidget.value = "Write a haiku about ComfyUI.";
|
||||
|
||||
client.connect(0, chat, 0);
|
||||
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:8080&backend=llamacpp");
|
||||
if (resp.ok) {
|
||||
const data = await resp.json();
|
||||
const models = data.models || [];
|
||||
if (models.length) {
|
||||
const modelWidget = chat.widgets.find(w => w.name === "model");
|
||||
if (modelWidget) modelWidget.value = models[0];
|
||||
}
|
||||
}
|
||||
} catch (e) {}
|
||||
|
||||
await new Promise(r => setTimeout(r, 800));
|
||||
window.app.canvas.setDirty(true, true);
|
||||
window.app.canvas.draw(true, true);
|
||||
|
||||
const nodes = [client, chat];
|
||||
const minX = Math.min(...nodes.map(n => n.pos[0])) - 20;
|
||||
const minY = Math.min(...nodes.map(n => n.pos[1])) - 20;
|
||||
const maxX = Math.max(...nodes.map(n => n.pos[0] + n.size[0])) + 20;
|
||||
const maxY = Math.max(...nodes.map(n => n.pos[1] + n.size[1])) + 20;
|
||||
return { pos: [minX, minY], size: [maxX - minX, maxY - minY] };
|
||||
}
|
||||
"""
|
||||
)
|
||||
|
||||
await asyncio.sleep(1.0)
|
||||
await _redraw(page)
|
||||
await _frame_node(page, info["pos"], info["size"], scale=1.0)
|
||||
await _capture(page, out, info["pos"], info["size"], scale=1.0)
|
||||
|
||||
|
||||
async def scene_ollama_workflow(page: Page, out: Path) -> None:
|
||||
"""Full mini-workflow: OllamaClient → OllamaChatCompletion + Temperature + Seed options."""
|
||||
"""Full mini-workflow: OllamaClient → ChatCompletion + Temperature + Seed options."""
|
||||
await _clear(page)
|
||||
|
||||
info = await page.evaluate(
|
||||
@@ -386,7 +524,7 @@ async def scene_ollama_workflow(page: Page, out: Path) -> None:
|
||||
if (twSeed) twSeed.value = 42;
|
||||
|
||||
// 4. ChatCompletion — right
|
||||
const chat = LiteGraph.createNode("OllamaChatCompletion");
|
||||
const chat = LiteGraph.createNode("ChatCompletion");
|
||||
chat.pos = [380, 100];
|
||||
graph.add(chat);
|
||||
const twPrompt = chat.widgets && chat.widgets.find(w => w.name === "prompt");
|
||||
@@ -409,7 +547,7 @@ async def scene_ollama_workflow(page: Page, out: Path) -> None:
|
||||
|
||||
// Refresh model dropdowns for chat node
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
|
||||
if (resp.ok) {
|
||||
const data = await resp.json();
|
||||
const models = data.models || [];
|
||||
@@ -466,25 +604,25 @@ async def scene_ollama_lifecycle(page: Page, out: Path) -> None:
|
||||
if (hostWidget) hostWidget.value = "http://localhost:11434";
|
||||
|
||||
// OllamaLoadModel — generous gap right of client
|
||||
const load = LiteGraph.createNode("OllamaLoadModel");
|
||||
const load = LiteGraph.createNode("LLMLoadModel");
|
||||
load.pos = [380, 80];
|
||||
graph.add(load);
|
||||
|
||||
// OllamaChatCompletion — wide node, plenty of space to the right of load
|
||||
const chat = LiteGraph.createNode("OllamaChatCompletion");
|
||||
const chat = LiteGraph.createNode("ChatCompletion");
|
||||
chat.pos = [720, 40];
|
||||
graph.add(chat);
|
||||
const twPrompt = chat.widgets && chat.widgets.find(w => w.name === "prompt");
|
||||
if (twPrompt) twPrompt.value = "Describe this image in one sentence.";
|
||||
|
||||
// OllamaUnloadModel — far right, vertically offset to match chat's outputs
|
||||
const unload = LiteGraph.createNode("OllamaUnloadModel");
|
||||
const unload = LiteGraph.createNode("LLMUnloadModel");
|
||||
unload.pos = [1200, 280];
|
||||
graph.add(unload);
|
||||
|
||||
// Populate model dropdowns from live Ollama
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
|
||||
if (resp.ok) {
|
||||
const data = await resp.json();
|
||||
const models = data.models || [];
|
||||
@@ -604,6 +742,11 @@ SCENES = [
|
||||
("ollama_workflow.png", scene_ollama_workflow),
|
||||
("ollama_options.png", scene_ollama_options),
|
||||
("ollama_lifecycle.png", scene_ollama_lifecycle),
|
||||
# Structured output (ADR-007 / pydantic-ai)
|
||||
("structured_output.png", scene_structured_output),
|
||||
# llama.cpp (spec 008)
|
||||
("llamacpp_client.png", scene_llamacpp_client),
|
||||
("llamacpp_workflow.png", scene_llamacpp_workflow),
|
||||
]
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
epic = "llm-provider-abstraction"
|
||||
@@ -0,0 +1,178 @@
|
||||
# Atomic Cutover Plan: Ollama Node Rename → Generic LLM Nodes
|
||||
|
||||
Companion to `tasks.md`'s "⚠️ Correction (2026-07-11)" section and
|
||||
[issue #16](https://github.com/darth-veitcher/comfydv/issues/16). This is
|
||||
the properly-specced version of that deferred work — based on a full,
|
||||
line-by-line inventory of every one of the ~125 affected references in
|
||||
`src/comfydv/ollama.py` and `tests/test_ollama.py` (1820 lines, read in
|
||||
full), not an estimate.
|
||||
|
||||
The inventory surfaced that this isn't one mechanical find-and-replace —
|
||||
several genuine design decisions were implicit in "rename it" and needed
|
||||
resolving before any code changes. Those decisions are below, followed by
|
||||
the sequenced task list that implements them.
|
||||
|
||||
## Resolved decisions
|
||||
|
||||
### D1 — Single source of truth for HTTP/cache infra
|
||||
|
||||
`comfydv._llm.ollama_provider` owns `_post_json`, `_run_async`,
|
||||
`_fetch_models`, `_TTLLRUCache`, `_cache_key`, `_MODEL_LIST_CACHE` (already
|
||||
ported there — see `ollama_provider.py`). `ollama.py` stops defining its
|
||||
own copies of these. The one remaining non-node use case —
|
||||
`_load_default_models()` (combo-widget population at import time, before
|
||||
any `OllamaClient` node exists) and the `/dv/ollama/models` refresh route —
|
||||
imports `_fetch_models`/`_run_async` from `comfydv._llm.ollama_provider`
|
||||
instead of duplicating them. `_CHAT_RESPONSE_CACHE` and `_post_json`
|
||||
disappear from `ollama.py` entirely — nothing there needs them once
|
||||
load/unload/chat delegate to `client.*`.
|
||||
|
||||
**Why not keep two copies:** they were already flagged as byte-identical by
|
||||
the inventory (§2.13) — the only reason `ollama.py` still has them is that
|
||||
nothing has repointed the imports yet. Keeping a second copy "just in case"
|
||||
is exactly the kind of duplication ADR-007 exists to eliminate.
|
||||
|
||||
### D2 — `client == "<host string>"` equality is removed, not preserved
|
||||
|
||||
`OllamaProvider` is a plain object, not a `str` subclass — this is a
|
||||
deliberate consequence of the adapter boundary (ADR-007), not an oversight
|
||||
to work around. Two tests assert string equality on `client`
|
||||
(`test_client_outputs_ollama_client_type:208-209`,
|
||||
`test_client_carries_headers:1233-1234`) — both get rewritten to assert
|
||||
`client.host == "..."` and `isinstance(client, OllamaProvider)`.
|
||||
`OllamaClientType` (the `str`-subclass, `ollama.py:97-109`) is left in
|
||||
place but becomes unused by `OllamaClient.create_client()` — not deleted in
|
||||
this cutover (no test depends on deleting it, and removing a class nothing
|
||||
references is a separate, lower-risk cleanup, not part of this bullet's
|
||||
scope).
|
||||
|
||||
### D3 — Bare-string `client` backward compatibility is removed
|
||||
|
||||
`test_plain_string_client_has_no_headers` (`test_ollama.py:1328-1344`)
|
||||
documents and tests that wiring a plain `STRING` node directly into
|
||||
`client` (skipping `OllamaClient` entirely) silently works, because
|
||||
`f"{client}/..."` succeeds on any string. This was never a documented,
|
||||
intended feature — it's a side effect of `OllamaClientType` being a `str`
|
||||
subclass, not mentioned in ADR-005 or the original Ollama epic's spec. Once
|
||||
node methods call `client.chat(...)`/`client.load_model(...)`, a bare
|
||||
string raises `AttributeError`. **This test is deleted**, not rewritten —
|
||||
its premise (bare-string clients are supported) is being intentionally
|
||||
removed, and asserting the new failure mode would just be testing that
|
||||
Python raises `AttributeError` on missing methods, which isn't
|
||||
comfydv-specific behavior worth a test.
|
||||
|
||||
The same bare-string pattern appears incidentally in ~8 other tests
|
||||
(`ollama_host` fixture returns a plain string, used as `client=ollama_host`
|
||||
in several integration tests — `test_ollama.py:220, 386, 407, 588, ...`).
|
||||
These need `client=OllamaClient().create_client(ollama_host)[0]` instead of
|
||||
`client=ollama_host` — a required edit, not optional, since they'll raise
|
||||
`AttributeError` otherwise. See task T-CUT-08 below.
|
||||
|
||||
### D4 — Test layer split (resolves all 35 relocated `_post_json` monkeypatches)
|
||||
|
||||
This is the biggest structural decision. Today, `test_ollama.py` tests
|
||||
ComfyUI node behavior by mocking `aiohttp` at the `ollama_mod._post_json`
|
||||
seam and asserting on the exact Ollama wire payload (`keep_alive`,
|
||||
`/api/generate`, tool-calling JSON shape, retry counts) *through* the node.
|
||||
That seam moves — nodes no longer call `_post_json` directly, they call
|
||||
`client.chat(...)` etc. Two options: (a) keep patching at whatever the new
|
||||
seam is, 1:1 per test, or (b) recognize this is an architectural boundary
|
||||
and split coverage accordingly. Going with **(b)**:
|
||||
|
||||
- **`tests/test_ollama.py`** — ComfyUI node **contract + delegation** only.
|
||||
A new `_FakeProvider` test double (implements `list_models`/
|
||||
`load_model`/`unload_model`/`chat`/`chat_structured`, records calls made
|
||||
to it) stands in for `client`. Tests assert: right method called, right
|
||||
arguments passed, return value flows through to the node's output tuple
|
||||
correctly. **No `aiohttp`/`_post_json` mocking at this layer anymore.**
|
||||
This directly matches the protocol contract's own rule ("generic nodes
|
||||
MUST NOT branch on which concrete provider type they received") — if the
|
||||
node tests don't need to know Ollama's wire format, they shouldn't mock
|
||||
it either.
|
||||
- **`tests/test_ollama_provider.py`** (new file) — `OllamaProvider`'s
|
||||
actual Ollama-wire-protocol behavior: `/api/generate`+`keep_alive` int
|
||||
shape, `/api/tags` parsing into `ModelInfo`, header/timeout forwarding,
|
||||
response caching, cache-key composition. This is where the *substance* of
|
||||
today's 35 `_post_json` monkeypatches lands — not 1:1, since several
|
||||
collapse or move (see D5).
|
||||
- **`tests/test_llm_chat_structured.py`** (exists, unchanged) — already
|
||||
covers the shared retry/validation/error-contract mechanism.
|
||||
|
||||
### D5 — Structured-output retry tests are not ported 1:1
|
||||
|
||||
`TestStructuredOutput` has 15 tests monkeypatching `_post_json` to assert
|
||||
exact retry-count behavior (`test_retries_on_invalid_json_then_succeeds`,
|
||||
`test_exhausts_retries_raises_runtime_error`, `test_max_retries_clamped_*`,
|
||||
etc.). Once `OllamaProvider.chat_structured()` delegates to the already-
|
||||
tested shared `chat_structured()` helper (`src/comfydv/_llm/chat.py`,
|
||||
covered by `tests/test_llm_chat_structured.py`'s 6 tests), re-asserting
|
||||
retry counts at the Ollama-node layer duplicates that coverage without
|
||||
adding confidence. Replaced with:
|
||||
- A handful of `test_ollama.py` delegation tests: `ChatCompletion` with
|
||||
`structured_output=True` calls `client.chat_structured(model, messages,
|
||||
schema, ...)` with the right schema/model/messages.
|
||||
- One `test_ollama_provider.py` test: `OllamaProvider.chat_structured()`
|
||||
builds `base_url=f"{self.host}/v1"` and forwards to the shared helper
|
||||
with the right arguments.
|
||||
- `test_structured_output_true_sends_tool_call_payload` (asserts the exact
|
||||
`tools`/`tool_choice` JSON shape) is **deleted** — that's `pydantic-ai`'s
|
||||
internal tool-calling mechanism now, not comfydv's; asserting on a
|
||||
third-party library's internals isn't a test worth keeping.
|
||||
- Pure schema-parsing tests that don't touch HTTP at all (`_parse_output_schema`
|
||||
fail-fast checks, `_coerce_structured_value`, dynamic-socket
|
||||
`RETURN_TYPES` mutation) are unaffected — they test code that stays in
|
||||
`ollama.py` unchanged, only need the class-name rename.
|
||||
|
||||
### D6 — `_client_headers()` is deleted
|
||||
|
||||
Dead code once `OllamaLoadModel`/`OllamaUnloadModel`/`OllamaChatCompletion`
|
||||
delegate to `client.*` (headers become internal to `OllamaProvider`,
|
||||
captured once at construction). No test calls it directly.
|
||||
|
||||
### D7 — Contract doc gets a small fix
|
||||
|
||||
`contracts/llm_provider_protocol.md`'s illustrative code sample is missing
|
||||
`timeout_secs` on `chat()`/`chat_structured()` — the actual `provider.py`
|
||||
(built after the doc) has it. Fix the doc to match the real protocol; docs
|
||||
follow code here, not the reverse.
|
||||
|
||||
### D8 — `conftest.py`'s `_clear_ollama_caches` fixture repoints
|
||||
|
||||
Per D1, there's now one cache source (`comfydv._llm.ollama_provider`).
|
||||
Fixture imports `_CHAT_RESPONSE_CACHE`/`_MODEL_LIST_CACHE` from there
|
||||
instead of `comfydv.ollama`, and `ChatCompletion` instead of
|
||||
`OllamaChatCompletion`. This is the single highest-priority fixture change
|
||||
— every `TestResponseCache` test (12) and every `structured_output`-
|
||||
toggling test depends on it for isolation.
|
||||
|
||||
## Sequenced task list
|
||||
|
||||
Replaces `tasks.md`'s Phase 3 (US1) + Phase 5 (US3) + the deferred T014.
|
||||
One coordinated PR/session, ordered so the codebase stays important at each
|
||||
step even though it can't be split across separate merges (per the
|
||||
2026-07-11 correction — this is genuinely atomic).
|
||||
|
||||
1. **T-CUT-01** — `ollama.py`: add `from comfydv._llm.ollama_provider import OllamaProvider, _fetch_models, _run_async` (drop the local `_post_json`, `_TTLLRUCache`, `_cache_key`, `_MODEL_LIST_CACHE`, `_CHAT_RESPONSE_CACHE`, `_run_async`, `_fetch_models`, `_post_json` definitions — lines 42-194 collapse to the import). Repoint `_load_default_models()` and the `/dv/ollama/models` route to the imported `_fetch_models`. (D1)
|
||||
2. **T-CUT-02** — `ollama_provider.py`: implement `OllamaProvider.list_models()` (port `_fetch_models`'s `/api/tags` logic, map to `ModelInfo`/`ModelStatus.UNLOADED`/`LOADED` — Ollama never emits `SLEEPING`/`DOWNLOADING`, per ADR-007's documented approximation), `load_model()` (port `/api/generate` + `keep_alive: -1`), `unload_model()` (port `/api/generate` + `keep_alive: 0`), `chat()` (port native `/api/chat` non-structured path), `chat_structured()` (build `base_url=f"{self.host}/v1"`, delegate to `comfydv._llm.chat.chat_structured()`).
|
||||
3. **T-CUT-03** — `tests/test_ollama_provider.py` (new): tests for T-CUT-02's method bodies, mocking at `ollama_provider_mod._post_json`/`aiohttp.ClientSession` — ports the *substance* of the 35 relocated monkeypatches per D4/D5 (not 1:1 — collapses redundant retry-count tests per D5).
|
||||
4. **T-CUT-04** — `ollama.py`: `OllamaClient.RETURN_TYPES = ("LLM_CLIENT",)`, `create_client()` returns `OllamaProvider(host, headers)`. (D2)
|
||||
5. **T-CUT-05** — `ollama.py`: rename `OllamaModelSelector`→`LLMModelSelector`, `OllamaLoadModel`→`LLMLoadModel`, `OllamaUnloadModel`→`LLMUnloadModel`, `OllamaChatCompletion`→`ChatCompletion`; every `"OLLAMA_CLIENT"` input socket → `"LLM_CLIENT"`; rewrite the 3 method bodies (`load_model`, `unload_model`, `chat`) to delegate to `client.*` instead of `_post_json`/f-string URLs; delete `_client_headers` (D6); update the 3 `OllamaChatCompletion.*` references in the `/dv/ollama/update_structured_outputs` route body.
|
||||
6. **T-CUT-06** — `src/comfydv/__init__.py`: update imports and `NODE_CLASS_MAPPINGS`/`NODE_DISPLAY_NAME_MAPPINGS` for the 4 renamed classes.
|
||||
7. **T-CUT-07** — `tests/conftest.py`: repoint `_clear_ollama_caches` (D8) and `first_generative_model`'s `_fetch_models` import (D1).
|
||||
8. **T-CUT-08** — `tests/test_ollama.py`: update the import block (4 class renames); add `_FakeProvider` test double; convert every `_post_json`-monkeypatched test to use `_FakeProvider` as `client` instead (D4); replace bare-string `client=ollama_host`/`client="http://..."` usages with a constructed provider (D3); rewrite the 2 `client == "<string>"` assertions (D2); delete `test_plain_string_client_has_no_headers` (D3) and `test_structured_output_true_sends_tool_call_payload` (D5); collapse the 15 `TestStructuredOutput` retry-count tests per D5; update `TestNodeContracts`'s `NODE_CLASSES` list (4 renames).
|
||||
9. **T-CUT-09** — `contracts/llm_provider_protocol.md`: add missing `timeout_secs` params (D7).
|
||||
10. **T-CUT-10** — Full suite green (`uv run pytest -m "not integration and not system"`), `ruff check --fix && ruff format`, `ty check`, `beacon doctor --strict`.
|
||||
11. **T-CUT-11** — `tasks.md`: mark T007-T010/T015-T018/T014 done, referencing this plan; migration mapping (old→new names, FR-009) as a module-level constant/docstring in `ollama.py`.
|
||||
12. **T-CUT-12** — Manual smoke test against a live local Ollama server per `quickstart.md`.
|
||||
|
||||
## What stays exactly as originally scoped
|
||||
|
||||
`OllamaHeader*`, `OllamaOption*`, `OllamaDebugHistory`, `OllamaHistoryLength`
|
||||
classes and the `OLLAMA_HEADERS`/`OLLAMA_OPTIONS`/`OLLAMA_HISTORY` socket
|
||||
types are **out of scope** — confirmed zero test dependencies force a
|
||||
change, and ADR-007 never proposed touching them (only the
|
||||
model-management/chat surface generalizes). `_parse_output_schema`,
|
||||
`_comfy_types_for_schema`, `_build_structured_model`,
|
||||
`_coerce_structured_value` stay in `ollama.py` unchanged — pure/local
|
||||
schema logic with no network dependency, still needed by the live-preview
|
||||
route.
|
||||
@@ -0,0 +1,38 @@
|
||||
# Specification Quality Checklist: LLM Provider Abstraction
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: 2026-07-11
|
||||
**Feature**: [spec.md](../spec.md)
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [x] No implementation details (languages, frameworks, APIs)
|
||||
- [x] Focused on user value and business needs
|
||||
- [x] Written for non-technical stakeholders
|
||||
- [x] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||
- [x] Requirements are testable and unambiguous
|
||||
- [x] Success criteria are measurable
|
||||
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||
- [x] All acceptance scenarios are defined
|
||||
- [x] Edge cases are identified
|
||||
- [x] Scope is clearly bounded
|
||||
- [x] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [x] All functional requirements have clear acceptance criteria
|
||||
- [x] User scenarios cover primary flows
|
||||
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [x] No implementation details leak into specification
|
||||
|
||||
## Notes
|
||||
|
||||
All items pass on first pass — no [NEEDS CLARIFICATION] markers were needed;
|
||||
scope boundaries (llama.cpp out of scope, no automatic workflow migration,
|
||||
no new tracing capability) came directly from the parent epic's Non-goals
|
||||
(`project-management/Roadmap/epics/llm-provider-abstraction.md`) and
|
||||
ADR-007, so no ambiguity required flagging back to the user.
|
||||
@@ -0,0 +1,81 @@
|
||||
# Contract: `LLMProvider` protocol
|
||||
|
||||
This is the interface the follow-on `llamacpp-integration` epic implements
|
||||
against (`LlamaCppProvider`) — it's the actual deliverable that makes ADR-007's
|
||||
adapter pattern real, not internal implementation detail. Treat changes to
|
||||
this contract as requiring epic-level sign-off (per ADR-007's own scope),
|
||||
not a routine refactor.
|
||||
|
||||
```python
|
||||
class ModelStatus(str, Enum):
|
||||
UNLOADED = "unloaded"
|
||||
LOADING = "loading"
|
||||
LOADED = "loaded"
|
||||
SLEEPING = "sleeping" # not all providers emit this
|
||||
DOWNLOADING = "downloading" # not all providers emit this
|
||||
|
||||
class ModelInfo(BaseModel):
|
||||
name: str
|
||||
status: ModelStatus
|
||||
size: int | None = None
|
||||
|
||||
class Message(BaseModel):
|
||||
role: Literal["system", "user", "assistant"]
|
||||
content: str
|
||||
|
||||
class LLMProvider(Protocol):
|
||||
async def list_models(self) -> list[ModelInfo]: ...
|
||||
async def load_model(self, model: str) -> None: ...
|
||||
async def unload_model(self, model: str) -> None: ...
|
||||
async def chat(
|
||||
self, model: str, messages: list[Message], options: dict | None = None,
|
||||
timeout_secs: float = 300.0,
|
||||
) -> str: ...
|
||||
async def chat_structured(
|
||||
self, model: str, messages: list[Message], schema: type[BaseModel],
|
||||
options: dict | None = None, timeout_secs: float = 300.0, max_retries: int = 2,
|
||||
) -> BaseModel: ...
|
||||
```
|
||||
|
||||
## Behavioral requirements (every implementation MUST satisfy)
|
||||
|
||||
- `load_model`/`unload_model` are **idempotent** — calling either on a model
|
||||
already in that state is not an error.
|
||||
- `chat_structured` **MUST NOT** return a `BaseModel` instance with a blank
|
||||
required `str` field — validate and retry (bounded, provider-internal)
|
||||
rather than pass through invalid data. On exhausted retries, raise
|
||||
`RuntimeError` naming the model, attempt count, and a truncated snippet of
|
||||
the last invalid response (FR-004 in `../spec.md`).
|
||||
- `list_models` MUST return every model the server currently knows about,
|
||||
including ones not currently loaded — this is a status listing, not a
|
||||
"loaded models only" filter.
|
||||
- A provider that cannot represent a given `ModelStatus` value (e.g. Ollama
|
||||
has no `sleeping`/`downloading` concept) MUST normalize to the closest
|
||||
applicable status rather than omit the model or invent a new status value
|
||||
outside this enum.
|
||||
- Connection state (host, auth headers, or equivalent) is captured once at
|
||||
provider-construction time; no method takes connection details as a
|
||||
parameter.
|
||||
|
||||
## Non-requirements (explicitly not part of this contract)
|
||||
|
||||
- No requirement that every provider support every `ModelStatus` value —
|
||||
see `data-model.md`'s per-provider emission notes.
|
||||
- No streaming contract — `chat`/`chat_structured` return a complete result,
|
||||
not a stream. (Not requested by the parent spec; a future contract change
|
||||
if ever needed.)
|
||||
- No multi-turn agent/tool-use contract beyond a single structured-output
|
||||
call — out of scope per the parent epic's Non-goals.
|
||||
|
||||
## ComfyUI-facing contract: `LLM_CLIENT` socket
|
||||
|
||||
An `LLMProvider`-implementing instance is the value carried by ComfyUI's
|
||||
`LLM_CLIENT` custom socket type. Any node that outputs `LLM_CLIENT` (e.g.
|
||||
`OllamaClient`, and later `LlamaCppClient`) is committing to have constructed
|
||||
a fully-configured provider instance — no partial/lazy construction that
|
||||
defers connection details to the consuming node.
|
||||
|
||||
Generic nodes (`LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`,
|
||||
`ChatCompletion`) accept `LLM_CLIENT` as their only connection-related input
|
||||
and MUST NOT branch on which concrete provider type they received — doing so
|
||||
would defeat the point of the protocol boundary (ADR-007).
|
||||
@@ -0,0 +1,69 @@
|
||||
# Data Model: LLM Provider Abstraction
|
||||
|
||||
## `ModelStatus` (enum)
|
||||
|
||||
Residency status of a model on a provider's server.
|
||||
|
||||
| Value | Meaning | Emitted by |
|
||||
|---|---|---|
|
||||
| `unloaded` | Known to the server, not resident in memory | all providers |
|
||||
| `loading` | Transitioning into memory | all providers |
|
||||
| `loaded` | Resident and ready to serve requests | all providers |
|
||||
| `sleeping` | Resident but idle-parked | llama.cpp only; Ollama has no distinct signal for this via its API and normalizes resident-and-idle to `loaded` (documented approximation, ADR-007) |
|
||||
| `downloading` | Server is fetching model weights | llama.cpp only; `OllamaProvider` never emits this (Ollama's pull/download flow is out of scope, per the original Ollama epic's non-goals) |
|
||||
|
||||
## `ModelInfo`
|
||||
|
||||
One entry returned by `list_models()`.
|
||||
|
||||
| Field | Type | Notes |
|
||||
|---|---|---|
|
||||
| `name` | `str` | Model identifier as the provider's server knows it |
|
||||
| `status` | `ModelStatus` | See above |
|
||||
| `size` | `int \| None` | Bytes, if the provider reports it; `None` otherwise |
|
||||
|
||||
## `LLMProvider` (Protocol)
|
||||
|
||||
The adapter boundary. Every backend (`OllamaProvider` now, `LlamaCppProvider`
|
||||
in the follow-on epic) implements this shape; ComfyUI nodes depend only on
|
||||
the protocol, never on a concrete provider class.
|
||||
|
||||
| Method | Signature | Notes |
|
||||
|---|---|---|
|
||||
| `list_models` | `async def list_models(self) -> list[ModelInfo]` | |
|
||||
| `load_model` | `async def load_model(self, model: str) -> None` | Idempotent: loading an already-loaded model is not an error |
|
||||
| `unload_model` | `async def unload_model(self, model: str) -> None` | Idempotent: unloading an already-unloaded model is not an error |
|
||||
| `chat` | `async def chat(self, model: str, messages: list[Message], options: dict) -> str` | Free-text response |
|
||||
| `chat_structured` | `async def chat_structured(self, model: str, messages: list[Message], schema: type[BaseModel], options: dict) -> BaseModel` | Validated response; raises on exhausted retries (see FR-004) |
|
||||
|
||||
A concrete provider instance is constructed once per ComfyUI client node with
|
||||
its connection's host/headers as instance state (Constitution Principle V
|
||||
justification — see `research.md`), and that instance is the value carried
|
||||
by the `LLM_CLIENT` ComfyUI socket type.
|
||||
|
||||
## `Message`
|
||||
|
||||
One turn in a chat request, matching the existing shape already sent to
|
||||
Ollama's `/api/chat`/`/v1/chat/completions` (`role` + `content`); unchanged
|
||||
by this feature, carried forward as-is.
|
||||
|
||||
| Field | Type | Notes |
|
||||
|---|---|---|
|
||||
| `role` | `Literal["system", "user", "assistant"]` | |
|
||||
| `content` | `str` | |
|
||||
|
||||
## Relationships
|
||||
|
||||
```
|
||||
ProviderConnection (ComfyUI client node)
|
||||
└─ produces → LLM_CLIENT socket value (an LLMProvider instance)
|
||||
└─ consumed by → LLMModelSelector, LLMLoadModel, LLMUnloadModel, ChatCompletion (ComfyUI nodes)
|
||||
├─ list_models() → ModelInfo[]
|
||||
├─ load_model()/unload_model() → mutates server-side residency, no return value
|
||||
└─ chat()/chat_structured() → str | BaseModel
|
||||
```
|
||||
|
||||
No new persistent storage is introduced — every entity above is
|
||||
constructed per-request or per-node-execution from the connected server's
|
||||
live state; the only caching is the existing in-memory TTL cache for model
|
||||
listing (`_TTLLRUCache`, unchanged, reused inside `OllamaProvider`).
|
||||
@@ -0,0 +1,16 @@
|
||||
Feature: US1 — Connect to a local inference server and get chat responses
|
||||
|
||||
Scenario: Client node feeds a chat node
|
||||
Given a running local inference server and a workflow with a client node wired into a chat node
|
||||
When the workflow executes
|
||||
Then the chat node returns the model's text response
|
||||
|
||||
Scenario: Unreachable server surfaces a clear error
|
||||
Given a client node configured with an unreachable server address
|
||||
When the workflow executes
|
||||
Then the chat node reports a clear connection error rather than hanging indefinitely or crashing the workflow
|
||||
|
||||
Scenario: One client node configures multiple chat nodes
|
||||
Given two chat nodes in the same workflow wired to the same client node
|
||||
When the host address is changed on the client node
|
||||
Then both chat nodes use the new address without being edited individually
|
||||
@@ -0,0 +1,16 @@
|
||||
Feature: US2 — Get structured, validated output instead of parsing raw text
|
||||
|
||||
Scenario: Valid structured response exposes typed fields
|
||||
Given a chat node with structured output enabled and a valid schema
|
||||
When the workflow executes and the model responds correctly
|
||||
Then each schema field is available as its own typed output, and no required field is blank
|
||||
|
||||
Scenario: Invalid response triggers automatic retry
|
||||
Given a model that returns invalid, incomplete, or empty-required-field output
|
||||
When the workflow executes
|
||||
Then the node automatically retries the request up to a configured limit
|
||||
|
||||
Scenario: Exhausted retries fail clearly instead of passing through bad data
|
||||
Given a model that continues to return invalid output after all retries are exhausted
|
||||
When the workflow executes
|
||||
Then the node fails with a clear, specific error rather than silently passing through invalid or partial data
|
||||
@@ -0,0 +1,16 @@
|
||||
Feature: US3 — Manage which models are resident in memory
|
||||
|
||||
Scenario: List models with current status
|
||||
Given a running local server with at least one available model
|
||||
When a workflow author uses the model-listing node
|
||||
Then they see each available model along with its current status
|
||||
|
||||
Scenario: Load a model into memory
|
||||
Given a model that is not currently loaded
|
||||
When a workflow author runs the load-model node against it
|
||||
Then the model becomes loaded and is then usable by the chat node
|
||||
|
||||
Scenario: Unload a model from memory
|
||||
Given a model that is loaded and idle
|
||||
When a workflow author runs the unload-model node against it
|
||||
Then the model is freed from memory and its reported status updates accordingly
|
||||
@@ -0,0 +1,12 @@
|
||||
Feature: US4 — Reconnect an existing workflow after upgrading
|
||||
|
||||
Scenario: Renamed nodes are reported with a documented replacement
|
||||
Given a saved workflow using the current Ollama-specific node and connection-socket names
|
||||
When it is opened after upgrading
|
||||
Then ComfyUI reports the now-missing node types
|
||||
And documentation identifies the replacement node for each one
|
||||
|
||||
Scenario: Reconnected workflow produces equivalent output
|
||||
Given a workflow that has been reconnected to the new generic nodes
|
||||
When it executes with the same inputs and model as before the upgrade
|
||||
Then it produces equivalent output
|
||||
@@ -0,0 +1,121 @@
|
||||
# Implementation Plan: LLM Provider Abstraction
|
||||
|
||||
**Branch**: `007-llm-provider-abstraction` | **Date**: 2026-07-11 | **Spec**: [spec.md](./spec.md)
|
||||
|
||||
**Input**: Feature specification from `/specs/007-llm-provider-abstraction/spec.md`
|
||||
|
||||
**Note**: This template is filled in by the `/speckit-plan` command. See `.specify/templates/plan-template.md` for the execution workflow.
|
||||
|
||||
## Summary
|
||||
|
||||
Define a shared `LLMProvider` protocol (list/load/unload/chat/structured-chat)
|
||||
and generic ComfyUI nodes so workflow authors can connect any supported local
|
||||
inference backend the same way. Migrate the existing Ollama integration onto
|
||||
it — `OllamaProvider` becomes the first (and, in this feature, only)
|
||||
implementation — including moving structured-output from ADR-006's
|
||||
hand-rolled tool-calling onto `pydantic-ai`, per ADR-007. This is a
|
||||
behavior-preserving mechanism swap for existing capability, plus the new
|
||||
protocol boundary that the follow-on `llamacpp-integration` epic builds a
|
||||
second provider against.
|
||||
|
||||
## Technical Context
|
||||
|
||||
**Language/Version**: Python ≥3.11 (per `pyproject.toml`)
|
||||
|
||||
**Primary Dependencies**: `aiohttp` (existing, unchanged — model-management
|
||||
REST calls), `pydantic` (existing, unchanged — validation), `pydantic-ai` +
|
||||
`openai` (new, per ADR-007 — powers `chat_structured()` only)
|
||||
|
||||
**Storage**: N/A — no persistent storage; the existing in-memory
|
||||
`_TTLLRUCache` for model listing is reused unchanged inside `OllamaProvider`
|
||||
|
||||
**Testing**: `pytest` via `uv run pytest`, following `tests/test_ollama.py`'s
|
||||
existing conventions (mocked `aiohttp`/`pydantic-ai` calls, no live server
|
||||
required for unit tests; the existing `integration` pytest marker — "requiring
|
||||
live Ollama at localhost:11434" — is reused for tests that exercise a real
|
||||
server)
|
||||
|
||||
**Target Platform**: ComfyUI custom-node runtime, cross-platform wherever
|
||||
ComfyUI runs; CPU-only dev harness per the project's stated vision
|
||||
|
||||
**Project Type**: Library / ComfyUI custom-node pack (single project,
|
||||
existing `src/comfydv/` layout — no new top-level project)
|
||||
|
||||
**Performance Goals**: No new numeric target; must not add latency beyond the
|
||||
existing bounded retry loop already in ADR-006 (`max_retries`, 0–5)
|
||||
|
||||
**Constraints**: Behavior-preserving for existing Ollama structured/
|
||||
non-structured chat and all model-management calls (FR-007, FR-008); no new
|
||||
dependency beyond what ADR-007 already accepted (`pydantic-ai`, `openai`,
|
||||
their transitive `httpx`/`tiktoken`); model-management stays on `aiohttp`
|
||||
|
||||
**Scale/Scope**: One new internal package (`src/comfydv/_llm/`), migration of
|
||||
the existing 1056-line `ollama.py` node/HTTP logic to consume it, five
|
||||
ComfyUI node classes renamed to generic names — no change to the project's
|
||||
single-repo, single-package scope
|
||||
|
||||
## Constitution Check
|
||||
|
||||
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
|
||||
|
||||
| Principle | Verdict | Notes |
|
||||
|---|---|---|
|
||||
| I. ComfyUI Contract First | PASS | Generic nodes (`ChatCompletion`, `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`, `OllamaClient`) still expose `INPUT_TYPES`/`RETURN_TYPES`/`RETURN_NAMES`/`FUNCTION`/`CATEGORY`; `NODE_CLASS_MAPPINGS` in `__init__.py` remains the only install-time interface. `LLMProvider` is internal, not a ComfyUI-facing contract change beyond the node/socket rename. |
|
||||
| II. Sandbox All User-Supplied Code | N/A | No template/expression evaluation in this feature — structured-output schemas are parsed as JSON Schema by `pydantic`, never `eval`/`exec`. |
|
||||
| III. Test-First | PASS (binding on tasks/implement phases) | `tests/test_ollama.py`'s existing assertions are the regression oracle (see `research.md`); new `_llm` package gets tests written before implementation, red→green→refactor. |
|
||||
| IV. Graceful Degradation Outside ComfyUI | PASS (binding on implementation) | `src/comfydv/_llm/` must not import `comfy`/`server` at module scope, matching `ollama.py`'s existing guarded-import pattern. |
|
||||
| V. Simplicity — Function Before Class | **Justified exception — see Complexity Tracking** | `LLMProvider` is a `Protocol` implemented by stateful provider classes, not module-level functions. |
|
||||
| VI. Fixed Output Positions | PASS (binding on implementation) | `ChatCompletion`'s (renamed from `OllamaChatCompletion`) `RETURN_TYPES`/`RETURN_NAMES` positions 0/1 carry forward unchanged — only the class/node name and internal mechanism change. |
|
||||
|
||||
Re-checked post-Phase 1 design (data-model.md, contracts/): unchanged — the
|
||||
`Protocol`-based design in `contracts/llm_provider_protocol.md` is exactly
|
||||
what was justified below, no new gate violations introduced by the detailed
|
||||
design.
|
||||
|
||||
## Project Structure
|
||||
|
||||
### Documentation (this feature)
|
||||
|
||||
```text
|
||||
specs/[###-feature]/
|
||||
├── plan.md # This file (/speckit-plan command output)
|
||||
├── research.md # Phase 0 output (/speckit-plan command)
|
||||
├── data-model.md # Phase 1 output (/speckit-plan command)
|
||||
├── quickstart.md # Phase 1 output (/speckit-plan command)
|
||||
├── contracts/ # Phase 1 output (/speckit-plan command)
|
||||
└── tasks.md # Phase 2 output (/speckit-tasks command - NOT created by /speckit-plan)
|
||||
```
|
||||
|
||||
### Source Code (repository root)
|
||||
|
||||
```text
|
||||
src/comfydv/
|
||||
├── ollama.py # existing — node classes renamed to generic names,
|
||||
│ # delegates HTTP/chat logic to _llm/ internally
|
||||
├── _llm/ # new internal package (not a ComfyUI node module)
|
||||
│ ├── __init__.py
|
||||
│ ├── provider.py # LLMProvider Protocol, ModelStatus, ModelInfo, Message
|
||||
│ ├── ollama_provider.py # OllamaProvider — wraps existing aiohttp REST logic
|
||||
│ └── chat.py # shared chat_structured() pydantic-ai helper
|
||||
└── __init__.py # NODE_CLASS_MAPPINGS updated for renamed nodes
|
||||
|
||||
tests/
|
||||
├── test_ollama.py # existing — updated for renamed nodes; behavior-
|
||||
│ # preserving assertions carried forward unchanged
|
||||
└── test_llm_provider.py # new — protocol conformance + OllamaProvider unit tests
|
||||
```
|
||||
|
||||
**Structure Decision**: Single project (existing `src/comfydv/` layout, no new
|
||||
top-level project). New internal package `src/comfydv/_llm/` (underscore
|
||||
prefix marks it as internal, consistent with existing internal helpers like
|
||||
`_TTLLRUCache` that already live inside `ollama.py`) hosts the protocol and
|
||||
shared chat logic; `ollama.py` keeps the actual ComfyUI-registered node
|
||||
classes and becomes a thin caller into `_llm`.
|
||||
|
||||
## Complexity Tracking
|
||||
|
||||
> **Fill ONLY if Constitution Check has violations that must be justified**
|
||||
|
||||
| Violation | Why Needed | Simpler Alternative Rejected Because |
|
||||
|-----------|------------|-------------------------------------|
|
||||
| `LLMProvider` as a `Protocol` implemented by stateful classes (Principle V: Function Before Class) | Every one of the five protocol methods (`list_models`/`load_model`/`unload_model`/`chat`/`chat_structured`) needs the same connection state (host, auth headers) — genuine shared state, the exact condition under which the constitution allows a class. A `Protocol` also lets ComfyUI's `LLM_CLIENT` socket carry one opaque object satisfying the shape, which is what makes the adapter pattern (ADR-007) work on the canvas. | Module-level functions taking host/headers as explicit parameters on every call were considered and rejected: they'd reintroduce the exact per-call-site repetition [ADR-005](../../project-management/ADRs/ADR-005-ollama-host-config-via-client-node.md)'s config-node pattern was built to eliminate, and a bare function can't be the typed payload of a ComfyUI socket the way an object implementing a `Protocol` can. |
|
||||
@@ -0,0 +1,49 @@
|
||||
# Quickstart: LLM Provider Abstraction
|
||||
|
||||
A minimal ComfyUI workflow using the generic nodes this feature introduces.
|
||||
|
||||
## 1. Connect to a local server
|
||||
|
||||
Add an **Ollama Client** node. Set its host widget (default
|
||||
`http://localhost:11434`). This is the only node that knows it's talking to
|
||||
Ollama specifically — everything downstream just sees `LLM_CLIENT`.
|
||||
|
||||
## 2. Chat
|
||||
|
||||
Add a **Chat Completion** node. Wire the client node's `LLM_CLIENT` output
|
||||
into it. Set a model name (or feed one from a model-selector node — see
|
||||
below) and a prompt. Run the workflow: the node returns the model's text
|
||||
response.
|
||||
|
||||
## 3. Get structured output instead of free text
|
||||
|
||||
On the same **Chat Completion** node, enable `structured_output` and supply
|
||||
a JSON Schema (e.g. `{"type": "object", "properties": {"summary": {"type": "string"}, "score": {"type": "number"}}, "required": ["summary", "score"]}`).
|
||||
Re-run: the node now exposes one typed output socket per schema property
|
||||
(`summary`, `score`) instead of a single text blob, and guarantees neither
|
||||
is blank.
|
||||
|
||||
## 4. Manage what's loaded in memory
|
||||
|
||||
Add an **LLM Model Selector** node wired to the same client, to see every
|
||||
model the server knows about and its current status (`unloaded` /
|
||||
`loading` / `loaded` / …). Add **LLM Load Model** / **LLM Unload Model**
|
||||
nodes, wired to the same client, to explicitly control residency before a
|
||||
chat node needs a model.
|
||||
|
||||
## 5. (Follow-on epic) Swap backends without touching downstream nodes
|
||||
|
||||
Once `llamacpp-integration` ships a **Llama.cpp Client** node, replacing the
|
||||
**Ollama Client** node in step 1 with it — and nothing else — is the whole
|
||||
migration: it emits the same `LLM_CLIENT` socket type, so every node from
|
||||
steps 2–4 keeps working unmodified. That's the point of this feature.
|
||||
|
||||
## Migrating an existing pre-upgrade workflow
|
||||
|
||||
If you have a saved workflow using the old node names (`OllamaClient`,
|
||||
`OllamaChatCompletion`, `OllamaModelSelector`, `OllamaLoadModel`,
|
||||
`OllamaUnloadModel`), ComfyUI will report those node types as missing on
|
||||
load. Replace each with its generic equivalent from the list above and
|
||||
reconnect — behavior is unchanged, only the node names and the
|
||||
`LLM_CLIENT` socket type (replacing `OLLAMA_CLIENT`) are different. See
|
||||
`spec.md`'s User Story 4 and Edge Cases for the full detail.
|
||||
@@ -0,0 +1,94 @@
|
||||
# Research: LLM Provider Abstraction
|
||||
|
||||
All unknowns below were already resolved during DESIGN-phase work on
|
||||
[ADR-007](../../project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md);
|
||||
this file consolidates that research for the plan gate rather than re-deriving it.
|
||||
|
||||
## Decision: `pydantic-ai` for `chat_structured()`, not hand-rolled tool-calling
|
||||
|
||||
**Decision**: Both `OllamaProvider` and the future `LlamaCppProvider` implement
|
||||
`chat_structured()` via `pydantic-ai`'s `Agent`/`output_type`, called through
|
||||
`OpenAIProvider(base_url=<host>/v1)`.
|
||||
|
||||
**Rationale**: A live research pass against current `pydantic-ai` docs/source
|
||||
found `httpx` is a base dependency of `pydantic-ai-slim` itself (not merely
|
||||
pulled in by an OpenAI extra), and `openai`+`tiktoken` are required for any
|
||||
OpenAI-compatible provider — a fixed, one-time dependency tax rather than a
|
||||
per-backend one. `pydantic.create_model()`-built `BaseModel` subclasses
|
||||
(comfydv's existing dynamic-schema pattern) work as `output_type` with no
|
||||
special-casing. `OpenAIProvider(base_url=...)` is one generic code path both
|
||||
Ollama's and llama.cpp's OpenAI-compatible `/v1/chat/completions` reach
|
||||
identically.
|
||||
|
||||
**Alternatives considered**: hand-roll llama.cpp's structured output too
|
||||
(duplicates [ADR-006](../../project-management/ADRs/ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md)'s
|
||||
mechanism — rejected, defeats the DRY goal); extract a shared aiohttp-based
|
||||
helper with no new dependencies (rejected — forces re-deriving pydantic-ai's
|
||||
retry/validation machinery by hand for no benefit now that two backends exist
|
||||
to amortize the dependency cost against).
|
||||
|
||||
## Decision: aiohttp stays authoritative for model-management REST calls
|
||||
|
||||
**Decision**: `list_models()` / `load_model()` / `unload_model()` on every
|
||||
provider use `aiohttp` — no dependency change from the existing Ollama
|
||||
integration for this surface.
|
||||
|
||||
**Rationale**: [ADR-004](../../project-management/ADRs/ADR-004-aiohttp-over-httpx-for-ollama.md)'s
|
||||
reasoning (ComfyUI's own server is aiohttp-based; httpx was an unjustified
|
||||
addition) still applies fully to REST calls that don't need pydantic-ai's
|
||||
machinery. ADR-007 narrows ADR-004's scope to exactly this surface, rather
|
||||
than superseding it.
|
||||
|
||||
**Alternatives considered**: route everything (including model management)
|
||||
through `pydantic-ai`/httpx for consistency — rejected, pydantic-ai has no
|
||||
model-lifecycle-management concept (it's a chat/agent framework, not a
|
||||
generic REST client) and would add no value over plain aiohttp calls that
|
||||
already exist and work.
|
||||
|
||||
## Decision: `LLMProvider` as a `Protocol` implemented by stateful provider classes
|
||||
|
||||
**Decision**: `list_models`/`load_model`/`unload_model`/`chat`/`chat_structured`
|
||||
are defined as a `typing.Protocol`, implemented by `OllamaProvider` (and later
|
||||
`LlamaCppProvider`) classes, each constructed once per ComfyUI client node
|
||||
with the connection's host/headers as instance state.
|
||||
|
||||
**Rationale**: This is a Constitution Principle V ("Function Before Class")
|
||||
gate — classes are only justified when there's shared state a group of
|
||||
functions would otherwise have to thread through every call. Here there is:
|
||||
every one of the five protocol methods needs the same host/headers, exactly
|
||||
the connection config [ADR-005](../../project-management/ADRs/ADR-005-ollama-host-config-via-client-node.md)'s
|
||||
config-node pattern centralizes. A `Protocol` (structural typing, no
|
||||
inheritance required) keeps this lightweight — `LlamaCppProvider` doesn't
|
||||
need to import or subclass `OllamaProvider`, it only needs to match the
|
||||
method shapes.
|
||||
|
||||
**Alternatives considered**: module-level functions taking host/headers as
|
||||
explicit parameters on every call — rejected, this reintroduces the exact
|
||||
per-call-site repetition ADR-005 eliminated, and loses the ability for a
|
||||
ComfyUI `LLM_CLIENT` socket to carry one opaque object implementing the
|
||||
protocol (functions can't be typed as a socket payload the way an object
|
||||
implementing a `Protocol` can).
|
||||
|
||||
## Decision: behavior-preserving migration, verified against existing tests
|
||||
|
||||
**Decision**: `tests/test_ollama.py`'s existing assertions (retry bounds,
|
||||
required-string validation, error messages) are the acceptance bar for the
|
||||
migrated `OllamaProvider.chat_structured()` — this is a mechanism swap, not a
|
||||
new capability, per the parent epic's Non-goals and spec FR-008.
|
||||
|
||||
**Rationale**: Constitution Principle III (Test-First) and the epic's
|
||||
explicit framing of this as the riskiest change in the whole llama.cpp
|
||||
proposal (touches a Done, shipped epic's code) both point the same way: the
|
||||
existing test suite is the regression oracle, not a new one written from
|
||||
scratch.
|
||||
|
||||
## Testing approach
|
||||
|
||||
Per Constitution Principle IV (Graceful Degradation Outside ComfyUI), the new
|
||||
`src/comfydv/_llm/` package must not import `comfy`/`server` at module scope,
|
||||
matching `ollama.py`'s existing runtime-guarded pattern. Unit tests mock
|
||||
`aiohttp`/`pydantic-ai` calls (no live server required, matching
|
||||
`tests/test_ollama.py`'s existing convention); the `integration` pytest
|
||||
marker (already defined in `pyproject.toml`, "requiring live Ollama at
|
||||
localhost:11434") is reused, not redefined, for tests that exercise a real
|
||||
local server.
|
||||
@@ -0,0 +1,119 @@
|
||||
# Feature Specification: LLM Provider Abstraction
|
||||
|
||||
**Feature Branch**: `007-llm-provider-abstraction`
|
||||
|
||||
**Created**: 2026-07-11
|
||||
|
||||
**Status**: Draft
|
||||
|
||||
**Input**: User description: "Introduce a shared LLMProvider protocol for comfydv's LLM backend nodes so ComfyUI workflow authors can swap between local inference servers (starting with Ollama, with llama.cpp planned next) without changing their chat/model-management nodes. Migrate the existing Ollama integration onto generic nodes (client config, model list, load, unload, chat completion with optional structured/validated output) backed by this shared interface, per ADR-007."
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Connect to a local inference server and get chat responses (Priority: P1)
|
||||
|
||||
As a ComfyUI workflow author, I want to point a single configuration node at my local LLM server and get chat responses through a generic chat node, so I can generate text without hardcoding a server address into every node that needs one.
|
||||
|
||||
**Why this priority**: This is the minimum viable path — without a working connection and a basic chat response, nothing else in this feature has value. It also directly replaces the most-used capability of the existing Ollama integration, so it carries the highest regression risk.
|
||||
|
||||
**Independent Test**: Wire a client configuration node into a chat node, run the workflow against a running local server, and confirm the chat node returns the model's text response.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a running local inference server and a workflow with a client node wired into a chat node, **When** the workflow executes, **Then** the chat node returns the model's text response.
|
||||
2. **Given** a client node configured with an unreachable server address, **When** the workflow executes, **Then** the chat node reports a clear connection error rather than hanging indefinitely or crashing the workflow.
|
||||
3. **Given** two chat nodes in the same workflow wired to the same client node, **When** the host address is changed on the client node, **Then** both chat nodes use the new address without being edited individually.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Get structured, validated output instead of parsing raw text (Priority: P1)
|
||||
|
||||
As a workflow author, I want to describe the shape of data I need and turn on structured output for a chat node, so downstream nodes receive individually typed fields I can trust are present and non-empty, instead of me parsing free text myself.
|
||||
|
||||
**Why this priority**: This is an existing, relied-upon capability of the current Ollama integration (structured output with retry-on-invalid-response). Preserving it exactly is required for this migration to be considered safe, so it's equal priority to basic chat.
|
||||
|
||||
**Independent Test**: Enable structured output on a chat node with a schema describing two or three fields, run the workflow against a model, and confirm each schema field is exposed as its own typed output socket with a valid value.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a chat node with structured output enabled and a valid schema, **When** the workflow executes and the model responds correctly, **Then** each schema field is available as its own typed output, and no required field is blank.
|
||||
2. **Given** a model that returns invalid, incomplete, or empty-required-field output, **When** the workflow executes, **Then** the node automatically retries the request up to a configured limit.
|
||||
3. **Given** a model that continues to return invalid output after all retries are exhausted, **When** the workflow executes, **Then** the node fails with a clear, specific error rather than silently passing through invalid or partial data.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Manage which models are resident in memory (Priority: P2)
|
||||
|
||||
As a workflow author running models locally, I want to see which models are currently loaded, loading, or unloaded, and explicitly load or unload a model, so I can control memory usage on my machine without leaving ComfyUI or using a separate terminal.
|
||||
|
||||
**Why this priority**: Valuable and already present in the current Ollama integration, but a workflow can still generate output without ever calling load/unload explicitly (servers can auto-load on first use) — so this is lower risk to defer than basic chat.
|
||||
|
||||
**Independent Test**: Use a model-listing node against a running server, confirm it shows each available model with a current status; use load/unload nodes against one model and confirm its reported status changes accordingly.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a running local server with at least one available model, **When** a workflow author uses the model-listing node, **Then** they see each available model along with its current status.
|
||||
2. **Given** a model that is not currently loaded, **When** a workflow author runs the load-model node against it, **Then** the model becomes loaded and is then usable by the chat node.
|
||||
3. **Given** a model that is loaded and idle, **When** a workflow author runs the unload-model node against it, **Then** the model is freed from memory and its reported status updates accordingly.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Reconnect an existing workflow after upgrading (Priority: P3)
|
||||
|
||||
As an existing user of the current Ollama nodes, when I open a workflow I saved before this change, I want it to be clear which new node replaces each renamed one, so I can reconnect my workflow with minimal effort and get the same results as before.
|
||||
|
||||
**Why this priority**: This is migration friction, not new capability — it matters for a good upgrade experience but doesn't block anyone building a new workflow from scratch, so it's the lowest priority of the four.
|
||||
|
||||
**Independent Test**: Open a workflow saved against the current Ollama-specific node names, follow the provided migration guidance to reconnect it to the new generic nodes, and confirm it produces the same output as before, given the same inputs and model.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a saved workflow using the current Ollama-specific node and connection-socket names, **When** it is opened after upgrading, **Then** ComfyUI reports the now-missing node types (standard ComfyUI behavior for renamed nodes), and documentation identifies the replacement node for each one.
|
||||
2. **Given** a workflow that has been reconnected to the new generic nodes, **When** it executes with the same inputs and model as before the upgrade, **Then** it produces equivalent output.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when the configured server address is unreachable at the moment a model-listing, load, or unload node runs (not just the chat node)?
|
||||
- What happens when a workflow author supplies an invalid or malformed schema to structured output, rather than an invalid model response?
|
||||
- What happens when the connected server does not support structured/validated output at all?
|
||||
- What happens to an in-flight chat request if the model it depends on is unloaded by another node in the same workflow run?
|
||||
- What happens when a workflow author tries to wire a pre-upgrade Ollama-specific node's output into a new generic node, or vice versa? (Expected: ComfyUI's own type-checking refuses the connection, since the socket types differ — this is the intended, safe failure mode, not a bug to work around.)
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: The system MUST allow a workflow author to configure a connection to a local inference server once and reuse that single configuration across multiple nodes in the same workflow.
|
||||
- **FR-002**: The system MUST allow a workflow author to request either free-text or schema-validated structured output from the same chat node, choosing per request.
|
||||
- **FR-003**: When structured output is requested, the system MUST validate the response against the supplied schema and MUST NOT deliver output to downstream nodes where a required field is missing or empty.
|
||||
- **FR-004**: When validation fails, the system MUST retry the request automatically up to a configurable limit before reporting a clear, actionable error that identifies the model, the number of attempts made, and a snippet of the last invalid response.
|
||||
- **FR-005**: The system MUST allow a workflow author to list available models on a connected server along with each model's current residency status.
|
||||
- **FR-006**: The system MUST allow a workflow author to explicitly load a model into memory and explicitly unload a model from memory.
|
||||
- **FR-007**: The system's chat and model-management nodes MUST behave identically regardless of which supported local inference server is connected, given equivalent inputs.
|
||||
- **FR-008**: The existing chat and structured-output behavior for the currently-supported local inference server (Ollama) MUST be unchanged in outcome after this migration — same retry limits, same validation rules, same error conditions — since this feature changes the underlying mechanism, not the capability.
|
||||
- **FR-009**: The system MUST document, for each node type renamed or removed by this change, which new node replaces it.
|
||||
|
||||
### Key Entities *(include if feature involves data)*
|
||||
|
||||
- **Provider connection**: A configured connection to one local inference server (address and any authentication), created once and reused by every model-management and chat node that needs it.
|
||||
- **Model**: An inference model known to a provider connection, identified by name, with a current residency status (e.g., unloaded, loading, loaded, and — on servers that support it — sleeping or downloading).
|
||||
- **Chat request/response**: A request for a model's output, optionally carrying a schema describing the required shape of a structured response, and the corresponding validated or free-text result.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: A workflow author can go from no nodes to a working chat response using no more than two nodes (one connection node, one chat node).
|
||||
- **SC-002**: Structured-output workflows never deliver a blank or missing required field to a downstream node — every request either produces fully valid data or a clear error, with zero silent partial results.
|
||||
- **SC-003**: Existing example/reference workflows built against the current Ollama nodes remain reproducible on the new nodes with equivalent output, after a workflow author reconnects the renamed nodes.
|
||||
- **SC-004**: Adding support for a second local inference server (planned as a follow-on feature) requires no visible change to chat or model-management node behavior — only a new connection node is needed.
|
||||
|
||||
## Assumptions
|
||||
|
||||
- Workflow authors run their own local inference server (e.g., Ollama) reachable over HTTP from the machine running ComfyUI; this feature does not host, install, or manage that server.
|
||||
- Users with workflows saved against the current Ollama-specific node and socket names will need to manually reconnect them after upgrading. This is an accepted, intentional breaking change (confirmed 2026-07-11), not a defect — see FR-009 for the mitigation (documented replacement mapping), not automatic migration.
|
||||
- Support for a second local inference server (llama.cpp) is planned as a separate, follow-on feature and is out of scope here — this feature only needs to prove the shared design works end-to-end for one real backend (Ollama).
|
||||
- Structured-output schemas remain limited to flat object shapes with typed properties, consistent with what the current Ollama integration already supports — deeper nested schemas are unaffected by (neither improved nor degraded by) this change.
|
||||
- No new observability/tracing capability is introduced for workflow authors as part of this feature, even though the underlying mechanism change makes it feasible to add later.
|
||||
@@ -0,0 +1,267 @@
|
||||
# Tasks: LLM Provider Abstraction
|
||||
|
||||
**Input**: Design documents from `/specs/007-llm-provider-abstraction/`
|
||||
|
||||
**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/llm_provider_protocol.md
|
||||
|
||||
**Tests**: First-class (spec carries Acceptance Scenarios) — every implementation task has a paired failing-test task (`-T`/`-I` suffix) per BEACON's test-first discipline.
|
||||
|
||||
**Organization**: Tasks are grouped by user story (spec.md priorities P1/P1/P2/P3) to enable independent implementation and testing of each.
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: Which user story this task belongs to (US1–US4)
|
||||
- **-T / -I**: paired test (red) / implementation (green) — a `-I` task is never parallel with its own `-T`
|
||||
|
||||
## Path Conventions
|
||||
|
||||
Single project: `src/comfydv/`, `tests/` at repository root (per plan.md's Project Structure).
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup
|
||||
|
||||
- [x] T001 Add `pydantic-ai` and `openai` to `pyproject.toml` dependencies; curate the addition into `requirements.txt` per [ADR-003](../../project-management/ADRs/ADR-003-requirements-txt-authoring-policy.md)
|
||||
- [x] T002 [P] Create `src/comfydv/_llm/__init__.py` (empty package init)
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**⚠️ CRITICAL**: No user story work can begin until this phase is complete.
|
||||
|
||||
- [x] T003 [P] Define `Message`, `ModelStatus`, `ModelInfo` in `src/comfydv/_llm/provider.py` per `data-model.md`
|
||||
- [x] T004 Define the `LLMProvider` `Protocol` in `src/comfydv/_llm/provider.py` per `contracts/llm_provider_protocol.md` (depends on T003)
|
||||
- [x] T005 [P] `LLM_CLIENT` is introduced as part of T007-I (`OllamaClient`'s output type) rather than as a standalone constant — ComfyUI socket types are plain string literals, not declared objects; folded in, not skipped
|
||||
- [x] T006 Scaffold `OllamaProvider.__init__(self, host, headers)` in `src/comfydv/_llm/ollama_provider.py`, porting the existing module-level `_post_json`/`_fetch_models`/`_run_async`/`_TTLLRUCache` helpers from `ollama.py` into it — behavior-preserving port, not a rewrite (depends on T004)
|
||||
|
||||
**Checkpoint**: protocol + provider skeleton exist; user story work can begin.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 — Connect to a local server and get chat responses (Priority: P1) 🎯 MVP
|
||||
|
||||
**Goal**: A workflow author wires a client node into a generic chat node and gets a text response.
|
||||
|
||||
**Independent Test**: Wire `OllamaClient` → `ChatCompletion`, run against a live server, confirm text output.
|
||||
|
||||
**Superseded 2026-07-11 — see `atomic-cutover-plan.md`.** T007–T010 as
|
||||
written below assumed US1 was independently deliverable; it isn't (see the
|
||||
correction at the bottom of this file). The actual work is now
|
||||
`atomic-cutover-plan.md`'s **T-CUT-04, T-CUT-05, T-CUT-06, T-CUT-08**
|
||||
(`OllamaClient` output-type change, `ChatCompletion` rename+delegation,
|
||||
`__init__.py` registration, and the corresponding `test_ollama.py`
|
||||
rewrite using a `_FakeProvider` double). Original text kept below for
|
||||
history, not as the active task list:
|
||||
|
||||
- [-] T007-T [US1] ~~Write FAILING test: `OllamaClient` node constructs and outputs an `OllamaProvider` via the `LLM_CLIENT` socket, in `tests/test_ollama.py`~~ _Superseded, see Phase 8 T-CUT-08._
|
||||
- [-] T007-I [US1] ~~Update `OllamaClient` in `src/comfydv/ollama.py` to construct and output an `OllamaProvider` via `LLM_CLIENT`~~ _Superseded, see Phase 8 T-CUT-04._
|
||||
- [-] T008-T [US1] ~~Write FAILING test: `OllamaProvider.chat()` returns model text via the existing `/api/chat` aiohttp call~~ _Superseded, see Phase 8 T-CUT-03._
|
||||
- [-] T008-I [US1] ~~Implement `OllamaProvider.chat()` in `src/comfydv/_llm/ollama_provider.py`~~ _Superseded, see Phase 8 T-CUT-02._
|
||||
- [-] T009-T [US1] ~~Write FAILING test: generic `ChatCompletion` node (non-structured) calls `provider.chat()` and surfaces a clear error on an unreachable host~~ _Superseded, see Phase 8 T-CUT-08._
|
||||
- [-] T009-I [US1] ~~Rename `OllamaChatCompletion` → `ChatCompletion`, delegate the non-structured path to `LLMProvider.chat()`, update `NODE_CLASS_MAPPINGS`~~ _Superseded, see Phase 8 T-CUT-05/T-CUT-06._
|
||||
- [-] T010-T [US1] ~~Write FAILING test: two `ChatCompletion` nodes sharing one `OllamaClient` both pick up a host change~~ _Superseded, see Phase 8 T-CUT-08._
|
||||
- [-] T010-I [US1] ~~Verify/adjust that `OllamaClient` → `OllamaProvider` construction happens per node execution~~ _Superseded, folded into Phase 8 T-CUT-04 (construction is already per-call in `create_client()`)._
|
||||
|
||||
**Checkpoint**: superseded — see `atomic-cutover-plan.md`'s checkpoint (T-CUT-10, full suite green).
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 — Structured, validated output (Priority: P1)
|
||||
|
||||
**Goal**: The chat node's `structured_output` toggle returns validated typed fields via the shared `pydantic-ai` mechanism.
|
||||
|
||||
**Independent Test**: Enable `structured_output` with a schema, run against a model, confirm typed sockets are populated and never blank.
|
||||
|
||||
- [x] T011-T [P] [US2] Write test: `chat_structured()` returns a validated schema instance on success, in `tests/test_llm_chat_structured.py` (witnesses `features/us2_structured_output.feature` scenario "Valid structured response exposes typed fields") — mocked at the `_build_agent` seam after an API-discovery spike into pydantic-ai's exact `Agent`/`OpenAIProvider`/`OpenAIChatModel` constructor and exception surface (not guessable from training data alone — verified live against the installed package); tests and implementation validated together rather than strictly red-first, noted honestly rather than presented as pure TDD
|
||||
- [x] T011-I [US2] Implement the shared `chat_structured()` helper in `src/comfydv/_llm/chat.py` using `pydantic-ai`'s `Agent`/`output_type` through `OpenAIProvider(base_url=<host>/v1)` + `OpenAIChatModel` — makes T011-T pass (depends on T004)
|
||||
- [x] T012-T [US2] Write test: invalid/failed-validation responses trigger automatic retry up to `max_retries` (clamped 0–5), in `tests/test_llm_chat_structured.py` (witnesses `features/us2_structured_output.feature` scenario "Invalid response triggers automatic retry")
|
||||
- [x] T012-I [US2] Implement the bounded retry loop (0–5, matching ADR-006's existing contract) around the `pydantic-ai` call in `src/comfydv/_llm/chat.py`, with the Agent's own internal retries disabled (`retries=0`) so the error contract is comfydv's — makes T012-T pass (depends on T011-I)
|
||||
- [x] T013-T [US2] Write test: exhausted retries raise `RuntimeError` naming the model, attempt count, and a truncated last-response snippet, in `tests/test_llm_chat_structured.py` (witnesses `features/us2_structured_output.feature` scenario "Exhausted retries fail clearly instead of passing through bad data")
|
||||
- [x] T013-I [US2] Implement the exhausted-retry error path in `src/comfydv/_llm/chat.py`, matching ADR-006's existing error message contract (model, attempt count, truncated last response) — makes T013-T pass (depends on T012-I)
|
||||
- [-] T014-T [US2] _Superseded — see `atomic-cutover-plan.md` D5 and T-CUT-08. Original: write FAILING test that `ChatCompletion`'s `structured_output=True` path wires a schema through `chat_structured()` to per-field dynamic ComfyUI output sockets. D5 replaces the originally-planned retry-count-style test with a delegation test against a `_FakeProvider`, since retry behavior is already covered by `tests/test_llm_chat_structured.py`._
|
||||
- [-] T014-I [US2] _Superseded — see `atomic-cutover-plan.md` T-CUT-05. Original: wire `ChatCompletion`'s `structured_output`/`output_schema` inputs to `LLMProvider.chat_structured()`, preserving the dynamic-socket UX from ADR-006 — still the right implementation shape, just executed as part of the coordinated T-CUT-05 rename, not standalone._
|
||||
|
||||
**Checkpoint**: US1 + US2 both independently functional — matches today's Ollama capability, now on the shared mechanism.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 — Manage model residency (Priority: P2)
|
||||
|
||||
**Goal**: List/load/unload models through generic nodes against any connected provider.
|
||||
|
||||
**Independent Test**: List models via `LLMModelSelector`; load/unload one via `LLMLoadModel`/`LLMUnloadModel`; confirm status changes.
|
||||
|
||||
**Superseded 2026-07-11 — see `atomic-cutover-plan.md`.** Maps to
|
||||
**T-CUT-02** (`OllamaProvider.list_models`/`load_model`/`unload_model`
|
||||
method bodies), **T-CUT-03** (`tests/test_ollama_provider.py`, new file),
|
||||
and **T-CUT-05/T-CUT-06/T-CUT-08** (the node renames + delegation +
|
||||
registration + test rewrite). Original text kept for history:
|
||||
|
||||
- [-] T015-T [US3] ~~Write FAILING test: `OllamaProvider.list_models()` returns `ModelInfo` entries with status normalized into `ModelStatus`~~ _Superseded, see Phase 8 T-CUT-03._
|
||||
- [-] T015-I [US3] ~~Implement `OllamaProvider.list_models()`~~ _Superseded, see Phase 8 T-CUT-02._
|
||||
- [-] T016-T [US3] ~~Write FAILING test: `OllamaProvider.load_model()`/`unload_model()` are idempotent~~ _Superseded, see Phase 8 T-CUT-03._
|
||||
- [-] T016-I [US3] ~~Implement `OllamaProvider.load_model()`/`unload_model()`~~ _Superseded, see Phase 8 T-CUT-02._
|
||||
- [-] T017-T [US3] ~~Write FAILING test: generic `LLMModelSelector` node returns model+status pairs~~ _Superseded, see Phase 8 T-CUT-08._
|
||||
- [-] T017-I [US3] ~~Rename `OllamaModelSelector` → `LLMModelSelector`, delegate to `LLMProvider.list_models()`~~ _Superseded, see Phase 8 T-CUT-05/T-CUT-06._
|
||||
- [-] T018-T [US3] ~~Write FAILING test: generic `LLMLoadModel`/`LLMUnloadModel` nodes call the protocol~~ _Superseded, see Phase 8 T-CUT-08._
|
||||
- [-] T018-I [US3] ~~Rename `OllamaLoadModel`/`OllamaUnloadModel` → `LLMLoadModel`/`LLMUnloadModel`~~ _Superseded, see Phase 8 T-CUT-05/T-CUT-06._
|
||||
|
||||
**Checkpoint**: superseded — see `atomic-cutover-plan.md`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 4 — Reconnect an existing workflow after upgrading (Priority: P3)
|
||||
|
||||
**Goal**: A clear old→new node mapping exists, and migrated workflows are output-equivalent.
|
||||
|
||||
**Independent Test**: Follow the mapping to reconnect a pre-upgrade workflow; confirm equivalent output.
|
||||
|
||||
- [-] T019 [US4] ~~Add a migration mapping constant~~ _Superseded, see Phase 8 T-CUT-11._
|
||||
- [-] T020 [US4] ~~Full suite green, SC-003 equivalence~~ _Superseded, see Phase 8 T-CUT-10._
|
||||
- [-] T021 [US4] ~~Update quickstart.md migration section~~ _Superseded, see Phase 8 T-CUT-12._
|
||||
|
||||
**Checkpoint**: all four user stories independently functional; migration path documented.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Atomic Node Cutover (supersedes Phases 3, 5, and T014/T019-T021)
|
||||
|
||||
**Goal**: execute `atomic-cutover-plan.md`'s 12-step sequenced plan as one
|
||||
coordinated change — this is the actual current work; Phases 3/5's `[-]`
|
||||
entries above are historical only.
|
||||
|
||||
**Not TDD-paired** the way earlier phases are — per the correction above,
|
||||
this genuinely can't be decomposed into independent red/green pairs (a
|
||||
class rename fails test *collection* for the whole file at once). Each
|
||||
T-CUT step is still verified incrementally during implementation; the
|
||||
suite only needs to be green as a whole at T-CUT-10, not after every step.
|
||||
|
||||
- [x] T-CUT-01 [P] `ollama.py`: import HTTP/cache infra from `comfydv._llm.ollama_provider` instead of duplicating it; repoint `_load_default_models()`/`/dv/ollama/models` route (plan D1)
|
||||
- [x] T-CUT-02 `ollama_provider.py`: implement `OllamaProvider.list_models()`/`load_model()`/`unload_model()`/`chat()`/`chat_structured()` method bodies (ports existing inline logic; `chat_structured()` delegates to `_llm/chat.py`; `list_models()` also queries `/api/ps` to distinguish loaded/unloaded, a genuinely new capability the old `OllamaModelSelector` never had)
|
||||
- [x] T-CUT-03 `tests/test_ollama_provider.py` (new file): tests for T-CUT-02, mocking at the `ollama_provider` seam (plan D4/D5)
|
||||
- [x] T-CUT-04 `ollama.py`: `OllamaClient.RETURN_TYPES` → `("LLM_CLIENT",)`, `create_client()` returns `OllamaProvider(host, headers)` (plan D2)
|
||||
- [x] T-CUT-05 `ollama.py`: rename the 4 classes, `"OLLAMA_CLIENT"`→`"LLM_CLIENT"` on every consumer, rewrite the 3 delegating method bodies, delete `_client_headers` (plan D6)
|
||||
- [x] T-CUT-06 `src/comfydv/__init__.py`: update imports and `NODE_CLASS_MAPPINGS`/`NODE_DISPLAY_NAME_MAPPINGS`
|
||||
- [x] T-CUT-07 `tests/conftest.py`: repoint `_clear_ollama_caches` and `first_generative_model`'s `_fetch_models` import (plan D8) — `first_generative_model`'s import needed no change (still re-exported from `comfydv.ollama`)
|
||||
- [x] T-CUT-08 `tests/test_ollama.py`: rewrote against a `_FakeProvider` double per plan D4/D5 — 98 unit tests, all passing. Also fixed a real gap the rename surfaced: `comfy-manager-entry.json`'s `nodename` list (and its matching test expectation) still had the old display names — updated both.
|
||||
- [x] T-CUT-09 [P] `contracts/llm_provider_protocol.md`: `timeout_secs` fix (commit `ef2464a`)
|
||||
- [x] T-CUT-10 Full suite green (218 passed, only the pre-existing unrelated Dockerfile-python-version test fails), `ruff check --fix && ruff format` clean, `ty check` clean (confirmed the `create_model`/`RandomChoice` diagnostics pre-date this cutover via `git stash` comparison), `beacon doctor --strict` shows only pre-existing/disclosed items (`tdd-commit-discipline` — already documented as an intentional deviation; `epic-gates` — `llamacpp-integration` correctly has no specs yet)
|
||||
- [x] T-CUT-11 [P] `tasks.md`/`ollama.py`: migration mapping constant (FR-009) — `MIGRATION_MAP` dict, `ollama.py`
|
||||
- [x] T-CUT-12 [P] Ollama was reachable in this environment — ran the real `@pytest.mark.integration` suite (not just a manual walkthrough). 6/8 passed, including the critical ones: unreachable-host error handling, real load/unload against the live server, structured-output retry-then-raise against the live server, temperature-determinism. 2 failures (`test_single_turn_returns_non_empty_response`, `test_multi_turn_receives_context`) — confirmed via direct `curl` to `/api/chat` (bypassing this codebase entirely) that the test model (`lukey03/qwen3.5-9b-abliterated-vision`) itself returns a degenerate empty response server-side; this is the exact pre-existing model unreliability ADR-006 already documented, not a cutover regression.
|
||||
|
||||
**Checkpoint**: T-CUT-10 green = all four user stories functional on the generic nodes; T-CUT-11/12 close out US4.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: Polish & Cross-Cutting Concerns
|
||||
|
||||
_Subsumed by T-CUT-10/T-CUT-12 (Phase 8) — same work, done together with the
|
||||
cutover rather than as a separate pass, since ruff/ty/doctor need to run
|
||||
against the final state anyway:_
|
||||
|
||||
- [x] T022 [P] `ruff check --fix && ruff format` — clean (T-CUT-10)
|
||||
- [x] T023 [P] `ty check` — clean, pre-existing diagnostics confirmed unrelated via `git stash` comparison (T-CUT-10)
|
||||
- [x] T024 Constitution Principle IV confirmed: `src/comfydv/_llm/*.py` import no `comfy`/`server`/`folder_paths` at module scope (verified via grep)
|
||||
- [x] T025 `beacon doctor --strict`: only `tdd-commit-discipline` (documented deviation, see Phase 4's US2 notes) and `epic-gates` (`llamacpp-integration` correctly has no specs yet) — both pre-disclosed, not new findings (T-CUT-10)
|
||||
- [x] T026 Live-server validation done via the real `@pytest.mark.integration` suite rather than a separate manual walkthrough — Ollama was reachable in this environment (T-CUT-12); a literal `quickstart.md` click-through in ComfyUI itself is still worth doing whenever this branch is reviewed in a real ComfyUI install, but the underlying behavior is now proven against a live server
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: no dependencies
|
||||
- **Foundational (Phase 2)**: depends on Setup — BLOCKS all user stories
|
||||
- **User Stories (Phase 3–6)**: all depend on Foundational; US1 has no dependency on US2/US3/US4; US2's `ChatCompletion` wiring (T014) depends on US1's node rename (T009-I); US3 is independent of US1/US2 except for sharing `OllamaProvider`'s constructor (T006); US4 depends on the node renames done in US1/US3 (T009-I, T017-I, T018-I) since it documents them
|
||||
- **Polish (Phase 7)**: depends on all four user stories
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- T002 (package init) can run alongside T001 (dependency addition)
|
||||
- T003 and T005 can run in parallel (different files/concerns) within Foundational
|
||||
- T008-T (provider-level test) can run in parallel with T007-T (node-level test) — different files
|
||||
- T011-T, T015-T, T016-T can each start as soon as Foundational is done, in parallel with US1 — different files, no shared dependency beyond T004/T006
|
||||
- T022/T023 (lint/type-check) can run in parallel in Polish
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First
|
||||
|
||||
1. Phase 1 (Setup) → Phase 2 (Foundational) → Phase 3 (US1) → **STOP and validate US1 independently** against a live local server.
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. Setup + Foundational → foundation ready.
|
||||
2. US1 → validate → this alone restores basic chat parity with today's Ollama integration, on the new mechanism.
|
||||
3. US2 → validate → restores structured-output parity (the ADR-006→ADR-007 migration is now complete in behavior).
|
||||
4. US3 → validate → restores model-management parity.
|
||||
5. US4 → validate → migration guidance ships; full regression pass (T020) confirms SC-003.
|
||||
6. Polish.
|
||||
|
||||
Each story adds value without breaking the previous one — this mirrors the epic's own framing: US1+US2 together are the risky "prove the migration is behavior-preserving" core; US3 and US4 round out parity and upgrade experience.
|
||||
|
||||
---
|
||||
|
||||
## ⚠️ Correction (2026-07-11) — US1/US3 independence claim was wrong
|
||||
|
||||
**Discovered mid-implementation, confirmed by independent product + engineering
|
||||
review (agent-trio deliberation, aligned verdicts):** the "US1 has no
|
||||
dependency on other stories" and "US3 is independent of US1/US2" claims above
|
||||
are **false**. `OllamaClient` is a single shared producer node — every
|
||||
downstream node (`OllamaModelSelector`, `OllamaLoadModel`,
|
||||
`OllamaUnloadModel`, `OllamaChatCompletion`) consumes its output via
|
||||
`f"{client}/api/..."` string interpolation (10 call sites in `ollama.py`).
|
||||
Changing `OllamaClient` to emit an `OllamaProvider` object instead of the
|
||||
current string-like `OllamaClientType` breaks **all four** consumers
|
||||
simultaneously — there is no way to migrate just `ChatCompletion` (US1)
|
||||
while leaving `OllamaModelSelector`/`OllamaLoadModel`/`OllamaUnloadModel`
|
||||
(US3) on the old string-based access pattern. Separately, renaming these
|
||||
classes breaks `tests/test_ollama.py`'s imports atomically (125 references
|
||||
across the file) — a class rename fails test *collection* for the whole
|
||||
file at once, not test-by-test.
|
||||
|
||||
**Rejected fix:** making `OllamaProvider` also subclass `str` (mirroring
|
||||
`OllamaClientType`'s trick) to preserve incremental per-node migration.
|
||||
Both reviewers rejected this — it reintroduces the exact hack ADR-007
|
||||
exists to eliminate into the new clean boundary, and would silently mask an
|
||||
incomplete cutover (un-migrated consumers keep working via the string
|
||||
trick, so T020's regression pass would go green for the wrong reason).
|
||||
|
||||
**Decision:** T007–T010 (US1) and T015–T018 (US3)'s *node-layer* work
|
||||
(everything that touches `OllamaClient`'s output type or renames a node
|
||||
class) must land as **one atomic cutover** — one coordinated change across
|
||||
`ollama.py` and `tests/test_ollama.py`, verified green as a whole, not as
|
||||
separable per-story TDD pairs. This is sized beyond a single 2–4h tracer
|
||||
bullet and is explicitly re-scoped as its own dedicated BUILD session
|
||||
(tracked in GitHub issue — see epic Notes for the link once filed), not
|
||||
attempted in the same session as the Foundational layer (T001–T006, already
|
||||
shipped safely — see git log). The *provider-layer* work that doesn't touch
|
||||
`ollama.py` (e.g. `OllamaProvider.chat()`/`list_models()`/`load_model()`/
|
||||
`unload_model()` method bodies, and the `pydantic-ai`-backed
|
||||
`chat_structured()` helper) remains genuinely independent and safe to build
|
||||
ahead of the cutover — only the ComfyUI node-layer rename is atomic.
|
||||
|
||||
ADR-007's own decision (breaking rename, no deprecated aliases) is
|
||||
**unaffected** — that call was about user-facing blast radius (small,
|
||||
Ollama integration shipped 2026-07-04), which this finding doesn't change.
|
||||
What's re-scoped is delivery sequencing, not the design decision.
|
||||
|
||||
---
|
||||
|
||||
## ✅ Properly specced (2026-07-11) — see `atomic-cutover-plan.md`
|
||||
|
||||
Full line-by-line inventory of every affected reference in `ollama.py` and
|
||||
`tests/test_ollama.py` (1820 lines, read in full), the design decisions it
|
||||
surfaced (cache-singleton duplication, `client == "<string>"` equality
|
||||
breaking, bare-string-client backward compat removal, and — the big one —
|
||||
a test-layer split so the 35 relocated `_post_json` monkeypatches land at
|
||||
the right architectural seam instead of being patched 1:1), and a
|
||||
12-step sequenced task list (T-CUT-01 … T-CUT-12) that supersedes the
|
||||
struck-through tasks above. That file is now the authoritative task list
|
||||
for this remaining work; this file's Phase 3/5/6 entries are kept only for
|
||||
history.
|
||||
@@ -0,0 +1 @@
|
||||
epic = "llamacpp-integration"
|
||||
@@ -0,0 +1,39 @@
|
||||
# Specification Quality Checklist: llama.cpp Model Integration
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: 2026-07-11
|
||||
**Feature**: [spec.md](../spec.md)
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [x] No implementation details (languages, frameworks, APIs)
|
||||
- [x] Focused on user value and business needs
|
||||
- [x] Written for non-technical stakeholders
|
||||
- [x] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||
- [x] Requirements are testable and unambiguous
|
||||
- [x] Success criteria are measurable
|
||||
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||
- [x] All acceptance scenarios are defined
|
||||
- [x] Edge cases are identified
|
||||
- [x] Scope is clearly bounded
|
||||
- [x] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [x] All functional requirements have clear acceptance criteria
|
||||
- [x] User scenarios cover primary flows
|
||||
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [x] No implementation details leak into specification
|
||||
|
||||
## Notes
|
||||
|
||||
No [NEEDS CLARIFICATION] markers needed — scope boundaries (router-mode-only,
|
||||
no GPU tuning, no auth/TLS, no Manager listing) came directly from the
|
||||
parent epic's Non-goals (`project-management/Roadmap/epics/llamacpp-integration.md`)
|
||||
and ADR-007. User Story 4 (swap backends without touching downstream nodes)
|
||||
is the adapter pattern's central promise made concrete and testable, not
|
||||
padding.
|
||||
@@ -0,0 +1,44 @@
|
||||
# Contract: `LlamaCppProvider` conforms to `LLMProvider`
|
||||
|
||||
This is the concrete proof of ADR-007's adapter pattern — the same protocol
|
||||
contract documented in
|
||||
`specs/007-llm-provider-abstraction/contracts/llm_provider_protocol.md`,
|
||||
now with a second implementation. Nothing in that contract changes; this
|
||||
file only documents `LlamaCppProvider`'s specific wire-format bindings.
|
||||
|
||||
```python
|
||||
class LlamaCppProvider:
|
||||
def __init__(self, host: str, headers: dict | None = None): ...
|
||||
|
||||
async def list_models(self) -> list[ModelInfo]:
|
||||
"""GET {host}/models → data[] → ModelInfo(name=m["id"], status=ModelStatus(m["status"]["value"]), size=None)"""
|
||||
|
||||
async def load_model(self, model: str) -> None:
|
||||
"""POST {host}/models/load {"model": model}"""
|
||||
|
||||
async def unload_model(self, model: str) -> None:
|
||||
"""POST {host}/models/unload {"model": model}"""
|
||||
|
||||
async def chat(self, model, messages, options=None, timeout_secs=300.0) -> str:
|
||||
"""POST {host}/v1/chat/completions → choices[0].message.content"""
|
||||
|
||||
async def chat_structured(self, model, messages, schema, options=None, timeout_secs=300.0, max_retries=2) -> BaseModel:
|
||||
"""Delegates to comfydv._llm.chat.chat_structured(base_url=f"{host}/v1", ...) — identical call OllamaProvider makes"""
|
||||
```
|
||||
|
||||
## Behavioral requirements (inherited from the protocol contract, restated for this implementation)
|
||||
|
||||
- `load_model`/`unload_model` MUST be idempotent. **Live-verified against a
|
||||
real router-mode server**: router mode's own endpoints are *not*
|
||||
idempotent — `/models/load` on an already-loaded model returns HTTP 400
|
||||
`"model is already running"`, and `/models/unload` on an already-unloaded
|
||||
model returns HTTP 400 `"model is not running"`, instead of `{"success": true}`.
|
||||
`LlamaCppProvider` absorbs this itself: these two specific error messages
|
||||
are treated as the desired end-state already reached, not a failure; any
|
||||
other error still propagates.
|
||||
- `list_models()` MUST NOT normalize away llama.cpp's `sleeping`/`downloading`
|
||||
states (unlike `OllamaProvider`, which has no choice but to normalize —
|
||||
see `research.md`).
|
||||
- A `llama-server` not running in router mode (missing endpoints) MUST
|
||||
surface a clear, specific error (spec.md FR-006) — not a generic
|
||||
connection failure indistinguishable from "server not running at all."
|
||||
@@ -0,0 +1,41 @@
|
||||
# Data Model: llama.cpp Model Integration
|
||||
|
||||
No new types — this feature is a second implementation of the existing
|
||||
`LLMProvider` protocol, `ModelStatus`, `ModelInfo`, and `Message` types
|
||||
(`src/comfydv/_llm/provider.py`, unchanged). This file documents
|
||||
`LlamaCppProvider`'s field mapping from llama-server's router-mode JSON onto
|
||||
those existing types (see `research.md` for the verified API shapes).
|
||||
|
||||
## `LlamaCppProvider.list_models()` → `ModelInfo` mapping
|
||||
|
||||
| `ModelInfo` field | Source (`GET /models` response, per model in `data[]`) |
|
||||
|---|---|
|
||||
| `name` | `id` — **not** `name` (llama.cpp's field name differs from Ollama's) |
|
||||
| `status` | `status.value` — nested object, not a flat string |
|
||||
| `size` | Not provided by this endpoint; `None` |
|
||||
|
||||
`status.value` maps directly onto `ModelStatus`'s five values
|
||||
(`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`) — llama.cpp's
|
||||
vocabulary is exactly `ModelStatus`'s full set, so unlike `OllamaProvider`
|
||||
(which normalizes into a narrower subset), `LlamaCppProvider` needs no
|
||||
approximation. A `"failed": true` state exists outside this vocabulary
|
||||
(model process crashed) — out of scope per spec.md's edge cases; treated as
|
||||
whatever `status.value` reports rather than added as a sixth enum value.
|
||||
|
||||
## `LlamaCppProvider.load_model()` / `unload_model()`
|
||||
|
||||
Both `POST /models/load` and `POST /models/unload` take `{"model": <id>}` —
|
||||
the same `id` string `list_models()` returns as `ModelInfo.name`. No mapping
|
||||
ambiguity here (unlike Ollama, where load/unload uses `/api/generate`'s
|
||||
`keep_alive` side effect rather than a dedicated endpoint).
|
||||
|
||||
## `LlamaCppProvider.chat()` / `chat_structured()`
|
||||
|
||||
Both reach `llama-server`'s OpenAI-compatible `/v1/chat/completions` —
|
||||
`chat_structured()` calls the existing shared `comfydv._llm.chat.chat_structured()`
|
||||
helper unchanged (`base_url=f"{self.host}/v1"`, matching `OllamaProvider`'s
|
||||
own call exactly). `chat()` parses the response as
|
||||
`choices[0].message.content` (OpenAI shape), not Ollama's native
|
||||
`message.content` — the two providers' non-structured paths differ here
|
||||
because llama-server doesn't have an Ollama-style native `/api/chat`
|
||||
endpoint to prefer instead.
|
||||
@@ -0,0 +1,11 @@
|
||||
Feature: US1 — Connect to a local llama.cpp server and get chat responses
|
||||
|
||||
Scenario: llama.cpp connection node feeds the existing chat node
|
||||
Given a running local llama-server (router mode) and a workflow with a llama.cpp connection node wired into the existing chat node
|
||||
When the workflow executes
|
||||
Then the chat node returns the model's text response
|
||||
|
||||
Scenario: Unreachable llama.cpp server surfaces a clear error
|
||||
Given the llama.cpp connection node configured with an unreachable server address
|
||||
When the workflow executes
|
||||
Then the chat node reports a clear connection error
|
||||
@@ -0,0 +1,11 @@
|
||||
Feature: US2 — Get structured, validated output from llama.cpp
|
||||
|
||||
Scenario: Valid structured response exposes typed fields, same as Ollama
|
||||
Given a chat node connected to llama.cpp with structured output enabled and a valid schema
|
||||
When the workflow executes and the model responds correctly
|
||||
Then each schema field is available as its own typed output, and no required field is blank
|
||||
|
||||
Scenario: Invalid response retries then fails clearly, same as Ollama
|
||||
Given a llama.cpp-hosted model that returns invalid or incomplete structured output
|
||||
When the workflow executes
|
||||
Then the node retries automatically and, if still unsuccessful, fails with a clear error
|
||||
@@ -0,0 +1,16 @@
|
||||
Feature: US3 — See and control which models are loaded on llama.cpp
|
||||
|
||||
Scenario: List models with full status vocabulary
|
||||
Given a running local llama-server with at least one available model
|
||||
When a workflow author uses the model-listing node
|
||||
Then they see each available model along with its current status, drawn from llama.cpp's full status vocabulary
|
||||
|
||||
Scenario: Load a model into memory
|
||||
Given a model that is not currently loaded
|
||||
When a workflow author runs the load-model node against it
|
||||
Then the model becomes loaded and is then usable by the chat node
|
||||
|
||||
Scenario: Unload a model from memory
|
||||
Given a model that is loaded and idle
|
||||
When a workflow author runs the unload-model node against it
|
||||
Then the model is freed from memory and its reported status updates accordingly
|
||||
@@ -0,0 +1,6 @@
|
||||
Feature: US4 — Swap from Ollama to llama.cpp without touching the rest of the workflow
|
||||
|
||||
Scenario: Replacing only the connection node preserves the workflow
|
||||
Given a workflow with chat/model-management nodes wired to an Ollama connection node
|
||||
When a workflow author replaces only the connection node with a llama.cpp one, pointed at a running llama-server
|
||||
Then the workflow runs successfully with no changes to any other node
|
||||
@@ -0,0 +1,117 @@
|
||||
# Implementation Plan: llama.cpp Model Integration
|
||||
|
||||
**Branch**: `008-llamacpp-integration` | **Date**: 2026-07-11 | **Spec**: [spec.md](./spec.md)
|
||||
|
||||
**Input**: Feature specification from `/specs/008-llamacpp-integration/spec.md`
|
||||
|
||||
**Note**: This template is filled in by the `/speckit-plan` command. See `.specify/templates/plan-template.md` for the execution workflow.
|
||||
|
||||
## Summary
|
||||
|
||||
Implement `LlamaCppProvider` as the second `LLMProvider` (ADR-007), backed by
|
||||
`llama-server`'s router mode (`GET /models`, `POST /models/load`,
|
||||
`POST /models/unload`, `/v1/chat/completions`). Add one new ComfyUI node
|
||||
(`LlamaCppClient`) emitting the existing `LLM_CLIENT` socket type — no other
|
||||
node classes change. This is the concrete proof the provider abstraction
|
||||
(prerequisite epic, PR #17) actually generalizes: a second backend, zero
|
||||
changes to `ChatCompletion`/`LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel`.
|
||||
|
||||
## Technical Context
|
||||
|
||||
**Language/Version**: Python ≥3.11 (unchanged, per `pyproject.toml`)
|
||||
|
||||
**Primary Dependencies**: `aiohttp` (existing — model-management REST calls),
|
||||
`pydantic-ai`/`openai` (existing, from the prerequisite epic — `chat_structured()`
|
||||
reuses the shared helper unchanged, zero new structured-output code)
|
||||
|
||||
**Storage**: N/A — no persistent storage; reuses the existing
|
||||
`_MODEL_LIST_CACHE`/`_CHAT_RESPONSE_CACHE` infra pattern from `OllamaProvider`
|
||||
|
||||
**Testing**: `pytest` via `uv run pytest`, following `tests/test_ollama_provider.py`'s
|
||||
established convention (mock at the provider's own `_post_json`/`_get_json`
|
||||
seam, no live server required for unit tests)
|
||||
|
||||
**Target Platform**: ComfyUI custom-node runtime, same as the existing Ollama
|
||||
integration
|
||||
|
||||
**Project Type**: Library / ComfyUI custom-node pack (single project, adds to
|
||||
existing `src/comfydv/` layout)
|
||||
|
||||
**Performance Goals**: No new numeric target; must not add latency beyond
|
||||
what `OllamaProvider`'s equivalent methods already accept
|
||||
|
||||
**Constraints**: Router-mode-only (spec.md Assumptions — a `llama-server`
|
||||
without `--models-dir`/`--models-preset` doesn't expose these endpoints at
|
||||
all, FR-006); model identifier field is `id` (llama.cpp) vs `name` (Ollama) —
|
||||
`LlamaCppProvider.list_models()` must map this correctly (see `research.md`);
|
||||
`status` is a nested object (`{"value": "..."}`), not a flat string
|
||||
|
||||
**Scale/Scope**: One new class (`LlamaCppProvider`, mirrors `OllamaProvider`'s
|
||||
shape), one new ComfyUI node (`LlamaCppClient`), one new test file — no
|
||||
changes to `ollama.py`, `_llm/provider.py`, `_llm/chat.py`, or any existing
|
||||
node class
|
||||
|
||||
## Constitution Check
|
||||
|
||||
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
|
||||
|
||||
| Principle | Verdict | Notes |
|
||||
|---|---|---|
|
||||
| I. ComfyUI Contract First | PASS | `LlamaCppClient` exposes the standard `INPUT_TYPES`/`RETURN_TYPES`/`FUNCTION`/`CATEGORY`; registered in `NODE_CLASS_MAPPINGS` like every other node. |
|
||||
| II. Sandbox All User-Supplied Code | N/A | No template/expression evaluation in this feature. |
|
||||
| III. Test-First | PASS (binding) | `tests/test_llamacpp_provider.py` written test-first, mirroring `test_ollama_provider.py`'s TDD-pair structure. |
|
||||
| IV. Graceful Degradation Outside ComfyUI | PASS (binding) | `LlamaCppProvider` lives in `src/comfydv/_llm/`, which already has no `comfy`/`server` imports at module scope (verified for the prerequisite epic; this feature adds no new module-scope imports of either). |
|
||||
| V. Simplicity — Function Before Class | PASS, same justification as `OllamaProvider` | `LlamaCppProvider` carries connection state (host, headers) across 5 methods — the same shared-state condition that already justified `OllamaProvider` as a class (research.md, prerequisite epic). No new gate — same precedent applies. |
|
||||
| VI. Fixed Output Positions | N/A | `LlamaCppClient`'s single output (`client`) isn't a multi-output node; no positional contract to preserve. |
|
||||
|
||||
Re-checked post-Phase 1 design (data-model.md): unchanged — no new gate
|
||||
violations. No Complexity Tracking entries needed (unlike the prerequisite
|
||||
epic, this feature introduces no new pattern, just a second instance of an
|
||||
already-justified one).
|
||||
|
||||
## Project Structure
|
||||
|
||||
### Documentation (this feature)
|
||||
|
||||
```text
|
||||
specs/008-llamacpp-integration/
|
||||
├── plan.md # This file
|
||||
├── research.md # Phase 0 — router-mode API shape, verified live
|
||||
├── data-model.md # Phase 1 — LlamaCppProvider field mapping
|
||||
├── quickstart.md # Phase 1 — minimal workflow walkthrough
|
||||
├── contracts/ # Phase 1 — LlamaCppProvider's protocol conformance
|
||||
└── tasks.md # Phase 2 (/speckit-tasks)
|
||||
```
|
||||
|
||||
### Source Code (repository root)
|
||||
|
||||
```text
|
||||
src/comfydv/
|
||||
├── ollama.py # unchanged — add LlamaCppClient node only via a new module
|
||||
├── llamacpp.py # new — LlamaCppClient node (mirrors OllamaClient's shape)
|
||||
├── _llm/
|
||||
│ ├── provider.py # unchanged — LLMProvider/ModelStatus/ModelInfo/Message
|
||||
│ ├── ollama_provider.py # unchanged
|
||||
│ ├── llamacpp_provider.py # new — LlamaCppProvider (mirrors ollama_provider.py's shape)
|
||||
│ └── chat.py # unchanged — chat_structured() reused as-is
|
||||
└── __init__.py # add LlamaCppClient import + NODE_CLASS_MAPPINGS entry
|
||||
|
||||
tests/
|
||||
├── test_ollama_provider.py # unchanged
|
||||
├── test_llamacpp_provider.py # new — mirrors test_ollama_provider.py's structure
|
||||
└── test_llamacpp.py # new — LlamaCppClient node contract test (small; mirrors
|
||||
# the OllamaClient-specific slice of test_ollama.py)
|
||||
```
|
||||
|
||||
**Structure Decision**: New `src/comfydv/llamacpp.py` module (not added into
|
||||
`ollama.py`) for the `LlamaCppClient` node, and a new `src/comfydv/_llm/llamacpp_provider.py`
|
||||
for `LlamaCppProvider` — mirroring the existing `ollama.py`/`ollama_provider.py`
|
||||
split exactly, so the two backends read as parallel, symmetric implementations
|
||||
rather than one growing to accommodate the other. No existing file is
|
||||
modified except `__init__.py`'s registration block.
|
||||
|
||||
## Complexity Tracking
|
||||
|
||||
> **Fill ONLY if Constitution Check has violations that must be justified**
|
||||
|
||||
None — see Constitution Check above.
|
||||
@@ -0,0 +1,26 @@
|
||||
# Quickstart: llama.cpp Model Integration
|
||||
|
||||
## Prerequisite
|
||||
|
||||
Launch `llama-server` in router mode:
|
||||
|
||||
```bash
|
||||
llama-server --models-dir ./models -c 8192
|
||||
```
|
||||
|
||||
## Minimal workflow
|
||||
|
||||
1. Add an **LlamaCpp Client** node. Set its host widget (default
|
||||
`http://localhost:8080`, llama-server's default port).
|
||||
2. Wire it into a **Chat Completion** node — the exact same node used for
|
||||
Ollama. Set a model and prompt, run.
|
||||
3. Structured output, model listing, and load/unload all work exactly as
|
||||
documented for Ollama in the main README/quickstart — swap the client
|
||||
node, nothing else changes.
|
||||
|
||||
## Swapping an existing Ollama workflow to llama.cpp
|
||||
|
||||
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed
|
||||
at your running `llama-server`. Every downstream node (Chat Completion, LLM
|
||||
Model Selector, LLM Load Model, LLM Unload Model) keeps working unmodified —
|
||||
this is the whole point of the provider abstraction (ADR-007).
|
||||
@@ -0,0 +1,78 @@
|
||||
# Research: llama.cpp Model Integration
|
||||
|
||||
## Decision: exact router-mode API shape (verified against `ggml-org/llama.cpp`'s live `tools/server/README.md`, not assumed)
|
||||
|
||||
llama.cpp's router mode postdates this session's training data — verified live
|
||||
against the authoritative source rather than guessed, since getting field
|
||||
names wrong here would silently produce broken code (wrong key = `KeyError`
|
||||
or silent `None`, not an obvious failure).
|
||||
|
||||
**`GET /models`** response:
|
||||
|
||||
```json
|
||||
{
|
||||
"data": [
|
||||
{
|
||||
"id": "ggml-org/gemma-3-4b-it-GGUF:Q4_K_M",
|
||||
"path": "/Users/.../gemma-3-4b-it-Q4_K_M.gguf",
|
||||
"status": {
|
||||
"value": "loaded",
|
||||
"args": ["llama-server", "-ctx", "4096"]
|
||||
},
|
||||
"architecture": {
|
||||
"input_modalities": ["text", "image"],
|
||||
"output_modalities": ["text"]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Two details that would have been wrong by assumption:**
|
||||
|
||||
1. The model identifier field is **`id`**, not `name` — different from Ollama's
|
||||
`/api/tags`, which uses `name`. `OllamaProvider.list_models()` maps
|
||||
`m["name"]`; `LlamaCppProvider.list_models()` must map `m["id"]` instead.
|
||||
2. **`status` is a nested object** (`{"value": "loaded", ...}`), not a flat
|
||||
string field. `LlamaCppProvider.list_models()` must read
|
||||
`m["status"]["value"]`, not `m["status"]` directly. A `"failed"` state
|
||||
also exists (`{"failed": true, "exit_code": ...}`) outside the five
|
||||
`ModelStatus` values the protocol defines — not handled by this feature
|
||||
(see Non-goals/edge cases in `spec.md`); a failed model is reported as
|
||||
whatever `status.value` degrades to rather than added as a sixth enum
|
||||
value, keeping `ModelStatus` unchanged across both providers.
|
||||
|
||||
**`POST /models/load`** and **`POST /models/unload`**: identical request
|
||||
shape, `{"model": "<id>"}` (using the same `id` string from `GET /models`,
|
||||
despite the request field being named `model` not `id`). Response:
|
||||
`{"success": true}`.
|
||||
|
||||
**CLI**: `--models-dir <path>` or `--models-preset <path>.ini` — a deployment
|
||||
prerequisite (spec.md Assumptions), not something comfydv configures.
|
||||
|
||||
## Decision: `chat_structured()` needs zero new code
|
||||
|
||||
`llama-server`'s `/v1/chat/completions` is OpenAI-compatible (the same
|
||||
assumption ADR-007 made when adopting `pydantic-ai`). `LlamaCppProvider.chat_structured()`
|
||||
calls the exact same `comfydv._llm.chat.chat_structured()` helper
|
||||
`OllamaProvider` already uses, with `base_url=f"{self.host}/v1"` — the only
|
||||
per-provider difference. This is the concrete proof the shared mechanism
|
||||
generalizes (spec.md User Story 2/FR-004), not just an assumption.
|
||||
|
||||
## Decision: `chat()` (non-structured) also reuses the OpenAI-compatible endpoint
|
||||
|
||||
Unlike Ollama (which has both a native `/api/chat` and an OpenAI-compat
|
||||
`/v1/chat/completions`), llama-server's primary chat endpoint is the
|
||||
OpenAI-compatible one. `LlamaCppProvider.chat()` POSTs to
|
||||
`{host}/v1/chat/completions` (via the existing `_post_json` helper, no new
|
||||
HTTP client) rather than mirroring Ollama's native-endpoint choice — the
|
||||
response shape (`choices[0].message.content`) differs from Ollama's native
|
||||
`message.content` and must be parsed accordingly.
|
||||
|
||||
## Decision: no protocol changes needed
|
||||
|
||||
`LLMProvider`'s five methods (`list_models`/`load_model`/`unload_model`/
|
||||
`chat`/`chat_structured`) already cover everything router mode needs — this
|
||||
was the actual point of designing the protocol at the operation level in
|
||||
ADR-007, and this research confirms it held up against llama.cpp's real API,
|
||||
not just Ollama's.
|
||||
@@ -0,0 +1,110 @@
|
||||
# Feature Specification: llama.cpp Model Integration
|
||||
|
||||
**Feature Branch**: `008-llamacpp-integration`
|
||||
|
||||
**Created**: 2026-07-11
|
||||
|
||||
**Status**: Draft
|
||||
|
||||
**Input**: User description: "Add ComfyUI nodes for llama.cpp local inference via llama-server's router mode, implementing the LlamaCppProvider as the second LLMProvider (ADR-007) alongside the existing OllamaProvider. Router mode exposes GET /models (with live status), POST /models/load, POST /models/unload, giving llama.cpp the same manual load/unload memory-management primitives as Ollama. No new ComfyUI node classes needed for chat/model-selection/load/unload — only a new LlamaCppClient config node; the existing generic ChatCompletion/LLMModelSelector/LLMLoadModel/LLMUnloadModel nodes work unchanged once wired to it."
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Connect to a local llama.cpp server and get chat responses (Priority: P1) 🎯 MVP
|
||||
|
||||
As a ComfyUI workflow author running `llama-server` locally, I want a connection node for it — just like the one I already use for Ollama — so I can get chat responses from a llama.cpp-hosted model using the same chat node I already know.
|
||||
|
||||
**Why this priority**: This is the entire point of the feature and the proof that the provider abstraction (shipped in the prerequisite epic) actually works: a second backend, zero changes to the chat node.
|
||||
|
||||
**Independent Test**: Wire a new llama.cpp connection node into the existing chat node, run against a local `llama-server` (router mode), confirm a text response.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a running local `llama-server` (router mode) and a workflow with a llama.cpp connection node wired into the existing chat node, **When** the workflow executes, **Then** the chat node returns the model's text response — using the exact same chat node a workflow author already uses for Ollama.
|
||||
2. **Given** the llama.cpp connection node configured with an unreachable server address, **When** the workflow executes, **Then** the chat node reports a clear connection error, matching the behavior workflow authors already know from the Ollama connection.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Get structured, validated output from llama.cpp (Priority: P1)
|
||||
|
||||
As a workflow author, I want structured output (a schema-validated response instead of free text) to work identically regardless of whether I'm connected to Ollama or llama.cpp, so I don't have to relearn or rebuild anything when switching backends.
|
||||
|
||||
**Why this priority**: Structured output is a core existing capability (already proven for Ollama); this story proves the shared mechanism genuinely generalizes rather than being Ollama-specific in practice, not just in name.
|
||||
|
||||
**Independent Test**: Enable structured output on the chat node with a schema, run against a llama.cpp-hosted model, confirm each schema field is populated and never blank — using the same steps as the equivalent Ollama test.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a chat node connected to llama.cpp with structured output enabled and a valid schema, **When** the workflow executes and the model responds correctly, **Then** each schema field is available as its own typed output, and no required field is blank.
|
||||
2. **Given** a llama.cpp-hosted model that returns invalid or incomplete structured output, **When** the workflow executes, **Then** the node retries automatically and, if still unsuccessful, fails with a clear error — identical behavior to the Ollama path.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - See and control which models are loaded on llama.cpp (Priority: P2)
|
||||
|
||||
As a workflow author running models locally, I want to see live model status (including whether a model is currently loading or being downloaded, not just loaded/unloaded) and explicitly load or unload a model on my llama.cpp server, so I can manage memory the same way I already do for Ollama — with more visibility, since llama.cpp's router mode reports richer status than Ollama does.
|
||||
|
||||
**Why this priority**: Valuable and proves the model-management path generalizes too, but a workflow can still run chat completions without ever calling load/unload explicitly (the server can load on first use), so it's lower risk to defer than basic chat.
|
||||
|
||||
**Independent Test**: Use the existing model-listing node against a running `llama-server`, confirm it shows each available model with its current status (including `loading`/`downloading` if applicable); use the existing load/unload nodes against one model and confirm its status changes.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a running local `llama-server` with at least one available model, **When** a workflow author uses the model-listing node, **Then** they see each available model along with its current status, drawn from llama.cpp's full status vocabulary (not just loaded/unloaded).
|
||||
2. **Given** a model that is not currently loaded, **When** a workflow author runs the load-model node against it, **Then** the model becomes loaded and is then usable by the chat node.
|
||||
3. **Given** a model that is loaded and idle, **When** a workflow author runs the unload-model node against it, **Then** the model is freed from memory and its reported status updates accordingly.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Swap from Ollama to llama.cpp without touching the rest of the workflow (Priority: P3)
|
||||
|
||||
As a workflow author with an existing Ollama-based workflow, I want to switch it to llama.cpp by changing only the connection node, so I don't have to rebuild my chat/model-management logic for a second backend.
|
||||
|
||||
**Why this priority**: This is the adapter pattern's actual promise made concrete for a user, but it's a validation/demonstration story rather than new capability — everything it depends on is already covered by User Stories 1–3.
|
||||
|
||||
**Independent Test**: Take a workflow using the Ollama connection node, replace it with the llama.cpp connection node (same downstream nodes, no other changes), run it, confirm it still works.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a workflow with chat/model-management nodes wired to an Ollama connection node, **When** a workflow author replaces only the connection node with a llama.cpp one (pointed at a running `llama-server`), **Then** the workflow runs successfully with no changes to any other node.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when `llama-server` is running but was launched without router mode (i.e. with `-m` instead of `--models-dir`)? The router-mode-only endpoints this feature depends on won't exist — the connection/model-management nodes should fail with a clear error, not hang or silently return empty results.
|
||||
- What happens when the configured server address is unreachable at the moment a model-listing, load, or unload node runs (not just the chat node)?
|
||||
- What happens when llama.cpp reports a model status this feature doesn't expect (a router-mode API change)? Should degrade gracefully (surface the status if recognized, don't crash on an unrecognized one), not silently misreport.
|
||||
- What happens to an in-flight chat request if the model it depends on is unloaded by another node in the same workflow run? (Same question already answered for Ollama — behavior should be consistent.)
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: The system MUST allow a workflow author to configure a connection to a local `llama-server` (router mode) the same way they already configure a connection to Ollama — a dedicated connection node, reusable across multiple nodes in a workflow.
|
||||
- **FR-002**: The system MUST NOT require any new or different node classes for chat, structured output, model listing, or load/unload when using llama.cpp — the existing generic nodes MUST work unchanged once connected to a llama.cpp connection node.
|
||||
- **FR-003**: The system MUST report each model's status using llama.cpp's full status vocabulary (unloaded, loading, loaded, sleeping, downloading) when connected to llama.cpp — not degraded to the narrower Ollama-compatible set.
|
||||
- **FR-004**: The system's chat and structured-output behavior MUST be identical between Ollama and llama.cpp connections, given equivalent inputs — same retry limits, same validation rules, same error conditions (this is the direct continuation of the prerequisite epic's own FR-007/FR-008).
|
||||
- **FR-005**: The system MUST allow a workflow author to explicitly load a model into memory and explicitly unload a model from memory on a connected llama.cpp server.
|
||||
- **FR-006**: The system MUST surface a clear, specific error when connected to a `llama-server` instance that isn't running in router mode (the endpoints this feature needs don't exist), rather than an unhelpful generic failure.
|
||||
|
||||
### Key Entities *(include if feature involves data)*
|
||||
|
||||
- **llama.cpp connection**: A configured connection to a local `llama-server` instance running in router mode (host + any authentication), implementing the same connection concept already established for Ollama.
|
||||
- **Model status**: Reuses the existing status concept from the prerequisite feature, now populated with llama.cpp's full vocabulary rather than a narrowed subset.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: A workflow author can connect to a llama.cpp server and get a chat response using the same node count and shape as connecting to Ollama (one connection node, one chat node) — no new nodes to learn for the chat path.
|
||||
- **SC-002**: An existing workflow can be repointed from Ollama to llama.cpp by changing exactly one node (the connection node) — zero edits to any chat or model-management node.
|
||||
- **SC-003**: Structured-output workflows behave identically (same validation guarantees, zero blank-required-field results) regardless of which backend is connected.
|
||||
- **SC-004**: Model status reporting for llama.cpp surfaces all five status values where applicable — a strictly richer view than what Ollama can report through the same interface.
|
||||
|
||||
## Assumptions
|
||||
|
||||
- Workflow authors run their own local `llama-server` instance, launched in router mode (`--models-dir` or `--models-preset`), reachable over HTTP from the machine running ComfyUI; this feature does not install, configure, or launch that server.
|
||||
- Non-router-mode `llama-server` usage (a single model launched with `-m`) is out of scope — router mode is required for the load/unload/status parity with Ollama that is this feature's whole point.
|
||||
- GPU inference optimisation, quantisation tuning, authentication/TLS, and ComfyUI Manager registry listing are out of scope, consistent with the prerequisite Ollama epic's own non-goals.
|
||||
- The `LLMProvider` protocol and generic nodes (`ChatCompletion`, `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`) already exist and are not modified by this feature — if llama.cpp's router mode needs a protocol capability that doesn't exist yet, that is a protocol change scoped as its own follow-up, not silently special-cased here.
|
||||
@@ -0,0 +1,224 @@
|
||||
# Tasks: llama.cpp Model Integration
|
||||
|
||||
**Input**: Design documents from `/specs/008-llamacpp-integration/`
|
||||
|
||||
**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/llamacpp_provider_conformance.md
|
||||
|
||||
**Tests**: First-class — every implementation task has a paired failing-test task (`-T`/`-I` suffix).
|
||||
|
||||
**Organization**: Grouped by user story (spec.md priorities P1/P1/P2/P3).
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: US1–US4
|
||||
- **-T / -I**: paired test (red) / implementation (green)
|
||||
|
||||
## Path Conventions
|
||||
|
||||
Single project: `src/comfydv/`, `tests/` at repository root, mirroring the
|
||||
`ollama.py`/`_llm/ollama_provider.py` split exactly (plan.md Structure
|
||||
Decision).
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup
|
||||
|
||||
- [x] T001 No new dependencies — `aiohttp`/`pydantic-ai` already present from the prerequisite epic (verified in `pyproject.toml`)
|
||||
- [x] T002 [P] Create `src/comfydv/_llm/llamacpp_provider.py` and `src/comfydv/llamacpp.py` (empty modules with docstrings, mirroring `ollama_provider.py`/`ollama.py`'s module docstring style)
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational
|
||||
|
||||
None — `LLMProvider`, `ModelStatus`, `ModelInfo`, `Message`, and the shared
|
||||
`chat_structured()` helper already exist from the prerequisite epic and are
|
||||
unmodified by this feature (plan.md Constitution Check, research.md).
|
||||
|
||||
**Checkpoint**: nothing blocks user story work — it can start immediately.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 — Connect to a local llama.cpp server and get chat responses (Priority: P1) 🎯 MVP
|
||||
|
||||
**Goal**: A workflow author wires an `LlamaCppClient` node into the existing `ChatCompletion` node and gets a text response.
|
||||
|
||||
**Independent Test**: Wire `LlamaCppClient` → `ChatCompletion`, run against a live `llama-server` (router mode), confirm text output.
|
||||
|
||||
- [x] T003-T [US1] Write FAILING test: `LlamaCppProvider.chat()` POSTs to `{host}/v1/chat/completions` and parses `choices[0].message.content`, in `tests/test_llamacpp_provider.py` (witnesses `features/us1_connect_and_chat.feature` scenario "llama.cpp connection node feeds the existing chat node")
|
||||
- [x] T003-I [US1] Implement `LlamaCppProvider.__init__`/`.chat()` in `src/comfydv/_llm/llamacpp_provider.py` (data-model.md — OpenAI-shape response parsing, not Ollama's native shape) — makes T003-T pass
|
||||
- [x] T004-T [US1] Write FAILING test: `LlamaCppClient` node's `INPUT_TYPES`/`RETURN_TYPES` match `OllamaClient`'s shape (`LLM_CLIENT` output), and `create_client()` constructs a `LlamaCppProvider`, in `tests/test_llamacpp.py`
|
||||
- [x] T004-I [US1] Implement `LlamaCppClient` node in `src/comfydv/llamacpp.py` (mirrors `OllamaClient` exactly, default host `http://localhost:8080` per llama-server's default port) — makes T004-T pass (depends on T003-I)
|
||||
- [x] T005-T [US1] Write FAILING test: `LlamaCppClient` registered in `NODE_CLASS_MAPPINGS`/`NODE_DISPLAY_NAME_MAPPINGS`, in `tests/test_llamacpp.py`
|
||||
- [x] T005-I [US1] Register `LlamaCppClient` in `src/comfydv/__init__.py` — makes T005-T pass (depends on T004-I)
|
||||
- [x] T006-T [US1] Write FAILING test: `LlamaCppProvider` connection error surfaces a clear message (mirrors `OllamaProvider`'s `_post_json` connection-error contract), in `tests/test_llamacpp_provider.py` (witnesses `features/us1_connect_and_chat.feature` scenario "Unreachable llama.cpp server surfaces a clear error")
|
||||
- [x] T006-I [US1] Confirm `LlamaCppProvider.chat()` reuses the shared `_post_json` connection-error handling unchanged (likely no code change needed — verify, don't assume) — makes T006-T pass
|
||||
|
||||
**Checkpoint**: US1 fully functional and independently testable (MVP) — proves the adapter pattern for the chat path.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 — Get structured, validated output from llama.cpp (Priority: P1)
|
||||
|
||||
**Goal**: `structured_output=True` on `ChatCompletion` works identically against llama.cpp.
|
||||
|
||||
**Independent Test**: Enable `structured_output` with a schema, run against a llama.cpp-hosted model, confirm typed sockets populate and are never blank.
|
||||
|
||||
- [x] T007-T [US2] Write FAILING test: `LlamaCppProvider.chat_structured()` builds `base_url=f"{host}/v1"` and delegates to the shared `comfydv._llm.chat.chat_structured()` helper unchanged, in `tests/test_llamacpp_provider.py` (witnesses `features/us2_structured_output.feature` scenario "Valid structured response exposes typed fields, same as Ollama")
|
||||
- [x] T007-I [US2] Implement `LlamaCppProvider.chat_structured()` in `src/comfydv/_llm/llamacpp_provider.py` — zero new structured-output logic, same call shape `OllamaProvider.chat_structured()` already makes — makes T007-T pass
|
||||
- [x] T008 [US2] No new test needed for the retry-then-fail path (witnesses `features/us2_structured_output.feature` scenario "Invalid response retries then fails clearly, same as Ollama") — already fully covered by `tests/test_llm_chat_structured.py`'s existing suite, since `LlamaCppProvider.chat_structured()` calls the identical shared helper `OllamaProvider` does; re-testing it here would duplicate coverage without adding confidence (same reasoning as the prerequisite epic's D5)
|
||||
|
||||
**Checkpoint**: US1 + US2 both independently functional — the chat surface is now backend-agnostic in practice, not just in name.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 — See and control which models are loaded on llama.cpp (Priority: P2)
|
||||
|
||||
**Goal**: `LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel` work against llama.cpp via `LlamaCppProvider`.
|
||||
|
||||
**Independent Test**: List models via `LLMModelSelector` wired to `LlamaCppClient`; load/unload one; confirm status changes, including `loading`/`downloading` states if triggered.
|
||||
|
||||
- [x] T009-T [P] [US3] Write FAILING test: `LlamaCppProvider.list_models()` maps `GET /models`'s `data[].id`→`ModelInfo.name` and `data[].status.value`→`ModelInfo.status`, surfacing all five `ModelStatus` values without normalization (data-model.md), in `tests/test_llamacpp_provider.py` (witnesses `features/us3_model_lifecycle.feature` scenario "List models with full status vocabulary")
|
||||
- [x] T009-I [US3] Implement `LlamaCppProvider.list_models()` in `src/comfydv/_llm/llamacpp_provider.py` — makes T009-T pass
|
||||
- [x] T010-T [P] [US3] Write FAILING test: `LlamaCppProvider.load_model()`/`unload_model()` POST `{"model": id}` to `/models/load`/`/models/unload` and are idempotent, in `tests/test_llamacpp_provider.py` (witnesses `features/us3_model_lifecycle.feature` scenarios "Load a model into memory" and "Unload a model from memory")
|
||||
- [x] T010-I [US3] Implement `LlamaCppProvider.load_model()`/`unload_model()` in `src/comfydv/_llm/llamacpp_provider.py` — makes T010-T pass
|
||||
- [x] T011 [US3] No new node-layer tests needed — `LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel` are untouched by this epic (plan.md Structure Decision) and already have delegation-test coverage against a generic `_FakeProvider` in `tests/test_ollama.py`; that coverage is provider-agnostic by construction (FR-002), so it already proves these nodes work with `LlamaCppProvider` too, not just `OllamaProvider`
|
||||
|
||||
**Checkpoint**: US1 + US2 + US3 independently functional.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 4 — Swap from Ollama to llama.cpp without touching the rest of the workflow (Priority: P3)
|
||||
|
||||
**Goal**: Demonstrate/prove the adapter pattern's actual promise end-to-end.
|
||||
|
||||
**Independent Test**: Same workflow, only the connection node changes.
|
||||
|
||||
- [x] T012-T [US4] Write FAILING test: a workflow-shaped test (client → `ChatCompletion` → `LLMModelSelector` → `LLMLoadModel` → `LLMUnloadModel`) runs identically whether `client` is an `OllamaProvider`-double or a `LlamaCppProvider`-double — i.e. no node branches on provider type, in `tests/test_llamacpp.py` (witnesses `features/us4_swap_backends.feature` scenario "Replacing only the connection node preserves the workflow")
|
||||
- [x] T012-I [US4] No implementation expected — this test should already pass given T003-T011 (it's a regression/integration proof, not new functionality); if it fails, that reveals a node secretly branching on provider type, which would be a real bug to fix, not a feature to add
|
||||
|
||||
**Checkpoint**: all four user stories independently functional; the adapter pattern is proven end-to-end, not just asserted.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: Polish & Cross-Cutting Concerns
|
||||
|
||||
- [x] T013 [P] `ruff check --fix && ruff format` — clean
|
||||
- [x] T014 [P] `ty check` — clean, same pre-existing diagnostics as the prerequisite epic (unrelated to this feature — `comfy`/`server`/`folder_paths` unresolved-import, `format_string.py`'s dynamic RETURN_TYPES, `create_model`/`RandomChoice` — none touch the new files)
|
||||
- [x] T015 Confirmed via grep: `llamacpp_provider.py`/`llamacpp.py` import no `comfy`/`server`/`folder_paths` at module scope
|
||||
- [x] T016 `beacon doctor --strict`: only the pre-existing `llm-provider-abstraction: all specs [complete]` epic-gates item (PR #18, the archive-bookkeeping PR for the *prerequisite* epic, not yet merged — unrelated to this feature) and `tdd-commit-discipline` (disclosed pattern, same reasoning as the prerequisite epic)
|
||||
- [x] T017 Live smoke test — run against a real router-mode `llama-server` (Homebrew-installed, already present in the dev environment; a prior pass wrongly assumed no server was reachable without actually checking). Full lifecycle exercised against a real 5.6GB local GGUF model: `list_models()` → `load_model()` → `chat()` → `unload_model()`, plus explicit idempotency checks (calling `load_model()`/`unload_model()` again in the already-satisfied state). Found and fixed a real gap not caught by the mocked suite — see the finding below.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: no dependencies
|
||||
- **Foundational (Phase 2)**: none — nothing blocks user story work
|
||||
- **US1**: no dependency on other stories — genuinely the MVP
|
||||
- **US2**: depends on US1's `LlamaCppProvider` skeleton existing (T003-I), but its own logic (T007) has no dependency on US1's chat() specifically
|
||||
- **US3**: independent of US1/US2 except sharing `LlamaCppProvider`'s constructor (T003-I) — unlike the prerequisite epic's atomic cutover, there is no shared "client output type" migration risk here, since `LlamaCppClient` is a brand-new node, not a changed one
|
||||
- **US4**: depends on US1–US3 all being done (it's a proof, not new functionality)
|
||||
- **Polish**: depends on all four user stories
|
||||
|
||||
### Parallel Opportunities
|
||||
|
||||
- T002 can start immediately
|
||||
- T009-T and T010-T can run in parallel (different methods, same file, no shared state)
|
||||
- T013/T014 can run in parallel in Polish
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First
|
||||
|
||||
1. Phase 1 (Setup, trivial) → Phase 3 (US1) → **STOP and validate US1 independently** against a live `llama-server`.
|
||||
|
||||
### Incremental Delivery
|
||||
|
||||
1. US1 → validate → basic chat parity with Ollama, on a second backend.
|
||||
2. US2 → validate → structured-output parity — the shared mechanism holds.
|
||||
3. US3 → validate → model-management parity, with richer status than Ollama can offer.
|
||||
4. US4 → validate → the adapter pattern is proven, not just asserted.
|
||||
5. Polish.
|
||||
|
||||
Unlike the prerequisite epic, **this decomposition genuinely holds** —
|
||||
there is no shared "output type" migration forcing an atomic cutover, because
|
||||
`LlamaCppClient` is new, not a change to an existing node. Each phase really
|
||||
can land independently.
|
||||
|
||||
---
|
||||
|
||||
## Post-implementation review finding (fixed)
|
||||
|
||||
A `beacon-reviewer` pass ahead of PR open found `LlamaCppProvider.list_models()`
|
||||
caught *every* exception and returned `[]`, silently indistinguishable from
|
||||
"no models installed" — violating FR-006 and `contracts/llamacpp_provider_conformance.md`'s
|
||||
explicit requirement that a non-router-mode `llama-server` (unreachable
|
||||
endpoints → HTTP error on `GET /models`) surface a clear, specific error.
|
||||
|
||||
Fixed: `_get_json` (shared with `OllamaProvider`, in `ollama_provider.py`) now
|
||||
raises `RuntimeError` on an HTTP error status, matching `_post_json`'s
|
||||
existing behavior — its docstring already claimed this, it just didn't do it.
|
||||
`LlamaCppProvider.list_models()` now distinguishes `OSError` (genuinely
|
||||
unreachable — connection refused, DNS failure, timeout; all aiohttp
|
||||
connection-level exceptions are `OSError` subclasses) from `RuntimeError`
|
||||
(server responded, but with an error): the former still degrades gracefully
|
||||
to `[]` (consistent with `OllamaProvider`'s existing UX), the latter is
|
||||
re-raised naming router mode as the likely cause. Regression test added:
|
||||
`test_list_models_non_router_mode_raises_clear_error` in
|
||||
`tests/test_llamacpp_provider.py`. `OllamaProvider`'s own `list_models()`/
|
||||
`_fetch_models()` still catch broadly and degrade to `[]` unchanged — no
|
||||
spec requirement asks Ollama to make this distinction, and this fix doesn't
|
||||
force it to.
|
||||
|
||||
---
|
||||
|
||||
## Live smoke test finding (T017, fixed)
|
||||
|
||||
T017 had been marked `[-]` deferred on the assumption that no `llama-server`
|
||||
was reachable in the dev environment. That assumption was never actually
|
||||
checked — `llama-server` was installed via Homebrew the whole time, and a
|
||||
router-mode server was launched against a real local GGUF model
|
||||
(`--models-dir` pointed at a symlinked model file) to run the smoke test for
|
||||
real.
|
||||
|
||||
This caught a genuine gap the mocked suite couldn't: `contracts/llamacpp_provider_conformance.md`
|
||||
claimed router mode's `/models/load`/`/models/unload` return `{"success": true}`
|
||||
on an already-loaded/unloaded model, satisfying the `LLMProvider` protocol's
|
||||
idempotency requirement "without extra handling." That claim was never
|
||||
live-verified — live testing showed the opposite: both endpoints return HTTP
|
||||
400 (`"model is already running"` / `"model is not running"`) instead.
|
||||
`LlamaCppProvider.load_model()`/`unload_model()` now absorb exactly those two
|
||||
error messages as the desired end-state already reached (any other error
|
||||
still propagates); the contract doc is corrected to describe the real
|
||||
behavior. Regression tests added (mocked, so they run in CI):
|
||||
`test_load_model_already_running_is_idempotent`,
|
||||
`test_unload_model_not_running_is_idempotent`, and their
|
||||
`_other_http_error_still_raises` counterparts confirming non-idempotency
|
||||
errors aren't over-broadly swallowed.
|
||||
|
||||
Also observed live (informational, no code change needed): `load_model()`/
|
||||
`unload_model()` return once the request is *accepted*, not once the state
|
||||
transition completes — a 5.6GB model reported `LOADING` for several seconds
|
||||
before `LOADED`. This matches `ModelStatus`'s documented vocabulary (`loading`
|
||||
is a real, intended state) and how a real UI would behave — fire the request,
|
||||
poll `list_models()` for the transition. No protocol change; noted here so
|
||||
it's not mistaken for a future bug report.
|
||||
|
||||
**`chat_structured()` live-verified separately** (T007/T008's mocked coverage
|
||||
only ever exercised the call-shape, never the real network path): ran a
|
||||
second live smoke test — `load_model()` → `chat_structured()` with a real
|
||||
`pydantic.BaseModel` schema — against the same router-mode server. Result
|
||||
validated correctly (`Color(name='Red', hex_code='#FF0000')`), confirming
|
||||
pydantic-ai's `Agent`/`OpenAIProvider(base_url=f"{host}/v1")` mechanism
|
||||
genuinely works against llama-server's OpenAI-compatible endpoint, not just
|
||||
Ollama's (which was the only one live-verified in the prerequisite epic).
|
||||
No gap found here — recorded as verification evidence, not a fix.
|
||||
|
||||
**Net result**: every `LlamaCppProvider` method (`list_models`, `load_model`,
|
||||
`unload_model`, `chat`, `chat_structured`) has now been exercised against a
|
||||
real router-mode `llama-server`, not just mocks. T017 is genuinely done.
|
||||
@@ -0,0 +1 @@
|
||||
epic = "vlm-image-input"
|
||||
@@ -0,0 +1,41 @@
|
||||
# Specification Quality Checklist: VLM Image Input for ChatCompletion
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: 2026-07-22
|
||||
**Feature**: [spec.md](../spec.md)
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [x] No implementation details (languages, frameworks, APIs)
|
||||
- [x] Focused on user value and business needs
|
||||
- [x] Written for non-technical stakeholders
|
||||
- [x] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||
- [x] Requirements are testable and unambiguous
|
||||
- [x] Success criteria are measurable
|
||||
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||
- [x] All acceptance scenarios are defined
|
||||
- [x] Edge cases are identified
|
||||
- [x] Scope is clearly bounded
|
||||
- [x] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [x] All functional requirements have clear acceptance criteria
|
||||
- [x] User scenarios cover primary flows
|
||||
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [x] No implementation details leak into specification
|
||||
|
||||
## Notes
|
||||
|
||||
- The cross-provider "where does the image live / who translates it" decision
|
||||
is intentionally kept out of the spec (WHAT/WHY) and recorded in
|
||||
ADR-008 (HOW). The spec references it via the epic, not inline.
|
||||
- Multi-image-per-turn is documented as an out-of-MVP extension in Assumptions,
|
||||
not a functional requirement — keeps scope bounded.
|
||||
- Items are validated by review, not by an automated gate (`beacon` CLI is not
|
||||
installed in this environment; placeholder/ADR-reference checks were run
|
||||
manually — see the session's validation step).
|
||||
@@ -0,0 +1,64 @@
|
||||
# Contract: Image Input across the LLMProvider boundary
|
||||
|
||||
**Spec**: [spec.md](../spec.md) · **Data model**: [data-model.md](../data-model.md) · **ADR**: [ADR-008](../../../project-management/ADRs/ADR-008-multimodal-image-input-across-llmprovider-boundary.md)
|
||||
|
||||
This feature adds no new protocol methods and no new socket types. The contract
|
||||
below is the **behavioural conformance** every `LLMProvider` must satisfy for
|
||||
the new `Message.images` field, plus the node's input contract.
|
||||
|
||||
---
|
||||
|
||||
## C1 — `Message.images` carrier
|
||||
|
||||
- `Message.images: list[str] | None = None`, base64 strings (no `data:` prefix).
|
||||
- `images=None` or `[]` ⇒ the turn is text-only and MUST produce a request
|
||||
**byte-for-byte identical** to the pre-feature behaviour.
|
||||
|
||||
## C2 — `LLMProvider.chat()` conformance (both providers)
|
||||
|
||||
Given `messages` where the last user turn carries `images`:
|
||||
1. The provider MUST transmit those images with that turn to its backend using
|
||||
its native shape (Ollama flat `images`; llama.cpp OpenAI `image_url` parts).
|
||||
2. The provider MUST NOT transmit an `images` field for turns that have none
|
||||
(empty key omitted).
|
||||
3. All existing behaviour is preserved: blank-retry-with-new-seed loop, response
|
||||
caching, timeout, and error surfacing are unchanged by the presence of images.
|
||||
4. A backend that cannot process images (non-vision model / no `--mmproj`) MUST
|
||||
have its error surfaced to the caller, not swallowed (FR-006).
|
||||
|
||||
## C3 — `chat_structured()` conformance (shared helper, both providers)
|
||||
|
||||
1. Images on the last user turn MUST be attached as `BinaryContent` on the
|
||||
`Agent.run` `user_prompt`; images on history user turns MUST be attached to
|
||||
their `UserPromptPart`.
|
||||
2. All existing structured guarantees hold unchanged: bounded retries (0–5),
|
||||
`RuntimeError` on exhaustion naming model/attempts/snippet, never returns a
|
||||
value that failed schema validation.
|
||||
3. A text-only structured call MUST be indistinguishable from today's.
|
||||
|
||||
## C4 — `ChatCompletion` node input contract
|
||||
|
||||
1. Adds exactly one **optional** `image: ("IMAGE",)` input. No required input
|
||||
added; `RETURN_TYPES`/`RETURN_NAMES` positions unchanged (Constitution VI).
|
||||
2. Un-wired ⇒ behaviour, request, and outputs identical to today.
|
||||
3. Wired ⇒ image(s) attached to the current user turn only; `history` turns
|
||||
unchanged (FR-007).
|
||||
4. Works with `structured_output=True` (C3) and free-text (C2) alike, on either
|
||||
backend, with no per-backend wiring difference (FR-004).
|
||||
|
||||
---
|
||||
|
||||
## Test contracts (test-first — Constitution III)
|
||||
|
||||
| ID | Level | Asserts |
|
||||
|---|---|---|
|
||||
| T1 | `Message` unit | `images` defaults `None`; round-trips base64 list; text-only dump omits the key |
|
||||
| T2 | `OllamaProvider.chat` | image turn → payload message has flat `images:[...]`; text-only payload byte-identical to today (regression) |
|
||||
| T3 | `LlamaCppProvider.chat` | image turn → `content` becomes text+`image_url` parts; text-only `content` stays a plain string (regression) |
|
||||
| T4 | `chat_structured` | image turn builds `BinaryContent` on the prompt; text-only path unchanged; retry/validation contract intact |
|
||||
| T5 | node encode helper | synthetic `[1,H,W,3]` tensor → decodable base64 PNG; `None`/empty → `[]` |
|
||||
| T6 | node contract | optional `image` in `INPUT_TYPES`; un-wired run == today; wired run attaches to last turn only |
|
||||
|
||||
All tests run without a live ComfyUI or a live backend (mock at each provider's
|
||||
own `_post_json`/`Agent.run` seam, per the `test_ollama_provider.py`
|
||||
convention). T5 uses a synthetic tensor + Pillow (dev dep), no ComfyUI.
|
||||
@@ -0,0 +1,83 @@
|
||||
# Phase 1 Data Model: VLM Image Input for ChatCompletion
|
||||
|
||||
**Spec**: [spec.md](./spec.md) · **Research**: [research.md](./research.md)
|
||||
|
||||
The feature adds **one optional field** to an existing model and defines how it
|
||||
maps into each backend's wire shape. No new entities, no new socket types.
|
||||
|
||||
---
|
||||
|
||||
## Modified entity — `Message` (`src/comfydv/_llm/provider.py`)
|
||||
|
||||
```python
|
||||
class Message(BaseModel):
|
||||
role: Literal["system", "user", "assistant"]
|
||||
content: str
|
||||
images: list[str] | None = None # NEW — base64-encoded images (no data: prefix)
|
||||
```
|
||||
|
||||
**Field: `images`**
|
||||
- **Type**: `list[str] | None`, default `None`.
|
||||
- **Meaning**: base64-encoded image payloads associated with this turn. `None`
|
||||
(or empty) means a text-only turn — **byte-for-byte identical to today**.
|
||||
- **Carrier form**: raw base64 string, no `data:` URI prefix. Chosen because
|
||||
every target adapts *from* it (Ollama `images` array, OpenAI data-URI,
|
||||
pydantic-ai `BinaryContent`) — ADR-008.
|
||||
- **Validation**: no format validation at the model layer (the model stays a
|
||||
dumb carrier); malformed data surfaces as a backend error (FR-006). A turn
|
||||
may carry ≥1 image; MVP exercises exactly one.
|
||||
- **Serialization invariant**: provider payload construction MUST omit the
|
||||
`images` key when `None`/empty so existing text-only requests are unchanged
|
||||
(research.md Decision 2; FR-003, SC-004).
|
||||
|
||||
---
|
||||
|
||||
## Mapping table — one carrier, three wire shapes
|
||||
|
||||
| Path | Code site | Transform |
|
||||
|---|---|---|
|
||||
| Ollama free-text | `ollama_provider.py::chat` | none — `model_dump()`'s flat `images` array already matches `/api/chat`; only drop the key when empty |
|
||||
| llama.cpp free-text | `llamacpp_provider.py::chat` | rebuild `content` as OpenAI parts: `[{"type":"text",...},{"type":"image_url","image_url":{"url":"data:image/png;base64,<b64>"}}]` |
|
||||
| Structured (both) | `chat.py::chat_structured` | build `BinaryContent(data=b64decode(img), media_type="image/png")`; attach to `user_prompt` (last turn) / `UserPromptPart` (history turns) as `[text, *images]` |
|
||||
|
||||
---
|
||||
|
||||
## Node input — `ChatCompletion` (`src/comfydv/ollama.py`)
|
||||
|
||||
Add to `INPUT_TYPES["optional"]`:
|
||||
|
||||
```python
|
||||
"image": ("IMAGE",),
|
||||
```
|
||||
|
||||
- **Optional** — un-wired ⇒ `image=None` ⇒ the node builds exactly today's
|
||||
text-only user message. No new required input; no output/socket change
|
||||
(Constitution VI untouched — `RETURN_TYPES` positions 0/1 unchanged).
|
||||
- When wired: encode the tensor to base64 PNG(s) (research.md Decision 4) and
|
||||
set them on the appended `Message(role="user", ...)`. History turns are not
|
||||
modified (FR-007).
|
||||
|
||||
### Encode helper (node-local, `comfy`/Pillow lazy)
|
||||
|
||||
```
|
||||
_encode_image_tensor(image) -> list[str]:
|
||||
# image: ComfyUI IMAGE, torch float tensor [B, H, W, C] in 0..1
|
||||
# → for each frame: *255 → uint8 → PIL.Image.fromarray → PNG bytes → base64
|
||||
# returns [] for None/empty so callers treat it as "no image"
|
||||
```
|
||||
|
||||
Lives in `ollama.py` (node module, already `comfy`-guarded). `src/comfydv/_llm/`
|
||||
never imports torch/numpy/Pillow — it deals only in the base64 strings this
|
||||
helper produces.
|
||||
|
||||
---
|
||||
|
||||
## State & relationships
|
||||
|
||||
- No persistent state; no new caching entity. Existing `_CHAT_RESPONSE_CACHE`
|
||||
keys already include the dumped messages, so an added `images` value
|
||||
participates in the cache key automatically (same image + prompt ⇒ cache
|
||||
hit), and a text-only turn's key is unchanged since the empty key is omitted.
|
||||
- Relationship: `images` belongs to exactly one `Message` (one turn) — this is
|
||||
why the carrier is a message field, not a side-channel parameter (ADR-008
|
||||
Alternative A rejected).
|
||||
@@ -0,0 +1,11 @@
|
||||
Feature: US1 — Describe an image with a chat node
|
||||
|
||||
Scenario: Describe a wired image
|
||||
Given a chat node connected to a backend with a vision-capable model loaded and an image wired into the node's image input
|
||||
When the workflow executes with a prompt like "describe this image"
|
||||
Then the node returns a text response that reflects the actual content of the wired image
|
||||
|
||||
Scenario: No image wired behaves exactly as today
|
||||
Given the same chat node with no image wired
|
||||
When the workflow executes
|
||||
Then the node behaves exactly as it does today — text-only chat, identical response for identical text input — with no new required inputs and no change in output
|
||||
@@ -0,0 +1,11 @@
|
||||
Feature: US2 — Same image input on either backend
|
||||
|
||||
Scenario: Swap Ollama for llama.cpp and the image path still works
|
||||
Given a workflow that describes an image via the chat node wired to Ollama
|
||||
When the connection node is swapped to a llama.cpp one (pointed at a server with a multimodal model) with no other change
|
||||
Then the workflow still returns a description of the same image
|
||||
|
||||
Scenario: Both backends produce an image-grounded response
|
||||
Given equivalent image + prompt inputs on both backends
|
||||
When each workflow executes
|
||||
Then both produce a coherent image-grounded text response — no backend requires a different node, input shape, or wiring for the image
|
||||
@@ -0,0 +1,11 @@
|
||||
Feature: US3 — Structured output about an image
|
||||
|
||||
Scenario: Structured fields populated from the image
|
||||
Given the chat node with an image wired and structured output enabled with a valid schema
|
||||
When the workflow executes against a vision-capable model
|
||||
Then each schema field is available as its own typed output, populated from the image, with no required field blank
|
||||
|
||||
Scenario: Invalid structured output retries then fails clearly
|
||||
Given the same setup where the model first returns invalid or incomplete structured output
|
||||
When the workflow executes
|
||||
Then the node retries and, if still unsuccessful, fails with a clear error — the same retry/validation behaviour the text-only structured path already guarantees
|
||||
@@ -0,0 +1,125 @@
|
||||
# Implementation Plan: VLM Image Input for ChatCompletion
|
||||
|
||||
**Branch**: `009-vlm-image-input` | **Date**: 2026-07-22 | **Spec**: [spec.md](./spec.md)
|
||||
|
||||
**Input**: Feature specification from `/specs/009-vlm-image-input/spec.md`
|
||||
|
||||
## Summary
|
||||
|
||||
Let a workflow author wire a ComfyUI `IMAGE` into the existing generic
|
||||
`ChatCompletion` node so a vision-capable model can describe or reason about it,
|
||||
on **either** backend. Per ADR-008 (extending ADR-007's adapter pattern to a
|
||||
second input modality), images ride on an optional `Message.images` carrier and
|
||||
each provider translates that carrier into its own wire shape: Ollama's flat
|
||||
`/api/chat` `images` array (passes through untouched), llama.cpp's OpenAI
|
||||
`image_url` content-parts, and — for structured output — pydantic-ai
|
||||
`BinaryContent`, shared by both backends through `OpenAIChatModel`. The node
|
||||
converts its `IMAGE` tensor to base64 PNG; everything below the node deals only
|
||||
in base64 strings. Text-only behaviour is byte-for-byte unchanged when no image
|
||||
is wired.
|
||||
|
||||
## Technical Context
|
||||
|
||||
**Language/Version**: Python ≥3.11 (unchanged, per `pyproject.toml`)
|
||||
|
||||
**Primary Dependencies**: existing only for runtime — `aiohttp` (Ollama/llama.cpp
|
||||
REST), `pydantic-ai-slim[openai]>=2.9.0` (structured path; its `BinaryContent`
|
||||
multimodal type was verified against the installed 2.9.0 source, see
|
||||
`research.md`). **No new core runtime dependency.** `pillow` is added to the
|
||||
**dev** group so the node's tensor→PNG encoder is unit-testable without a live
|
||||
ComfyUI; at runtime Pillow/numpy are ComfyUI-provided (same stance the repo
|
||||
already takes for torch).
|
||||
|
||||
**Storage**: N/A — no persistent storage; reuses the existing
|
||||
`_CHAT_RESPONSE_CACHE` (an `images` value participates in the cache key
|
||||
automatically).
|
||||
|
||||
**Testing**: `pytest` via `uv run pytest`, following the
|
||||
`tests/test_ollama_provider.py` convention (mock at each provider's own
|
||||
`_post_json` / `Agent.run` seam, no live server or ComfyUI required). Test-first
|
||||
per Constitution III; the tensor-encode test uses a synthetic tensor + Pillow.
|
||||
|
||||
**Target Platform**: ComfyUI custom-node runtime, same as the existing LLM nodes.
|
||||
|
||||
**Project Type**: Library / ComfyUI custom-node pack (single project).
|
||||
|
||||
**Performance Goals**: No new numeric target; image encoding is a one-shot
|
||||
per-execution PNG encode, negligible against inference latency.
|
||||
|
||||
**Constraints**: Text-only requests MUST stay byte-identical (FR-003/SC-004) —
|
||||
providers omit an empty `images` key. `src/comfydv/_llm/` must not import
|
||||
torch/numpy/Pillow (Constitution IV) — tensor handling stays in the node.
|
||||
llama.cpp image support requires a server launched with `--mmproj` (deployment
|
||||
prerequisite, surfaced as a clear error when absent, not configured by comfydv).
|
||||
|
||||
**Scale/Scope**: One new `Message` field; a per-provider mapping in each
|
||||
`chat()` plus the shared `chat_structured()`; one optional node input + a
|
||||
node-local encode helper. No new node classes, no new socket types, no protocol
|
||||
method changes.
|
||||
|
||||
## Constitution Check
|
||||
|
||||
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
|
||||
|
||||
| Principle | Verdict | Notes |
|
||||
|---|---|---|
|
||||
| I. ComfyUI Contract First | PASS | `ChatCompletion` keeps its `INPUT_TYPES`/`RETURN_TYPES`/`FUNCTION`/`CATEGORY`; only an optional input is added. No new registration, no ComfyUI changes. |
|
||||
| II. Sandbox All User-Supplied Code | N/A | No template/expression evaluation in this feature. |
|
||||
| III. Test-First | PASS (binding) | Test contracts T1–T6 (`contracts/image-input-contract.md`) written test-first, mirroring `test_ollama_provider.py`. Each runs without a live ComfyUI/backend. |
|
||||
| IV. Graceful Degradation Outside ComfyUI | PASS (binding) | `_llm/` stays torch/numpy/Pillow-free — pure base64 carrier + mapping, unit-testable. Tensor→PNG lives in `ollama.py` (already `comfy`-guarded) with lazy Pillow/numpy import, so module import outside ComfyUI is unaffected. |
|
||||
| V. Simplicity — Function Before Class | PASS | No new class. New logic is a `Message` field, two small per-provider transforms, one shared helper edit, and one module-level encode function. |
|
||||
| VI. Fixed Output Positions | PASS | Outputs are untouched — only an optional **input** is added; `RETURN_TYPES`/`RETURN_NAMES` positions 0/1 and the structured extra-outputs contract are unchanged. |
|
||||
|
||||
Re-checked post-Phase 1 design (data-model.md, contracts/): unchanged — no new
|
||||
gate violations. No Complexity Tracking entries needed.
|
||||
|
||||
## Project Structure
|
||||
|
||||
### Documentation (this feature)
|
||||
|
||||
```text
|
||||
specs/009-vlm-image-input/
|
||||
├── plan.md # This file
|
||||
├── research.md # Phase 0 — wire shapes verified against installed deps
|
||||
├── data-model.md # Phase 1 — Message.images + per-provider mapping
|
||||
├── quickstart.md # Phase 1 — minimal describe-an-image workflow
|
||||
├── contracts/
|
||||
│ └── image-input-contract.md # Phase 1 — behavioural + test contracts (T1–T6)
|
||||
├── checklists/requirements.md # Spec quality checklist (from /speckit-specify)
|
||||
└── tasks.md # Phase 2 (/speckit-tasks) — not created here
|
||||
```
|
||||
|
||||
### Source Code (repository root)
|
||||
|
||||
```text
|
||||
src/comfydv/
|
||||
├── ollama.py # MODIFIED — ChatCompletion: optional `image` input +
|
||||
│ # node-local _encode_image_tensor() (lazy Pillow/numpy)
|
||||
├── _llm/
|
||||
│ ├── provider.py # MODIFIED — Message gains `images: list[str] | None = None`
|
||||
│ ├── ollama_provider.py # MODIFIED — chat(): pass flat images through; omit empty key
|
||||
│ ├── llamacpp_provider.py # MODIFIED — chat(): map images → OpenAI image_url parts
|
||||
│ └── chat.py # MODIFIED — chat_structured(): images → BinaryContent on prompt
|
||||
└── __init__.py # unchanged — no new node class or mapping
|
||||
|
||||
tests/
|
||||
├── test_provider.py (or test_ollama_provider.py) # T1 Message carrier + regression
|
||||
├── test_ollama_provider.py # MODIFIED — T2 Ollama image mapping + text regression
|
||||
├── test_llamacpp_provider.py # MODIFIED — T3 llama.cpp content-parts + text regression
|
||||
├── test_llm_chat.py / chat tests # T4 chat_structured multimodal + regression
|
||||
└── test_ollama.py # MODIFIED — T5 encode helper, T6 node input contract
|
||||
|
||||
pyproject.toml # MODIFIED — add `pillow` to [dependency-groups].dev only
|
||||
```
|
||||
|
||||
**Structure Decision**: Purely additive edits to the four existing `_llm`/node
|
||||
files that ADR-007 established — no new module, because there is no new class or
|
||||
node (contrast 008, which added a provider + node). The image path threads
|
||||
through the exact seams the text path already uses, which is the whole point of
|
||||
ADR-008: a second modality on the same adapter, not a parallel structure.
|
||||
|
||||
## Complexity Tracking
|
||||
|
||||
> **Fill ONLY if Constitution Check has violations that must be justified**
|
||||
|
||||
None — see Constitution Check above.
|
||||
@@ -0,0 +1,48 @@
|
||||
# Quickstart: Describe an image with ChatCompletion
|
||||
|
||||
**Spec**: [spec.md](./spec.md)
|
||||
|
||||
Minimal end-to-end walkthrough of the feature once shipped.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- A running backend with a **vision-capable** model:
|
||||
- **Ollama** — a multimodal model pulled and available (e.g. a llava-class model), or
|
||||
- **llama.cpp** — `llama-server` launched in router mode **with a multimodal projector**: `--mmproj <projector.gguf>` alongside the model.
|
||||
- comfydv installed in ComfyUI.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Add an image source to the canvas (e.g. **Load Image**) → gives an `IMAGE`.
|
||||
2. Add a client node (**OllamaClient** or **LlamaCppClient**) → gives `LLM_CLIENT`.
|
||||
3. Add **ChatCompletion**. Wire:
|
||||
- `client` ← the client node
|
||||
- `model` ← a vision-capable model name (typed or wired)
|
||||
- `prompt` ← `"Describe this image in one sentence."`
|
||||
- `image` ← the `IMAGE` from step 1 ← **the only new wire**
|
||||
4. Queue the prompt. The `response` output is a text description of the image.
|
||||
|
||||
## Structured variant (optional)
|
||||
|
||||
- On **ChatCompletion**, set `structured_output = True` and provide a schema, e.g.:
|
||||
```json
|
||||
{"type":"object","properties":{"caption":{"type":"string"},"has_text":{"type":"boolean"}},"required":["caption","has_text"]}
|
||||
```
|
||||
- Run: each field (`caption`, `has_text`) appears as its own typed output,
|
||||
populated from the image, with no required field blank.
|
||||
|
||||
## Swap backends (proves FR-004 / SC-002)
|
||||
|
||||
- Replace **OllamaClient** with **LlamaCppClient** (pointed at an `--mmproj`
|
||||
server) — **change nothing else**. Re-queue: same image description path.
|
||||
|
||||
## What stays the same
|
||||
|
||||
- Leave `image` un-wired and ChatCompletion behaves exactly as before — text-only,
|
||||
identical results. No existing workflow changes.
|
||||
|
||||
## Expected failure (proves FR-006 / SC-005)
|
||||
|
||||
- Wire an image but select a **non-vision** model (or a `llama-server` started
|
||||
without `--mmproj`): the node reports a clear error that the model/server
|
||||
can't process images — it does not silently answer as if no image was sent.
|
||||
@@ -0,0 +1,155 @@
|
||||
# Phase 0 Research: VLM Image Input for ChatCompletion
|
||||
|
||||
**Spec**: [spec.md](./spec.md) · **Plan**: [plan.md](./plan.md) · **ADR**: [ADR-008](../../project-management/ADRs/ADR-008-multimodal-image-input-across-llmprovider-boundary.md)
|
||||
|
||||
ADR-008 recorded the boundary decision (images on `Message.images`, translated
|
||||
per-provider) but deferred the exact wire shapes for live verification. This
|
||||
file resolves them against the **installed** dependency versions and the
|
||||
current provider code, not from memory.
|
||||
|
||||
---
|
||||
|
||||
## Decision 1 — pydantic-ai multimodal vehicle (structured-output path)
|
||||
|
||||
**Decision**: In the shared `chat_structured()` helper (`src/comfydv/_llm/chat.py`),
|
||||
attach images as `pydantic_ai.messages.BinaryContent(data=<png bytes>,
|
||||
media_type="image/png")` inside a `Sequence[UserContent]`. The current-turn
|
||||
image rides on `Agent.run(user_prompt=[text, BinaryContent(...)])`; a
|
||||
history turn's image rides on `UserPromptPart(content=[text, BinaryContent(...)])`.
|
||||
|
||||
**Rationale / verified**: Read directly from the pinned
|
||||
`pydantic_ai_slim==2.9.0` source in this environment:
|
||||
|
||||
- `BinaryContent` (`messages.py:521`) — `__init__(self, data: bytes, *,
|
||||
media_type: ..., identifier=None, ...)`; exposes a `.base64` helper. `data`
|
||||
is **bytes**, so `chat.py` must `base64.b64decode()` the `Message.images`
|
||||
string into bytes when building it.
|
||||
- `UserPromptPart.content: str | Sequence[UserContent]` (`messages.py:1022`)
|
||||
and `user_prompt` on `Agent.run` accept the same. `UserContent = str |
|
||||
TextContent | MultiModalContent | CachePoint` (`messages.py:899`), and
|
||||
`MultiModalContent` includes `BinaryContent`/`ImageUrl` — so a `[text,
|
||||
image]` list is the supported shape.
|
||||
- `OpenAIChatModel` renders `BinaryContent` for images as an OpenAI
|
||||
`image_url` data-URI, and reads `BinaryContent.vendor_metadata['detail']`
|
||||
for the `detail` setting (documented in the field's own docstring).
|
||||
|
||||
**Consequence**: the structured path is provider-agnostic *for free* — both
|
||||
Ollama and llama.cpp reach `/v1/chat/completions` through the same
|
||||
`OpenAIChatModel`, so one change in `chat.py` covers structured output on both
|
||||
backends. No per-provider structured code.
|
||||
|
||||
**Alternatives considered**: `ImageUrl(url="data:image/png;base64,...")` — also
|
||||
supported, but requires assembling a data URI string; `BinaryContent` from raw
|
||||
bytes + media type is the more direct representation of what we hold and lets
|
||||
pydantic-ai own the data-URI formatting.
|
||||
|
||||
---
|
||||
|
||||
## Decision 2 — Ollama free-text path (`/api/chat`)
|
||||
|
||||
**Decision**: `OllamaProvider.chat()` sends each message's images as a flat
|
||||
`images` array of **base64 strings** (no `data:` prefix) alongside `content`,
|
||||
which is exactly Ollama's native `/api/chat` message schema. Because
|
||||
`Message.images` already holds base64 strings, `Message.model_dump()` produces
|
||||
the correct shape with **no transform** — the field flows straight through.
|
||||
|
||||
**Rationale / verified**: `OllamaProvider.chat()`
|
||||
(`src/comfydv/_llm/ollama_provider.py:280`) already builds
|
||||
`payload_messages = [m.model_dump() for m in messages]` and POSTs to
|
||||
`/api/chat`. Ollama's documented `/api/chat` message object is
|
||||
`{"role", "content", "images": [<base64>, ...]}` — the flat sibling field this
|
||||
carrier maps onto directly. This is the reason base64 is the neutral carrier
|
||||
form (ADR-008).
|
||||
|
||||
**Constraint discovered — byte-identical text path (FR-003/SC-004)**: adding
|
||||
`images: list[str] | None = None` to `Message` means a text-only message would
|
||||
dump as `{"role","content","images":null}`, changing today's request body.
|
||||
Providers MUST drop a `None`/empty `images` before sending. Resolution:
|
||||
serialize provider payload messages with the images key omitted when empty
|
||||
(e.g. `model_dump(exclude_none=True)`, or drop the key explicitly). Guarded by
|
||||
the existing Ollama provider tests, which assert the exact payload.
|
||||
|
||||
---
|
||||
|
||||
## Decision 3 — llama.cpp free-text path (`/v1/chat/completions`)
|
||||
|
||||
**Decision**: `LlamaCppProvider.chat()` maps a message carrying images into
|
||||
OpenAI-style multimodal `content` **parts** before POSTing:
|
||||
`content: [{"type":"text","text":<content>},
|
||||
{"type":"image_url","image_url":{"url":"data:image/png;base64,<b64>"}}]`.
|
||||
Messages with no images keep the plain-string `content` unchanged.
|
||||
|
||||
**Rationale / verified**: `LlamaCppProvider.chat()`
|
||||
(`src/comfydv/_llm/llamacpp_provider.py:163`) builds
|
||||
`payload_messages = [m.model_dump() for m in messages]` and POSTs to
|
||||
`/v1/chat/completions`. Unlike Ollama, a flat `images` sibling is **not**
|
||||
understood there — OpenAI's vision schema requires images inside `content` as
|
||||
typed parts. `llama-server` implements this OpenAI-compatible multimodal
|
||||
format **only when launched with a multimodal projector (`--mmproj`)**; without
|
||||
it, image parts yield a server error (surfaced per FR-006, not crashed on).
|
||||
This is the single point where the two providers genuinely diverge — exactly
|
||||
the leakage ADR-008 localizes inside each provider.
|
||||
|
||||
**Alternatives considered**: normalizing Ollama *up* to content-parts too (one
|
||||
shared mapper) — rejected in ADR-008 Alternative C: it forces the
|
||||
currently-simpler Ollama path to do extra work and inverts "each provider owns
|
||||
its wire format."
|
||||
|
||||
---
|
||||
|
||||
## Decision 4 — ComfyUI IMAGE tensor → base64 PNG (node layer)
|
||||
|
||||
**Decision**: The `ChatCompletion` node converts its optional `IMAGE` input to
|
||||
base64 PNG(s) via Pillow: ComfyUI IMAGE is a float tensor `[B, H, W, C]` in
|
||||
`0..1`; scale to `uint8`, `PIL.Image.fromarray(...)`, save PNG to an in-memory
|
||||
buffer, base64-encode. A batch of `B` frames becomes `B` base64 strings in the
|
||||
turn's `images` list (natural multi-image; MVP exercises `B=1`). The import of
|
||||
Pillow/numpy is **lazy** (inside the encode function), so the module still
|
||||
imports cleanly outside ComfyUI (Constitution IV).
|
||||
|
||||
**Rationale**: Pillow is the ComfyUI-ecosystem standard for IMAGE tensor ↔
|
||||
file and is present in every ComfyUI install; numpy comes with torch. Neither
|
||||
is added to comfydv's **core** runtime deps — they are ComfyUI-provided, the
|
||||
same stance the repo already takes for torch (dev-only in `pyproject.toml`).
|
||||
To keep the encoder **test-first** (Constitution III) without a live ComfyUI,
|
||||
add `pillow` to the **dev** dependency group so a unit test can feed a
|
||||
synthetic `numpy`/`torch` tensor through the pure encode function and assert a
|
||||
decodable PNG.
|
||||
|
||||
**Boundary kept clean**: only the node (`src/comfydv/ollama.py`, already
|
||||
`comfy`-guarded) touches tensors/Pillow. Everything in `src/comfydv/_llm/`
|
||||
deals purely in base64 strings and stays unit-testable with hand-crafted
|
||||
strings — no torch, numpy, or Pillow import there.
|
||||
|
||||
**Edge cases (FR-006, Edge Cases)**: an un-wired optional input arrives as
|
||||
`None` → node builds today's exact text-only message. A zero-size / empty batch
|
||||
tensor → treated as "no image". A non-vision model or non-`mmproj` server
|
||||
returns a backend error → surfaced with a clear message, never a silent
|
||||
image-less answer.
|
||||
|
||||
---
|
||||
|
||||
## Decision 5 — where the image attaches on the turn (FR-007)
|
||||
|
||||
**Decision**: The node attaches images to the **current user turn only** — the
|
||||
`Message(role="user", content=prompt, images=[...])` it already appends. Prior
|
||||
`history` turns are untouched. The structured helper likewise only lifts images
|
||||
onto the final user turn (and any history turn that already carried them),
|
||||
matching its existing "last message is the prompt" contract
|
||||
(`chat.py:106`, which requires `messages[-1].role == "user"`).
|
||||
|
||||
---
|
||||
|
||||
## Summary of resolved unknowns
|
||||
|
||||
| Unknown (from ADR-008) | Resolved to |
|
||||
|---|---|
|
||||
| pydantic-ai multimodal type | `BinaryContent(data=bytes, media_type="image/png")` — verified in installed 2.9.0 |
|
||||
| Structured path per-provider? | No — shared via `OpenAIChatModel`; one change in `chat.py` |
|
||||
| Ollama wire shape | flat `images: [base64]` on the message; passes through `model_dump()` |
|
||||
| llama.cpp wire shape | OpenAI `image_url` content-parts; requires `--mmproj` |
|
||||
| Text-path byte-identity | drop empty `images` key in provider payloads (guarded by existing tests) |
|
||||
| Tensor → base64 | Pillow, lazy import in node; `pillow` added to dev deps for testability |
|
||||
| No new runtime deps | Confirmed — Pillow/numpy are ComfyUI-provided, dev-only here |
|
||||
|
||||
No `NEEDS CLARIFICATION` remain.
|
||||
@@ -0,0 +1,148 @@
|
||||
# Feature Specification: VLM Image Input for ChatCompletion
|
||||
|
||||
**Feature Branch**: `009-vlm-image-input`
|
||||
|
||||
**Created**: 2026-07-22
|
||||
|
||||
**Status**: Draft
|
||||
|
||||
**Input**: User description: "Let a workflow author wire a ComfyUI IMAGE into the existing generic ChatCompletion node so a vision-capable model (VLM) on either backend (Ollama multimodal models, llama.cpp multimodal via mmproj) can describe or understand the image. Provider-agnostic per ADR-007/ADR-008: the node attaches the image to the user message; each provider maps it to its own wire format. The Message carrier gains an optional image field; text-only behaviour is unchanged when no image is wired."
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Describe an image with a chat node (Priority: P1) 🎯 MVP
|
||||
|
||||
As a ComfyUI workflow author with a vision-capable model available, I want to
|
||||
wire an image into the chat node I already use and get back a text description
|
||||
or answer about that image, so I can add image understanding to a workflow
|
||||
without learning a new node.
|
||||
|
||||
**Why this priority**: This is the entire point of the feature — a picture in,
|
||||
a text understanding out — and the proof that image input works through the
|
||||
existing generic node on at least one backend.
|
||||
|
||||
**Independent Test**: Wire any image source into the chat node's image input,
|
||||
point the node at a loaded vision-capable model, run the workflow, and confirm
|
||||
the response text describes the wired image.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a chat node connected to a backend with a vision-capable model loaded and an image wired into the node's image input, **When** the workflow executes with a prompt like "describe this image", **Then** the node returns a text response that reflects the actual content of the wired image.
|
||||
2. **Given** the same chat node with **no** image wired, **When** the workflow executes, **Then** the node behaves exactly as it does today — text-only chat, identical response for identical text input — with no new required inputs and no change in output.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Same image input on either backend (Priority: P1)
|
||||
|
||||
As a workflow author, I want image input to work the same way whether my chat
|
||||
node is connected to Ollama or to llama.cpp, so I don't have to rebuild or
|
||||
relearn the image path when I switch backends — exactly as text and structured
|
||||
output already behave identically across the two.
|
||||
|
||||
**Why this priority**: The generic-node promise (ADR-007) is the reason this
|
||||
feature is small; this story is what proves the image path honours it rather
|
||||
than quietly becoming backend-specific.
|
||||
|
||||
**Independent Test**: Run User Story 1 unchanged against an Ollama connection
|
||||
and against a llama.cpp connection (each with a vision-capable model), and
|
||||
confirm both return a description of the wired image using the identical node
|
||||
setup.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a workflow that describes an image via the chat node wired to Ollama, **When** the connection node is swapped to a llama.cpp one (pointed at a server with a multimodal model) with no other change, **Then** the workflow still returns a description of the same image.
|
||||
2. **Given** equivalent image + prompt inputs on both backends, **When** each workflow executes, **Then** both produce a coherent image-grounded text response — no backend requires a different node, input shape, or wiring for the image.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Structured output about an image (Priority: P2)
|
||||
|
||||
As a workflow author, I want to combine image input with the node's existing
|
||||
structured-output mode, so a VLM can return schema-validated fields extracted
|
||||
from an image (for example a caption, a list of detected objects, or a
|
||||
yes/no), not just free text.
|
||||
|
||||
**Why this priority**: Structured output is an existing, valued capability;
|
||||
making it work with images turns "describe this" into usable, wired,
|
||||
downstream-typed data. It builds on User Story 1 and is lower risk to defer
|
||||
than getting basic image chat working at all.
|
||||
|
||||
**Independent Test**: Enable structured output on the chat node with a schema,
|
||||
wire an image, run against a vision-capable model, and confirm each schema
|
||||
field is populated from the image and no required field is blank.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** the chat node with an image wired and structured output enabled with a valid schema, **When** the workflow executes against a vision-capable model, **Then** each schema field is available as its own typed output, populated from the image, with no required field blank.
|
||||
2. **Given** the same setup where the model first returns invalid or incomplete structured output, **When** the workflow executes, **Then** the node retries and, if still unsuccessful, fails with a clear error — the same retry/validation behaviour the text-only structured path already guarantees.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- What happens when an image is wired but the selected model is **not**
|
||||
vision-capable? The node must surface a clear error attributable to the model
|
||||
lacking image support, not crash and not silently drop the image and answer
|
||||
as if none was sent.
|
||||
- What happens when the backend server is reachable but was not started with
|
||||
multimodal support (e.g. a llama.cpp server launched without an `mmproj`
|
||||
projector)? The node should report a clear, specific error rather than an
|
||||
unhelpful generic failure.
|
||||
- What happens with an empty or zero-size image input, or an image input that
|
||||
is wired but carries no actual image data? The node should treat it as "no
|
||||
image" or report a clear error — never send a malformed request.
|
||||
- What happens when both an image and a multi-turn history are present? The
|
||||
image must be associated with the current user turn, and prior turns must
|
||||
remain unaffected.
|
||||
- What happens to the node's text-only path for a model/backend that does not
|
||||
understand images at all — does an un-wired image input leave the request
|
||||
byte-for-byte identical to today's? (It must.)
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: The system MUST let a workflow author provide an image to the existing chat node through a single, **optional** image input — no new node and no new required input.
|
||||
- **FR-002**: When an image is provided, the system MUST include it with the current user turn sent to the connected model, so a vision-capable model can ground its response in that image.
|
||||
- **FR-003**: When **no** image is provided, the system MUST send exactly the request it sends today — text-only behaviour, inputs, and outputs unchanged, with no regression for existing workflows.
|
||||
- **FR-004**: Image input MUST work identically across both supported backends from the workflow author's perspective — same node, same wiring, same input shape — with each backend's differing native image format handled internally, not exposed on the graph.
|
||||
- **FR-005**: Image input MUST be compatible with the node's existing structured-output mode: an image-grounded response can be schema-validated with the same retry and validation guarantees as the text-only structured path.
|
||||
- **FR-006**: The system MUST surface a clear, specific error when an image is provided but the target model or backend cannot process images (non-vision model, or a server without multimodal support), rather than crashing or silently discarding the image.
|
||||
- **FR-007**: The system MUST associate a provided image with the current user turn only, leaving any prior conversation history unchanged.
|
||||
|
||||
### Key Entities *(include if feature involves data)*
|
||||
|
||||
- **Chat message**: The existing per-turn unit of a chat request. Extended so a
|
||||
turn can optionally carry one or more images in addition to its text; a
|
||||
turn with no image is unchanged from today.
|
||||
- **Image input**: An image supplied on the workflow canvas (the standard
|
||||
ComfyUI image type) and attached to the current user turn; provider-neutral
|
||||
at the boundary, translated to each backend's native shape internally.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: A workflow author can make an existing chat node describe a wired image by adding exactly one connection (the image), with no new node and no other node changes.
|
||||
- **SC-002**: The same image-describing workflow runs unchanged when repointed from one backend to the other — zero edits beyond swapping the connection node.
|
||||
- **SC-003**: Structured-output workflows with an image populate every required schema field from the image content, with zero blank-required-field results, matching the text-only structured guarantee.
|
||||
- **SC-004**: Every existing text-only workflow produces identical results after this feature ships — no observable change when no image is wired (existing backend behaviour tests remain green).
|
||||
- **SC-005**: Providing an image to a non-vision model or a non-multimodal server yields a clear, specific error in 100% of such cases — never a crash and never a silently image-less answer presented as if the image was seen.
|
||||
|
||||
## Assumptions
|
||||
|
||||
- Workflow authors run their own backend (Ollama or llama.cpp) with a
|
||||
vision-capable model available and loaded; for llama.cpp this means the
|
||||
server was launched with a multimodal projector (`mmproj`). This feature does
|
||||
not install, configure, download, or launch vision models.
|
||||
- Scope is still **images only** — no video, audio, or document modalities; and
|
||||
image **input** only — no image generation or output.
|
||||
- A single image per turn is the primary target; carrying more than one image
|
||||
per turn is a natural extension of the same carrier but is not a required
|
||||
acceptance criterion of the MVP.
|
||||
- The generic `ChatCompletion` node, the `LLMProvider` protocol, and both
|
||||
providers already exist (ADR-007) and are extended, not replaced; the
|
||||
cross-provider image-carrier decision is recorded in ADR-008.
|
||||
- The standard ComfyUI image type is the input; converting it to the neutral
|
||||
form each backend consumes is an internal concern of this feature, not
|
||||
something the workflow author sees.
|
||||
@@ -0,0 +1,146 @@
|
||||
# Tasks: VLM Image Input for ChatCompletion
|
||||
|
||||
**Input**: Design documents from `/specs/009-vlm-image-input/`
|
||||
|
||||
**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/image-input-contract.md
|
||||
|
||||
**Tests**: First-class — every implementation task has a paired failing-test task (`-T`/`-I` suffix). Contracts T1–T6 in `contracts/image-input-contract.md` map to the pairs below.
|
||||
|
||||
**Organization**: Grouped by user story (spec.md priorities: US1 P1 🎯 MVP, US2 P1, US3 P2).
|
||||
|
||||
## Format: `[ID] [P?] [Story] Description`
|
||||
|
||||
- **[P]**: Can run in parallel (different files, no dependencies)
|
||||
- **[Story]**: US1–US3
|
||||
- **-T / -I**: paired test (red) / implementation (green) — the `-T` is committed failing before its `-I` partner (no test + impl in one commit)
|
||||
|
||||
## Path Conventions
|
||||
|
||||
Single project: `src/comfydv/`, `tests/` at repo root. Purely additive edits to
|
||||
the existing `_llm`/node files (plan.md Structure Decision) — no new module.
|
||||
`src/comfydv/_llm/` stays torch/numpy/Pillow-free (Constitution IV); tensor
|
||||
handling lives only in the `comfy`-guarded `ollama.py`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup
|
||||
|
||||
- [x] T001 Add `pillow` to `[dependency-groups].dev` in `pyproject.toml` — lets the node's tensor→PNG encoder be unit-tested without a live ComfyUI; runtime Pillow/numpy are ComfyUI-provided, so **no core runtime dependency is added** (research.md Decision 4)
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: the image carrier every path depends on. **⚠️ No user story work can begin until this is complete.**
|
||||
|
||||
- [x] T002-T Write FAILING test: `Message.images` defaults to `None`, round-trips a base64 list, and a text-only message's transport dump **omits** the `images` key (byte-identical to today), in `tests/test_llm_provider.py` (contract T1)
|
||||
- [x] T002-I Add `images: list[str] | None = None` to `Message` in `src/comfydv/_llm/provider.py` — makes T002-T pass
|
||||
|
||||
**Checkpoint**: carrier ready — user stories can begin.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 — Describe an image with a chat node (Priority: P1) 🎯 MVP
|
||||
|
||||
**Goal**: A workflow author wires a ComfyUI `IMAGE` into the existing `ChatCompletion` node and gets back a text description via an Ollama vision model; the text-only path is unchanged when no image is wired.
|
||||
|
||||
**Independent Test**: Wire an image → `ChatCompletion` → Ollama (vision model), confirm the response describes the image; un-wire the image and confirm behaviour/output identical to today.
|
||||
|
||||
- [x] T003-T [P] [US1] Write FAILING test: node-local `_encode_image_tensor()` converts a synthetic `[1,H,W,3]` float tensor (0..1) into a **decodable** base64 PNG, encodes a `B>1` batch to a list of that length, and returns `[]` for `None`/empty, in `tests/test_ollama.py` (contract T5; witnesses `features/us1_describe_image.feature` scenario "Describe a wired image")
|
||||
- [x] T003-I [US1] Implement `_encode_image_tensor()` in `src/comfydv/ollama.py` — lazy `PIL`/`numpy` import so module import stays clean outside ComfyUI (Constitution IV); batch → one base64 string per frame — makes T003-T pass
|
||||
- [x] T004-T [US1] Write FAILING test: `ChatCompletion.INPUT_TYPES` exposes an **optional** `image: ("IMAGE",)`; `RETURN_TYPES`/`RETURN_NAMES` positions are unchanged; an un-wired run builds the same text-only messages as today; a wired run attaches images to the **last user turn only** (history untouched), in `tests/test_ollama.py` (contract T6; witnesses both `features/us1_describe_image.feature` scenarios)
|
||||
- [x] T004-I [US1] Add the optional `image` input (with a tooltip noting a vision-capable model is required; llama.cpp needs `--mmproj`) and attach encoded images to the appended user `Message` in `ChatCompletion.chat()` in `src/comfydv/ollama.py` — makes T004-T pass (depends on T003-I, T002-I)
|
||||
- [x] T005-T [P] [US1] Write FAILING test: `OllamaProvider.chat()` forwards a message's images as a flat `images:[...]` array to `/api/chat`, and a text-only call's payload is **byte-identical to today** (regression), in `tests/test_ollama_provider.py` (contract T2; witnesses `features/us1_describe_image.feature` scenario "Describe a wired image")
|
||||
- [x] T005-I [US1] Ensure `OllamaProvider.chat()` passes images through and omits the empty `images` key (e.g. `model_dump(exclude_none=True)`) in `src/comfydv/_llm/ollama_provider.py` — makes T005-T pass (depends on T002-I)
|
||||
|
||||
**Checkpoint**: describe-an-image works end-to-end on Ollama (MVP); every existing text-only test stays green.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 — Same image input on either backend (Priority: P1)
|
||||
|
||||
**Goal**: The same node and wiring drive image input on llama.cpp too, via its OpenAI-compatible content-parts shape — proving the generic-node promise (ADR-007/008) holds for the image path.
|
||||
|
||||
**Independent Test**: Run the US1 workflow unchanged against a llama.cpp server (launched with `--mmproj`); swap the Ollama client node for the llama.cpp one with no other change and confirm the image is still described.
|
||||
|
||||
- [x] T006-T [P] [US2] Write FAILING test: `LlamaCppProvider.chat()` maps a message's images into OpenAI `content` parts (`{"type":"text",...}` + `{"type":"image_url","image_url":{"url":"data:image/png;base64,..."}}`) for `/v1/chat/completions`, and a text-only message keeps a **plain-string** `content` (regression), in `tests/test_llamacpp_provider.py` (contract T3; witnesses both `features/us2_both_backends.feature` scenarios)
|
||||
- [x] T006-I [US2] Implement the images→content-parts mapping in `LlamaCppProvider.chat()` in `src/comfydv/_llm/llamacpp_provider.py`; leave text-only messages untouched — makes T006-T pass (depends on T002-I)
|
||||
|
||||
**Checkpoint**: parity proven — the identical node/wiring describes an image on both backends; swapping the client node is the only change.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 — Structured output about an image (Priority: P2)
|
||||
|
||||
**Goal**: Image input works with the node's existing structured-output mode, via the shared `chat_structured()` helper (pydantic-ai `BinaryContent`) — one implementation covering both backends through `OpenAIChatModel`.
|
||||
|
||||
**Independent Test**: Enable structured output with a schema, wire an image, run against a vision model, confirm each field is populated from the image with no required field blank; a first-invalid response retries then fails clearly.
|
||||
|
||||
- [x] T007-T [US3] Write FAILING test: `chat_structured()` attaches a message's images as `BinaryContent(data=b64decode(img), media_type="image/png")` onto the run's `user_prompt` (last turn) and onto history `UserPromptPart`s, a text-only structured call is unchanged, and the retry/validation contract is intact, in `tests/test_llm_chat_structured.py` (contract T4; witnesses both `features/us3_structured_image.feature` scenarios) — mock at the `Agent.run`/`_build_agent` seam per the established convention
|
||||
- [x] T007-I [US3] Implement image→`BinaryContent` handling in `chat_structured()` and `_history_to_messages()` in `src/comfydv/_llm/chat.py` — makes T007-T pass (depends on T002-I)
|
||||
|
||||
**Checkpoint**: structured image output works on both backends via the one shared helper; all prior stories remain green.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: Polish & Cross-Cutting Concerns
|
||||
|
||||
- [x] T008 [P] Document image input on `ChatCompletion` in `README.md` and add a `CHANGELOG.md` Unreleased entry — note the vision-model / llama.cpp `--mmproj` prerequisite (quickstart.md)
|
||||
- [x] T009 Run the full quality gate: `ruff check` ✓, `ruff format` ✓, `pytest` ✓ (289 passed, +21 new; all spec-009 code green), `beacon doctor --strict` ✓ for this spec (bullet + BDD + backlinks pass). _Pre-existing, out of scope: `ty check` has 36 diagnostics repo-wide (0 from spec-009 code — verified), one Docker packaging test (`test_dockerfile_uses_python_311_base`) fails at baseline, and `spec-task-alignment` flags 007's deferred tasks under --strict._
|
||||
- [-] T010 End-to-end `quickstart.md` validation against a live vision backend (Ollama multimodal model and `llama-server --mmproj`) _Deferred — requires a live vision-capable backend not available in CI/this environment; validate manually before release._
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
- **Setup (T001)** → no dependencies; start immediately.
|
||||
- **Foundational (T002-T/I)** → depends on nothing; **blocks all user stories** (every path reads `Message.images`).
|
||||
- **US1 (T003–T005)** → after T002-I. `T004-I` depends on `T003-I`; `T005-I` depends on `T002-I`. MVP.
|
||||
- **US2 (T006)** → after T002-I. Independent of US1's files; independently testable.
|
||||
- **US3 (T007)** → after T002-I. Independent of US1/US2's files; independently testable.
|
||||
- **Polish (T008–T010)** → after the stories you intend to ship.
|
||||
|
||||
### Within each story
|
||||
|
||||
- The `-T` task is written and committed **failing** before its `-I` partner (`tdd-commit-discipline`).
|
||||
- `-I` is never `[P]` with its own `-T`.
|
||||
|
||||
### Parallel opportunities
|
||||
|
||||
- US1: `T003-T` (`tests/test_ollama.py`) and `T005-T` (`tests/test_ollama_provider.py`) are different files → `[P]`.
|
||||
- Across stories: US1, US2, US3 touch different provider/helper files and can proceed in parallel once T002-I lands.
|
||||
|
||||
---
|
||||
|
||||
## Parallel Example: User Story 1
|
||||
|
||||
```bash
|
||||
# Different test files, no shared deps — write both failing tests together:
|
||||
Task: "T003-T encode-helper test in tests/test_ollama.py"
|
||||
Task: "T005-T Ollama image-passthrough test in tests/test_ollama_provider.py"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP first (US1 only)
|
||||
|
||||
1. T001 Setup → T002 carrier → T003–T005 US1.
|
||||
2. **STOP and VALIDATE**: an Ollama vision model describes a wired image; every text-only test stays green.
|
||||
3. Demoable as-is.
|
||||
|
||||
### Incremental delivery
|
||||
|
||||
1. Foundation + US1 → describe-an-image on Ollama (MVP).
|
||||
2. + US2 → same node works on llama.cpp (parity).
|
||||
3. + US3 → structured output about an image (both backends).
|
||||
4. Polish → docs, quality gate, manual live validation (T010).
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- `[-]` (T010) is a **known-deferred** follow-up — `beacon bullet finish` skips it rather than flipping to `[x]`; `beacon doctor` reports it as deferred, held under `--strict`.
|
||||
- `beacon doctor` runs two gates against this discipline: `spec-bdd-coverage` (every acceptance scenario has a `.feature` witness — 6 scenarios across 3 features here) and `tdd-commit-discipline` (no test + implementation in the same commit). Both FAIL under `--strict`.
|
||||
- Commit after each task or `-T`/`-I` pair; keep existing Ollama/llama.cpp/text tests green throughout (FR-003/SC-004 regression guard).
|
||||
@@ -2,24 +2,27 @@ import logging
|
||||
|
||||
from .circuit_breaker import CircuitBreaker
|
||||
from .format_string import FormatString
|
||||
from .llamacpp import LlamaCppClient
|
||||
from .ollama import (
|
||||
ChatCompletion,
|
||||
LLMLoadModel,
|
||||
LLMModelSelector,
|
||||
LLMUnloadModel,
|
||||
OllamaClient,
|
||||
OllamaChatCompletion,
|
||||
OllamaDebugHistory,
|
||||
OllamaHeaderBasicAuth,
|
||||
OllamaHeaderBearerToken,
|
||||
OllamaHeaderCustom,
|
||||
OllamaHistoryLength,
|
||||
OllamaLoadModel,
|
||||
OllamaModelSelector,
|
||||
OllamaOptionDisableThinking,
|
||||
OllamaOptionExtraBody,
|
||||
OllamaOptionMaxTokens,
|
||||
OllamaOptionRefusalRetry,
|
||||
OllamaOptionRepeatPenalty,
|
||||
OllamaOptionSeed,
|
||||
OllamaOptionTemperature,
|
||||
OllamaOptionTopK,
|
||||
OllamaOptionTopP,
|
||||
OllamaUnloadModel,
|
||||
)
|
||||
from .random_choice import RandomChoice
|
||||
|
||||
@@ -31,18 +34,22 @@ NODE_CLASS_MAPPINGS = {
|
||||
"RandomChoice": RandomChoice,
|
||||
"CircuitBreaker": CircuitBreaker,
|
||||
"FormatString": FormatString,
|
||||
# Ollama nodes
|
||||
# LLM nodes (generic, ADR-007) — see comfydv.ollama.MIGRATION_MAP for
|
||||
# the pre-cutover Ollama-specific names these replace
|
||||
"OllamaClient": OllamaClient,
|
||||
"OllamaModelSelector": OllamaModelSelector,
|
||||
"OllamaLoadModel": OllamaLoadModel,
|
||||
"OllamaUnloadModel": OllamaUnloadModel,
|
||||
"OllamaChatCompletion": OllamaChatCompletion,
|
||||
"LlamaCppClient": LlamaCppClient,
|
||||
"LLMModelSelector": LLMModelSelector,
|
||||
"LLMLoadModel": LLMLoadModel,
|
||||
"LLMUnloadModel": LLMUnloadModel,
|
||||
"ChatCompletion": ChatCompletion,
|
||||
"OllamaOptionTemperature": OllamaOptionTemperature,
|
||||
"OllamaOptionSeed": OllamaOptionSeed,
|
||||
"OllamaOptionMaxTokens": OllamaOptionMaxTokens,
|
||||
"OllamaOptionTopP": OllamaOptionTopP,
|
||||
"OllamaOptionTopK": OllamaOptionTopK,
|
||||
"OllamaOptionRepeatPenalty": OllamaOptionRepeatPenalty,
|
||||
"OllamaOptionDisableThinking": OllamaOptionDisableThinking,
|
||||
"OllamaOptionRefusalRetry": OllamaOptionRefusalRetry,
|
||||
"OllamaOptionExtraBody": OllamaOptionExtraBody,
|
||||
"OllamaDebugHistory": OllamaDebugHistory,
|
||||
"OllamaHistoryLength": OllamaHistoryLength,
|
||||
@@ -56,18 +63,21 @@ NODE_DISPLAY_NAME_MAPPINGS = {
|
||||
"RandomChoice": "Random Choice",
|
||||
"CircuitBreaker": "Circuit Breaker",
|
||||
"FormatString": "Format String (Python f-strings)",
|
||||
# Ollama nodes
|
||||
# LLM nodes (generic, ADR-007)
|
||||
"OllamaClient": "Ollama Client",
|
||||
"OllamaModelSelector": "Ollama Model Selector",
|
||||
"OllamaLoadModel": "Ollama Load Model",
|
||||
"OllamaUnloadModel": "Ollama Unload Model",
|
||||
"OllamaChatCompletion": "Ollama Chat Completion",
|
||||
"LlamaCppClient": "LlamaCpp Client",
|
||||
"LLMModelSelector": "LLM Model Selector",
|
||||
"LLMLoadModel": "LLM Load Model",
|
||||
"LLMUnloadModel": "LLM Unload Model",
|
||||
"ChatCompletion": "Chat Completion",
|
||||
"OllamaOptionTemperature": "Ollama Option — Temperature",
|
||||
"OllamaOptionSeed": "Ollama Option — Seed",
|
||||
"OllamaOptionMaxTokens": "Ollama Option — Max Tokens",
|
||||
"OllamaOptionTopP": "Ollama Option — Top P",
|
||||
"OllamaOptionTopK": "Ollama Option — Top K",
|
||||
"OllamaOptionRepeatPenalty": "Ollama Option — Repeat Penalty",
|
||||
"OllamaOptionDisableThinking": "Ollama Option — Disable Thinking",
|
||||
"OllamaOptionRefusalRetry": "Ollama Option — Refusal Retry",
|
||||
"OllamaOptionExtraBody": "Ollama Option — Extra Body",
|
||||
"OllamaDebugHistory": "Ollama Debug History",
|
||||
"OllamaHistoryLength": "Ollama History Length",
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
"""Internal package: shared LLM provider abstraction (ADR-007).
|
||||
|
||||
Not a ComfyUI node module — nothing here is registered in
|
||||
``NODE_CLASS_MAPPINGS``. ``comfydv.ollama`` (and, in a follow-on epic,
|
||||
``comfydv.llamacpp``) import from here.
|
||||
"""
|
||||
@@ -0,0 +1,335 @@
|
||||
"""Shared chat_structured() helper — pydantic-ai backed structured output.
|
||||
|
||||
Used by ``LlamaCppProvider.chat_structured()`` (ADR-007) over llama-server's
|
||||
OpenAI-compatible ``/v1/chat/completions``. ``OllamaProvider`` no longer uses
|
||||
this module (ADR-009): Ollama's OpenAI-compatible endpoint was found to
|
||||
silently reload the model at its default context size on every call,
|
||||
discarding any ``options.num_ctx`` override even when included in that same
|
||||
request — a behavior specific to Ollama's compat layer, not llama-server's.
|
||||
``OllamaProvider.chat_structured()`` now hand-rolls its own structured-output
|
||||
call over Ollama's *native* ``/api/chat`` + ``"format"``, which doesn't have
|
||||
that problem.
|
||||
|
||||
ADR-009: the Agent uses ``NativeOutput`` (``response_format``/JSON-schema
|
||||
constrained decoding), not pydantic-ai's default tool-calling. Live-tested
|
||||
against a "thinking"-capable model: tool-calling let the model spend its
|
||||
whole token budget on chain-of-thought reasoning and never emit the tool
|
||||
call; native output keeps reasoning in a separate response field and the
|
||||
constrained ``content`` always comes back as schema-valid JSON. This benefit
|
||||
still applies to llama.cpp, which is why this module (and its NativeOutput
|
||||
choice) is kept for that provider.
|
||||
|
||||
Ports ADR-006's retry/validation contract exactly: bounded retries (0-5,
|
||||
clamped), and a ``RuntimeError`` naming the model, attempt count, and a
|
||||
truncated snippet of the last invalid response on exhausted retries. The
|
||||
Agent's own internal retries are disabled (``retries=0``) — this helper
|
||||
drives its own retry loop so the error contract is comfydv's, not
|
||||
pydantic-ai's internal one.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
from collections.abc import Callable
|
||||
from typing import cast
|
||||
|
||||
from pydantic import BaseModel, ValidationError
|
||||
from pydantic_ai import Agent, NativeOutput
|
||||
from pydantic_ai.exceptions import ModelRetry, UnexpectedModelBehavior
|
||||
from pydantic_ai.messages import (
|
||||
BinaryContent,
|
||||
ModelRequest,
|
||||
ModelResponse,
|
||||
SystemPromptPart,
|
||||
TextPart,
|
||||
UserPromptPart,
|
||||
)
|
||||
from pydantic_ai.models.openai import OpenAIChatModel
|
||||
from pydantic_ai.providers.openai import OpenAIProvider
|
||||
from pydantic_ai.settings import ModelSettings
|
||||
|
||||
from .provider import Message
|
||||
from .retry import (
|
||||
RETRY_BACKOFF_SECS,
|
||||
EmbedFn,
|
||||
format_recovered_status,
|
||||
format_retry_status,
|
||||
is_refusal,
|
||||
next_seed,
|
||||
next_timeout_secs,
|
||||
record_attempt_info,
|
||||
)
|
||||
|
||||
_STRUCTURED_OUTPUT_FAILURE_EXCEPTIONS = (
|
||||
UnexpectedModelBehavior,
|
||||
ModelRetry,
|
||||
ValidationError,
|
||||
)
|
||||
|
||||
|
||||
def _build_agent(
|
||||
*,
|
||||
base_url: str,
|
||||
model: str,
|
||||
schema: type[BaseModel],
|
||||
headers: dict | None,
|
||||
timeout_secs: float,
|
||||
) -> Agent:
|
||||
import httpx
|
||||
|
||||
http_client = httpx.AsyncClient(
|
||||
headers=headers or None, timeout=httpx.Timeout(timeout_secs)
|
||||
)
|
||||
provider = OpenAIProvider(
|
||||
base_url=base_url, api_key="not-needed", http_client=http_client
|
||||
)
|
||||
chat_model = OpenAIChatModel(model, provider=provider)
|
||||
return Agent(chat_model, output_type=NativeOutput(schema), retries=0)
|
||||
|
||||
|
||||
def _user_prompt_content(msg: Message):
|
||||
"""Render a user turn as pydantic-ai user-prompt content.
|
||||
|
||||
Text-only ``msg`` → the plain ``content`` string, byte-identical to the
|
||||
pre-009 path (FR-003). A turn carrying images → ``[content, *images]``
|
||||
where each image is a ``BinaryContent`` PNG (ADR-008 / research.md
|
||||
Decision 1); ``OpenAIChatModel`` renders these as OpenAI ``image_url``
|
||||
parts, so both backends reach the same multimodal request through one
|
||||
shared code path.
|
||||
"""
|
||||
if not msg.images:
|
||||
return msg.content
|
||||
import base64
|
||||
|
||||
content: list = [msg.content]
|
||||
for image in msg.images:
|
||||
content.append(
|
||||
BinaryContent(data=base64.b64decode(image), media_type="image/png")
|
||||
)
|
||||
return content
|
||||
|
||||
|
||||
def _history_to_messages(messages: list[Message]) -> list:
|
||||
"""Convert all but the last message into pydantic-ai's typed history.
|
||||
|
||||
The last message (the current turn) is passed separately as
|
||||
``Agent.run()``'s ``user_prompt`` — see ``chat_structured()``.
|
||||
"""
|
||||
history: list = []
|
||||
for msg in messages[:-1]:
|
||||
if msg.role == "assistant":
|
||||
history.append(ModelResponse(parts=[TextPart(msg.content)]))
|
||||
elif msg.role == "system":
|
||||
history.append(ModelRequest(parts=[SystemPromptPart(msg.content)]))
|
||||
else:
|
||||
history.append(
|
||||
ModelRequest(parts=[UserPromptPart(_user_prompt_content(msg))])
|
||||
)
|
||||
return history
|
||||
|
||||
|
||||
async def chat_structured(
|
||||
*,
|
||||
base_url: str,
|
||||
model: str,
|
||||
messages: list[Message],
|
||||
schema: type[BaseModel],
|
||||
headers: dict | None = None,
|
||||
options: dict | None = None,
|
||||
max_retries: int = 2,
|
||||
timeout_secs: float = 300.0,
|
||||
embed_fn: EmbedFn | None = None,
|
||||
attempt_info: dict | None = None,
|
||||
on_status: Callable[[str], None] | None = None,
|
||||
) -> BaseModel:
|
||||
"""Call ``model`` at ``base_url`` (an OpenAI-compatible ``/v1`` root) and
|
||||
return a validated instance of ``schema``.
|
||||
|
||||
``options`` is forwarded verbatim as a top-level ``"options"`` field in
|
||||
the request body via pydantic-ai's ``extra_body`` — the same shape the
|
||||
pre-ADR-007 hand-rolled implementation sent, so provider-native sampling
|
||||
params (Ollama's ``num_predict``/``repeat_penalty``/etc., set via the
|
||||
``OllamaOption*`` nodes) keep working unchanged rather than being
|
||||
lossily remapped onto pydantic-ai's own standardized ``ModelSettings``
|
||||
fields.
|
||||
|
||||
ADR-010: ``options`` may also carry a ``"think"`` key (bool), popped out
|
||||
here rather than forwarded inside the nested ``options`` object —
|
||||
llama-server's OpenAI-compatible endpoint doesn't recognize a literal
|
||||
``"think"`` key there. Translated to its own two documented
|
||||
request-body toggles instead: ``chat_template_kwargs:
|
||||
{"enable_thinking": ...}`` (Qwen3-style models) and, when disabling,
|
||||
``reasoning_effort: "none"`` (the more model-agnostic OpenAI convention
|
||||
llama-server also honors) — both via ``extra_body`` the same way
|
||||
``options`` is. This provider only serves ``LlamaCppProvider`` — see
|
||||
``OllamaProvider``'s own hand-rolled ``chat_structured`` for why Ollama
|
||||
needed a different mechanism entirely. Sourced from llama.cpp's server
|
||||
docs, not live-verified against a running llama-server (no instance
|
||||
available at implementation time) — verify against your own deployment.
|
||||
|
||||
Retries up to ``max_retries`` times (clamped 0-5) on validation failure
|
||||
before raising ``RuntimeError``. Never returns a value that failed
|
||||
validation against ``schema``.
|
||||
|
||||
``options`` may also carry a ``"refusal_retry"`` config dict (same
|
||||
comfydv-level convention as ``"think"``, emitted by
|
||||
``OllamaOptionRefusalRetry``) — a detected refusal/deflection (see
|
||||
``_llm/retry.py``) is treated exactly like a validation failure: retried
|
||||
with a bumped seed rather than returned to the caller. ``embed_fn`` is
|
||||
``LlamaCppProvider``'s own ``embed()``, bound to whatever embedding
|
||||
model the config names — passed in rather than looked up here since
|
||||
this module has no provider instance of its own to call.
|
||||
"""
|
||||
if not messages or messages[-1].role != "user":
|
||||
raise ValueError(
|
||||
"chat_structured requires the last message to have role='user'"
|
||||
)
|
||||
|
||||
history = _history_to_messages(messages)
|
||||
prompt = _user_prompt_content(messages[-1])
|
||||
think = None
|
||||
if options and "think" in options:
|
||||
options = dict(options)
|
||||
think = options.pop("think")
|
||||
options = options or None
|
||||
refusal_cfg = None
|
||||
if options and "refusal_retry" in options:
|
||||
options = dict(options)
|
||||
refusal_cfg = options.pop("refusal_retry")
|
||||
options = options or None
|
||||
extra_body: dict = {}
|
||||
if options:
|
||||
extra_body["options"] = options
|
||||
if think is not None:
|
||||
extra_body["chat_template_kwargs"] = {"enable_thinking": think}
|
||||
if not think:
|
||||
extra_body["reasoning_effort"] = "none"
|
||||
model_settings: ModelSettings | None = (
|
||||
{"extra_body": extra_body} if extra_body else None
|
||||
)
|
||||
|
||||
total_attempts = max(0, min(int(max_retries), 5)) + 1
|
||||
last_error: Exception | None = None
|
||||
last_invalid_text = ""
|
||||
refusal_count = 0
|
||||
attempt_seed = (options or {}).get("seed", 0) if isinstance(options, dict) else 0
|
||||
attempt_timeout = timeout_secs
|
||||
|
||||
def _emit_retry_status(reason: str, attempt: int) -> None:
|
||||
if on_status is None or attempt >= total_attempts:
|
||||
return
|
||||
upcoming_seed = next_seed(options, attempt + 1)
|
||||
upcoming_timeout = next_timeout_secs(timeout_secs, attempt + 1)
|
||||
on_status(
|
||||
format_retry_status(
|
||||
reason, attempt, total_attempts, upcoming_seed, upcoming_timeout
|
||||
)
|
||||
)
|
||||
|
||||
for attempt in range(1, total_attempts + 1):
|
||||
attempt_timeout = next_timeout_secs(timeout_secs, attempt)
|
||||
# Rebuilt each attempt so the escalated timeout actually takes
|
||||
# effect — httpx.AsyncClient's timeout is fixed at construction,
|
||||
# not mutable per-request.
|
||||
agent = _build_agent(
|
||||
base_url=base_url,
|
||||
model=model,
|
||||
schema=schema,
|
||||
headers=headers,
|
||||
timeout_secs=attempt_timeout,
|
||||
)
|
||||
attempt_settings = dict(model_settings) if model_settings else {}
|
||||
if attempt > 1:
|
||||
# Confirmed live: a freshly-loaded model's first structured-output
|
||||
# attempt can fail outright (no valid tool call at all) and then
|
||||
# behave normally on the very next call. Retrying with the exact
|
||||
# same request reproduces the same failure if the model is
|
||||
# genuinely stuck rather than just unlucky, so force a new seed
|
||||
# (pydantic-ai maps ModelSettings["seed"] to the OpenAI API's
|
||||
# top-level "seed" param, which works against both Ollama's and
|
||||
# llama-server's OpenAI-compatible endpoints) and give it a beat
|
||||
# via RETRY_BACKOFF_SECS in case it's still finishing loading.
|
||||
seed = next_seed(options, attempt)
|
||||
attempt_seed = seed
|
||||
attempt_settings["seed"] = seed
|
||||
if "extra_body" in attempt_settings:
|
||||
# beacon-reviewer caught this: if a caller pinned options["seed"],
|
||||
# it's also sitting in extra_body.options.seed (the Ollama-native
|
||||
# passthrough). Left untouched, a backend that honors that nested
|
||||
# field over the top-level OpenAI "seed" above would keep sending
|
||||
# the same old seed on every retry — silently defeating this fix
|
||||
# for exactly the pinned-seed case. Copy rather than mutate in
|
||||
# place: extra_body/options here are the caller's own dicts,
|
||||
# shared across every attempt (and possibly other calls).
|
||||
# ModelSettings declares extra_body as `object` (it's an
|
||||
# opaque passthrough field), so a plain dict() call on it
|
||||
# doesn't type-check — cast first, this module always builds
|
||||
# it as a dict (see model_settings above).
|
||||
extra_body = dict(cast(dict, attempt_settings["extra_body"]))
|
||||
nested_options = dict(extra_body.get("options") or {})
|
||||
nested_options["seed"] = seed
|
||||
extra_body["options"] = nested_options
|
||||
attempt_settings["extra_body"] = extra_body
|
||||
try:
|
||||
result = await agent.run(
|
||||
prompt,
|
||||
message_history=history,
|
||||
model_settings=cast(ModelSettings, attempt_settings)
|
||||
if attempt_settings
|
||||
else None,
|
||||
)
|
||||
# agent's output_type is the caller's `schema` (a runtime value,
|
||||
# not a static type parameter), so the checker can't narrow
|
||||
# result.output past Agent's default `str` — cast to the
|
||||
# function's declared return type, which schema is a subtype of.
|
||||
output = cast(BaseModel, result.output)
|
||||
if refusal_cfg and refusal_cfg.get("enabled"):
|
||||
# Re-serialized, not the original wire text — pydantic-ai's
|
||||
# NativeOutput doesn't expose that separately, and the
|
||||
# regex/embedding check works the same either way (same
|
||||
# textual content, just re-encoded).
|
||||
content = output.model_dump_json()
|
||||
refused = await is_refusal(
|
||||
content,
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key=refusal_cfg.get("embedding_model", ""),
|
||||
threshold=refusal_cfg.get("threshold", 0.82),
|
||||
custom_phrases=tuple(refusal_cfg.get("custom_phrases") or ()),
|
||||
)
|
||||
if refused:
|
||||
refusal_count += 1
|
||||
last_error = RuntimeError("refusal/deflection detected")
|
||||
last_invalid_text = content
|
||||
_emit_retry_status("Refusal/deflection detected", attempt)
|
||||
if attempt < total_attempts:
|
||||
await asyncio.sleep(RETRY_BACKOFF_SECS)
|
||||
continue
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=attempt,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
if on_status is not None and attempt > 1:
|
||||
on_status(
|
||||
format_recovered_status(attempt, total_attempts, attempt_seed)
|
||||
)
|
||||
return output
|
||||
except _STRUCTURED_OUTPUT_FAILURE_EXCEPTIONS as exc:
|
||||
last_error = exc
|
||||
last_invalid_text = str(exc)
|
||||
_emit_retry_status("Structured output failed", attempt)
|
||||
if attempt < total_attempts:
|
||||
await asyncio.sleep(RETRY_BACKOFF_SECS)
|
||||
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=total_attempts,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
raise RuntimeError(
|
||||
f"chat_structured: response failed validation against schema after "
|
||||
f"{total_attempts} attempt(s) (model={model!r}). Last error: "
|
||||
f"{last_error}. Last response (truncated): {last_invalid_text[:300]!r}"
|
||||
)
|
||||
@@ -0,0 +1,440 @@
|
||||
"""LlamaCppProvider — LLMProvider implementation backed by llama-server's
|
||||
router mode.
|
||||
|
||||
Mirrors comfydv._llm.ollama_provider's structure exactly (ADR-007's parallel-
|
||||
implementation pattern). Router-mode API shape verified live against
|
||||
ggml-org/llama.cpp's tools/server/README.md (postdates training data) — see
|
||||
specs/008-llamacpp-integration/research.md. Two details differ from Ollama:
|
||||
the model identifier field is "id" (not "name"), and "status" is a nested
|
||||
object ({"value": "..."}), not a flat string.
|
||||
|
||||
Deployment prerequisite: llama-server must be launched with --models-dir or
|
||||
--models-preset (router mode) — the endpoints this provider calls don't
|
||||
exist otherwise (spec.md FR-006).
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
from collections.abc import Callable
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
from .ollama_provider import (
|
||||
_TTLLRUCache,
|
||||
_cache_key,
|
||||
_get_json,
|
||||
_pop_refusal_retry,
|
||||
_pop_think,
|
||||
_post_json,
|
||||
)
|
||||
from .provider import Message, ModelInfo, ModelStatus
|
||||
from .retry import (
|
||||
RETRY_BACKOFF_SECS,
|
||||
format_recovered_status,
|
||||
format_retry_status,
|
||||
is_refusal,
|
||||
next_seed,
|
||||
next_timeout_secs,
|
||||
record_attempt_info,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Own cache pool, not shared with OllamaProvider's — see plan.md's Structure
|
||||
# Decision (parallel, symmetric, independent implementations). ChatCompletion
|
||||
# is OUTPUT_NODE=True and re-executes every queue run regardless of which
|
||||
# provider is wired in, so caching parity matters for llama.cpp too, not
|
||||
# just Ollama.
|
||||
_MODEL_LIST_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=20.0)
|
||||
_CHAT_RESPONSE_CACHE = _TTLLRUCache(maxsize=64, ttl_seconds=None)
|
||||
|
||||
|
||||
async def _fetch_models(host: str, headers: dict | None = None) -> list[str]:
|
||||
"""Name-only view for ComfyUI's combo-widget population (the JS refresh
|
||||
button and node-creation auto-populate) — mirrors
|
||||
ollama_provider._fetch_models's narrower, gracefully-degrading contract.
|
||||
|
||||
Deliberately more forgiving than LlamaCppProvider.list_models(): that
|
||||
method raises on a non-router-mode server (FR-006, for real workflow
|
||||
execution, where a silent empty result would be misleading). This
|
||||
combo-population use case wants the same quiet "just show an empty
|
||||
dropdown" degradation Ollama's nodes already give for *any* failure —
|
||||
consistent UX across backends for this specific, lower-stakes path.
|
||||
"""
|
||||
try:
|
||||
models = await LlamaCppProvider(host, headers).list_models()
|
||||
except Exception as exc:
|
||||
logger.warning("Could not fetch llama.cpp models from %s: %s", host, exc)
|
||||
return []
|
||||
return [m.name for m in models]
|
||||
|
||||
|
||||
def _to_openai_message(message: Message) -> dict:
|
||||
"""Render a ``Message`` in llama.cpp's OpenAI-compatible shape.
|
||||
|
||||
A text-only turn stays ``{"role", "content": <str>}`` — byte-identical to
|
||||
the pre-009 payload (FR-003). A turn carrying images becomes OpenAI
|
||||
multimodal ``content`` parts: the text followed by one ``image_url`` part
|
||||
per base64 image, as a ``data:`` URI (ADR-008). ``llama-server`` only
|
||||
honours these parts when launched with a multimodal projector
|
||||
(``--mmproj``); without it the server errors, surfaced to the caller
|
||||
rather than crashed on (FR-006).
|
||||
"""
|
||||
if not message.images:
|
||||
return {"role": message.role, "content": message.content}
|
||||
parts: list[dict] = [{"type": "text", "text": message.content}]
|
||||
for image in message.images:
|
||||
parts.append(
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {"url": f"data:image/png;base64,{image}"},
|
||||
}
|
||||
)
|
||||
return {"role": message.role, "content": parts}
|
||||
|
||||
|
||||
class LlamaCppProvider:
|
||||
"""LLMProvider implementation backed by llama-server's router mode.
|
||||
|
||||
Host and headers are captured once at construction — every method
|
||||
reuses them, the same pattern OllamaProvider already established.
|
||||
"""
|
||||
|
||||
def __init__(self, host: str, headers: dict | None = None):
|
||||
self.host = host
|
||||
self.headers = dict(headers) if headers else None
|
||||
|
||||
async def list_models(self) -> list[ModelInfo]:
|
||||
"""GET {host}/models — every model llama-server's router knows
|
||||
about, with its live status. Unlike OllamaProvider, no
|
||||
normalization is needed: llama.cpp's status vocabulary is exactly
|
||||
ModelStatus's full set.
|
||||
"""
|
||||
cache_key = _cache_key("llamacpp_list_models", self.host, self.headers or {})
|
||||
cached, hit = _MODEL_LIST_CACHE.get(cache_key)
|
||||
if hit:
|
||||
return cached
|
||||
|
||||
try:
|
||||
data = await _get_json(f"{self.host}/models", headers=self.headers)
|
||||
except OSError as exc:
|
||||
# Genuinely unreachable (connection refused, DNS failure, timed
|
||||
# out — aiohttp's connection-level exceptions are all OSError
|
||||
# subclasses) — degrade gracefully like OllamaProvider does, so
|
||||
# a not-yet-started server just shows an empty dropdown rather
|
||||
# than a hard error.
|
||||
logger.warning(
|
||||
"Could not fetch llama.cpp models from %s: %s", self.host, exc
|
||||
)
|
||||
return []
|
||||
except RuntimeError as exc:
|
||||
# The server answered but with an HTTP error status — GET
|
||||
# /models only exists in router mode, so this is almost always
|
||||
# a llama-server launched without --models-dir/--models-preset.
|
||||
# Surfacing this distinctly (FR-006) matters: silently returning
|
||||
# [] here would be indistinguishable from "no models installed".
|
||||
raise RuntimeError(
|
||||
f"llama-server at {self.host} did not return a model list from "
|
||||
f"GET {self.host}/models — is it running in router mode "
|
||||
f"(--models-dir or --models-preset)? Underlying error: {exc}"
|
||||
) from exc
|
||||
|
||||
models = []
|
||||
for m in data.get("data", []):
|
||||
status_value = m.get("status", {}).get("value")
|
||||
try:
|
||||
status = ModelStatus(status_value)
|
||||
except ValueError:
|
||||
logger.warning(
|
||||
"llama.cpp reported an unrecognized model status %r for %r — "
|
||||
"skipping status normalization, this model will be omitted",
|
||||
status_value,
|
||||
m.get("id"),
|
||||
)
|
||||
continue
|
||||
models.append(ModelInfo(name=m["id"], status=status, size=None))
|
||||
if models:
|
||||
_MODEL_LIST_CACHE.set(cache_key, models)
|
||||
return models
|
||||
|
||||
async def load_model(self, model: str) -> None:
|
||||
if not model.strip():
|
||||
raise ValueError("model name cannot be empty")
|
||||
try:
|
||||
await _post_json(
|
||||
f"{self.host}/models/load",
|
||||
{"model": model},
|
||||
headers=self.headers,
|
||||
)
|
||||
except RuntimeError as exc:
|
||||
# Confirmed live: router mode's own /models/load is NOT
|
||||
# idempotent — it 400s "model is already running" rather than
|
||||
# the {"success": true} the contract assumed. The LLMProvider
|
||||
# protocol requires load_model() to be idempotent, so this
|
||||
# error is the desired end-state, not a failure — absorb it
|
||||
# here rather than leaking the wire-level quirk to callers.
|
||||
if "model is already running" not in str(exc):
|
||||
raise
|
||||
|
||||
async def unload_model(self, model: str) -> None:
|
||||
if not model.strip():
|
||||
raise ValueError("model name cannot be empty")
|
||||
try:
|
||||
await _post_json(
|
||||
f"{self.host}/models/unload",
|
||||
{"model": model},
|
||||
headers=self.headers,
|
||||
)
|
||||
except RuntimeError as exc:
|
||||
# Mirror of load_model()'s non-idempotency above, confirmed live:
|
||||
# /models/unload 400s "model is not running" on an already-
|
||||
# unloaded model instead of {"success": true}.
|
||||
if "model is not running" not in str(exc):
|
||||
raise
|
||||
|
||||
async def chat(
|
||||
self,
|
||||
model: str,
|
||||
messages: list[Message],
|
||||
options: dict | None = None,
|
||||
timeout_secs: float = 300.0,
|
||||
max_retries: int = 2,
|
||||
attempt_info: dict | None = None,
|
||||
on_status: Callable[[str], None] | None = None,
|
||||
) -> str:
|
||||
payload_messages = [_to_openai_message(m) for m in messages]
|
||||
options, think = _pop_think(options)
|
||||
options, refusal_cfg = _pop_refusal_retry(options)
|
||||
embed_fn = None
|
||||
custom_phrases: tuple[str, ...] = ()
|
||||
if refusal_cfg and refusal_cfg.get("enabled"):
|
||||
custom_phrases = tuple(refusal_cfg.get("custom_phrases") or ())
|
||||
if refusal_cfg.get("embedding_model"):
|
||||
embedding_model = refusal_cfg["embedding_model"]
|
||||
embed_fn = lambda t: self.embed(embedding_model, t) # noqa: E731
|
||||
total_attempts = max(0, min(int(max_retries), 5)) + 1
|
||||
response_text = ""
|
||||
refusal_count = 0
|
||||
attempt_seed = 0
|
||||
attempt_timeout = timeout_secs
|
||||
|
||||
for attempt in range(1, total_attempts + 1):
|
||||
attempt_timeout = next_timeout_secs(timeout_secs, attempt)
|
||||
payload: dict = {
|
||||
"model": model,
|
||||
"messages": payload_messages,
|
||||
"stream": False,
|
||||
}
|
||||
if options:
|
||||
# Passed through verbatim, same nesting OllamaProvider.chat()
|
||||
# uses (payload["options"] = options) — the OllamaOption*
|
||||
# nodes emit Ollama-native parameter names (num_predict,
|
||||
# repeat_penalty, ...), which llama-server's OpenAI-compatible
|
||||
# endpoint won't recognize either way; translating them is
|
||||
# out of scope for this epic (plan.md Non-goals — no changes
|
||||
# to the generic nodes). This keeps the two providers'
|
||||
# handling consistent rather than silently special-casing
|
||||
# one of them.
|
||||
payload["options"] = options
|
||||
if think is not None:
|
||||
# ADR-010: llama-server's two documented reasoning toggles —
|
||||
# sourced from server docs, not live-verified (no instance
|
||||
# available at implementation time).
|
||||
payload["chat_template_kwargs"] = {"enable_thinking": think}
|
||||
if not think:
|
||||
payload["reasoning_effort"] = "none"
|
||||
if attempt > 1:
|
||||
# Unlike the options-passthrough above, this IS the OpenAI
|
||||
# spec's actual top-level "seed" field, so it takes effect
|
||||
# against llama-server's /v1/chat/completions.
|
||||
payload["seed"] = next_seed(options, attempt)
|
||||
attempt_seed = payload["seed"]
|
||||
else:
|
||||
# attempt 1 never sets the top-level "seed" field above (only
|
||||
# retries do) — fall back to whatever the caller pinned in
|
||||
# options, so attempt_info/seed_used reports the real seed in
|
||||
# play even on a first-attempt success, not a stale 0.
|
||||
attempt_seed = (options or {}).get("seed", 0)
|
||||
|
||||
cache_key = _cache_key(
|
||||
"llamacpp_chat",
|
||||
self.host,
|
||||
self.headers or {},
|
||||
model,
|
||||
payload_messages,
|
||||
options or {},
|
||||
think,
|
||||
payload.get("seed"),
|
||||
)
|
||||
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
|
||||
if hit:
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=attempt,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
return cached
|
||||
|
||||
result = await _post_json(
|
||||
f"{self.host}/v1/chat/completions",
|
||||
payload,
|
||||
timeout=attempt_timeout,
|
||||
headers=self.headers,
|
||||
)
|
||||
choices = result.get("choices") or []
|
||||
response_text = (
|
||||
choices[0].get("message", {}).get("content", "") or ""
|
||||
if choices
|
||||
else ""
|
||||
)
|
||||
retry_reason: str | None = None
|
||||
if response_text.strip():
|
||||
refused = False
|
||||
if refusal_cfg and refusal_cfg.get("enabled"):
|
||||
refused = await is_refusal(
|
||||
response_text,
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key=refusal_cfg.get("embedding_model", ""),
|
||||
threshold=refusal_cfg.get("threshold", 0.82),
|
||||
custom_phrases=custom_phrases,
|
||||
)
|
||||
if not refused:
|
||||
_CHAT_RESPONSE_CACHE.set(cache_key, response_text)
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=attempt,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
if on_status is not None and attempt > 1:
|
||||
on_status(
|
||||
format_recovered_status(
|
||||
attempt, total_attempts, attempt_seed
|
||||
)
|
||||
)
|
||||
return response_text
|
||||
refusal_count += 1
|
||||
retry_reason = "Refusal/deflection detected"
|
||||
else:
|
||||
retry_reason = "Blank response"
|
||||
|
||||
if attempt < total_attempts:
|
||||
if on_status is not None:
|
||||
upcoming_seed = next_seed(options, attempt + 1)
|
||||
upcoming_timeout = next_timeout_secs(timeout_secs, attempt + 1)
|
||||
on_status(
|
||||
format_retry_status(
|
||||
retry_reason,
|
||||
attempt,
|
||||
total_attempts,
|
||||
upcoming_seed,
|
||||
upcoming_timeout,
|
||||
)
|
||||
)
|
||||
await asyncio.sleep(RETRY_BACKOFF_SECS)
|
||||
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=total_attempts,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
# Every attempt came back blank — never raises here (chat() has
|
||||
# never validated its output, unlike chat_structured()); return the
|
||||
# last (blank) attempt uncached so the next queue run tries fresh.
|
||||
return response_text
|
||||
|
||||
async def chat_structured(
|
||||
self,
|
||||
model: str,
|
||||
messages: list[Message],
|
||||
schema: type[BaseModel],
|
||||
options: dict | None = None,
|
||||
timeout_secs: float = 300.0,
|
||||
max_retries: int = 2,
|
||||
attempt_info: dict | None = None,
|
||||
on_status: Callable[[str], None] | None = None,
|
||||
) -> BaseModel:
|
||||
from .chat import chat_structured as _chat_structured_impl
|
||||
|
||||
payload_messages = [m.model_dump() for m in messages]
|
||||
cache_key = _cache_key(
|
||||
"llamacpp_chat_structured",
|
||||
self.host,
|
||||
self.headers or {},
|
||||
model,
|
||||
payload_messages,
|
||||
options or {},
|
||||
schema.model_json_schema(),
|
||||
)
|
||||
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
|
||||
if hit:
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=(options or {}).get("seed", 0),
|
||||
attempts=1,
|
||||
timeout_secs=timeout_secs,
|
||||
refusals=0,
|
||||
)
|
||||
return schema.model_validate(cached)
|
||||
|
||||
embed_fn = None
|
||||
refusal_cfg = (options or {}).get("refusal_retry")
|
||||
if (
|
||||
refusal_cfg
|
||||
and refusal_cfg.get("enabled")
|
||||
and refusal_cfg.get("embedding_model")
|
||||
):
|
||||
embedding_model = refusal_cfg["embedding_model"]
|
||||
embed_fn = lambda t: self.embed(embedding_model, t) # noqa: E731
|
||||
|
||||
result = await _chat_structured_impl(
|
||||
base_url=f"{self.host}/v1",
|
||||
model=model,
|
||||
messages=messages,
|
||||
schema=schema,
|
||||
headers=self.headers,
|
||||
options=options,
|
||||
max_retries=max_retries,
|
||||
timeout_secs=timeout_secs,
|
||||
embed_fn=embed_fn,
|
||||
attempt_info=attempt_info,
|
||||
on_status=on_status,
|
||||
)
|
||||
_CHAT_RESPONSE_CACHE.set(cache_key, result.model_dump())
|
||||
return result
|
||||
|
||||
async def embed(self, model: str, text: str) -> list[float] | None:
|
||||
"""POST {host}/v1/embeddings — llama-server's OpenAI-compatible
|
||||
embeddings endpoint.
|
||||
|
||||
Requires an embedding-capable model to be loaded in the router
|
||||
(typically a *different* model from whatever's answering chat
|
||||
requests) — not live-verified against a running llama-server (no
|
||||
instance available at implementation time), mirroring this
|
||||
provider's other sourced-from-docs-not-verified caveats. Returns
|
||||
``None`` rather than raising on any failure, same contract as
|
||||
``OllamaProvider.embed()``.
|
||||
"""
|
||||
if not model.strip() or not text.strip():
|
||||
return None
|
||||
try:
|
||||
result = await _post_json(
|
||||
f"{self.host}/v1/embeddings",
|
||||
{"model": model, "input": text},
|
||||
timeout=30.0,
|
||||
headers=self.headers,
|
||||
)
|
||||
except Exception:
|
||||
return None
|
||||
data = result.get("data")
|
||||
if not isinstance(data, list) or not data:
|
||||
return None
|
||||
vec = data[0].get("embedding")
|
||||
if not isinstance(vec, list) or not vec:
|
||||
return None
|
||||
return vec
|
||||
@@ -0,0 +1,731 @@
|
||||
"""OllamaProvider — LLMProvider implementation backed by Ollama's REST API.
|
||||
|
||||
Ported from comfydv.ollama's original module-level HTTP/cache helpers
|
||||
(_post_json, _fetch_models, _run_async, _TTLLRUCache) — behavior-preserving,
|
||||
not a rewrite. See ADR-007 and
|
||||
specs/007-llm-provider-abstraction/research.md.
|
||||
|
||||
OllamaProvider implements the LLMProvider Protocol structurally (no explicit
|
||||
inheritance — that's the point of typing.Protocol); conformance is checked
|
||||
by ``ty check``, not the runtime.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import logging
|
||||
import threading
|
||||
import time
|
||||
from collections.abc import Callable
|
||||
|
||||
from pydantic import BaseModel, ValidationError
|
||||
|
||||
from .provider import Message, ModelInfo, ModelStatus
|
||||
from .retry import (
|
||||
RETRY_BACKOFF_SECS,
|
||||
format_recovered_status,
|
||||
format_retry_status,
|
||||
is_refusal,
|
||||
next_seed,
|
||||
next_timeout_secs,
|
||||
record_attempt_info,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Local response cache (ported from comfydv.ollama)
|
||||
# ---------------------------------------------------------------------------
|
||||
#
|
||||
# See comfydv.ollama's original module docstring for why this exists:
|
||||
# OUTPUT_NODE=True chat nodes re-execute every queue run even when inputs
|
||||
# are unchanged; this cache absorbs the redundant round-trips.
|
||||
|
||||
|
||||
class _TTLLRUCache:
|
||||
"""Bounded cache, LRU-evicted, with an optional per-entry TTL."""
|
||||
|
||||
def __init__(self, maxsize: int, ttl_seconds: float | None = None):
|
||||
self.maxsize = maxsize
|
||||
self.ttl_seconds = ttl_seconds
|
||||
self._data: dict = {}
|
||||
self._lock = threading.Lock()
|
||||
|
||||
def get(self, key):
|
||||
with self._lock:
|
||||
entry = self._data.get(key)
|
||||
if entry is None:
|
||||
return None, False
|
||||
expires_at, value = entry
|
||||
if self.ttl_seconds is not None and time.monotonic() > expires_at:
|
||||
del self._data[key]
|
||||
return None, False
|
||||
# Re-insert to mark as most-recently-used (dicts preserve insertion order).
|
||||
del self._data[key]
|
||||
self._data[key] = (expires_at, value)
|
||||
return value, True
|
||||
|
||||
def set(self, key, value):
|
||||
with self._lock:
|
||||
expires_at = (
|
||||
time.monotonic() + self.ttl_seconds
|
||||
if self.ttl_seconds is not None
|
||||
else float("inf")
|
||||
)
|
||||
self._data.pop(key, None)
|
||||
self._data[key] = (expires_at, value)
|
||||
while len(self._data) > self.maxsize:
|
||||
oldest_key = next(iter(self._data))
|
||||
del self._data[oldest_key]
|
||||
|
||||
def clear(self):
|
||||
with self._lock:
|
||||
self._data.clear()
|
||||
|
||||
|
||||
def _cache_key(*parts) -> str:
|
||||
"""Deterministic, hashable key from arbitrary JSON-serializable parts."""
|
||||
return json.dumps(parts, sort_keys=True, default=str)
|
||||
|
||||
|
||||
def _pop_think(options: dict | None) -> tuple[dict | None, bool | None]:
|
||||
"""Split a ``"think"`` toggle out of a generic ``options`` dict.
|
||||
|
||||
ADR-010: ``OllamaOptionDisableThinking`` merges a ``"think": bool`` key
|
||||
into the same composable ``OLLAMA_OPTIONS`` chain every other
|
||||
``OllamaOption*`` node feeds into ``ChatCompletion``'s ``options``
|
||||
input — but unlike those (Ollama-native sampling params, passed through
|
||||
verbatim), ``"think"`` needs real per-provider translation: neither
|
||||
Ollama's native ``/api/chat`` nor llama-server's OpenAI-compatible
|
||||
endpoint recognizes a literal ``"think"`` key nested inside their own
|
||||
``options``/sampling-params object, so every provider pops it out here
|
||||
(or in ``LlamaCppProvider``'s own copy) before building its request.
|
||||
Returns ``options`` with ``"think"`` removed (unchanged if absent, so a
|
||||
falsy/empty result stays falsy) and the popped value, or ``None`` if the
|
||||
caller didn't set it — never touches the caller's own dict in place.
|
||||
"""
|
||||
if not options or "think" not in options:
|
||||
return options, None
|
||||
remaining = dict(options)
|
||||
think = remaining.pop("think")
|
||||
return (remaining or None), think
|
||||
|
||||
|
||||
def _pop_refusal_retry(options: dict | None) -> tuple[dict | None, dict | None]:
|
||||
"""Split a ``"refusal_retry"`` config dict out of a generic ``options``
|
||||
dict — same convention as ``_pop_think``: ``OllamaOptionRefusalRetry``
|
||||
merges ``{"refusal_retry": {"enabled", "embedding_model", "threshold"}}``
|
||||
into the same composable ``OLLAMA_OPTIONS`` chain every other
|
||||
``OllamaOption*`` node feeds into ``ChatCompletion``'s ``options``
|
||||
input, and neither Ollama's nor llama.cpp's own API recognizes this key,
|
||||
so every provider pops it out here before building its request.
|
||||
"""
|
||||
if not options or "refusal_retry" not in options:
|
||||
return options, None
|
||||
remaining = dict(options)
|
||||
cfg = remaining.pop("refusal_retry")
|
||||
return (remaining or None), cfg
|
||||
|
||||
|
||||
_MODEL_LIST_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=20.0)
|
||||
_CHAT_RESPONSE_CACHE = _TTLLRUCache(maxsize=64, ttl_seconds=None)
|
||||
_CAPABILITY_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=300.0)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Async infrastructure (ported from comfydv.ollama)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _run_async(coro):
|
||||
"""Run an async coroutine synchronously in an isolated worker thread.
|
||||
|
||||
Always spins up a fresh thread rather than conditionally checking
|
||||
asyncio.get_running_loop() first: live-verified against a real running
|
||||
ComfyUI instance (its execution engine runs its own event loop, in
|
||||
Python 3.13, on the same process) that the conditional version — try
|
||||
get_running_loop(), spin up a thread only if it succeeds, otherwise
|
||||
call asyncio.run(coro) directly — is unreliable there. Under real
|
||||
ComfyUI, get_running_loop() sometimes raised inside that try block
|
||||
(unlike under pytest or a standalone script, where it never does),
|
||||
which routed straight into `asyncio.run(coro)` on the *current* thread
|
||||
— the one thread guaranteed to already have ComfyUI's own loop running
|
||||
— reproducing exactly the "asyncio.run() cannot be called from a
|
||||
running event loop" crash this function exists to prevent. Always
|
||||
using a dedicated thread sidesteps the detection entirely: a freshly
|
||||
spawned thread never has an ambient loop, so asyncio.run() is safe
|
||||
there unconditionally, regardless of what the calling thread's loop
|
||||
state actually is.
|
||||
"""
|
||||
import concurrent.futures
|
||||
|
||||
with concurrent.futures.ThreadPoolExecutor(max_workers=1) as pool:
|
||||
return pool.submit(asyncio.run, coro).result()
|
||||
|
||||
|
||||
async def _post_json(
|
||||
url: str,
|
||||
payload: dict,
|
||||
*,
|
||||
timeout: float = 120.0,
|
||||
headers: dict | None = None,
|
||||
) -> dict:
|
||||
"""POST JSON to url, return parsed response dict."""
|
||||
import aiohttp
|
||||
|
||||
try:
|
||||
async with aiohttp.ClientSession() as session:
|
||||
async with session.post(
|
||||
url,
|
||||
json=payload,
|
||||
headers=headers or None,
|
||||
timeout=aiohttp.ClientTimeout(total=timeout),
|
||||
) as resp:
|
||||
if resp.status >= 400:
|
||||
body = await resp.text()
|
||||
raise RuntimeError(
|
||||
f"Ollama returned HTTP {resp.status} for {url}: {body[:300]}"
|
||||
)
|
||||
return await resp.json()
|
||||
except aiohttp.ClientConnectionError as exc:
|
||||
raise RuntimeError(f"Cannot reach Ollama at {url}: {exc}") from exc
|
||||
|
||||
|
||||
async def _get_json(
|
||||
url: str, *, timeout: float = 5.0, headers: dict | None = None
|
||||
) -> dict:
|
||||
"""GET url, return parsed response dict.
|
||||
|
||||
Raises RuntimeError on an HTTP error status (distinct message, so callers
|
||||
can tell "server responded with an error" from "couldn't reach it at
|
||||
all" — aiohttp connection/timeout errors propagate unwrapped for that
|
||||
reason). Message is generic, not backend-branded: this helper is shared
|
||||
by every LLMProvider implementation.
|
||||
"""
|
||||
import aiohttp
|
||||
|
||||
async with aiohttp.ClientSession() as session:
|
||||
async with session.get(
|
||||
url,
|
||||
headers=headers or None,
|
||||
timeout=aiohttp.ClientTimeout(total=timeout),
|
||||
) as resp:
|
||||
if resp.status >= 400:
|
||||
body = await resp.text()
|
||||
raise RuntimeError(
|
||||
f"Server returned HTTP {resp.status} for {url}: {body[:300]}"
|
||||
)
|
||||
return await resp.json()
|
||||
|
||||
|
||||
async def _fetch_models(host: str, headers: dict | None = None) -> list[str]:
|
||||
"""GET {host}/api/tags — return list of model name strings.
|
||||
|
||||
Used by comfydv.ollama's combo-widget population (_load_default_models,
|
||||
the /dv/ollama/models route) — a narrower, name-only view than
|
||||
OllamaProvider.list_models(), which returns full ModelInfo with status.
|
||||
Cached for _MODEL_LIST_CACHE.ttl_seconds per (host, headers) pair.
|
||||
"""
|
||||
cache_key = _cache_key("models", host, headers or {})
|
||||
cached, hit = _MODEL_LIST_CACHE.get(cache_key)
|
||||
if hit:
|
||||
return cached
|
||||
|
||||
try:
|
||||
data = await _get_json(f"{host}/api/tags", headers=headers)
|
||||
models = [m["name"] for m in data.get("models", [])]
|
||||
except Exception as exc:
|
||||
logger.warning("Could not fetch Ollama models from %s: %s", host, exc)
|
||||
return []
|
||||
|
||||
if models:
|
||||
_MODEL_LIST_CACHE.set(cache_key, models)
|
||||
return models
|
||||
|
||||
|
||||
async def _require_vision_capability(
|
||||
host: str, model: str, headers: dict | None
|
||||
) -> None:
|
||||
"""Raise a clear error if ``model`` lacks Ollama's ``vision`` capability.
|
||||
|
||||
Only called when a request carries at least one image (spec 009 FR-006):
|
||||
Ollama's /api/chat silently accepts an unsupported ``images`` field and
|
||||
answers with a blank/malformed HTTP 200 instead of an error — which
|
||||
would otherwise be indistinguishable from an ordinary blank generation
|
||||
and get swallowed by chat()'s existing blank-response retry. /api/show's
|
||||
``capabilities`` list is the only place Ollama states support explicitly,
|
||||
so a request carrying an image is checked against it up front.
|
||||
|
||||
Fails open on any lookup problem (older Ollama without ``capabilities``,
|
||||
unreachable host, unexpected shape) — a lookup failure must not block a
|
||||
request that would otherwise have worked; the real request surfaces its
|
||||
own clear error if the host is genuinely unreachable.
|
||||
"""
|
||||
cache_key = _cache_key("capabilities", host, headers or {}, model)
|
||||
cached, hit = _CAPABILITY_CACHE.get(cache_key)
|
||||
if hit:
|
||||
capabilities = cached
|
||||
else:
|
||||
try:
|
||||
data = await _post_json(
|
||||
f"{host}/api/show", {"model": model}, timeout=10.0, headers=headers
|
||||
)
|
||||
except Exception:
|
||||
return
|
||||
capabilities = data.get("capabilities")
|
||||
if capabilities is None:
|
||||
return
|
||||
_CAPABILITY_CACHE.set(cache_key, capabilities)
|
||||
|
||||
if "vision" not in capabilities:
|
||||
raise ValueError(
|
||||
f"Model '{model}' does not support image input — Ollama reports "
|
||||
f"capabilities {capabilities!r} for it, no 'vision'. Wire a "
|
||||
"vision-capable model, or disconnect the image input for "
|
||||
"text-only chat."
|
||||
)
|
||||
|
||||
|
||||
class OllamaProvider:
|
||||
"""LLMProvider implementation backed by Ollama's REST API.
|
||||
|
||||
Host and headers are captured once at construction — every method
|
||||
reuses them, matching the ADR-005 config-node pattern (one
|
||||
``OllamaClient`` node's output is one ``OllamaProvider`` instance).
|
||||
"""
|
||||
|
||||
def __init__(self, host: str, headers: dict | None = None):
|
||||
self.host = host
|
||||
self.headers = dict(headers) if headers else None
|
||||
|
||||
async def list_models(self) -> list[ModelInfo]:
|
||||
"""Every installed model, with live loaded/unloaded status.
|
||||
|
||||
`/api/tags` lists installed models; `/api/ps` lists currently-loaded
|
||||
ones. Ollama has no `sleeping`/`downloading` concept via this API —
|
||||
never emitted here (ADR-007's documented approximation).
|
||||
"""
|
||||
cache_key = _cache_key("list_models", self.host, self.headers or {})
|
||||
cached, hit = _MODEL_LIST_CACHE.get(cache_key)
|
||||
if hit:
|
||||
return cached
|
||||
|
||||
try:
|
||||
tags = await _get_json(f"{self.host}/api/tags", headers=self.headers)
|
||||
except Exception as exc:
|
||||
logger.warning("Could not fetch Ollama models from %s: %s", self.host, exc)
|
||||
return []
|
||||
|
||||
loaded_names: set[str] = set()
|
||||
try:
|
||||
ps = await _get_json(f"{self.host}/api/ps", headers=self.headers)
|
||||
loaded_names = {m["name"] for m in ps.get("models", [])}
|
||||
except Exception as exc:
|
||||
logger.warning(
|
||||
"Could not fetch Ollama running models from %s: %s", self.host, exc
|
||||
)
|
||||
|
||||
models = [
|
||||
ModelInfo(
|
||||
name=m["name"],
|
||||
status=(
|
||||
ModelStatus.LOADED
|
||||
if m["name"] in loaded_names
|
||||
else ModelStatus.UNLOADED
|
||||
),
|
||||
size=m.get("size"),
|
||||
)
|
||||
for m in tags.get("models", [])
|
||||
]
|
||||
if models:
|
||||
_MODEL_LIST_CACHE.set(cache_key, models)
|
||||
return models
|
||||
|
||||
async def load_model(self, model: str) -> None:
|
||||
if not model.strip():
|
||||
raise ValueError("model name cannot be empty")
|
||||
await _post_json(
|
||||
f"{self.host}/api/generate",
|
||||
{"model": model, "keep_alive": -1, "stream": False},
|
||||
timeout=300.0,
|
||||
headers=self.headers,
|
||||
)
|
||||
|
||||
async def unload_model(self, model: str) -> None:
|
||||
if not model.strip():
|
||||
raise ValueError("model name cannot be empty")
|
||||
await _post_json(
|
||||
f"{self.host}/api/generate",
|
||||
{"model": model, "keep_alive": 0, "stream": False},
|
||||
timeout=30.0,
|
||||
headers=self.headers,
|
||||
)
|
||||
|
||||
async def chat(
|
||||
self,
|
||||
model: str,
|
||||
messages: list[Message],
|
||||
options: dict | None = None,
|
||||
timeout_secs: float = 300.0,
|
||||
max_retries: int = 2,
|
||||
attempt_info: dict | None = None,
|
||||
on_status: Callable[[str], None] | None = None,
|
||||
) -> str:
|
||||
if any(m.images for m in messages):
|
||||
await _require_vision_capability(self.host, model, self.headers)
|
||||
|
||||
# exclude_none drops the images key for text-only turns so an
|
||||
# image-less request is byte-identical to the pre-009 payload
|
||||
# (FR-003); a turn with images keeps Ollama's native flat images
|
||||
# array (ADR-008 — no transform needed for /api/chat).
|
||||
payload_messages = [m.model_dump(exclude_none=True) for m in messages]
|
||||
options, think = _pop_think(options)
|
||||
options, refusal_cfg = _pop_refusal_retry(options)
|
||||
embed_fn = None
|
||||
custom_phrases: tuple[str, ...] = ()
|
||||
if refusal_cfg and refusal_cfg.get("enabled"):
|
||||
custom_phrases = tuple(refusal_cfg.get("custom_phrases") or ())
|
||||
if refusal_cfg.get("embedding_model"):
|
||||
embedding_model = refusal_cfg["embedding_model"]
|
||||
embed_fn = lambda t: self.embed(embedding_model, t) # noqa: E731
|
||||
total_attempts = max(0, min(int(max_retries), 5)) + 1
|
||||
response_text = ""
|
||||
incomplete = False
|
||||
refusal_count = 0
|
||||
attempt_seed = 0
|
||||
attempt_timeout = timeout_secs
|
||||
|
||||
for attempt in range(1, total_attempts + 1):
|
||||
attempt_options = dict(options) if options else {}
|
||||
if attempt > 1:
|
||||
attempt_options["seed"] = next_seed(options, attempt)
|
||||
attempt_seed = attempt_options.get("seed", 0)
|
||||
attempt_timeout = next_timeout_secs(timeout_secs, attempt)
|
||||
|
||||
payload: dict = {
|
||||
"model": model,
|
||||
"messages": payload_messages,
|
||||
"stream": False,
|
||||
}
|
||||
if attempt_options:
|
||||
payload["options"] = attempt_options
|
||||
if think is not None:
|
||||
# ADR-010: confirmed live this must be a top-level field —
|
||||
# Ollama silently ignores "think" nested inside "options".
|
||||
payload["think"] = think
|
||||
|
||||
cache_key = _cache_key(
|
||||
"chat",
|
||||
self.host,
|
||||
self.headers or {},
|
||||
model,
|
||||
payload_messages,
|
||||
attempt_options,
|
||||
think,
|
||||
)
|
||||
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
|
||||
if hit:
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=attempt,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
return cached
|
||||
|
||||
result = await _post_json(
|
||||
f"{self.host}/api/chat",
|
||||
payload,
|
||||
timeout=attempt_timeout,
|
||||
headers=self.headers,
|
||||
)
|
||||
response_text = result.get("message", {}).get("content", "")
|
||||
retry_reason: str | None = None
|
||||
if response_text.strip():
|
||||
refused = False
|
||||
if refusal_cfg and refusal_cfg.get("enabled"):
|
||||
refused = await is_refusal(
|
||||
response_text,
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key=refusal_cfg.get("embedding_model", ""),
|
||||
threshold=refusal_cfg.get("threshold", 0.82),
|
||||
custom_phrases=custom_phrases,
|
||||
)
|
||||
if not refused:
|
||||
_CHAT_RESPONSE_CACHE.set(cache_key, response_text)
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=attempt,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
if on_status is not None and attempt > 1:
|
||||
on_status(
|
||||
format_recovered_status(
|
||||
attempt, total_attempts, attempt_seed
|
||||
)
|
||||
)
|
||||
return response_text
|
||||
# A detected refusal is handled exactly like a blank
|
||||
# response below: fall through to the backoff/retry with a
|
||||
# bumped seed (next_seed), rather than returning the refusal
|
||||
# text to the caller.
|
||||
refusal_count += 1
|
||||
retry_reason = "Refusal/deflection detected"
|
||||
|
||||
# done: false alongside blank content is a distinct signal from
|
||||
# an ordinary blank generation — it's Ollama answering before
|
||||
# the model has actually finished loading/swapping in, observed
|
||||
# live under model-swap load (issue #27), not the model having
|
||||
# genuinely generated nothing. Tracked separately so it can be
|
||||
# raised on below instead of silently returned like a real
|
||||
# blank generation would be.
|
||||
incomplete = result.get("done") is False
|
||||
if retry_reason is None and not response_text.strip():
|
||||
retry_reason = (
|
||||
"Model still loading/swapping" if incomplete else "Blank response"
|
||||
)
|
||||
|
||||
if attempt < total_attempts:
|
||||
if on_status is not None and retry_reason is not None:
|
||||
upcoming_seed = next_seed(options, attempt + 1)
|
||||
upcoming_timeout = next_timeout_secs(timeout_secs, attempt + 1)
|
||||
on_status(
|
||||
format_retry_status(
|
||||
retry_reason,
|
||||
attempt,
|
||||
total_attempts,
|
||||
upcoming_seed,
|
||||
upcoming_timeout,
|
||||
)
|
||||
)
|
||||
await asyncio.sleep(RETRY_BACKOFF_SECS)
|
||||
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=total_attempts,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
|
||||
if incomplete:
|
||||
raise RuntimeError(
|
||||
f"Ollama returned an incomplete response after "
|
||||
f"{total_attempts} attempt(s) for model '{model}' — it may "
|
||||
"still be loading or swapping in memory. Try again in a "
|
||||
"few seconds."
|
||||
)
|
||||
|
||||
# Every attempt came back blank (and complete) — never raises here
|
||||
# (chat() has never validated its output, unlike chat_structured());
|
||||
# return the last (blank) attempt uncached so the next queue run
|
||||
# tries fresh.
|
||||
return response_text
|
||||
|
||||
async def chat_structured(
|
||||
self,
|
||||
model: str,
|
||||
messages: list[Message],
|
||||
schema: type[BaseModel],
|
||||
options: dict | None = None,
|
||||
timeout_secs: float = 300.0,
|
||||
max_retries: int = 2,
|
||||
attempt_info: dict | None = None,
|
||||
on_status: Callable[[str], None] | None = None,
|
||||
) -> BaseModel:
|
||||
"""Native ``/api/chat`` + ``"format"`` (grammar-constrained JSON
|
||||
decoding), not the shared pydantic-ai ``chat.py`` helper.
|
||||
|
||||
ADR-009 originally routed this through the OpenAI-compatible
|
||||
``/v1/chat/completions`` endpoint via pydantic-ai's ``NativeOutput``.
|
||||
Confirmed live that endpoint silently *reloads the model at its
|
||||
default context size on every call*, discarding any prior
|
||||
``options.num_ctx`` — even when the same ``options`` are included in
|
||||
that very request. Priming with a separate native call first
|
||||
(the original fix) didn't help: the very next OpenAI-compat call
|
||||
undid it immediately. The native ``/api/chat`` endpoint doesn't
|
||||
have this problem — confirmed live it preserves an already-primed
|
||||
context, and it supports structured output directly via
|
||||
``"format"``, so ``options`` and structured output now apply
|
||||
atomically in one request. ``LlamaCppProvider`` is unaffected — it
|
||||
keeps using the shared pydantic-ai path, since llama-server's
|
||||
context is fixed at process launch, not a per-request concern.
|
||||
"""
|
||||
if any(m.images for m in messages):
|
||||
await _require_vision_capability(self.host, model, self.headers)
|
||||
|
||||
payload_messages = [m.model_dump(exclude_none=True) for m in messages]
|
||||
json_schema = schema.model_json_schema()
|
||||
options, think = _pop_think(options)
|
||||
options, refusal_cfg = _pop_refusal_retry(options)
|
||||
embed_fn = None
|
||||
custom_phrases: tuple[str, ...] = ()
|
||||
if refusal_cfg and refusal_cfg.get("enabled"):
|
||||
custom_phrases = tuple(refusal_cfg.get("custom_phrases") or ())
|
||||
if refusal_cfg.get("embedding_model"):
|
||||
embedding_model = refusal_cfg["embedding_model"]
|
||||
embed_fn = lambda t: self.embed(embedding_model, t) # noqa: E731
|
||||
cache_key = _cache_key(
|
||||
"chat_structured",
|
||||
self.host,
|
||||
self.headers or {},
|
||||
model,
|
||||
payload_messages,
|
||||
options or {},
|
||||
json_schema,
|
||||
think,
|
||||
)
|
||||
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
|
||||
if hit:
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=(options or {}).get("seed", 0),
|
||||
attempts=1,
|
||||
timeout_secs=timeout_secs,
|
||||
refusals=0,
|
||||
)
|
||||
return schema.model_validate(cached)
|
||||
|
||||
total_attempts = max(0, min(int(max_retries), 5)) + 1
|
||||
last_error: Exception | None = None
|
||||
last_invalid_text = ""
|
||||
refusal_count = 0
|
||||
attempt_seed = 0
|
||||
attempt_timeout = timeout_secs
|
||||
|
||||
def _emit_retry_status(reason: str, attempt: int) -> None:
|
||||
if on_status is None or attempt >= total_attempts:
|
||||
return
|
||||
upcoming_seed = next_seed(options, attempt + 1)
|
||||
upcoming_timeout = next_timeout_secs(timeout_secs, attempt + 1)
|
||||
on_status(
|
||||
format_retry_status(
|
||||
reason, attempt, total_attempts, upcoming_seed, upcoming_timeout
|
||||
)
|
||||
)
|
||||
|
||||
for attempt in range(1, total_attempts + 1):
|
||||
attempt_options = dict(options) if options else {}
|
||||
if attempt > 1:
|
||||
attempt_options["seed"] = next_seed(options, attempt)
|
||||
attempt_seed = attempt_options.get("seed", 0)
|
||||
attempt_timeout = next_timeout_secs(timeout_secs, attempt)
|
||||
|
||||
payload: dict = {
|
||||
"model": model,
|
||||
"messages": payload_messages,
|
||||
"format": json_schema,
|
||||
"stream": False,
|
||||
}
|
||||
if attempt_options:
|
||||
payload["options"] = attempt_options
|
||||
if think is not None:
|
||||
# ADR-010: confirmed live this must be a top-level field —
|
||||
# Ollama silently ignores "think" nested inside "options".
|
||||
payload["think"] = think
|
||||
|
||||
try:
|
||||
result = await _post_json(
|
||||
f"{self.host}/api/chat",
|
||||
payload,
|
||||
timeout=attempt_timeout,
|
||||
headers=self.headers,
|
||||
)
|
||||
except RuntimeError as exc:
|
||||
last_error = exc
|
||||
last_invalid_text = str(exc)
|
||||
_emit_retry_status("Request failed", attempt)
|
||||
if attempt < total_attempts:
|
||||
await asyncio.sleep(RETRY_BACKOFF_SECS)
|
||||
continue
|
||||
|
||||
content = result.get("message", {}).get("content", "")
|
||||
try:
|
||||
parsed = schema.model_validate_json(content)
|
||||
except ValidationError as exc:
|
||||
last_error = exc
|
||||
last_invalid_text = content
|
||||
_emit_retry_status("Schema validation failed", attempt)
|
||||
if attempt < total_attempts:
|
||||
await asyncio.sleep(RETRY_BACKOFF_SECS)
|
||||
continue
|
||||
|
||||
if refusal_cfg and refusal_cfg.get("enabled"):
|
||||
# Checked against the raw JSON text, not a specific parsed
|
||||
# field: ChatCompletion's schema is caller-defined and this
|
||||
# provider has no idea which field would carry refusal
|
||||
# language — the regex/embedding check still matches text
|
||||
# sitting inside a JSON string value either way.
|
||||
refused = await is_refusal(
|
||||
content,
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key=refusal_cfg.get("embedding_model", ""),
|
||||
threshold=refusal_cfg.get("threshold", 0.82),
|
||||
custom_phrases=custom_phrases,
|
||||
)
|
||||
if refused:
|
||||
refusal_count += 1
|
||||
last_error = RuntimeError("refusal/deflection detected")
|
||||
last_invalid_text = content
|
||||
_emit_retry_status("Refusal/deflection detected", attempt)
|
||||
if attempt < total_attempts:
|
||||
await asyncio.sleep(RETRY_BACKOFF_SECS)
|
||||
continue
|
||||
|
||||
_CHAT_RESPONSE_CACHE.set(cache_key, parsed.model_dump())
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=attempt,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
if on_status is not None and attempt > 1:
|
||||
on_status(
|
||||
format_recovered_status(attempt, total_attempts, attempt_seed)
|
||||
)
|
||||
return parsed
|
||||
|
||||
record_attempt_info(
|
||||
attempt_info,
|
||||
seed=attempt_seed,
|
||||
attempts=total_attempts,
|
||||
timeout_secs=attempt_timeout,
|
||||
refusals=refusal_count,
|
||||
)
|
||||
raise RuntimeError(
|
||||
f"chat_structured: response failed validation against schema after "
|
||||
f"{total_attempts} attempt(s) (model={model!r}). Last error: "
|
||||
f"{last_error}. Last response (truncated): {last_invalid_text[:300]!r}"
|
||||
)
|
||||
|
||||
async def embed(self, model: str, text: str) -> list[float] | None:
|
||||
"""POST {host}/api/embed — Ollama's native embeddings endpoint.
|
||||
|
||||
Returns ``None`` rather than raising on any failure (wrong/missing
|
||||
embedding model, unreachable server, malformed response) — this is
|
||||
a best-effort capability per the ``LLMProvider`` protocol, and its
|
||||
one current caller (refusal-retry detection) already treats
|
||||
``None`` as "skip the embedding check", not an error.
|
||||
"""
|
||||
if not model.strip() or not text.strip():
|
||||
return None
|
||||
try:
|
||||
result = await _post_json(
|
||||
f"{self.host}/api/embed",
|
||||
{"model": model, "input": text},
|
||||
timeout=30.0,
|
||||
headers=self.headers,
|
||||
)
|
||||
except Exception:
|
||||
return None
|
||||
embeddings = result.get("embeddings")
|
||||
if not isinstance(embeddings, list) or not embeddings:
|
||||
return None
|
||||
vec = embeddings[0]
|
||||
if not isinstance(vec, list) or not vec:
|
||||
return None
|
||||
return vec
|
||||
@@ -0,0 +1,161 @@
|
||||
"""LLMProvider protocol — the adapter boundary between ComfyUI nodes and
|
||||
specific local inference backends.
|
||||
|
||||
ADR-007: every backend (OllamaProvider now, LlamaCppProvider in a follow-on
|
||||
epic) implements this shape; ComfyUI nodes depend only on the protocol,
|
||||
never on a concrete provider class. See
|
||||
project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md and
|
||||
specs/007-llm-provider-abstraction/contracts/llm_provider_protocol.md.
|
||||
"""
|
||||
|
||||
from collections.abc import Callable
|
||||
from enum import Enum
|
||||
from typing import Literal, Protocol
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
|
||||
class ModelStatus(str, Enum):
|
||||
"""Residency status of a model on a provider's server.
|
||||
|
||||
Not every provider emits every value — e.g. Ollama has no distinct
|
||||
signal for SLEEPING or DOWNLOADING via its API and normalizes to the
|
||||
closest applicable status rather than omitting the model (documented
|
||||
approximation, ADR-007).
|
||||
"""
|
||||
|
||||
UNLOADED = "unloaded"
|
||||
LOADING = "loading"
|
||||
LOADED = "loaded"
|
||||
SLEEPING = "sleeping"
|
||||
DOWNLOADING = "downloading"
|
||||
|
||||
|
||||
class ModelInfo(BaseModel):
|
||||
"""One entry returned by ``LLMProvider.list_models()``."""
|
||||
|
||||
name: str
|
||||
status: ModelStatus
|
||||
size: int | None = None
|
||||
|
||||
|
||||
class Message(BaseModel):
|
||||
"""One turn in a chat request.
|
||||
|
||||
``images`` carries optional base64-encoded image payloads (no ``data:``
|
||||
prefix) associated with this turn, for vision-capable models. ``None``
|
||||
(the default) means a text-only turn that serializes byte-for-byte as
|
||||
before — providers dump with ``exclude_none=True`` so no ``images`` key
|
||||
reaches the wire for image-less turns. Each provider translates this
|
||||
neutral carrier into its own native shape (ADR-008): Ollama's flat
|
||||
per-message ``images`` array, llama.cpp's OpenAI ``image_url`` content
|
||||
parts, and pydantic-ai ``BinaryContent`` on the structured path.
|
||||
"""
|
||||
|
||||
role: Literal["system", "user", "assistant"]
|
||||
content: str
|
||||
images: list[str] | None = None
|
||||
|
||||
|
||||
class LLMProvider(Protocol):
|
||||
"""Adapter boundary every backend implements.
|
||||
|
||||
Connection state (host, auth headers, or equivalent) is captured once
|
||||
at provider-construction time — no method takes connection details as
|
||||
a parameter.
|
||||
"""
|
||||
|
||||
async def list_models(self) -> list[ModelInfo]:
|
||||
"""Every model the server currently knows about, loaded or not."""
|
||||
...
|
||||
|
||||
async def load_model(self, model: str) -> None:
|
||||
"""Load a model into memory. Idempotent — already-loaded is not an error."""
|
||||
...
|
||||
|
||||
async def unload_model(self, model: str) -> None:
|
||||
"""Unload a model from memory. Idempotent — already-unloaded is not an error."""
|
||||
...
|
||||
|
||||
async def chat(
|
||||
self,
|
||||
model: str,
|
||||
messages: list[Message],
|
||||
options: dict | None = None,
|
||||
timeout_secs: float = 300.0,
|
||||
max_retries: int = 2,
|
||||
attempt_info: dict | None = None,
|
||||
on_status: Callable[[str], None] | None = None,
|
||||
) -> str:
|
||||
"""Free-text chat response.
|
||||
|
||||
Retries up to ``max_retries`` times (clamped 0-5) with a new seed if
|
||||
the response comes back blank — confirmed live on a freshly-loaded
|
||||
model, whose first response is sometimes empty before it settles
|
||||
into normal behavior. Still returns the (possibly blank) last
|
||||
attempt's text rather than raising if every retry comes back blank —
|
||||
this method has never validated its output, unlike
|
||||
``chat_structured()``. Each retry's request timeout also escalates
|
||||
(``_llm/retry.py``'s ``next_timeout_secs``) rather than reusing the
|
||||
same budget that just ran out.
|
||||
|
||||
ADR-010: ``options`` may carry a ``"think"`` key (bool) to disable a
|
||||
"thinking"-capable model's chain-of-thought reasoning — every
|
||||
implementation pops it out of ``options`` and translates it to its
|
||||
own wire shape (Ollama: a top-level ``think`` field; llama.cpp:
|
||||
``chat_template_kwargs``/``reasoning_effort`` in the request body),
|
||||
since neither backend recognizes a literal ``"think"`` key nested
|
||||
inside a generic options object.
|
||||
|
||||
``attempt_info``, if given, is populated in place with the retry
|
||||
loop's final outcome (seed/timeout used, attempt count, refusal
|
||||
count) via ``_llm/retry.py``'s ``record_attempt_info`` — an optional
|
||||
out-param, not a return-type change, so existing callers that don't
|
||||
pass it see no behavior change.
|
||||
|
||||
``on_status``, if given, is called synchronously at each retry
|
||||
boundary with a one-line human-readable status (see
|
||||
``_llm/retry.py``'s ``format_retry_status``/``format_recovered_status``)
|
||||
— a live counterpart to ``attempt_info``, which only reports the
|
||||
final outcome after the call returns.
|
||||
"""
|
||||
...
|
||||
|
||||
async def chat_structured(
|
||||
self,
|
||||
model: str,
|
||||
messages: list[Message],
|
||||
schema: type[BaseModel],
|
||||
options: dict | None = None,
|
||||
timeout_secs: float = 300.0,
|
||||
max_retries: int = 2,
|
||||
attempt_info: dict | None = None,
|
||||
on_status: Callable[[str], None] | None = None,
|
||||
) -> BaseModel:
|
||||
"""Schema-validated chat response.
|
||||
|
||||
Raises ``RuntimeError`` (naming the model, attempt count, and a
|
||||
truncated snippet of the last invalid response) if every retry is
|
||||
exhausted — never returns a value with a missing/blank required
|
||||
field.
|
||||
|
||||
ADR-010: see ``chat()`` — same ``options["think"]`` convention,
|
||||
same per-provider translation, same escalating per-attempt timeout,
|
||||
and the same ``attempt_info``/``on_status`` conventions.
|
||||
"""
|
||||
...
|
||||
|
||||
async def embed(self, model: str, text: str) -> list[float] | None:
|
||||
"""Embedding vector for ``text``, or ``None`` if unavailable.
|
||||
|
||||
Best-effort, not a core capability every deployment has configured:
|
||||
``model`` must itself be embedding-capable, which is typically a
|
||||
*different* model from whatever's answering chat requests (e.g.
|
||||
``nomic-embed-text``, not the model passed to ``chat()``). Returns
|
||||
``None`` rather than raising when embeddings aren't usable right now
|
||||
(wrong/missing model, unreachable server) — the one current caller,
|
||||
refusal-retry detection (see ``_llm/retry.py``), degrades gracefully
|
||||
to lexical-only detection when this returns ``None``, so a provider
|
||||
with no embedding model configured is never a hard failure.
|
||||
"""
|
||||
...
|
||||
@@ -0,0 +1,333 @@
|
||||
"""Shared retry-on-empty-output helpers for chat()/chat_structured().
|
||||
|
||||
Both providers' chat() calls (ADR-007) and the shared chat_structured()
|
||||
helper (_llm/chat.py) hit the same class of failure, confirmed live against
|
||||
a freshly-started Ollama instance on a fresh runpod: the model's first
|
||||
response after loading is sometimes blank or fails structured-output
|
||||
validation outright, then behaves normally on the very next call. Centralized
|
||||
here so both providers and both chat modes retry the same way rather than
|
||||
each re-deriving the policy.
|
||||
|
||||
Only blank/whitespace-only responses trigger a retry for plain chat() —
|
||||
not merely "short" ones — because a fixed length threshold would misfire on
|
||||
legitimately short, valid answers (single-word replies, labels, "yes"/"no").
|
||||
|
||||
Refusal/deflection detection (below) is a separate, opt-in trigger for the
|
||||
same retry-with-a-new-seed mechanism: some models (observed with an
|
||||
abliterated Qwen variant) answer with a soft refusal on a topic they judge
|
||||
"sensitive" instead of erroring or returning blank, so neither of the above
|
||||
checks catches it. This is deliberately a model-behavior concern, not a
|
||||
backend one — every ``LLMProvider`` implementation (Ollama, llama.cpp, and
|
||||
whatever comes next) wires the same detector into its own retry loop via its
|
||||
own ``embed()``, rather than each backend inventing its own heuristic.
|
||||
"""
|
||||
|
||||
import math
|
||||
import re
|
||||
from collections.abc import Awaitable, Callable
|
||||
|
||||
RETRY_BACKOFF_SECS = 1.5
|
||||
"""Flat delay between retries — gives a still-loading model time to finish
|
||||
before the next attempt, rather than hammering it with identical requests
|
||||
back-to-back."""
|
||||
|
||||
|
||||
def next_seed(options: dict | None, attempt: int) -> int:
|
||||
"""Deterministic seed for retry ``attempt`` (1-indexed).
|
||||
|
||||
Attempt 1 is the caller's original request and is never touched by this
|
||||
function — callers only call it for attempt >= 2. Starts from
|
||||
``options["seed"]`` if the caller pinned one, else 0, and increments by
|
||||
``attempt - 1`` so each retry is a new, reproducible value instead of
|
||||
repeating the exact same request that just failed.
|
||||
"""
|
||||
base = 0
|
||||
if options and isinstance(options.get("seed"), int):
|
||||
base = options["seed"]
|
||||
return base + (attempt - 1)
|
||||
|
||||
|
||||
def next_timeout_secs(base_timeout: float, attempt: int) -> float:
|
||||
"""Escalating per-attempt timeout for retries (1-indexed ``attempt``).
|
||||
|
||||
Attempt 1 gets the caller's own ``timeout_secs`` unchanged; each retry
|
||||
multiplies it by the attempt number. A request that timed out may
|
||||
genuinely need more time — a slow-to-load or heavily-loaded model, a
|
||||
large prompt — not just an identical retry under the same budget it
|
||||
just failed to meet.
|
||||
"""
|
||||
return base_timeout * attempt
|
||||
|
||||
|
||||
def record_attempt_info(
|
||||
attempt_info: dict | None,
|
||||
*,
|
||||
seed: int,
|
||||
attempts: int,
|
||||
timeout_secs: float,
|
||||
refusals: int,
|
||||
) -> None:
|
||||
"""Populate an optional caller-supplied dict with the retry loop's
|
||||
final outcome — the seed/timeout actually used, how many attempts it
|
||||
took, and how many were refusal-triggered.
|
||||
|
||||
A plain out-param rather than a return-type change, so it's fully
|
||||
backward compatible: a caller that doesn't pass ``attempt_info`` sees
|
||||
no change in behavior at all. ``ChatCompletion`` uses this to expose
|
||||
the seed actually used as a node output and to build a UI status line
|
||||
when a retry/refusal happened.
|
||||
"""
|
||||
if attempt_info is None:
|
||||
return
|
||||
attempt_info.update(
|
||||
{
|
||||
"seed": seed,
|
||||
"attempts": attempts,
|
||||
"timeout_secs": timeout_secs,
|
||||
"refusals": refusals,
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
OnStatus = Callable[[str], None]
|
||||
"""A caller-supplied, synchronous, best-effort progress callback — see
|
||||
``format_retry_status``/``format_recovered_status``. Not async: providers
|
||||
call it inline mid-retry-loop, and the one real implementation
|
||||
(``ChatCompletion``'s closure over ``PromptServer.send_progress_text``) is
|
||||
itself synchronous, so there's nothing to await."""
|
||||
|
||||
|
||||
def format_retry_status(
|
||||
reason: str, attempt: int, total_attempts: int, seed: int, timeout_secs: float
|
||||
) -> str:
|
||||
"""One-line, human-readable status for ``on_status()`` callers — shown
|
||||
live on the node via ComfyUI's ``PromptServer.send_progress_text``
|
||||
(see ``ChatCompletion.chat()``). Centralized so every provider's retry
|
||||
loop describes a retry the same way rather than each inventing its own
|
||||
wording.
|
||||
"""
|
||||
return (
|
||||
f"⚠ {reason} on attempt {attempt}/{total_attempts} — "
|
||||
f"retrying with seed={seed}, timeout={timeout_secs:.0f}s"
|
||||
)
|
||||
|
||||
|
||||
def format_recovered_status(attempt: int, total_attempts: int, seed: int) -> str:
|
||||
"""Final status shown once a retry loop succeeds after >1 attempt —
|
||||
lets a live status left over from ``format_retry_status`` resolve to
|
||||
something other than a stale "retrying..." message."""
|
||||
return f"✅ Recovered on attempt {attempt}/{total_attempts} (seed={seed})"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Refusal/deflection detection
|
||||
# ---------------------------------------------------------------------------
|
||||
#
|
||||
# Hybrid, cheapest-check-first: a fast, free lexical pass catches the blatant
|
||||
# majority ("I cannot generate...") without ever touching the network; only
|
||||
# a response that's short and/or hedge-y enough to be genuinely ambiguous
|
||||
# pays for an embedding call. A long, on-topic response never reaches the
|
||||
# embedding step at all.
|
||||
|
||||
REFUSAL_LEXICAL_PATTERNS: tuple[re.Pattern, ...] = tuple(
|
||||
re.compile(p, re.IGNORECASE)
|
||||
for p in (
|
||||
r"\b(?:I\s*(?:'m|\s+am)?\s*)?(?:cannot|can't|won't|will not)\b[^.]{0,60}?\b"
|
||||
r"(?:generate|create|produce|write|provide|help|assist|describe|depict|continue)\b",
|
||||
r"\bI(?:'m|\s+am) (?:(?:not able|unable) to|restricted from)\b",
|
||||
r"\bI don't feel comfortable\b",
|
||||
r"\bI'm sorry,?\s*(?:but\s+)?I\s*(?:can't|cannot)\b",
|
||||
r"\bas an AI\b[^.]{0,60}?\b(?:cannot|can't|unable|not able)\b",
|
||||
r"\bI must (?:decline|refuse)\b",
|
||||
r"\bthis (?:request|content|image|prompt) (?:violates|goes against)\b",
|
||||
r"\bnot something I(?:'m| am)? (?:able to|going to) (?:help with|generate|create|provide)\b",
|
||||
r"\bI (?:won't|will not) be (?:able to )?(?:generat|creat|provid)",
|
||||
)
|
||||
)
|
||||
|
||||
_SOFT_HEDGE_KEYWORDS: tuple[str, ...] = (
|
||||
"sensitive",
|
||||
"unfortunately",
|
||||
"i must",
|
||||
"policy",
|
||||
"guideline",
|
||||
"responsible ai",
|
||||
"not appropriate",
|
||||
"inappropriate",
|
||||
"instead, i",
|
||||
"i'd rather",
|
||||
"i would rather",
|
||||
"controversial",
|
||||
"harmful",
|
||||
"offensive",
|
||||
"disturbing",
|
||||
"explicit content",
|
||||
"nsfw",
|
||||
"consider an alternative",
|
||||
)
|
||||
|
||||
_AMBIGUOUS_LENGTH_THRESHOLD = 600
|
||||
"""Below this many characters, a response is short enough that a soft
|
||||
refusal is plausible and worth the extra embedding check — chosen well
|
||||
under this pipeline's normal structured-JSON response sizes (typically
|
||||
1000+ characters), not a tuned/validated threshold.
|
||||
|
||||
Note: ``is_refusal`` no longer applies this gate when the caller supplies
|
||||
``embed_fn`` (see below) — an explicitly configured embedding model always
|
||||
runs the check regardless of length/keywords. ``is_ambiguous`` is kept as a
|
||||
standalone, independently tested heuristic for other callers (e.g. logging
|
||||
"why was this worth embedding-checking") rather than wired into the hybrid
|
||||
detector itself."""
|
||||
|
||||
REFUSAL_EXEMPLARS: tuple[str, ...] = (
|
||||
"I cannot generate an image or description involving this topic.",
|
||||
"I'm not able to help with this request as it involves a politically sensitive subject.",
|
||||
"I don't feel comfortable creating content related to this figure or event.",
|
||||
"This request touches on a sensitive political topic that I must decline to engage with.",
|
||||
"As an AI, I'm unable to produce content depicting this scenario.",
|
||||
"I'd rather not generate this — it involves content I'm not comfortable with.",
|
||||
)
|
||||
|
||||
|
||||
def is_lexical_refusal(text: str) -> bool:
|
||||
"""Cheap, free regex pass — catches the blatant majority of refusals."""
|
||||
return any(p.search(text) for p in REFUSAL_LEXICAL_PATTERNS)
|
||||
|
||||
|
||||
def is_ambiguous(text: str) -> bool:
|
||||
"""Whether ``text`` is short/hedge-y enough to be worth the pricier
|
||||
embedding check, having already failed the free lexical pass.
|
||||
|
||||
Deliberately cheap and approximate — false positives here only cost one
|
||||
extra embedding call, false negatives skip a refusal that a real
|
||||
similarity check might have caught. Not meant to be a precise signal on
|
||||
its own, just a gate on when the more expensive check runs at all.
|
||||
"""
|
||||
stripped = text.strip()
|
||||
if not stripped:
|
||||
return False
|
||||
if len(stripped) < _AMBIGUOUS_LENGTH_THRESHOLD:
|
||||
return True
|
||||
lowered = stripped.lower()
|
||||
return any(keyword in lowered for keyword in _SOFT_HEDGE_KEYWORDS)
|
||||
|
||||
|
||||
def cosine_similarity(a: list[float], b: list[float]) -> float:
|
||||
"""Standard cosine similarity, no numpy dependency (comfydv has none)."""
|
||||
if not a or not b or len(a) != len(b):
|
||||
return 0.0
|
||||
dot = sum(x * y for x, y in zip(a, b))
|
||||
norm_a = math.sqrt(sum(x * x for x in a))
|
||||
norm_b = math.sqrt(sum(y * y for y in b))
|
||||
if norm_a == 0.0 or norm_b == 0.0:
|
||||
return 0.0
|
||||
return dot / (norm_a * norm_b)
|
||||
|
||||
|
||||
EmbedFn = Callable[[str], Awaitable[list[float] | None]]
|
||||
|
||||
_exemplar_embedding_cache: dict[str, list[list[float]]] = {}
|
||||
|
||||
|
||||
async def _exemplar_embeddings(
|
||||
embed_fn: EmbedFn, cache_key: str, exemplars: tuple[str, ...] = REFUSAL_EXEMPLARS
|
||||
) -> list[list[float]]:
|
||||
"""Embed ``exemplars`` once per ``cache_key`` and reuse — the exemplar
|
||||
set only changes if the caller's custom phrases change (folded into
|
||||
``cache_key`` by the caller), or the embedding space (i.e. which model
|
||||
produced the vectors) does.
|
||||
"""
|
||||
cached = _exemplar_embedding_cache.get(cache_key)
|
||||
if cached is not None:
|
||||
return cached
|
||||
embeddings = []
|
||||
for exemplar in exemplars:
|
||||
vec = await embed_fn(exemplar)
|
||||
if not vec:
|
||||
# An embedding call failing for one exemplar almost certainly
|
||||
# means embeddings aren't usable at all right now (wrong/missing
|
||||
# embedding model, unreachable server) — bail out rather than
|
||||
# caching a partial, unusable exemplar set.
|
||||
return []
|
||||
embeddings.append(vec)
|
||||
_exemplar_embedding_cache[cache_key] = embeddings
|
||||
return embeddings
|
||||
|
||||
|
||||
def _matches_custom_phrase(text: str, custom_phrases: tuple[str, ...]) -> bool:
|
||||
"""Case-insensitive substring match against user-supplied phrases."""
|
||||
if not custom_phrases:
|
||||
return False
|
||||
lowered = text.lower()
|
||||
return any(phrase.lower() in lowered for phrase in custom_phrases)
|
||||
|
||||
|
||||
async def is_refusal(
|
||||
text: str,
|
||||
*,
|
||||
embed_fn: EmbedFn | None = None,
|
||||
embed_cache_key: str = "",
|
||||
threshold: float = 0.82,
|
||||
custom_phrases: tuple[str, ...] = (),
|
||||
) -> bool:
|
||||
"""Hybrid refusal/deflection detector: free lexical pass first, then an
|
||||
embedding-similarity fallback whenever the caller has configured one.
|
||||
|
||||
``embed_fn`` is supplied by the caller's own ``LLMProvider.embed()`` —
|
||||
this function has no idea which backend or model produced ``text``, by
|
||||
design (ADR: refusal detection is a model-behavior concern, not a
|
||||
backend one). ``embed_fn=None`` (no embedding model configured) degrades
|
||||
to lexical-only detection rather than erroring; ``embed_fn`` present
|
||||
means the caller already opted in to the extra cost, so every non-blank,
|
||||
non-lexically-caught response gets checked — no further length/keyword
|
||||
gating. Any failure while embedding (unreachable server, no
|
||||
embedding-capable model loaded) is swallowed the same way — an optional
|
||||
enhancement failing shouldn't take down the retry loop it's assisting.
|
||||
|
||||
``custom_phrases`` lets a caller extend detection at runtime — e.g. a
|
||||
ComfyUI node field the user edits directly — without touching the
|
||||
shipped patterns/exemplars. Each phrase is checked two ways: a free
|
||||
case-insensitive substring match (same cost tier as the lexical pass,
|
||||
so it runs even with no ``embed_fn`` configured), and, when ``embed_fn``
|
||||
is present, folded in as additional exemplars for the similarity check
|
||||
so near-matches (not just exact substrings) of the user's phrases count
|
||||
too.
|
||||
"""
|
||||
if not text or not text.strip():
|
||||
return False # blank responses are the *other* retry trigger, not this one
|
||||
if is_lexical_refusal(text):
|
||||
return True
|
||||
custom_phrases = tuple(p.strip() for p in custom_phrases if p and p.strip())
|
||||
if _matches_custom_phrase(text, custom_phrases):
|
||||
return True
|
||||
if embed_fn is None:
|
||||
return False
|
||||
# embed_fn only exists when the caller explicitly configured an
|
||||
# embedding_model — that's an opt-in to pay for the check, so run it on
|
||||
# every non-blank, non-lexically-caught response rather than gating
|
||||
# further on is_ambiguous. The length/keyword heuristic exists to avoid
|
||||
# *unwanted* embedding calls when no embedding model is configured (see
|
||||
# the embed_fn is None branch above); it has no reason to also suppress
|
||||
# calls once the caller has already asked for them, and doing so was
|
||||
# exactly what let the subtle/on-topic-looking deflections this feature
|
||||
# targets slip through undetected.
|
||||
exemplars = (
|
||||
REFUSAL_EXEMPLARS + custom_phrases if custom_phrases else REFUSAL_EXEMPLARS
|
||||
)
|
||||
exemplar_cache_key = (
|
||||
f"{embed_cache_key}|custom:{','.join(custom_phrases)}"
|
||||
if custom_phrases
|
||||
else embed_cache_key
|
||||
)
|
||||
try:
|
||||
exemplar_vecs = await _exemplar_embeddings(
|
||||
embed_fn, exemplar_cache_key, exemplars
|
||||
)
|
||||
if not exemplar_vecs:
|
||||
return False
|
||||
text_vec = await embed_fn(text)
|
||||
if not text_vec:
|
||||
return False
|
||||
except Exception:
|
||||
return False
|
||||
return max(cosine_similarity(text_vec, vec) for vec in exemplar_vecs) >= threshold
|
||||
@@ -21,7 +21,7 @@ import sys
|
||||
from typing import Any, Dict, List
|
||||
|
||||
from aiohttp import web
|
||||
from jinja2 import exceptions, sandbox
|
||||
from jinja2 import exceptions, meta, sandbox
|
||||
|
||||
# Set up logger for this module
|
||||
logger = logging.getLogger(__name__)
|
||||
@@ -75,6 +75,10 @@ class FormatString:
|
||||
|
||||
# Create a sandboxed Jinja2 environment for security
|
||||
jinja_env = sandbox.SandboxedEnvironment()
|
||||
# Jinja2 ships `tojson` but not its inverse; add one so STRING inputs
|
||||
# carrying a JSON array/object (ComfyUI has no native list socket type)
|
||||
# can be parsed back into real Python data, e.g. `{{ hints | fromjson }}`.
|
||||
jinja_env.filters["fromjson"] = json.loads
|
||||
|
||||
# Define additional context
|
||||
@staticmethod
|
||||
@@ -252,38 +256,48 @@ class FormatString:
|
||||
>>> # Test with additional context (should be excluded)
|
||||
>>> keys = FormatString._extract_keys("Time: {{ datetime.now() }}")
|
||||
>>> assert keys == []
|
||||
>>> # Test {% for %} control structures: the loop variable is bound by the
|
||||
>>> # template itself and must not be treated as a required input, while the
|
||||
>>> # iterable it draws from must be.
|
||||
>>> keys = FormatString._extract_keys(
|
||||
... "{% for hint in extraction_hints %}{{ hint }}{% endfor %}"
|
||||
... )
|
||||
>>> assert keys == ['extraction_hints']
|
||||
-->
|
||||
"""
|
||||
variables = []
|
||||
seen = set()
|
||||
|
||||
def add_var(var):
|
||||
var = var.split("|")[0].split(".")[0].strip()
|
||||
if var not in seen and var not in FormatString.additional_context:
|
||||
seen.add(var)
|
||||
variables.append(var)
|
||||
|
||||
# Extract variables from Jinja2 expressions {{ }}
|
||||
for match in re.finditer(
|
||||
r"\{\{\s*([\w.]+)(?:\s*\|[\w\s]+)?(?:\.[^\(\)]+\(\))?\s*\}\}", template
|
||||
):
|
||||
add_var(match.group(1))
|
||||
|
||||
# Extract variables from f-string style { }
|
||||
# Extract variables from Python str.format() style { }
|
||||
for match in re.finditer(r"\{(\w+)\}", template):
|
||||
add_var(match.group(1))
|
||||
|
||||
# Extract variables from Jinja2 control structures {% %}
|
||||
for structure in re.finditer(r"\{%.*?%\}", template):
|
||||
for var in re.findall(r"\b(\w+)\|\b", structure.group(0)):
|
||||
if not var.startswith("end") and var not in {
|
||||
"if",
|
||||
"else",
|
||||
"elif",
|
||||
"for",
|
||||
"in",
|
||||
}:
|
||||
add_var(var)
|
||||
# Extract variables referenced anywhere in Jinja2 syntax ({{ }} expressions
|
||||
# and {% %} control structures) by parsing the template with Jinja2 itself
|
||||
# rather than approximating it with regexes. This is what correctly excludes
|
||||
# names bound within the template (e.g. the `hint` loop variable in
|
||||
# `{% for hint in extraction_hints %}`) while still surfacing names the
|
||||
# template expects the caller to supply (e.g. `extraction_hints`).
|
||||
try:
|
||||
template_ast = FormatString.jinja_env.parse(template)
|
||||
except exceptions.TemplateSyntaxError:
|
||||
pass
|
||||
else:
|
||||
# find_undeclared_variables returns an unordered set; sort by first
|
||||
# textual occurrence so extraction order is deterministic and matches
|
||||
# the order the template reads left to right (callers rely on this
|
||||
# for positional outputs, e.g. two {{ }} variables in sequence).
|
||||
undeclared = sorted(
|
||||
meta.find_undeclared_variables(template_ast),
|
||||
key=lambda name: template.find(name),
|
||||
)
|
||||
for var in undeclared:
|
||||
add_var(var)
|
||||
|
||||
return variables
|
||||
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
"""llama.cpp connection node for ComfyUI.
|
||||
|
||||
Mirrors comfydv.ollama's OllamaClient exactly (ADR-007's parallel-
|
||||
implementation pattern) — LlamaCppClient is the only new node this feature
|
||||
introduces. Every other generic node (ChatCompletion, LLMModelSelector,
|
||||
LLMLoadModel, LLMUnloadModel) already works with any LLM_CLIENT-typed
|
||||
provider unchanged.
|
||||
|
||||
Deployment prerequisite: llama-server must be launched in router mode
|
||||
(--models-dir or --models-preset) — see specs/008-llamacpp-integration/quickstart.md.
|
||||
"""
|
||||
|
||||
from ._llm.llamacpp_provider import LlamaCppProvider
|
||||
|
||||
|
||||
class LlamaCppClient:
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {
|
||||
"required": {
|
||||
"host": ("STRING", {"default": "http://localhost:8080"}),
|
||||
},
|
||||
"optional": {
|
||||
"headers": ("OLLAMA_HEADERS",),
|
||||
},
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("LLM_CLIENT",)
|
||||
RETURN_NAMES = ("client",)
|
||||
FUNCTION = "create_client"
|
||||
CATEGORY = "dv/llamacpp"
|
||||
|
||||
def create_client(self, host: str, headers: dict | None = None):
|
||||
return (LlamaCppProvider(host, headers),)
|
||||
@@ -1,3 +1,4 @@
|
||||
import json
|
||||
import logging
|
||||
import random
|
||||
import sys
|
||||
@@ -7,6 +8,28 @@ from .utils import any_type
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def _preview_text(value) -> str:
|
||||
"""Best-effort text preview for RandomChoice's arbitrary-typed output.
|
||||
|
||||
Mirrors ComfyUI core's own ``PreviewAny`` node's value handling (str/
|
||||
number passthrough, else JSON, else ``str()``) rather than inventing a
|
||||
new convention — RandomChoice's output can be anything (an IMAGE
|
||||
tensor, a LATENT, a plain string), so this only needs to be "good
|
||||
enough to glance at," not a faithful repr of every type.
|
||||
"""
|
||||
if isinstance(value, str):
|
||||
return value
|
||||
if isinstance(value, (int, float, bool)):
|
||||
return str(value)
|
||||
try:
|
||||
return json.dumps(value, default=str, indent=2)
|
||||
except Exception:
|
||||
try:
|
||||
return str(value)
|
||||
except Exception:
|
||||
return "<value could not be serialized>"
|
||||
|
||||
|
||||
class RandomChoice:
|
||||
def __init__(self):
|
||||
pass
|
||||
@@ -25,15 +48,20 @@ class RandomChoice:
|
||||
|
||||
FUNCTION = "random_choice"
|
||||
|
||||
OUTPUT_NODE = False
|
||||
OUTPUT_NODE = True
|
||||
|
||||
CATEGORY = "dv/utils"
|
||||
|
||||
@classmethod
|
||||
def IS_CHANGED(s, **kwargs):
|
||||
return s.random_choice(s, **kwargs)
|
||||
# Unchanged from before the UI-preview addition: returns the raw
|
||||
# picked value (not the ui-wrapped dict random_choice() now returns)
|
||||
# so ComfyUI's change-detection comparison keeps working exactly as
|
||||
# it did previously.
|
||||
return s._pick(**kwargs)
|
||||
|
||||
def random_choice(self, **kwargs):
|
||||
@staticmethod
|
||||
def _pick(**kwargs):
|
||||
(
|
||||
random.seed(kwargs.get("seed"))
|
||||
if kwargs.get("seed")
|
||||
@@ -41,10 +69,13 @@ class RandomChoice:
|
||||
)
|
||||
input = [i for i in kwargs.items() if i[0] != "seed"]
|
||||
logger.debug("RandomChoice inputs: %s", input)
|
||||
return random.choice(input)[1]
|
||||
|
||||
def random_choice(self, **kwargs):
|
||||
try:
|
||||
choice = random.choice(input)[1]
|
||||
choice = self._pick(**kwargs)
|
||||
logger.debug("RandomChoice chose: %s", choice)
|
||||
return (choice,)
|
||||
return {"ui": {"text": [_preview_text(choice)]}, "result": (choice,)}
|
||||
except Exception as e:
|
||||
logger.error("RandomChoice: unexpected error: %s", e)
|
||||
raise
|
||||
|
||||
@@ -1,11 +1,15 @@
|
||||
/**
|
||||
* ollama.js — ComfyUI frontend extension for comfydv Ollama nodes.
|
||||
* ollama.js — ComfyUI frontend extension for comfydv's generic LLM nodes.
|
||||
*
|
||||
* Populates the model widget on Ollama nodes from a live call to
|
||||
* GET /dv/ollama/models?host=<url>.
|
||||
* Populates the model widget on LLM nodes from a live call to
|
||||
* GET /dv/ollama/models?host=<url>&backend=<ollama|llamacpp>. Despite the
|
||||
* file/route name (kept for historical reasons — see MIGRATION_MAP in
|
||||
* comfydv.ollama), this now serves both backends: which one a given node's
|
||||
* upstream client is determines the `backend` param (see
|
||||
* getHostAndBackendFromNode below).
|
||||
*
|
||||
* OllamaModelSelector and OllamaLoadModel use a COMBO widget (dropdown).
|
||||
* OllamaChatCompletion uses a plain STRING widget (accepts wired values).
|
||||
* LLMModelSelector and LLMLoadModel use a COMBO widget (dropdown).
|
||||
* ChatCompletion uses a plain STRING widget (accepts wired values).
|
||||
* The Refresh button works the same way for both: it fetches the live list
|
||||
* and sets the widget value / updates COMBO options as appropriate.
|
||||
*/
|
||||
@@ -13,12 +17,15 @@
|
||||
import { app } from "../../scripts/app.js";
|
||||
|
||||
/** Nodes whose model widget is a COMBO dropdown. */
|
||||
const OLLAMA_COMBO_NODES = new Set(["OllamaModelSelector", "OllamaLoadModel"]);
|
||||
const LLM_COMBO_NODES = new Set(["LLMModelSelector", "LLMLoadModel"]);
|
||||
|
||||
/** Nodes whose model widget is a plain STRING (accepts wired input). */
|
||||
const OLLAMA_STRING_MODEL_NODES = new Set(["OllamaChatCompletion"]);
|
||||
const LLM_STRING_MODEL_NODES = new Set(["ChatCompletion"]);
|
||||
|
||||
const OLLAMA_ALL_NODES = new Set([...OLLAMA_COMBO_NODES, ...OLLAMA_STRING_MODEL_NODES]);
|
||||
const LLM_ALL_NODES = new Set([...LLM_COMBO_NODES, ...LLM_STRING_MODEL_NODES]);
|
||||
|
||||
/** Registered client node type -> backend param the /dv/ollama/models route expects. */
|
||||
const CLIENT_NODE_BACKENDS = { OllamaClient: "ollama", LlamaCppClient: "llamacpp" };
|
||||
|
||||
/**
|
||||
* Fetch model list and update the node's model widget.
|
||||
@@ -27,9 +34,11 @@ const OLLAMA_ALL_NODES = new Set([...OLLAMA_COMBO_NODES, ...OLLAMA_STRING_MODEL_
|
||||
* - STRING: sets the value to the first model; keeps existing value if it
|
||||
* still appears in the live list (user may have typed a valid name).
|
||||
*/
|
||||
async function refreshModelWidget(node, host) {
|
||||
async function refreshModelWidget(node, host, backend) {
|
||||
try {
|
||||
const resp = await fetch(`/dv/ollama/models?host=${encodeURIComponent(host)}`);
|
||||
const resp = await fetch(
|
||||
`/dv/ollama/models?host=${encodeURIComponent(host)}&backend=${encodeURIComponent(backend)}`
|
||||
);
|
||||
if (!resp.ok) return;
|
||||
const data = await resp.json();
|
||||
const models = data.models ?? [];
|
||||
@@ -53,54 +62,115 @@ async function refreshModelWidget(node, host) {
|
||||
|
||||
node.setDirtyCanvas(true, false);
|
||||
} catch (_) {
|
||||
// Ollama unreachable — leave widget unchanged
|
||||
// Server unreachable — leave widget unchanged
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Locate the host string for a node.
|
||||
* Locate the host and backend for a node's connected LLM client.
|
||||
*
|
||||
* First checks the node's own widgets (OllamaClient has a "host" widget).
|
||||
* Otherwise traverses graph links to find a connected OllamaClient node and
|
||||
* reads its "host" widget — this is the common case for downstream nodes.
|
||||
* Traverses graph links to find a connected OllamaClient or LlamaCppClient
|
||||
* node and reads its "host" widget. Falls back to Ollama's default if
|
||||
* nothing is wired yet, matching the pre-existing fallback behavior.
|
||||
*/
|
||||
function getHostFromNode(node) {
|
||||
const ownHostWidget = node.widgets?.find(w => w.name === "host");
|
||||
if (ownHostWidget) return ownHostWidget.value;
|
||||
|
||||
function getHostAndBackendFromNode(node) {
|
||||
for (const input of node.inputs ?? []) {
|
||||
if (!input.link) continue;
|
||||
const link = node.graph?.links[input.link];
|
||||
if (!link) continue;
|
||||
const sourceNode = node.graph?.getNodeById(link.origin_id);
|
||||
if (sourceNode?.type === "OllamaClient") {
|
||||
const backend = sourceNode ? CLIENT_NODE_BACKENDS[sourceNode.type] : undefined;
|
||||
if (backend) {
|
||||
const hostWidget = sourceNode.widgets?.find(w => w.name === "host");
|
||||
if (hostWidget?.value) return hostWidget.value;
|
||||
if (hostWidget?.value) return { host: hostWidget.value, backend };
|
||||
}
|
||||
}
|
||||
|
||||
return "http://localhost:11434";
|
||||
return { host: "http://localhost:11434", backend: "ollama" };
|
||||
}
|
||||
|
||||
app.registerExtension({
|
||||
name: "comfydv.ollama",
|
||||
|
||||
async beforeRegisterNodeDef(nodeType, nodeData) {
|
||||
if (!OLLAMA_ALL_NODES.has(nodeData.name)) return;
|
||||
if (!LLM_ALL_NODES.has(nodeData.name)) return;
|
||||
|
||||
const onNodeCreated = nodeType.prototype.onNodeCreated;
|
||||
nodeType.prototype.onNodeCreated = function () {
|
||||
const result = onNodeCreated?.apply(this, arguments);
|
||||
|
||||
const refresh = () => {
|
||||
const { host, backend } = getHostAndBackendFromNode(this);
|
||||
refreshModelWidget(this, host, backend);
|
||||
};
|
||||
|
||||
// Add a Refresh button below the model widget
|
||||
this.addWidget("button", "⟳ Refresh models", null, () => {
|
||||
const host = getHostFromNode(this);
|
||||
refreshModelWidget(this, host);
|
||||
});
|
||||
this.addWidget("button", "⟳ Refresh models", null, refresh);
|
||||
|
||||
// Initial population on node creation
|
||||
const host = getHostFromNode(this);
|
||||
refreshModelWidget(this, host);
|
||||
refresh();
|
||||
|
||||
return result;
|
||||
};
|
||||
},
|
||||
});
|
||||
|
||||
/**
|
||||
* Live structured-output dynamic sockets for ChatCompletion.
|
||||
*
|
||||
* Mirrors FormatString's live dynamic-output pattern (see format_string.js):
|
||||
* editing structured_output or output_schema posts to a backend route that
|
||||
* recomputes ChatCompletion.RETURN_TYPES/RETURN_NAMES (the same
|
||||
* update_outputs() path chat() itself uses at execution time) and returns
|
||||
* the resulting output list — applied to this node's sockets immediately,
|
||||
* so you see the extracted fields appear without having to run the graph
|
||||
* first. Backend-agnostic: ChatCompletion is the one generic node both
|
||||
* OllamaProvider and LlamaCppProvider feed.
|
||||
*/
|
||||
async function updateStructuredOutputs(node, structuredOutput, outputSchema) {
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/update_structured_outputs", {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify({
|
||||
unique_id: String(node.id),
|
||||
structured_output: structuredOutput,
|
||||
output_schema: outputSchema,
|
||||
}),
|
||||
});
|
||||
if (!resp.ok) return;
|
||||
const data = await resp.json();
|
||||
applyOutputs(node, data.outputs ?? []);
|
||||
} catch (_) {
|
||||
// Backend unreachable — leave sockets unchanged.
|
||||
}
|
||||
}
|
||||
|
||||
function applyOutputs(node, outputs) {
|
||||
node.outputs.length = 0;
|
||||
outputs.forEach(o => node.addOutput(o.name, o.type));
|
||||
node.setDirtyCanvas(true, true);
|
||||
node.graph?.setDirtyCanvas(true, true);
|
||||
}
|
||||
|
||||
app.registerExtension({
|
||||
name: "comfydv.ollama.structuredOutput",
|
||||
|
||||
async beforeRegisterNodeDef(nodeType, nodeData) {
|
||||
if (nodeData.name !== "ChatCompletion") return;
|
||||
|
||||
const onNodeCreated = nodeType.prototype.onNodeCreated;
|
||||
nodeType.prototype.onNodeCreated = function () {
|
||||
const result = onNodeCreated?.apply(this, arguments);
|
||||
|
||||
const structuredWidget = this.widgets?.find(w => w.name === "structured_output");
|
||||
const schemaWidget = this.widgets?.find(w => w.name === "output_schema");
|
||||
if (!structuredWidget || !schemaWidget) return result;
|
||||
|
||||
const update = () =>
|
||||
updateStructuredOutputs(this, structuredWidget.value, schemaWidget.value);
|
||||
structuredWidget.callback = update;
|
||||
schemaWidget.callback = update;
|
||||
|
||||
return result;
|
||||
};
|
||||
|
||||
@@ -0,0 +1,57 @@
|
||||
/**
|
||||
* preview_text.js — read-only output preview for comfydv's OUTPUT_NODE=True
|
||||
* nodes that return a ComfyUI "ui": {"text": [...]} payload
|
||||
* (ChatCompletion, FormatString, RandomChoice).
|
||||
*
|
||||
* ComfyUI does NOT auto-render an arbitrary node's ui.text — each node type
|
||||
* that wants one implements its own onExecuted handler. This mirrors core's
|
||||
* own ``PreviewAny`` node (comfy_extras/nodes_preview_any.py +
|
||||
* "Comfy.PreviewAny" in the frontend bundle) minus its Markdown/Plaintext
|
||||
* toggle, which none of these three nodes need.
|
||||
*/
|
||||
|
||||
import { app } from "../../scripts/app.js";
|
||||
import { ComfyWidgets } from "../../scripts/widgets.js";
|
||||
|
||||
const PREVIEW_NODES = new Set(["ChatCompletion", "FormatString", "RandomChoice"]);
|
||||
|
||||
app.registerExtension({
|
||||
name: "comfydv.previewText",
|
||||
|
||||
async beforeRegisterNodeDef(nodeType, nodeData) {
|
||||
if (!PREVIEW_NODES.has(nodeData.name)) return;
|
||||
|
||||
const onNodeCreated = nodeType.prototype.onNodeCreated;
|
||||
nodeType.prototype.onNodeCreated = function () {
|
||||
const result = onNodeCreated?.apply(this, arguments);
|
||||
|
||||
const widget = ComfyWidgets.STRING(
|
||||
this,
|
||||
"comfydv_preview_text",
|
||||
["STRING", { multiline: true }],
|
||||
app
|
||||
).widget;
|
||||
widget.label = "Preview";
|
||||
widget.options.read_only = true;
|
||||
// Not a real input — nothing to save/replay in the saved
|
||||
// workflow JSON, and read-only anyway.
|
||||
widget.options.serialize = false;
|
||||
widget.serialize = false;
|
||||
widget.inputEl.readOnly = true;
|
||||
|
||||
return result;
|
||||
};
|
||||
|
||||
const onExecuted = nodeType.prototype.onExecuted;
|
||||
nodeType.prototype.onExecuted = function (message) {
|
||||
onExecuted?.apply(this, arguments);
|
||||
|
||||
const widget = this.widgets?.find(w => w.name === "comfydv_preview_text");
|
||||
if (!widget) return;
|
||||
|
||||
const text = message?.text ?? "";
|
||||
widget.value = Array.isArray(text) ? (text.join("\n\n") ?? "") : text;
|
||||
this.setDirtyCanvas(true, true);
|
||||
};
|
||||
},
|
||||
});
|
||||
@@ -60,6 +60,15 @@ def pytest_configure(config):
|
||||
|
||||
sys.modules["folder_paths"] = MockFolderPaths
|
||||
|
||||
# Force "comfydv" to resolve to src/comfydv and get cached in sys.modules now,
|
||||
# while our sys.path.insert(0, ...) above is still the definitive answer. The
|
||||
# repo root's own __init__.py (ComfyUI's custom-node entry point) is also a
|
||||
# valid "comfydv" package from certain sys.path states pytest transiently
|
||||
# constructs during fixture setup; without this, a later bare `import comfydv`
|
||||
# (e.g. in the _clear_ollama_caches fixture) can resolve to that root package
|
||||
# instead, which lacks the _llm submodule and fails with ModuleNotFoundError.
|
||||
import comfydv # noqa: F401
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Ollama fixtures (used by @pytest.mark.integration tests)
|
||||
@@ -68,20 +77,35 @@ def pytest_configure(config):
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_ollama_caches():
|
||||
"""Reset comfydv.ollama's module-level LRU caches around every test.
|
||||
"""Reset the shared LLM provider caches and ChatCompletion's dynamic
|
||||
RETURN_TYPES/RETURN_NAMES around every test.
|
||||
|
||||
Several tests reuse identical client/model/prompt inputs across cases
|
||||
with different monkeypatched responses — without this, a later test would
|
||||
silently get an earlier test's cached result instead of exercising its
|
||||
own fake.
|
||||
"""
|
||||
from comfydv.ollama import _CHAT_RESPONSE_CACHE, _MODEL_LIST_CACHE
|
||||
own fake. RETURN_TYPES/RETURN_NAMES are class-level mutable state (set by
|
||||
ChatCompletion.update_outputs for structured_output mode) shared across
|
||||
every test in the module — without resetting them, a structured-output
|
||||
test would leak its dynamic outputs into unrelated tests that assert the
|
||||
fixed 3-tuple.
|
||||
|
||||
_MODEL_LIST_CACHE.clear()
|
||||
_CHAT_RESPONSE_CACHE.clear()
|
||||
Caches live in comfydv._llm.ollama_provider (ADR-007's single source of
|
||||
truth) — comfydv.ollama's combo-widget helpers (_fetch_models) share the
|
||||
same cache instance, not a separate copy.
|
||||
"""
|
||||
from comfydv._llm.ollama_provider import _CHAT_RESPONSE_CACHE, _MODEL_LIST_CACHE
|
||||
from comfydv.ollama import ChatCompletion
|
||||
|
||||
def _reset():
|
||||
_MODEL_LIST_CACHE.clear()
|
||||
_CHAT_RESPONSE_CACHE.clear()
|
||||
ChatCompletion.RETURN_TYPES = ChatCompletion._BASE_RETURN_TYPES
|
||||
ChatCompletion.RETURN_NAMES = ChatCompletion._BASE_RETURN_NAMES
|
||||
ChatCompletion.node_configs.clear()
|
||||
|
||||
_reset()
|
||||
yield
|
||||
_MODEL_LIST_CACHE.clear()
|
||||
_CHAT_RESPONSE_CACHE.clear()
|
||||
_reset()
|
||||
|
||||
|
||||
@pytest.fixture(scope="session")
|
||||
|
||||
@@ -0,0 +1,148 @@
|
||||
"""Guards against the exact bug found while validating spec 008 against a
|
||||
real ComfyUI dev harness (docker-compose): every `from comfydv._llm.X
|
||||
import Y`-style absolute self-import inside src/comfydv/ silently broke the
|
||||
*entire* plugin (every node, not just LLM ones) as soon as ComfyUI actually
|
||||
loaded it.
|
||||
|
||||
ComfyUI's custom_nodes loader imports the plugin via a *relative* chain —
|
||||
the repo-root __init__.py does `from .src.comfydv import ...`, nesting
|
||||
comfydv under whatever top-level name the folder has (never `comfydv`
|
||||
itself). An absolute `from comfydv...` self-import only resolves if `src/`
|
||||
has separately been placed on sys.path — which conftest.py does for every
|
||||
other test file in this suite, masking the bug completely. This file
|
||||
deliberately does NOT rely on that sys.path insertion: it reproduces
|
||||
ComfyUI's actual nested-relative-import shape in a subprocess.
|
||||
|
||||
Confirmed via git history: this predates spec 008 entirely — it was already
|
||||
broken immediately after PR #17 merged (spec 007), well before llamacpp.py
|
||||
existed. No test caught it because none exercised this exact loading shape
|
||||
until the docker harness was run by hand.
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import sys
|
||||
import textwrap
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
REPO_ROOT = Path(__file__).parent.parent
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_ollama_caches():
|
||||
"""Shadow conftest.py's autouse fixture of the same name for this module
|
||||
only. That fixture's own setup does `from comfydv._llm.ollama_provider
|
||||
import ...` — this file's tests are the exact reproduction of an
|
||||
environment where `comfydv` resolving correctly can't be assumed (that's
|
||||
the point of the file), so depending on it for an unrelated cache-reset
|
||||
would make these tests order-dependent on whichever other test file
|
||||
happens to import `comfydv` "the normal way" first in the session. These
|
||||
tests touch no OllamaProvider/ChatCompletion state, so there is nothing
|
||||
to reset."""
|
||||
yield
|
||||
|
||||
|
||||
_SUBPROCESS_SCRIPT = textwrap.dedent(
|
||||
"""
|
||||
import sys
|
||||
import types
|
||||
|
||||
# Minimal ComfyUI stubs — same shape as conftest.py's pytest_configure,
|
||||
# but this script intentionally runs outside pytest so it isn't reusing
|
||||
# (or accidentally validated by) that fixture's sys.path setup.
|
||||
class _InterruptProcessingException(Exception):
|
||||
pass
|
||||
|
||||
comfy_module = types.ModuleType("comfy")
|
||||
comfy_module.model_management = types.SimpleNamespace(
|
||||
InterruptProcessingException=_InterruptProcessingException
|
||||
)
|
||||
sys.modules["comfy"] = comfy_module
|
||||
sys.modules["comfy.model_management"] = comfy_module.model_management
|
||||
|
||||
class _Routes:
|
||||
def post(self, path):
|
||||
return lambda fn: fn
|
||||
|
||||
def get(self, path):
|
||||
return lambda fn: fn
|
||||
|
||||
class _PromptServer:
|
||||
pass
|
||||
|
||||
_PromptServer.instance = _PromptServer()
|
||||
_PromptServer.instance.routes = _Routes()
|
||||
server_module = types.ModuleType("server")
|
||||
server_module.PromptServer = _PromptServer
|
||||
sys.modules["server"] = server_module
|
||||
|
||||
folder_paths_module = types.ModuleType("folder_paths")
|
||||
folder_paths_module.get_output_directory = lambda: "/tmp/comfydv_test"
|
||||
sys.modules["folder_paths"] = folder_paths_module
|
||||
|
||||
# The critical part: put the repo's *parent* directory on sys.path, so
|
||||
# `import comfydv` resolves to the repo-root __init__.py — exactly how
|
||||
# ComfyUI resolves a folder under custom_nodes/ — NOT to src/comfydv
|
||||
# directly (that's what conftest.py's sys.path.insert(0, ".../src")
|
||||
# does for the rest of this test suite, and why it never caught this).
|
||||
sys.path.insert(0, sys.argv[1])
|
||||
|
||||
import comfydv
|
||||
|
||||
required = {"FormatString", "RandomChoice", "CircuitBreaker",
|
||||
"OllamaClient", "LlamaCppClient", "ChatCompletion"}
|
||||
missing = required - set(comfydv.NODE_CLASS_MAPPINGS)
|
||||
if missing:
|
||||
print(f"MISSING_NODES:{missing}")
|
||||
sys.exit(1)
|
||||
print("OK")
|
||||
"""
|
||||
)
|
||||
|
||||
|
||||
def test_package_imports_under_comfyui_style_relative_nesting():
|
||||
"""Reproduces ComfyUI's real loading shape and fails loudly — with the
|
||||
actual traceback — if any internal module reverts to an absolute
|
||||
`from comfydv...` self-import."""
|
||||
result = subprocess.run(
|
||||
[sys.executable, "-c", _SUBPROCESS_SCRIPT, str(REPO_ROOT.parent)],
|
||||
cwd=REPO_ROOT,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=30,
|
||||
)
|
||||
assert result.returncode == 0, (
|
||||
"comfydv failed to import the way ComfyUI actually loads it "
|
||||
"(relative nesting, not a top-level `comfydv` on sys.path). "
|
||||
f"This means every node in the plugin would fail to register.\n"
|
||||
f"--- stdout ---\n{result.stdout}\n--- stderr ---\n{result.stderr}"
|
||||
)
|
||||
assert "OK" in result.stdout
|
||||
|
||||
|
||||
def test_no_absolute_self_imports_in_package():
|
||||
"""Cheap, fast static guard alongside the dynamic test above: no file
|
||||
under src/comfydv/ should import itself as `comfydv.X` — internal
|
||||
imports must be relative (`.X` / `..X`) so they resolve regardless of
|
||||
what the outer package happens to be named at load time."""
|
||||
import ast
|
||||
|
||||
offenders = []
|
||||
for path in (REPO_ROOT / "src" / "comfydv").rglob("*.py"):
|
||||
tree = ast.parse(path.read_text(), filename=str(path))
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom):
|
||||
if node.module and (
|
||||
node.module == "comfydv" or node.module.startswith("comfydv.")
|
||||
):
|
||||
offenders.append(f"{path.relative_to(REPO_ROOT)}:{node.lineno}")
|
||||
elif isinstance(node, ast.Import):
|
||||
for alias in node.names:
|
||||
if alias.name == "comfydv" or alias.name.startswith("comfydv."):
|
||||
offenders.append(f"{path.relative_to(REPO_ROOT)}:{node.lineno}")
|
||||
|
||||
assert not offenders, (
|
||||
"Absolute self-imports found — use relative imports instead "
|
||||
f"(they break under ComfyUI's actual loader): {offenders}"
|
||||
)
|
||||
@@ -35,9 +35,32 @@ class TestVariableExtraction:
|
||||
def test_extract_jinja2_with_multiple_filters(self, format_string_class):
|
||||
"""Test extraction of variables with multiple Jinja2 filters."""
|
||||
keys = format_string_class._extract_keys("{{ name | upper | trim }}")
|
||||
# Multiple chained filters may not extract - that's a limitation of the regex
|
||||
# Just test that it doesn't crash
|
||||
assert isinstance(keys, list)
|
||||
assert keys == ["name"]
|
||||
|
||||
def test_extract_jinja2_for_loop_excludes_loop_variable(self, format_string_class):
|
||||
"""The for-loop target (e.g. `hint`) is bound by the template and must
|
||||
not be treated as a required input; the iterable it draws from must be."""
|
||||
keys = format_string_class._extract_keys(
|
||||
"{% for hint in extraction_hints %}- {{ hint }}\n{% endfor %}"
|
||||
)
|
||||
assert keys == ["extraction_hints"]
|
||||
|
||||
def test_extract_jinja2_if_condition_variable(self, format_string_class):
|
||||
"""A variable referenced only in an {% if %} condition must still be
|
||||
detected, even without a matching {{ }} expression elsewhere."""
|
||||
keys = format_string_class._extract_keys(
|
||||
"{% if extraction_hints is defined and extraction_hints %}yes{% endif %}"
|
||||
)
|
||||
assert keys == ["extraction_hints"]
|
||||
|
||||
def test_extract_jinja2_filter_with_arguments(self, format_string_class):
|
||||
"""A filter called with arguments (e.g. tojson(indent=2)) has parens
|
||||
in the way of the old regex's anchor to the closing }} — the variable
|
||||
must still be detected."""
|
||||
keys = format_string_class._extract_keys(
|
||||
"{{ scene_manifest | tojson(indent=2) }}"
|
||||
)
|
||||
assert keys == ["scene_manifest"]
|
||||
|
||||
def test_extract_jinja2_multiple_variables(self, format_string_class):
|
||||
"""Test extraction of multiple variables from Jinja2 template."""
|
||||
@@ -178,6 +201,18 @@ class TestJinja2Formatting:
|
||||
assert result[2] == "John"
|
||||
assert result[3] == "Doe"
|
||||
|
||||
def test_jinja2_fromjson_filter_parses_array(self, format_string_class):
|
||||
"""ComfyUI has no native list socket, so a STRING input carrying a
|
||||
JSON array must be parseable back into a real list for iteration."""
|
||||
result = format_string_class.format_string(
|
||||
template_type="Jinja2",
|
||||
template="{% for hint in hints | fromjson %}- {{ hint }}\n{% endfor %}",
|
||||
save_path="",
|
||||
unique_id="test-fromjson",
|
||||
hints='["motion", "camera pan"]',
|
||||
)["result"]
|
||||
assert result[0] == "- motion\n- camera pan\n"
|
||||
|
||||
def test_jinja2_with_datetime(self, format_string_class):
|
||||
"""Test Jinja2 formatting with datetime context."""
|
||||
result = format_string_class.format_string(
|
||||
@@ -199,10 +234,11 @@ class TestJinja2Formatting:
|
||||
unique_id="test9",
|
||||
value=sample_data["value"],
|
||||
)["result"]
|
||||
# value is not extracted as a variable because it's used in an expression
|
||||
assert len(result) == 2 # Just formatted_string, saved_file_path
|
||||
# value is extracted even though it's used in an expression
|
||||
assert len(result) == 3 # formatted_string, saved_file_path, value
|
||||
assert result[0] == "Result: 10"
|
||||
assert result[1] == ""
|
||||
assert result[2] == str(sample_data["value"])
|
||||
|
||||
|
||||
class TestInlineDisplay:
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
"""
|
||||
Tests for comfydv.llamacpp.LlamaCppClient — the one new ComfyUI node this
|
||||
feature introduces. Also proves the adapter pattern end-to-end (US4): the
|
||||
same generic nodes work unmodified against either provider.
|
||||
|
||||
BDD coverage:
|
||||
../specs/008-llamacpp-integration/features/us1_connect_and_chat.feature
|
||||
../specs/008-llamacpp-integration/features/us4_swap_backends.feature
|
||||
"""
|
||||
|
||||
from comfydv._llm.llamacpp_provider import LlamaCppProvider
|
||||
from comfydv._llm.ollama_provider import OllamaProvider
|
||||
from comfydv.llamacpp import LlamaCppClient
|
||||
from comfydv.ollama import (
|
||||
ChatCompletion,
|
||||
LLMLoadModel,
|
||||
LLMModelSelector,
|
||||
LLMUnloadModel,
|
||||
)
|
||||
|
||||
|
||||
class _FakeProvider:
|
||||
"""Mirrors tests/test_ollama.py's _FakeProvider — reused here for US4's
|
||||
swap-backends proof rather than duplicated, since the whole point is
|
||||
that node behavior doesn't depend on which concrete provider it gets."""
|
||||
|
||||
def __init__(self, chat_response="ok"):
|
||||
self.chat_response = chat_response
|
||||
self.models = [{"name": "m", "status": "loaded"}]
|
||||
self.calls: list[tuple] = []
|
||||
|
||||
async def list_models(self):
|
||||
self.calls.append(("list_models",))
|
||||
return self.models
|
||||
|
||||
async def load_model(self, model):
|
||||
self.calls.append(("load_model", model))
|
||||
|
||||
async def unload_model(self, model):
|
||||
self.calls.append(("unload_model", model))
|
||||
|
||||
async def chat(
|
||||
self,
|
||||
model,
|
||||
messages,
|
||||
options=None,
|
||||
timeout_secs=300.0,
|
||||
max_retries=2,
|
||||
attempt_info=None,
|
||||
on_status=None,
|
||||
):
|
||||
self.calls.append(("chat", model))
|
||||
if attempt_info is not None:
|
||||
attempt_info.update(
|
||||
{
|
||||
"seed": (options or {}).get("seed", 0),
|
||||
"attempts": 1,
|
||||
"timeout_secs": timeout_secs,
|
||||
"refusals": 0,
|
||||
}
|
||||
)
|
||||
return self.chat_response
|
||||
|
||||
|
||||
def test_client_outputs_llamacpp_provider():
|
||||
(client,) = LlamaCppClient().create_client("http://localhost:8080")
|
||||
assert isinstance(client, LlamaCppProvider)
|
||||
assert client.host == "http://localhost:8080"
|
||||
|
||||
|
||||
def test_client_default_host_matches_llama_server_default_port():
|
||||
input_types = LlamaCppClient.INPUT_TYPES()
|
||||
assert input_types["required"]["host"][1]["default"] == "http://localhost:8080"
|
||||
|
||||
|
||||
def test_client_output_type_is_generic_llm_client():
|
||||
assert LlamaCppClient.RETURN_TYPES == ("LLM_CLIENT",)
|
||||
|
||||
|
||||
def test_client_carries_headers():
|
||||
(client,) = LlamaCppClient().create_client(
|
||||
"http://localhost:8080", headers={"Authorization": "Bearer abc"}
|
||||
)
|
||||
assert client.headers == {"Authorization": "Bearer abc"}
|
||||
|
||||
|
||||
def test_node_contract():
|
||||
assert hasattr(LlamaCppClient, "INPUT_TYPES")
|
||||
assert hasattr(LlamaCppClient, "RETURN_TYPES")
|
||||
assert hasattr(LlamaCppClient, "FUNCTION")
|
||||
assert hasattr(LlamaCppClient, "CATEGORY")
|
||||
assert hasattr(LlamaCppClient, LlamaCppClient.FUNCTION)
|
||||
|
||||
|
||||
def test_registered_in_node_class_mappings():
|
||||
from comfydv import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS
|
||||
|
||||
assert NODE_CLASS_MAPPINGS["LlamaCppClient"] is LlamaCppClient
|
||||
assert "LlamaCppClient" in NODE_DISPLAY_NAME_MAPPINGS
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# US4 — swap backends without touching downstream nodes
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _run_workflow(client) -> None:
|
||||
"""The same node sequence a workflow author would wire up, regardless
|
||||
of which provider `client` is."""
|
||||
ChatCompletion().chat(client=client, model="m", prompt="hi")
|
||||
LLMModelSelector().select_model(client=client, model="m")
|
||||
LLMLoadModel().load_model(client=client, model="m")
|
||||
LLMUnloadModel().unload_model(client=client, model="m")
|
||||
|
||||
|
||||
def test_same_workflow_runs_against_either_fake_provider():
|
||||
"""No node branches on provider type — the same call sequence succeeds
|
||||
whether client looks like an Ollama-shaped or llama.cpp-shaped provider."""
|
||||
ollama_like = _FakeProvider(chat_response="ollama says hi")
|
||||
llamacpp_like = _FakeProvider(chat_response="llamacpp says hi")
|
||||
|
||||
# Neither call raises — that's the actual assertion. If ChatCompletion/
|
||||
# LLMModelSelector/LLMLoadModel/LLMUnloadModel secretly special-cased a
|
||||
# concrete provider type (isinstance checks, attribute probing beyond
|
||||
# the protocol), one of these would fail.
|
||||
_run_workflow(ollama_like)
|
||||
_run_workflow(llamacpp_like)
|
||||
|
||||
# LLMModelSelector is pure passthrough (client is accepted only for
|
||||
# wiring/typing, never dereferenced), so it makes no provider call.
|
||||
expected = ["chat", "load_model", "unload_model"]
|
||||
assert [c[0] for c in ollama_like.calls] == expected
|
||||
assert [c[0] for c in llamacpp_like.calls] == expected
|
||||
|
||||
|
||||
def test_real_providers_are_interchangeable_client_output():
|
||||
"""OllamaClient and LlamaCppClient both emit LLM_CLIENT — a workflow
|
||||
author can wire either one into the same downstream nodes."""
|
||||
from comfydv.ollama import OllamaClient
|
||||
|
||||
(ollama_client,) = OllamaClient().create_client("http://localhost:11434")
|
||||
(llamacpp_client,) = LlamaCppClient().create_client("http://localhost:8080")
|
||||
|
||||
assert isinstance(ollama_client, OllamaProvider)
|
||||
assert isinstance(llamacpp_client, LlamaCppProvider)
|
||||
# Both satisfy the same protocol shape — same method names available.
|
||||
for method in (
|
||||
"list_models",
|
||||
"load_model",
|
||||
"unload_model",
|
||||
"chat",
|
||||
"chat_structured",
|
||||
):
|
||||
assert callable(getattr(ollama_client, method))
|
||||
assert callable(getattr(llamacpp_client, method))
|
||||
@@ -0,0 +1,777 @@
|
||||
"""
|
||||
Tests for comfydv._llm.llamacpp_provider.LlamaCppProvider — mirrors
|
||||
test_ollama_provider.py's structure exactly (ADR-007's parallel-
|
||||
implementation pattern). Mocks at the provider's own _post_json/_get_json
|
||||
seam.
|
||||
|
||||
BDD coverage:
|
||||
../specs/008-llamacpp-integration/features/us1_connect_and_chat.feature
|
||||
../specs/008-llamacpp-integration/features/us2_structured_output.feature
|
||||
../specs/008-llamacpp-integration/features/us3_model_lifecycle.feature
|
||||
"""
|
||||
|
||||
import pytest
|
||||
|
||||
import comfydv._llm.llamacpp_provider as provider_mod
|
||||
from comfydv._llm.llamacpp_provider import LlamaCppProvider, _fetch_models
|
||||
from comfydv._llm.ollama_provider import _run_async
|
||||
from comfydv._llm.provider import Message, ModelStatus
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_provider_caches():
|
||||
provider_mod._MODEL_LIST_CACHE.clear()
|
||||
provider_mod._CHAT_RESPONSE_CACHE.clear()
|
||||
yield
|
||||
provider_mod._MODEL_LIST_CACHE.clear()
|
||||
provider_mod._CHAT_RESPONSE_CACHE.clear()
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# list_models — the "id" field name and nested "status.value" are the two
|
||||
# details research.md flagged as easy to get wrong by assumption.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_list_models_maps_id_field_to_name(monkeypatch):
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
return {"data": [{"id": "gemma-3-4b:Q4_K_M", "status": {"value": "loaded"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
(model,) = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
|
||||
|
||||
assert model.name == "gemma-3-4b:Q4_K_M"
|
||||
|
||||
|
||||
def test_list_models_reads_nested_status_value(monkeypatch):
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
return {
|
||||
"data": [
|
||||
{"id": "a", "status": {"value": "sleeping"}},
|
||||
{"id": "b", "status": {"value": "downloading", "progress": {}}},
|
||||
]
|
||||
}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
|
||||
|
||||
by_name = {m.name: m for m in models}
|
||||
assert by_name["a"].status == ModelStatus.SLEEPING
|
||||
assert by_name["b"].status == ModelStatus.DOWNLOADING
|
||||
|
||||
|
||||
def test_list_models_no_normalization_needed_full_vocabulary(monkeypatch):
|
||||
"""Unlike OllamaProvider, llama.cpp's status vocabulary is exactly
|
||||
ModelStatus's full set — every value should pass through untouched."""
|
||||
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
return {
|
||||
"data": [
|
||||
{"id": v, "status": {"value": v}}
|
||||
for v in ["unloaded", "loading", "loaded", "sleeping", "downloading"]
|
||||
]
|
||||
}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
|
||||
|
||||
assert {m.status for m in models} == set(ModelStatus)
|
||||
|
||||
|
||||
def test_list_models_skips_unrecognized_status(monkeypatch):
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
return {
|
||||
"data": [
|
||||
{"id": "crashed", "status": {"value": "failed", "exit_code": 1}},
|
||||
{"id": "ok", "status": {"value": "loaded"}},
|
||||
]
|
||||
}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
|
||||
|
||||
assert [m.name for m in models] == ["ok"]
|
||||
|
||||
|
||||
def test_list_models_unreachable_returns_empty(monkeypatch):
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
raise ConnectionError("no route to host")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
models = _run_async(LlamaCppProvider("http://localhost:19999").list_models())
|
||||
|
||||
assert models == []
|
||||
|
||||
|
||||
def test_list_models_non_router_mode_raises_clear_error(monkeypatch):
|
||||
"""FR-006: a llama-server that IS reachable but wasn't launched with
|
||||
--models-dir/--models-preset answers GET /models with an HTTP error
|
||||
(the endpoint doesn't exist outside router mode). That must surface as
|
||||
a specific, actionable error — not silently degrade to an empty list,
|
||||
which would be indistinguishable from "server has no models"."""
|
||||
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
raise RuntimeError("Server returned HTTP 404 for http://x/models: not found")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
|
||||
with pytest.raises(RuntimeError, match="router mode"):
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").list_models())
|
||||
|
||||
|
||||
def test_list_models_cached_second_call(monkeypatch):
|
||||
calls = {"n": 0}
|
||||
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
calls["n"] += 1
|
||||
return {"data": [{"id": "a", "status": {"value": "loaded"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
provider = LlamaCppProvider("http://localhost:8080")
|
||||
_run_async(provider.list_models())
|
||||
_run_async(provider.list_models())
|
||||
|
||||
assert calls["n"] == 1
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# load_model / unload_model
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_load_model_posts_to_models_load_with_model_field(monkeypatch):
|
||||
captured = {}
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
captured["url"] = url
|
||||
captured["payload"] = payload
|
||||
return {"success": True}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
|
||||
|
||||
assert captured["url"] == "http://localhost:8080/models/load"
|
||||
assert captured["payload"] == {"model": "gemma-3-4b"}
|
||||
|
||||
|
||||
def test_unload_model_posts_to_models_unload_with_model_field(monkeypatch):
|
||||
captured = {}
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
captured["url"] = url
|
||||
captured["payload"] = payload
|
||||
return {"success": True}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
|
||||
|
||||
assert captured["url"] == "http://localhost:8080/models/unload"
|
||||
assert captured["payload"] == {"model": "gemma-3-4b"}
|
||||
|
||||
|
||||
def test_load_model_already_running_is_idempotent(monkeypatch):
|
||||
"""Confirmed live: router mode's /models/load is NOT idempotent at the
|
||||
wire level — it 400s "model is already running" rather than returning
|
||||
{"success": true}. The LLMProvider protocol requires load_model() to be
|
||||
idempotent, so LlamaCppProvider must absorb this itself."""
|
||||
calls = []
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls.append((url, payload))
|
||||
raise RuntimeError(
|
||||
"Server returned HTTP 400 for "
|
||||
f'{url}: {{"error":{{"code":400,"message":"model is already '
|
||||
'running","type":"invalid_request_error"}}}}'
|
||||
)
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
|
||||
|
||||
# The absence of a raised exception is only meaningful if the request
|
||||
# actually happened and hit the "already running" branch — assert that
|
||||
# directly rather than trusting silence alone.
|
||||
assert calls == [("http://localhost:8080/models/load", {"model": "gemma-3-4b"})]
|
||||
|
||||
|
||||
def test_load_model_other_http_error_still_raises(monkeypatch):
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
raise RuntimeError(f"Server returned HTTP 500 for {url}: internal error")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
with pytest.raises(RuntimeError, match="500"):
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
|
||||
|
||||
|
||||
def test_unload_model_not_running_is_idempotent(monkeypatch):
|
||||
"""Mirror of the load_model case, confirmed live: /models/unload 400s
|
||||
"model is not running" on an already-unloaded model."""
|
||||
calls = []
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls.append((url, payload))
|
||||
raise RuntimeError(
|
||||
"Server returned HTTP 400 for "
|
||||
f'{url}: {{"error":{{"code":400,"message":"model is not '
|
||||
'running","type":"invalid_request_error"}}}}'
|
||||
)
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
|
||||
|
||||
assert calls == [("http://localhost:8080/models/unload", {"model": "gemma-3-4b"})]
|
||||
|
||||
|
||||
def test_unload_model_other_http_error_still_raises(monkeypatch):
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
raise RuntimeError(f"Server returned HTTP 500 for {url}: internal error")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
with pytest.raises(RuntimeError, match="500"):
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
|
||||
|
||||
|
||||
def test_load_model_empty_raises_before_network(monkeypatch):
|
||||
def fail_post(*a, **k):
|
||||
raise AssertionError("must not call _post_json for an empty model name")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fail_post)
|
||||
with pytest.raises(ValueError, match="cannot be empty"):
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").load_model(""))
|
||||
|
||||
|
||||
def test_unload_model_empty_raises_before_network(monkeypatch):
|
||||
def fail_post(*a, **k):
|
||||
raise AssertionError("must not call _post_json for an empty model name")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fail_post)
|
||||
with pytest.raises(ValueError, match="cannot be empty"):
|
||||
_run_async(LlamaCppProvider("http://localhost:8080").unload_model(" "))
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# chat — OpenAI response shape (choices[0].message.content), not Ollama's
|
||||
# native shape
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_chat_posts_to_v1_chat_completions_and_parses_openai_shape(monkeypatch):
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
assert url == "http://localhost:8080/v1/chat/completions"
|
||||
return {"choices": [{"message": {"role": "assistant", "content": "hello"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b", [Message(role="user", content="hi")]
|
||||
)
|
||||
)
|
||||
|
||||
assert result == "hello"
|
||||
|
||||
|
||||
def test_chat_no_choices_returns_empty_string(monkeypatch):
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
return {"choices": []}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b", [Message(role="user", content="hi")], max_retries=0
|
||||
)
|
||||
)
|
||||
|
||||
assert result == ""
|
||||
|
||||
|
||||
def test_chat_second_identical_call_is_cached(monkeypatch):
|
||||
calls = {"n": 0}
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls["n"] += 1
|
||||
return {"choices": [{"message": {"content": "cached"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
provider = LlamaCppProvider("http://localhost:8080")
|
||||
messages = [Message(role="user", content="hi")]
|
||||
|
||||
r1 = _run_async(provider.chat("m", messages))
|
||||
r2 = _run_async(provider.chat("m", messages))
|
||||
|
||||
assert r1 == r2 == "cached"
|
||||
assert calls["n"] == 1
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# chat — retry-on-blank-output. Mirrors test_ollama_provider.py's coverage;
|
||||
# the one llama.cpp-specific detail is that the retry seed must land in the
|
||||
# top-level OpenAI "seed" field, not nested under "options" (see chat()'s
|
||||
# comment on why the options passthrough doesn't reach llama-server at all).
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
async def _fake_sleep(_secs):
|
||||
"""No-op stand-in for asyncio.sleep — keeps retry tests instant."""
|
||||
|
||||
|
||||
def test_chat_retries_on_blank_response_and_returns_second_attempt(monkeypatch):
|
||||
calls = []
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls.append(payload)
|
||||
if len(calls) == 1:
|
||||
return {"choices": [{"message": {"content": ""}}]}
|
||||
return {"choices": [{"message": {"content": "real answer"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
|
||||
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b", [Message(role="user", content="hi")]
|
||||
)
|
||||
)
|
||||
|
||||
assert result == "real answer"
|
||||
assert len(calls) == 2
|
||||
|
||||
|
||||
def test_chat_retry_seed_is_top_level_not_nested_in_options(monkeypatch):
|
||||
calls = []
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls.append(payload)
|
||||
if len(calls) < 3:
|
||||
return {"choices": [{"message": {"content": ""}}]}
|
||||
return {"choices": [{"message": {"content": "ok"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
|
||||
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b", [Message(role="user", content="hi")], max_retries=2
|
||||
)
|
||||
)
|
||||
|
||||
assert "seed" not in calls[0]
|
||||
assert calls[1]["seed"] == 1
|
||||
assert calls[2]["seed"] == 2
|
||||
|
||||
|
||||
def test_chat_disable_thinking_sets_chat_template_kwargs_and_reasoning_effort(
|
||||
monkeypatch,
|
||||
):
|
||||
"""ADR-010: llama-server doesn't recognize a "think" key nested inside
|
||||
"options" (that's an Ollama-native convention) — it needs its own two
|
||||
documented request-body toggles instead, and "think" must not leak into
|
||||
the nested options object llama-server actually does understand."""
|
||||
captured = {}
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
captured.update(payload)
|
||||
return {"choices": [{"message": {"content": "ok"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b",
|
||||
[Message(role="user", content="hi")],
|
||||
options={"temperature": 0.5, "think": False},
|
||||
)
|
||||
)
|
||||
|
||||
assert captured["chat_template_kwargs"] == {"enable_thinking": False}
|
||||
assert captured["reasoning_effort"] == "none"
|
||||
assert captured["options"] == {"temperature": 0.5} # "think" popped out
|
||||
|
||||
|
||||
def test_chat_exhausted_retries_returns_blank_without_raising(monkeypatch):
|
||||
calls = {"n": 0}
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls["n"] += 1
|
||||
return {"choices": [{"message": {"content": ""}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
|
||||
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b", [Message(role="user", content="hi")], max_retries=2
|
||||
)
|
||||
)
|
||||
|
||||
assert result == ""
|
||||
assert calls["n"] == 3 # original + 2 retries, per max_retries=2
|
||||
|
||||
|
||||
def test_chat_timeout_escalates_per_retry_attempt(monkeypatch):
|
||||
timeouts = []
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
timeouts.append(timeout)
|
||||
if len(timeouts) < 3:
|
||||
return {"choices": [{"message": {"content": ""}}]}
|
||||
return {"choices": [{"message": {"content": "done"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
|
||||
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b",
|
||||
[Message(role="user", content="hi")],
|
||||
timeout_secs=50.0,
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
assert timeouts == [50.0, 100.0, 150.0]
|
||||
|
||||
|
||||
def test_chat_attempt_info_populated_on_success(monkeypatch):
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
return {"choices": [{"message": {"content": "ok"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
|
||||
attempt_info: dict = {}
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b",
|
||||
[Message(role="user", content="hi")],
|
||||
options={"seed": 9},
|
||||
attempt_info=attempt_info,
|
||||
)
|
||||
)
|
||||
|
||||
assert attempt_info == {
|
||||
"seed": 9,
|
||||
"attempts": 1,
|
||||
"timeout_secs": 300.0,
|
||||
"refusals": 0,
|
||||
}
|
||||
|
||||
|
||||
def test_chat_on_status_called_on_retry_and_recovery(monkeypatch):
|
||||
calls = []
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls.append(payload)
|
||||
if len(calls) == 1:
|
||||
return {"choices": [{"message": {"content": "I cannot generate that."}}]}
|
||||
return {"choices": [{"message": {"content": "a real, on-topic answer"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
|
||||
|
||||
statuses = []
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b",
|
||||
[Message(role="user", content="hi")],
|
||||
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
|
||||
on_status=statuses.append,
|
||||
)
|
||||
)
|
||||
|
||||
assert len(statuses) == 2
|
||||
assert "Refusal/deflection detected" in statuses[0]
|
||||
assert "Recovered" in statuses[1]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# chat_structured — zero new logic, delegates to the shared helper unchanged
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_chat_structured_builds_v1_base_url_and_delegates(monkeypatch):
|
||||
from pydantic import BaseModel
|
||||
|
||||
class Widget(BaseModel):
|
||||
name: str
|
||||
|
||||
captured = {}
|
||||
|
||||
async def fake_chat_structured(**kwargs):
|
||||
captured.update(kwargs)
|
||||
return Widget(name="x")
|
||||
|
||||
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
|
||||
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat_structured(
|
||||
"gemma-3-4b", [Message(role="user", content="hi")], Widget
|
||||
)
|
||||
)
|
||||
|
||||
assert result == Widget(name="x")
|
||||
assert captured["base_url"] == "http://localhost:8080/v1"
|
||||
assert captured["model"] == "gemma-3-4b"
|
||||
|
||||
|
||||
def test_chat_structured_forwards_options(monkeypatch):
|
||||
from pydantic import BaseModel
|
||||
|
||||
class Widget(BaseModel):
|
||||
name: str
|
||||
|
||||
captured = {}
|
||||
|
||||
async def fake_chat_structured(**kwargs):
|
||||
captured.update(kwargs)
|
||||
return Widget(name="x")
|
||||
|
||||
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
|
||||
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat_structured(
|
||||
"gemma-3-4b",
|
||||
[Message(role="user", content="hi")],
|
||||
Widget,
|
||||
options={"temperature": 0.0},
|
||||
)
|
||||
)
|
||||
|
||||
assert captured["options"] == {"temperature": 0.0}
|
||||
|
||||
|
||||
def test_chat_structured_forwards_attempt_info(monkeypatch):
|
||||
from pydantic import BaseModel
|
||||
|
||||
class Widget(BaseModel):
|
||||
name: str
|
||||
|
||||
captured = {}
|
||||
|
||||
async def fake_chat_structured(**kwargs):
|
||||
captured.update(kwargs)
|
||||
return Widget(name="x")
|
||||
|
||||
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
|
||||
|
||||
attempt_info: dict = {}
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat_structured(
|
||||
"gemma-3-4b",
|
||||
[Message(role="user", content="hi")],
|
||||
Widget,
|
||||
attempt_info=attempt_info,
|
||||
)
|
||||
)
|
||||
|
||||
# Same object passed straight through — the shared helper populates it,
|
||||
# this provider doesn't need to know its shape.
|
||||
assert captured["attempt_info"] is attempt_info
|
||||
|
||||
|
||||
def test_chat_structured_forwards_on_status(monkeypatch):
|
||||
from pydantic import BaseModel
|
||||
|
||||
class Widget(BaseModel):
|
||||
name: str
|
||||
|
||||
captured = {}
|
||||
|
||||
async def fake_chat_structured(**kwargs):
|
||||
captured.update(kwargs)
|
||||
return Widget(name="x")
|
||||
|
||||
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
|
||||
|
||||
def on_status(msg):
|
||||
pass
|
||||
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat_structured(
|
||||
"gemma-3-4b",
|
||||
[Message(role="user", content="hi")],
|
||||
Widget,
|
||||
on_status=on_status,
|
||||
)
|
||||
)
|
||||
|
||||
assert captured["on_status"] is on_status
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# _fetch_models — name-only view used by ComfyUI's /dv/ollama/models?backend=
|
||||
# llamacpp route (the JS refresh button / node-creation auto-populate).
|
||||
# Deliberately more forgiving than list_models(): degrades to [] on any
|
||||
# failure rather than raising on a non-router-mode server, matching the
|
||||
# combo-widget UX OllamaProvider's own _fetch_models already gives.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_fetch_models_returns_name_only_list(monkeypatch):
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
return {
|
||||
"data": [
|
||||
{"id": "a", "status": {"value": "loaded"}},
|
||||
{"id": "b", "status": {"value": "unloaded"}},
|
||||
]
|
||||
}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
names = _run_async(_fetch_models("http://localhost:8080"))
|
||||
|
||||
assert names == ["a", "b"]
|
||||
|
||||
|
||||
def test_fetch_models_degrades_to_empty_on_non_router_mode(monkeypatch):
|
||||
"""Unlike list_models() (FR-006), this combo-population view swallows
|
||||
even the non-router-mode error — a quiet empty dropdown, not a toast."""
|
||||
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
raise RuntimeError("Server returned HTTP 404 for http://x/models: not found")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
names = _run_async(_fetch_models("http://localhost:8080"))
|
||||
|
||||
assert names == []
|
||||
|
||||
|
||||
def test_fetch_models_degrades_to_empty_when_unreachable(monkeypatch):
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
raise ConnectionError("no route to host")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
names = _run_async(_fetch_models("http://localhost:19999"))
|
||||
|
||||
assert names == []
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# chat — image input (spec 009, US2; features/us2_both_backends.feature)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_chat_maps_images_to_openai_content_parts(monkeypatch):
|
||||
"""llama.cpp's OpenAI-compatible endpoint takes images as image_url parts
|
||||
inside content, not a flat images field (ADR-008)."""
|
||||
captured = {}
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
captured["payload"] = payload
|
||||
return {"choices": [{"message": {"content": "a red square"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"m", [Message(role="user", content="describe", images=["QUJD"])]
|
||||
)
|
||||
)
|
||||
|
||||
assert captured["payload"]["messages"][-1] == {
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "text", "text": "describe"},
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {"url": "data:image/png;base64,QUJD"},
|
||||
},
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
def test_chat_text_only_content_stays_plain_string(monkeypatch):
|
||||
"""FR-003/SC-004: an image-less message keeps a plain string content,
|
||||
byte-identical to today (no content-parts, no images key)."""
|
||||
captured = {}
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
captured["payload"] = payload
|
||||
return {"choices": [{"message": {"content": "ok"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
_run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"m", [Message(role="user", content="hi")]
|
||||
)
|
||||
)
|
||||
|
||||
assert captured["payload"]["messages"] == [{"role": "user", "content": "hi"}]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# refusal-retry — parity with test_ollama_provider.py's coverage (ADR: this
|
||||
# is a model-behavior concern, not a backend one — both providers wire the
|
||||
# same comfydv._llm.retry.is_refusal() into their own retry loop).
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_chat_retries_on_lexical_refusal_and_returns_clean_second_attempt(
|
||||
monkeypatch,
|
||||
):
|
||||
calls = []
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls.append(payload)
|
||||
if len(calls) == 1:
|
||||
return {"choices": [{"message": {"content": "I cannot generate that."}}]}
|
||||
return {"choices": [{"message": {"content": "a real, on-topic answer"}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
|
||||
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b",
|
||||
[Message(role="user", content="hi")],
|
||||
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
|
||||
)
|
||||
)
|
||||
|
||||
assert result == "a real, on-topic answer"
|
||||
assert len(calls) == 2
|
||||
|
||||
|
||||
def test_chat_refusal_retry_disabled_returns_refusal_text_unchanged(monkeypatch):
|
||||
calls = []
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
calls.append(payload)
|
||||
return {"choices": [{"message": {"content": "I cannot generate that."}}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").chat(
|
||||
"gemma-3-4b", [Message(role="user", content="hi")]
|
||||
)
|
||||
)
|
||||
|
||||
assert result == "I cannot generate that."
|
||||
assert len(calls) == 1
|
||||
|
||||
|
||||
def test_embed_returns_vector_from_v1_embeddings(monkeypatch):
|
||||
captured = {}
|
||||
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
captured["url"] = url
|
||||
captured["payload"] = payload
|
||||
return {"data": [{"embedding": [0.4, 0.5, 0.6], "index": 0}]}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").embed("nomic-embed-text", "hello")
|
||||
)
|
||||
|
||||
assert result == [0.4, 0.5, 0.6]
|
||||
assert captured["url"] == "http://localhost:8080/v1/embeddings"
|
||||
assert captured["payload"] == {"model": "nomic-embed-text", "input": "hello"}
|
||||
|
||||
|
||||
def test_embed_returns_none_when_no_embedding_model_configured(monkeypatch):
|
||||
async def fake_post(url, payload, *, timeout=120.0, headers=None):
|
||||
raise RuntimeError("llama-server returned HTTP 404 for /v1/embeddings")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
|
||||
|
||||
result = _run_async(
|
||||
LlamaCppProvider("http://localhost:8080").embed("nomic-embed-text", "hello")
|
||||
)
|
||||
|
||||
assert result is None
|
||||
@@ -0,0 +1,704 @@
|
||||
"""
|
||||
Tests for comfydv._llm.chat.chat_structured — shared pydantic-ai backed
|
||||
structured output, used by every LLMProvider implementation (ADR-007).
|
||||
|
||||
Mocks at the comfydv._llm.chat._build_agent seam (returns a fake agent
|
||||
exposing an async .run()), mirroring test_ollama.py's existing convention
|
||||
of monkeypatching the module-level HTTP seam rather than the network itself.
|
||||
Uses _run_async (same helper comfydv.ollama uses) to drive the coroutine
|
||||
synchronously, matching this project's existing test style rather than
|
||||
introducing a pytest-asyncio dependency.
|
||||
|
||||
BDD coverage:
|
||||
../specs/007-llm-provider-abstraction/features/us2_structured_output.feature
|
||||
"""
|
||||
|
||||
from dataclasses import dataclass
|
||||
|
||||
import pytest
|
||||
from pydantic import BaseModel, ValidationError
|
||||
|
||||
import comfydv._llm.chat as chat_mod
|
||||
from comfydv._llm.ollama_provider import _run_async
|
||||
from comfydv._llm.provider import Message
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _no_retry_backoff(monkeypatch):
|
||||
"""Keep the real RETRY_BACKOFF_SECS delay out of this suite's wall-clock
|
||||
time for every test except the ones that specifically assert on it
|
||||
(which re-monkeypatch locally, overriding this)."""
|
||||
|
||||
async def _instant_sleep(_secs):
|
||||
pass
|
||||
|
||||
monkeypatch.setattr(chat_mod.asyncio, "sleep", _instant_sleep)
|
||||
|
||||
|
||||
class _Widget(BaseModel):
|
||||
name: str
|
||||
count: int
|
||||
|
||||
|
||||
@dataclass
|
||||
class _FakeResult:
|
||||
output: object
|
||||
|
||||
|
||||
class _FakeAgent:
|
||||
"""Stand-in for pydantic_ai.Agent — .run() is scripted per test."""
|
||||
|
||||
def __init__(self, responses):
|
||||
self._responses = list(responses)
|
||||
self.calls = []
|
||||
|
||||
async def run(self, prompt, *, message_history=None, model_settings=None):
|
||||
self.calls.append((prompt, message_history, model_settings))
|
||||
outcome = self._responses.pop(0)
|
||||
if isinstance(outcome, Exception):
|
||||
raise outcome
|
||||
return _FakeResult(output=outcome)
|
||||
|
||||
|
||||
def _messages(*, system=None, history=None, prompt="hi"):
|
||||
msgs = []
|
||||
if system:
|
||||
msgs.append(Message(role="system", content=system))
|
||||
for role, content in history or []:
|
||||
msgs.append(Message(role=role, content=content))
|
||||
msgs.append(Message(role="user", content=prompt))
|
||||
return msgs
|
||||
|
||||
|
||||
def test_chat_structured_returns_validated_output(monkeypatch):
|
||||
fake = _FakeAgent([_Widget(name="a", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
result = _run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(prompt="describe a widget"),
|
||||
schema=_Widget,
|
||||
)
|
||||
)
|
||||
|
||||
assert result == _Widget(name="a", count=1)
|
||||
assert fake.calls[0][0] == "describe a widget"
|
||||
|
||||
|
||||
def test_build_agent_uses_native_output_not_tool_calling(monkeypatch):
|
||||
"""ADR-009: the Agent must be built with NativeOutput (response_format /
|
||||
JSON-schema constrained decoding), not pydantic-ai's tool-calling
|
||||
default. Regression guard against reverting to a bare `output_type=schema`,
|
||||
which let a thinking-capable model exhaust its token budget on reasoning
|
||||
and never emit a tool call (confirmed live against a real Ollama server).
|
||||
|
||||
Spies on the real pydantic_ai.Agent constructor (only _build_agent's own
|
||||
seam is mocked in every other test in this file) and asserts on its
|
||||
output_type argument via the public NativeOutput marker class, rather
|
||||
than pydantic-ai's private internal schema representation."""
|
||||
from pydantic_ai import NativeOutput
|
||||
|
||||
captured = {}
|
||||
real_agent_cls = chat_mod.Agent
|
||||
|
||||
class _SpyAgent(real_agent_cls):
|
||||
def __init__(self, *args, **kwargs):
|
||||
captured.update(kwargs)
|
||||
super().__init__(*args, **kwargs)
|
||||
|
||||
monkeypatch.setattr(chat_mod, "Agent", _SpyAgent)
|
||||
|
||||
chat_mod._build_agent(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
schema=_Widget,
|
||||
headers=None,
|
||||
timeout_secs=300.0,
|
||||
)
|
||||
|
||||
output_type = captured["output_type"]
|
||||
assert isinstance(output_type, NativeOutput)
|
||||
assert output_type.outputs == _Widget
|
||||
|
||||
|
||||
def test_chat_structured_retries_on_validation_failure(monkeypatch):
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, _Widget(name="b", count=2)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
result = _run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
assert result == _Widget(name="b", count=2)
|
||||
assert len(fake.calls) == 2
|
||||
|
||||
|
||||
def test_chat_structured_timeout_escalates_per_retry_attempt(monkeypatch):
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, bad, _Widget(name="b", count=2)])
|
||||
build_calls = []
|
||||
|
||||
def fake_build_agent(**kw):
|
||||
build_calls.append(kw["timeout_secs"])
|
||||
return fake
|
||||
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", fake_build_agent)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
timeout_secs=100.0,
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
assert build_calls == [100.0, 200.0, 300.0]
|
||||
|
||||
|
||||
def test_chat_structured_attempt_info_populated_on_success(monkeypatch):
|
||||
fake = _FakeAgent([_Widget(name="a", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
attempt_info: dict = {}
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={"seed": 5},
|
||||
attempt_info=attempt_info,
|
||||
)
|
||||
)
|
||||
|
||||
assert attempt_info == {
|
||||
"seed": 5,
|
||||
"attempts": 1,
|
||||
"timeout_secs": 300.0,
|
||||
"refusals": 0,
|
||||
}
|
||||
|
||||
|
||||
def test_chat_structured_attempt_info_populated_on_exhaustion(monkeypatch):
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, bad])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
attempt_info: dict = {}
|
||||
with pytest.raises(RuntimeError):
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
max_retries=1,
|
||||
attempt_info=attempt_info,
|
||||
)
|
||||
)
|
||||
|
||||
assert attempt_info["attempts"] == 2
|
||||
|
||||
|
||||
def test_chat_structured_on_status_called_on_retry_and_recovery(monkeypatch):
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, _Widget(name="b", count=2)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
statuses = []
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
max_retries=2,
|
||||
on_status=statuses.append,
|
||||
)
|
||||
)
|
||||
|
||||
assert len(statuses) == 2
|
||||
assert "Structured output failed" in statuses[0]
|
||||
assert "attempt 1/3" in statuses[0]
|
||||
assert "Recovered" in statuses[1]
|
||||
assert "attempt 2/3" in statuses[1]
|
||||
|
||||
|
||||
def test_chat_structured_exhausted_retries_raises_runtime_error(monkeypatch):
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, bad, bad]) # max_retries=2 -> 3 total attempts
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
with pytest.raises(RuntimeError) as exc_info:
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
message = str(exc_info.value)
|
||||
assert "llama3" in message
|
||||
assert "3 attempt(s)" in message
|
||||
assert len(fake.calls) == 3
|
||||
|
||||
|
||||
def test_chat_structured_max_retries_clamped_to_five(monkeypatch):
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad] * 6)
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
with pytest.raises(RuntimeError, match=r"6 attempt\(s\)"):
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
max_retries=999, # clamped to 5 -> 6 total attempts
|
||||
)
|
||||
)
|
||||
assert len(fake.calls) == 6
|
||||
|
||||
|
||||
def test_chat_structured_forwards_options_as_extra_body(monkeypatch):
|
||||
"""Regression guard: options (Ollama-native sampling params set via the
|
||||
OllamaOption* nodes — temperature, seed, num_predict, repeat_penalty,
|
||||
etc.) must reach the request, not be silently dropped in structured
|
||||
mode. Forwarded verbatim via pydantic-ai's model_settings.extra_body,
|
||||
matching the pre-ADR-007 payload shape exactly (no lossy remapping onto
|
||||
ModelSettings' own standardized field names)."""
|
||||
fake = _FakeAgent([_Widget(name="a", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={"temperature": 0.0, "seed": 42, "num_predict": 128},
|
||||
)
|
||||
)
|
||||
|
||||
assert fake.calls[0][2] == {
|
||||
"extra_body": {"options": {"temperature": 0.0, "seed": 42, "num_predict": 128}}
|
||||
}
|
||||
|
||||
|
||||
def test_chat_structured_no_options_means_no_model_settings(monkeypatch):
|
||||
fake = _FakeAgent([_Widget(name="a", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
)
|
||||
)
|
||||
|
||||
assert fake.calls[0][2] is None
|
||||
|
||||
|
||||
def test_chat_structured_disable_thinking_sets_chat_template_kwargs(monkeypatch):
|
||||
"""ADR-010: llama-server's two documented reasoning toggles, applied via
|
||||
extra_body the same way options is — "think" must not leak into the
|
||||
nested options.extra_body.options object llama-server's native sampling
|
||||
params live in."""
|
||||
fake = _FakeAgent([_Widget(name="a", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:8080/v1",
|
||||
model="gemma-3-4b",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={"temperature": 0.0, "think": False},
|
||||
)
|
||||
)
|
||||
|
||||
assert fake.calls[0][2] == {
|
||||
"extra_body": {
|
||||
"options": {"temperature": 0.0},
|
||||
"chat_template_kwargs": {"enable_thinking": False},
|
||||
"reasoning_effort": "none",
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
def test_chat_structured_enable_thinking_skips_reasoning_effort(monkeypatch):
|
||||
fake = _FakeAgent([_Widget(name="a", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:8080/v1",
|
||||
model="gemma-3-4b",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={"think": True},
|
||||
)
|
||||
)
|
||||
|
||||
extra_body = fake.calls[0][2]["extra_body"]
|
||||
assert extra_body["chat_template_kwargs"] == {"enable_thinking": True}
|
||||
assert "reasoning_effort" not in extra_body
|
||||
assert "options" not in extra_body # only "think" was in options
|
||||
|
||||
|
||||
def test_chat_structured_requires_last_message_user_role():
|
||||
with pytest.raises(ValueError, match="role='user'"):
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=[Message(role="system", content="only a system message")],
|
||||
schema=_Widget,
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# retry-on-failure seed/backoff — live-verified: a freshly-loaded model's
|
||||
# first structured-output attempt can fail outright, then behave normally on
|
||||
# the very next call.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_chat_structured_retry_injects_incrementing_seed(monkeypatch):
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, bad, _Widget(name="c", count=3)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
assert fake.calls[0][2] is None # attempt 1 untouched — no options set
|
||||
assert fake.calls[1][2]["seed"] == 1
|
||||
assert fake.calls[2][2]["seed"] == 2
|
||||
|
||||
|
||||
def test_chat_structured_retry_seed_starts_from_pinned_base(monkeypatch):
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, _Widget(name="c", count=3)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={"seed": 42},
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
assert fake.calls[0][2] == {"extra_body": {"options": {"seed": 42}}}
|
||||
assert fake.calls[1][2]["seed"] == 43 # base(42) + (attempt 2 - 1)
|
||||
# The nested Ollama-native options.seed must track the same retry seed as
|
||||
# the top-level one — a backend that honors the nested field over the
|
||||
# top-level OpenAI "seed" must not keep seeing the stale pinned value.
|
||||
assert fake.calls[1][2]["extra_body"] == {"options": {"seed": 43}}
|
||||
|
||||
|
||||
def test_chat_structured_retry_does_not_mutate_callers_options_dict(monkeypatch):
|
||||
"""Regression guard for the fix above: syncing the nested seed must copy,
|
||||
not mutate, the caller's options dict — otherwise a second call reusing
|
||||
the same options object would start from the wrong base seed."""
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, _Widget(name="c", count=3)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
caller_options = {"seed": 42}
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options=caller_options,
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
assert caller_options == {"seed": 42}
|
||||
|
||||
|
||||
def test_chat_structured_retry_sleeps_between_attempts(monkeypatch):
|
||||
sleep_calls = []
|
||||
|
||||
async def fake_sleep(secs):
|
||||
sleep_calls.append(secs)
|
||||
|
||||
monkeypatch.setattr(chat_mod.asyncio, "sleep", fake_sleep)
|
||||
|
||||
bad = ValidationError.from_exception_data("Widget", [])
|
||||
fake = _FakeAgent([bad, _Widget(name="c", count=3)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
assert sleep_calls == [chat_mod.RETRY_BACKOFF_SECS]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# refusal-retry — parity with test_ollama_provider.py's coverage (this
|
||||
# module serves LlamaCppProvider.chat_structured() — see
|
||||
# comfydv._llm.retry.is_refusal and OllamaOptionRefusalRetry).
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_chat_structured_retries_on_refusal_and_returns_clean_second_attempt(
|
||||
monkeypatch,
|
||||
):
|
||||
refused = _Widget(name="I cannot generate that content.", count=1)
|
||||
clean = _Widget(name="clean", count=2)
|
||||
fake = _FakeAgent([refused, clean])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
result = _run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
|
||||
max_retries=2,
|
||||
)
|
||||
)
|
||||
|
||||
assert result == clean
|
||||
assert len(fake.calls) == 2
|
||||
|
||||
|
||||
def test_chat_structured_on_status_reports_refusal_reason(monkeypatch):
|
||||
refused = _Widget(name="I cannot generate that content.", count=1)
|
||||
clean = _Widget(name="clean", count=2)
|
||||
fake = _FakeAgent([refused, clean])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
statuses = []
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
|
||||
max_retries=2,
|
||||
on_status=statuses.append,
|
||||
)
|
||||
)
|
||||
|
||||
assert len(statuses) == 2
|
||||
assert "Refusal/deflection detected" in statuses[0]
|
||||
assert "Recovered" in statuses[1]
|
||||
|
||||
|
||||
def test_chat_structured_refusal_retry_disabled_returns_refusal_unchanged(
|
||||
monkeypatch,
|
||||
):
|
||||
refused = _Widget(name="I cannot generate that content.", count=1)
|
||||
fake = _FakeAgent([refused])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
result = _run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
)
|
||||
)
|
||||
|
||||
assert result == refused
|
||||
assert len(fake.calls) == 1
|
||||
|
||||
|
||||
def test_chat_structured_refusal_retry_exhausted_raises(monkeypatch):
|
||||
refused = _Widget(name="I cannot help with this.", count=1)
|
||||
fake = _FakeAgent([refused, refused])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
with pytest.raises(RuntimeError, match="failed validation"):
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
|
||||
max_retries=1,
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def test_chat_structured_refusal_retry_uses_embed_fn(monkeypatch):
|
||||
from comfydv._llm.retry import REFUSAL_EXEMPLARS
|
||||
|
||||
refused = _Widget(name="not today, sorry", count=1)
|
||||
clean = _Widget(name="a clean value", count=2)
|
||||
fake = _FakeAgent([refused, clean])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
embed_calls = []
|
||||
|
||||
async def fake_embed(text):
|
||||
embed_calls.append(text)
|
||||
if "not today" in text or text in REFUSAL_EXEMPLARS:
|
||||
return [1.0, 0.0]
|
||||
return [0.0, 1.0]
|
||||
|
||||
result = _run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://localhost:11434/v1",
|
||||
model="llama3",
|
||||
messages=_messages(),
|
||||
schema=_Widget,
|
||||
options={
|
||||
"refusal_retry": {
|
||||
"enabled": True,
|
||||
"embedding_model": "nomic-embed-text",
|
||||
"threshold": 0.5,
|
||||
}
|
||||
},
|
||||
max_retries=2,
|
||||
embed_fn=fake_embed,
|
||||
)
|
||||
)
|
||||
|
||||
assert result == clean
|
||||
assert embed_calls # embedding path was actually exercised
|
||||
|
||||
|
||||
def test_history_to_messages_preserves_order_and_roles():
|
||||
from pydantic_ai.messages import ModelRequest, ModelResponse
|
||||
|
||||
msgs = _messages(
|
||||
system="be terse",
|
||||
history=[("user", "first"), ("assistant", "reply")],
|
||||
prompt="second",
|
||||
)
|
||||
history = chat_mod._history_to_messages(msgs)
|
||||
|
||||
# system, user(first), assistant(reply) — "second" is excluded (it's the
|
||||
# current turn, passed separately as Agent.run()'s user_prompt).
|
||||
assert len(history) == 3
|
||||
assert isinstance(history[0], ModelRequest) # system
|
||||
assert isinstance(history[1], ModelRequest) # user
|
||||
assert isinstance(history[2], ModelResponse) # assistant
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Image input (spec 009, US3; features/us3_structured_image.feature)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_chat_structured_attaches_image_to_user_prompt(monkeypatch):
|
||||
"""The current turn's image rides on Agent.run()'s user_prompt as a
|
||||
pydantic-ai BinaryContent (ADR-008, research.md Decision 1)."""
|
||||
import base64
|
||||
|
||||
from pydantic_ai.messages import BinaryContent
|
||||
|
||||
fake = _FakeAgent([_Widget(name="sq", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
b64 = base64.b64encode(b"PNGDATA").decode()
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://x/v1",
|
||||
model="m",
|
||||
schema=_Widget,
|
||||
messages=[Message(role="user", content="describe", images=[b64])],
|
||||
)
|
||||
)
|
||||
|
||||
prompt = fake.calls[0][0]
|
||||
assert isinstance(prompt, list)
|
||||
assert prompt[0] == "describe"
|
||||
assert isinstance(prompt[1], BinaryContent)
|
||||
assert prompt[1].data == b"PNGDATA"
|
||||
assert prompt[1].media_type == "image/png"
|
||||
|
||||
|
||||
def test_chat_structured_text_only_prompt_is_plain_string(monkeypatch):
|
||||
"""FR-003: an image-less structured call is unchanged — plain-string
|
||||
user_prompt, exactly as before spec 009."""
|
||||
fake = _FakeAgent([_Widget(name="a", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://x/v1",
|
||||
model="m",
|
||||
schema=_Widget,
|
||||
messages=[Message(role="user", content="hi")],
|
||||
)
|
||||
)
|
||||
|
||||
assert fake.calls[0][0] == "hi"
|
||||
|
||||
|
||||
def test_chat_structured_attaches_image_to_history_user_turn(monkeypatch):
|
||||
"""A prior user turn that carried an image keeps it in message_history."""
|
||||
import base64
|
||||
|
||||
from pydantic_ai.messages import BinaryContent, UserPromptPart
|
||||
|
||||
fake = _FakeAgent([_Widget(name="a", count=1)])
|
||||
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
|
||||
b64 = base64.b64encode(b"IMG").decode()
|
||||
msgs = [
|
||||
Message(role="user", content="earlier", images=[b64]),
|
||||
Message(role="assistant", content="ok"),
|
||||
Message(role="user", content="now"),
|
||||
]
|
||||
|
||||
_run_async(
|
||||
chat_mod.chat_structured(
|
||||
base_url="http://x/v1", model="m", schema=_Widget, messages=msgs
|
||||
)
|
||||
)
|
||||
|
||||
history = fake.calls[0][1]
|
||||
part = history[0].parts[0]
|
||||
assert isinstance(part, UserPromptPart)
|
||||
assert isinstance(part.content, list)
|
||||
assert part.content[0] == "earlier"
|
||||
assert isinstance(part.content[1], BinaryContent)
|
||||
assert part.content[1].data == b"IMG"
|
||||
@@ -0,0 +1,74 @@
|
||||
"""
|
||||
Tests for comfydv._llm — shared LLMProvider protocol and OllamaProvider.
|
||||
|
||||
Test layers:
|
||||
Unit (no marker) — pure Python, no live services
|
||||
Integration (-m integration) — requires live Ollama at localhost:11434
|
||||
|
||||
BDD coverage:
|
||||
../specs/007-llm-provider-abstraction/features/us1_connect_and_chat.feature
|
||||
../specs/007-llm-provider-abstraction/features/us2_structured_output.feature
|
||||
../specs/007-llm-provider-abstraction/features/us3_model_lifecycle.feature
|
||||
"""
|
||||
|
||||
from comfydv._llm.ollama_provider import OllamaProvider
|
||||
from comfydv._llm.provider import Message, ModelInfo, ModelStatus
|
||||
|
||||
|
||||
def test_ollama_provider_captures_connection_state():
|
||||
provider = OllamaProvider("http://localhost:11434", headers={"X-Test": "1"})
|
||||
assert provider.host == "http://localhost:11434"
|
||||
assert provider.headers == {"X-Test": "1"}
|
||||
|
||||
|
||||
def test_ollama_provider_headers_default_to_none():
|
||||
provider = OllamaProvider("http://localhost:11434")
|
||||
assert provider.headers is None
|
||||
|
||||
|
||||
def test_model_status_values():
|
||||
assert ModelStatus.LOADED == "loaded"
|
||||
assert ModelStatus.SLEEPING == "sleeping"
|
||||
assert ModelStatus.DOWNLOADING == "downloading"
|
||||
|
||||
|
||||
def test_model_info_optional_size():
|
||||
info = ModelInfo(name="llama3", status=ModelStatus.UNLOADED)
|
||||
assert info.size is None
|
||||
|
||||
|
||||
def test_message_roles():
|
||||
Message(role="system", content="be terse")
|
||||
Message(role="user", content="hi")
|
||||
Message(role="assistant", content="hello")
|
||||
|
||||
|
||||
# --- US1 foundational: Message.images carrier (spec 009, contract T1) ---
|
||||
|
||||
|
||||
def test_message_images_defaults_to_none():
|
||||
"""A text-only turn carries no images."""
|
||||
msg = Message(role="user", content="hi")
|
||||
assert msg.images is None
|
||||
|
||||
|
||||
def test_message_images_round_trips_base64_list():
|
||||
msg = Message(role="user", content="describe", images=["aGVsbG8=", "d29ybGQ="])
|
||||
assert msg.images == ["aGVsbG8=", "d29ybGQ="]
|
||||
|
||||
|
||||
def test_message_text_only_dump_omits_images_key():
|
||||
"""FR-003/SC-004: an image-less message must serialize byte-identically to
|
||||
today — no stray ``images`` key in the transport payload."""
|
||||
msg = Message(role="user", content="hi")
|
||||
assert msg.model_dump(exclude_none=True) == {"role": "user", "content": "hi"}
|
||||
|
||||
|
||||
def test_message_with_images_dump_includes_images_key():
|
||||
msg = Message(role="user", content="describe", images=["aGVsbG8="])
|
||||
dumped = msg.model_dump(exclude_none=True)
|
||||
assert dumped == {
|
||||
"role": "user",
|
||||
"content": "describe",
|
||||
"images": ["aGVsbG8="],
|
||||
}
|
||||
@@ -0,0 +1,426 @@
|
||||
"""Tests for comfydv._llm.retry — shared retry-on-blank-output helpers used
|
||||
by both providers' chat() and the shared chat_structured() helper, plus the
|
||||
refusal/deflection detector that rides the same retry-with-a-new-seed
|
||||
mechanism.
|
||||
|
||||
Refusal-detection tests here are pure-logic only — no provider/HTTP
|
||||
involved. See test_ollama_provider.py, test_llamacpp_provider.py, and
|
||||
test_llm_chat_structured.py for the retry-loop integration (does a detected
|
||||
refusal actually trigger a reseeded retry).
|
||||
"""
|
||||
|
||||
import pytest
|
||||
|
||||
from comfydv._llm.ollama_provider import _run_async
|
||||
from comfydv._llm.retry import (
|
||||
REFUSAL_EXEMPLARS,
|
||||
cosine_similarity,
|
||||
format_recovered_status,
|
||||
format_retry_status,
|
||||
is_ambiguous,
|
||||
is_lexical_refusal,
|
||||
is_refusal,
|
||||
next_seed,
|
||||
next_timeout_secs,
|
||||
record_attempt_info,
|
||||
)
|
||||
|
||||
|
||||
def test_next_seed_attempt_one_is_zero_by_default():
|
||||
assert next_seed(None, 1) == 0
|
||||
|
||||
|
||||
def test_next_seed_increments_from_zero_when_unset():
|
||||
assert next_seed(None, 2) == 1
|
||||
assert next_seed({}, 3) == 2
|
||||
|
||||
|
||||
def test_next_seed_starts_from_pinned_base():
|
||||
assert next_seed({"seed": 42}, 1) == 42
|
||||
assert next_seed({"seed": 42}, 2) == 43
|
||||
assert next_seed({"seed": 42}, 3) == 44
|
||||
|
||||
|
||||
def test_next_seed_ignores_non_int_seed():
|
||||
assert next_seed({"seed": "not-an-int"}, 2) == 1
|
||||
|
||||
|
||||
class TestNextTimeoutSecs:
|
||||
def test_attempt_one_returns_base_timeout_unchanged(self):
|
||||
assert next_timeout_secs(300.0, 1) == 300.0
|
||||
|
||||
def test_escalates_multiplicatively_per_attempt(self):
|
||||
assert next_timeout_secs(300.0, 2) == 600.0
|
||||
assert next_timeout_secs(300.0, 3) == 900.0
|
||||
|
||||
|
||||
class TestRecordAttemptInfo:
|
||||
def test_none_attempt_info_is_a_no_op(self):
|
||||
# Must not raise — callers that don't care about this metadata pass
|
||||
# None and should see no behavior change at all.
|
||||
record_attempt_info(None, seed=1, attempts=2, timeout_secs=600.0, refusals=1)
|
||||
|
||||
def test_populates_dict_in_place(self):
|
||||
info: dict = {}
|
||||
record_attempt_info(info, seed=7, attempts=3, timeout_secs=900.0, refusals=2)
|
||||
assert info == {
|
||||
"seed": 7,
|
||||
"attempts": 3,
|
||||
"timeout_secs": 900.0,
|
||||
"refusals": 2,
|
||||
}
|
||||
|
||||
def test_overwrites_previous_values(self):
|
||||
# Callers call this once per attempt (or once per return path) —
|
||||
# a later call must replace, not merge with, an earlier one.
|
||||
info: dict = {"seed": 1, "attempts": 1, "timeout_secs": 300.0, "refusals": 0}
|
||||
record_attempt_info(info, seed=2, attempts=2, timeout_secs=600.0, refusals=1)
|
||||
assert info == {
|
||||
"seed": 2,
|
||||
"attempts": 2,
|
||||
"timeout_secs": 600.0,
|
||||
"refusals": 1,
|
||||
}
|
||||
|
||||
|
||||
class TestFormatStatus:
|
||||
def test_retry_status_includes_reason_attempt_seed_and_timeout(self):
|
||||
msg = format_retry_status("Blank response", 1, 3, seed=1, timeout_secs=200.0)
|
||||
assert "Blank response" in msg
|
||||
assert "attempt 1/3" in msg
|
||||
assert "seed=1" in msg
|
||||
assert "timeout=200s" in msg
|
||||
|
||||
def test_recovered_status_includes_attempt_and_seed(self):
|
||||
msg = format_recovered_status(2, 3, seed=1)
|
||||
assert "attempt 2/3" in msg
|
||||
assert "seed=1" in msg
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Refusal/deflection detection
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
class TestLexicalDetection:
|
||||
def test_blatant_refusal_openers_match(self):
|
||||
blatant = [
|
||||
"I cannot generate that image for you.",
|
||||
"I'm not able to help with this request.",
|
||||
"I am unable to comply with your instruction.",
|
||||
"Cannot generate this image due to the sensitive content.",
|
||||
'{"error": "cannot generate this content"}',
|
||||
"I don't feel comfortable creating this content.",
|
||||
"I'm sorry, but I can't produce that description.",
|
||||
"As an AI, I'm unable to depict this scenario.",
|
||||
"I must decline this request.",
|
||||
"This request violates my usage policies.",
|
||||
"That's not something I'm able to help with right now.",
|
||||
]
|
||||
for text in blatant:
|
||||
assert is_lexical_refusal(text), f"expected refusal match: {text!r}"
|
||||
|
||||
def test_ordinary_content_does_not_match(self):
|
||||
ordinary = [
|
||||
"The subject turns to face the camera and smiles warmly.",
|
||||
"A person cannot simply walk into Mordor, the guide joked.",
|
||||
"I can help you plan a birthday party for your dog.",
|
||||
"",
|
||||
]
|
||||
for text in ordinary:
|
||||
assert not is_lexical_refusal(text), f"unexpected match: {text!r}"
|
||||
|
||||
|
||||
class TestAmbiguityHeuristic:
|
||||
def test_short_response_is_ambiguous(self):
|
||||
assert is_ambiguous("Sorry, can't do that one.")
|
||||
|
||||
def test_long_response_without_hedge_keywords_is_not_ambiguous(self):
|
||||
long_text = "The subject rotates smoothly toward the lens. " * 20
|
||||
assert len(long_text) >= 400
|
||||
assert not is_ambiguous(long_text)
|
||||
|
||||
def test_long_response_with_hedge_keyword_is_ambiguous(self):
|
||||
long_text = "Unfortunately, " + "this touches on a sensitive area. " * 20
|
||||
assert len(long_text) >= 400
|
||||
assert is_ambiguous(long_text)
|
||||
|
||||
def test_blank_text_is_not_ambiguous(self):
|
||||
assert not is_ambiguous(" ")
|
||||
|
||||
|
||||
class TestCosineSimilarity:
|
||||
def test_identical_vectors_score_one(self):
|
||||
assert cosine_similarity([1.0, 0.0], [1.0, 0.0]) == pytest.approx(1.0)
|
||||
|
||||
def test_orthogonal_vectors_score_zero(self):
|
||||
assert cosine_similarity([1.0, 0.0], [0.0, 1.0]) == pytest.approx(0.0)
|
||||
|
||||
def test_opposite_vectors_score_negative_one(self):
|
||||
assert cosine_similarity([1.0, 0.0], [-1.0, 0.0]) == pytest.approx(-1.0)
|
||||
|
||||
def test_mismatched_lengths_return_zero(self):
|
||||
assert cosine_similarity([1.0, 0.0], [1.0, 0.0, 0.0]) == 0.0
|
||||
|
||||
def test_empty_vectors_return_zero(self):
|
||||
assert cosine_similarity([], []) == 0.0
|
||||
|
||||
|
||||
class TestIsRefusalHybrid:
|
||||
def test_blank_text_is_never_a_refusal(self):
|
||||
assert not _run_async(is_refusal(""))
|
||||
assert not _run_async(is_refusal(" "))
|
||||
|
||||
def test_lexical_match_short_circuits_without_embedding(self):
|
||||
calls = []
|
||||
|
||||
async def embed_fn(text):
|
||||
calls.append(text)
|
||||
return [1.0, 0.0]
|
||||
|
||||
result = _run_async(
|
||||
is_refusal("I cannot generate that image for you.", embed_fn=embed_fn)
|
||||
)
|
||||
assert result is True
|
||||
assert calls == [] # never reached the embedding step
|
||||
|
||||
def test_long_clean_response_is_still_embedding_checked_but_not_a_refusal(self):
|
||||
# embed_fn present means the caller opted in to the embedding check
|
||||
# regardless of length/hedge-keywords (is_ambiguous no longer gates
|
||||
# this) — a long, on-topic response should still be embedding-
|
||||
# checked, it just shouldn't score as similar to the refusal
|
||||
# exemplars.
|
||||
calls = []
|
||||
|
||||
async def embed_fn(text):
|
||||
calls.append(text)
|
||||
if text == long_text:
|
||||
return [0.0, 1.0] # orthogonal to the exemplar vector below
|
||||
return [1.0, 0.0] # exemplars
|
||||
|
||||
long_text = "The subject rotates smoothly toward the lens. " * 20
|
||||
result = _run_async(is_refusal(long_text, embed_fn=embed_fn))
|
||||
assert result is False
|
||||
assert long_text in calls # embedding check DID run, just scored low
|
||||
|
||||
def test_no_embed_fn_degrades_to_lexical_only(self):
|
||||
# Ambiguous (short), no lexical match, no embed_fn -> can't check further
|
||||
assert not _run_async(is_refusal("Not today, sorry.", embed_fn=None))
|
||||
|
||||
def test_ambiguous_response_above_threshold_is_refusal(self):
|
||||
async def embed_fn(text):
|
||||
# Exemplars and a near-identical short "refusal-ish" probe get a
|
||||
# high similarity score; distinguish by a marker substring.
|
||||
if "PROBE" in text:
|
||||
return [1.0, 0.0]
|
||||
return [0.99, 0.14] # cos-sim with [1,0] is ~0.99
|
||||
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"PROBE: not comfortable with this one",
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key="test-model",
|
||||
threshold=0.8,
|
||||
)
|
||||
)
|
||||
assert result is True
|
||||
|
||||
def test_ambiguous_response_below_threshold_is_not_refusal(self):
|
||||
async def embed_fn(text):
|
||||
if "PROBE" in text:
|
||||
return [1.0, 0.0]
|
||||
return [0.0, 1.0] # orthogonal -> cos-sim 0.0
|
||||
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"PROBE: a short reply",
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key="test-model-2",
|
||||
threshold=0.8,
|
||||
)
|
||||
)
|
||||
assert result is False
|
||||
|
||||
def test_long_json_shaped_soft_refusal_without_hedge_keywords_is_caught(self):
|
||||
# Regression case: a structured-output-shaped response (>600 chars
|
||||
# once you count JSON braces/field names) whose deflection doesn't
|
||||
# use any of the canned hedge keywords used to be invisible to the
|
||||
# embedding check entirely, because is_ambiguous gated on length
|
||||
# and keywords. embed_fn now runs unconditionally once configured.
|
||||
soft_refusal = (
|
||||
'{"prompt": "'
|
||||
+ "Let's take this in a different creative direction that everyone can enjoy. "
|
||||
* 8
|
||||
+ '"}'
|
||||
)
|
||||
assert len(soft_refusal) >= 600
|
||||
assert not is_lexical_refusal(soft_refusal)
|
||||
|
||||
async def embed_fn(text):
|
||||
if text == soft_refusal:
|
||||
return [1.0, 0.0]
|
||||
return [0.97, 0.24] # exemplars: cos-sim with [1,0] is ~0.97
|
||||
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
soft_refusal,
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key="regression-model",
|
||||
threshold=0.8,
|
||||
)
|
||||
)
|
||||
assert result is True
|
||||
|
||||
def test_exemplar_embeddings_cached_across_calls(self):
|
||||
exemplar_calls = {"n": 0}
|
||||
|
||||
async def embed_fn(text):
|
||||
if "PROBE" in text:
|
||||
return [1.0, 0.0]
|
||||
exemplar_calls["n"] += 1
|
||||
return [1.0, 0.0]
|
||||
|
||||
_run_async(
|
||||
is_refusal(
|
||||
"PROBE: first ambiguous call",
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key="cache-key-shared",
|
||||
threshold=0.5,
|
||||
)
|
||||
)
|
||||
first_count = exemplar_calls["n"]
|
||||
assert first_count > 0
|
||||
|
||||
_run_async(
|
||||
is_refusal(
|
||||
"PROBE: second ambiguous call",
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key="cache-key-shared",
|
||||
threshold=0.5,
|
||||
)
|
||||
)
|
||||
# Exemplar embeddings reused from cache -> no additional exemplar calls
|
||||
assert exemplar_calls["n"] == first_count
|
||||
|
||||
def test_embed_fn_failure_degrades_to_not_refused(self):
|
||||
async def failing_embed_fn(text):
|
||||
raise RuntimeError("server unreachable")
|
||||
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"Not comfortable with this one, sorry.",
|
||||
embed_fn=failing_embed_fn,
|
||||
embed_cache_key="unreachable-model",
|
||||
)
|
||||
)
|
||||
assert result is False
|
||||
|
||||
def test_embed_fn_returning_none_degrades_to_not_refused(self):
|
||||
async def none_embed_fn(text):
|
||||
return None
|
||||
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"Not comfortable with this one, sorry.",
|
||||
embed_fn=none_embed_fn,
|
||||
embed_cache_key="no-embeddings-model",
|
||||
)
|
||||
)
|
||||
assert result is False
|
||||
|
||||
|
||||
class TestCustomPhrases:
|
||||
def test_custom_phrase_substring_match_needs_no_embed_fn(self):
|
||||
# A phrase the user added at runtime that the shipped lexical
|
||||
# patterns don't cover — should be caught for free, no embedding
|
||||
# model required.
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"I am restricted from producing that kind of content.",
|
||||
custom_phrases=("restricted from",),
|
||||
)
|
||||
)
|
||||
assert result is True
|
||||
|
||||
def test_custom_phrase_match_is_case_insensitive(self):
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"SORRY, THAT'S OFF LIMITS FOR ME.",
|
||||
custom_phrases=("off limits",),
|
||||
)
|
||||
)
|
||||
assert result is True
|
||||
|
||||
def test_unrelated_custom_phrase_does_not_match(self):
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"The subject walks calmly toward the horizon.",
|
||||
custom_phrases=("restricted from", "off limits"),
|
||||
)
|
||||
)
|
||||
assert result is False
|
||||
|
||||
def test_blank_and_whitespace_custom_phrases_are_ignored(self):
|
||||
# A stray empty entry must never become a universal substring match.
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"The subject walks calmly toward the horizon.",
|
||||
custom_phrases=("", " "),
|
||||
)
|
||||
)
|
||||
assert result is False
|
||||
|
||||
def test_custom_phrase_folded_into_embedding_exemplars(self):
|
||||
# No exact substring match, but embed_fn scores the response as
|
||||
# similar to the custom phrase (not one of the shipped exemplars).
|
||||
custom = "my creators have limited what I can show you"
|
||||
|
||||
async def embed_fn(text):
|
||||
if text == custom:
|
||||
return [1.0, 0.0]
|
||||
if text in REFUSAL_EXEMPLARS:
|
||||
return [0.0, 1.0] # shipped exemplars score orthogonal
|
||||
return [0.99, 0.14] # the probe response is near the custom one
|
||||
|
||||
result = _run_async(
|
||||
is_refusal(
|
||||
"There are limits my creators placed on what I can show.",
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key="custom-exemplar-model",
|
||||
threshold=0.8,
|
||||
custom_phrases=(custom,),
|
||||
)
|
||||
)
|
||||
assert result is True
|
||||
|
||||
def test_different_custom_phrase_sets_do_not_share_exemplar_cache(self):
|
||||
# Regression guard: if the exemplar cache key ignored custom_phrases,
|
||||
# a second call with a different custom phrase set would incorrectly
|
||||
# reuse the first call's cached (and now stale) exemplar vectors.
|
||||
calls = []
|
||||
|
||||
async def embed_fn(text):
|
||||
calls.append(text)
|
||||
return [1.0, 0.0]
|
||||
|
||||
_run_async(
|
||||
is_refusal(
|
||||
"short reply",
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key="shared-model",
|
||||
custom_phrases=("phrase one",),
|
||||
)
|
||||
)
|
||||
first_call_count = len(calls)
|
||||
|
||||
_run_async(
|
||||
is_refusal(
|
||||
"short reply",
|
||||
embed_fn=embed_fn,
|
||||
embed_cache_key="shared-model",
|
||||
custom_phrases=("phrase two",),
|
||||
)
|
||||
)
|
||||
# A fresh custom phrase set re-embeds the exemplars (including the
|
||||
# new phrase) rather than reusing the first set's cached vectors.
|
||||
assert len(calls) > first_call_count
|
||||