Merge pull request #21 from darth-veitcher/fix/comfyui-relative-import-loading

This commit is contained in:
James Veitch
2026-07-11 18:09:56 +01:00
committed by GitHub
24 changed files with 684 additions and 257 deletions
+84 -105
View File
@@ -1,29 +1,37 @@
# comfydv
A collection of workflow efficiency and quality-of-life nodes built out of necessity for personal ComfyUI use.
**Quality-of-life nodes for ComfyUI, built to disappear into your workflow.**
## What is this?
`comfydv` fills the gaps ComfyUI's built-in library leaves on the table: string templates that build their own sockets as you type, seed-controlled randomisation, graceful mid-queue interruption, and a local-LLM integration that doesn't care whether you're running Ollama or llama.cpp. No Python required — install it, drop the nodes on your canvas, wire them up.
`comfydv` fills gaps in ComfyUI's built-in node library: dynamic string formatting, seed-controlled random selection, graceful workflow interruption, and local LLM integration (Ollama and llama.cpp). Install it once and connect the nodes like any other — no Python knowledge required.
![Chat Completion in action](docs/assets/ollama_chat.png)
## What is comfydv?
A small, focused ComfyUI utility pack. It exists because:
- **It reads your intent, not just your syntax.** Format String detects `{variables}` in a template and adds/removes input sockets live, as you type — no manual socket wrangling.
- **One LLM integration, any local backend.** Wire a Chat Completion node once; swap between Ollama and llama.cpp by changing a single upstream client node. Structured output, multi-turn history, and model load/unload work identically on both.
- **It fails politely.** Circuit Breaker halts a queue run cleanly instead of throwing a stack trace at you; a disconnected LLM server gets a specific, actionable error instead of a silent empty dropdown.
- **Small, tested, boring in the best way.** Every node is unit-tested and the local-LLM nodes are verified against real running servers, not just mocks.
## What's inside
| Node | What it does |
|------|-------------|
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the connection through the graph as an `LLM_CLIENT` socket — a generic connection type any backend's client node emits. |
| **LlamaCpp Client** | Configures a connection to a `llama-server` instance running in router mode (default: `http://localhost:8080`). Emits the same `LLM_CLIENT` socket as Ollama Client — every node below works with either. |
| **LLM Model Selector** | Fetches the live model list from the connected server and presents it as a dropdown. Outputs the selected model name. |
| **LLM Load Model** | Loads a model into memory on the connected server. |
| **LLM Unload Model** | Evicts a model from memory on the connected server. |
| **Chat Completion** | Sends a prompt (and optional conversation history) to the connected server. Response and history are shown inline in the node body and available as output sockets. |
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
| **Ollama History Length** | Returns the number of messages in an `OLLAMA_HISTORY` list as an integer. |
|------|---------------|
| **Format String** | Renders a Python f-string or Jinja2 template. Sockets appear and disappear automatically as you type variables. |
| **Random Choice** | Accepts any number of typed inputs and returns one at random, with a seed for reproducibility. |
| **Circuit Breaker** | Halts the current queue run gracefully — no crash, just a clean stop — when a condition isn't met. |
| **Ollama Client** / **LlamaCpp Client** | Configure a connection to a local Ollama or llama.cpp server. Both emit the same `LLM_CLIENT` socket — every node below works with either. |
| **LLM Model Selector** | Live dropdown of models available on the connected server. |
| **LLM Load Model** / **LLM Unload Model** | Explicit VRAM management — pin a model in memory before inference, evict it after. |
| **Chat Completion** | Send a prompt (optionally with history) to the connected server; response shown inline and as an output socket. |
| **Ollama Option — \*** | Seven composable parameter nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into Chat Completion's `options` input. |
| **Ollama Debug History** / **Ollama History Length** | Inspect an `OLLAMA_HISTORY` conversation list — pretty-print it or count its messages. |
## Install
**Via ComfyUI Manager** (recommended): search for `comfydv` and click Install.
**Via ComfyUI Manager** (recommended): search for `comfydv`, click Install.
**Manual:**
@@ -32,134 +40,125 @@ cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/darth-veitcher/comfydv.git
```
Restart ComfyUI. The nodes appear under the **dv/**, **dv/ollama**, and **dv/llamacpp** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) are installed automatically via `requirements.txt`.
Restart ComfyUI. Nodes appear under **dv/**, **dv/ollama**, and **dv/llamacpp** in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) install automatically via `requirements.txt`.
For local LLM nodes, pick one backend (or both):
For local-LLM nodes, bring your own backend:
- **Ollama**: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`).
- **llama.cpp**: [build/install `llama-server`](https://github.com/ggml-org/llama.cpp) and launch it in router mode (`llama-server --models-dir ./models`) — see the [llama.cpp section](#llamacpp) below.
- **Ollama** — [install Ollama](https://ollama.com/download), pull a model (`ollama pull qwen2.5:latest`).
- **llama.cpp** — [build/install `llama-server`](https://github.com/ggml-org/llama.cpp), launch it in [router mode](#llamacpp).
## Quickstart
1. Install via ComfyUI Manager (search `comfydv`) or clone manually into `custom_nodes/`.
2. Right-click the canvas → Add Node → **dv/** to find Format String, Random Choice, and Circuit Breaker.
3. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
1. Right-click the canvas → Add Node → **dv/** for Format String, Random Choice, and Circuit Breaker.
2. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
## Documentation
Full documentation: [darth-veitcher.github.io/comfydv](https://darth-veitcher.github.io/comfydv/stable/)
Full documentation, including every node's inputs/outputs: **[darth-veitcher.github.io/comfydv](https://darth-veitcher.github.io/comfydv/stable/)**
---
## Format String
Formats text from a Python f-string or Jinja2 template. As you type the template, input sockets appear and disappear automatically — one per variable detected.
### Python f-strings
Type `{variable_name}` and a socket appears. Wire it to any string output in your workflow.
Type a template, get sockets. `{variable_name}` in f-string mode, `{{ variable_name }}` in Jinja2 mode — either way, comfydv watches what you type and keeps the node's inputs in sync automatically.
![Format String — f-string mode](docs/assets/fstring.png)
Outputs are always in a stable order:
| Output | Content |
|--------|---------|
| `formatted_string` | The rendered result |
| `saved_file_path` | Path written to disk (if `save_path` is set) |
| `<var>` … | Pass-through of each input value, for easy chaining |
| `saved_file_path` | Where it was written, if `save_path` is set |
| `<var>` … | Each input passed through unchanged, for easy chaining |
### Jinja2 templates
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`, …), conditionals (`{% if %}…{% endif %}`), and loops.
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`), conditionals, and loops:
![Format String — Jinja2 mode](docs/assets/jinja2.png)
Variables detected in `{{ }}` expressions become input sockets exactly as in Simple mode. See the [Jinja2 documentation](https://jinja.palletsprojects.com/en/latest/) for the full filter/test reference.
---
## Random Choice
Connect any number of inputs of the same type. Each run picks one at random. Set `seed` for reproducibility.
Wire in any number of same-typed inputs — images, strings, conditioning, anything ComfyUI can carry over a socket — and get one back at random.
![Random Choice](docs/assets/random.png)
- Accepts any ComfyUI type (STRING, IMAGE, CONDITIONING, …)
- Add as many inputs as you like; unused slots are removed automatically when disconnected
- `seed = 0` randomises on every run; any other value locks the selection
`seed = 0` randomises every run; any other value locks the selection. Unused input slots vanish automatically when you disconnect them.
---
## Circuit Breaker
Stops the queue gracefully when a condition isn't met — no crash, no error, just a clean halt.
Stop a queue run cleanly when a condition isn't met, instead of letting a downstream node crash on bad input.
![Circuit Breaker](docs/assets/circuit_breaker.png)
Wire an image (or any trigger) into `trigger` and a boolean into `status`. When `status` is **false** the node raises `InterruptProcessingException`, which tells ComfyUI to stop the current run cleanly. When `status` is **true** the image passes through unchanged.
Wire a trigger (an image, or anything) into `trigger` and a boolean into `status`. `status = false` raises `InterruptProcessingException` — ComfyUI stops the run without an error dialog. `status = true` passes the trigger straight through.
Typical use: skip an expensive upscale step when a quality-check node says the draft is already good enough.
Typical use: skip an expensive upscale pass when an upstream quality-check node says the draft's already good enough.
---
## Ollama
## Local LLMs
Nodes for integrating a local Ollama LLM into your ComfyUI workflow. The host is configured once in **Ollama Client** and threaded through the graph as an `LLM_CLIENT` socket — a generic connection type any future backend's client node can also emit, so the chat/model-management nodes below aren't Ollama-specific.
One set of nodes, two interchangeable backends. Configure a connection once with **Ollama Client** or **LlamaCpp Client** — both output the same `LLM_CLIENT` socket — and every downstream node (model selection, load/unload, chat, structured output, multi-turn history) works exactly the same way regardless of which one you picked. Swapping backends means rewiring one node, not rebuilding your graph.
### Ollama Client node
Configure the server address once; all downstream nodes inherit it automatically.
### Connect and chat
![Ollama Client](docs/assets/ollama_client.png)
### Model lifecycle (load and unload)
On memory-constrained machines and single-GPU setups, explicitly loading and unloading the model before and after inference is critical. **LLM Load Model** pins the model into VRAM (`keep_alive=-1`); **LLM Unload Model** evicts it immediately (`keep_alive=0`), freeing memory for image generation or other models.
![Ollama Load / Unload](docs/assets/ollama_lifecycle.png)
The correct chain is **Load → Chat → Unload**, enforced through data dependencies:
1. Wire `LLMLoadModel.model_name` → `ChatCompletion.model`. This creates the data dependency that guarantees Load runs before Chat and passes the model name into the Chat node's plain-string `model` input.
2. Wire `ChatCompletion.model_name` → `LLMUnloadModel.model`. This guarantees Unload runs after Chat completes.
3. Optionally wire `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged so the rest of your workflow can still consume it.
### Minimal chat workflow
1. **Ollama Client** → set host (default `http://localhost:11434`)
2. **LLM Model Selector** → pick a model from the live dropdown (or type/wire a model name directly into Chat Completion's `model` input)
3. **Chat Completion** → wire client + model + prompt; the response appears inline in the node body and is also available as an output socket
1. **Ollama Client** (default `http://localhost:11434`) or **LlamaCpp Client** (default `http://localhost:8080`) — set the host.
2. **LLM Model Selector** — pick a model from the live dropdown, or wire a model name straight into Chat Completion.
3. **Chat Completion** — wire in client, model, and prompt. The response renders inline in the node and is also available as an output socket.
![Chat Completion](docs/assets/ollama_chat.png)
Wire multiple nodes together for a complete end-to-end workflow:
A complete graph looks like this:
![Ollama Full Workflow](docs/assets/ollama_workflow.png)
![Full LLM workflow](docs/assets/ollama_workflow.png)
### Option nodes
### Manual memory management
Chain any combination of **Ollama Option —** nodes before Chat Completion to override inference parameters:
Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline.
| Option node | Ollama param |
|-------------|-------------|
![Load / Unload lifecycle](docs/assets/ollama_lifecycle.png)
The **Load → Chat → Unload** chain is enforced by data dependencies, not by convention:
1. `LLMLoadModel.model_name` → `ChatCompletion.model` — guarantees Load runs before Chat, and feeds the model name straight in.
2. `ChatCompletion.model_name` → `LLMUnloadModel.model` — guarantees Unload runs after Chat completes.
3. *(Optional)* `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged, so the rest of your workflow can still consume it.
### Tuning generation
Chain any combination of **Ollama Option —** nodes ahead of Chat Completion to override inference parameters:
![Ollama Option nodes](docs/assets/ollama_options.png)
| Option node | Parameter |
|-------------|-----------|
| Temperature | `temperature` |
| Seed | `seed` |
| Max Tokens | `num_predict` |
| Top P | `top_p` |
| Top K | `top_k` |
| Repeat Penalty | `repeat_penalty` |
| Extra Body | arbitrary JSON merged into options |
![Ollama Option Nodes](docs/assets/ollama_options.png)
| Extra Body | arbitrary JSON, merged into `options` |
### Multi-turn conversations
`OLLAMA_HISTORY` flows out of Chat Completion as a list of `{"role", "content"}` dicts. Wire it back into the next Chat Completion for multi-turn conversations, or inspect it with **Ollama Debug History** / **Ollama History Length**.
`OLLAMA_HISTORY` flows out of Chat Completion as a `{"role", "content"}` list. Feed it back into the next Chat Completion call for multi-turn context, or inspect it with **Ollama Debug History** / **Ollama History Length**.
### Upgrading an older workflow
### llama.cpp
If you saved a workflow before this rename, ComfyUI will report the old node types as missing when you reopen it. Reconnect using this mapping, then re-run — behavior is unchanged, only the names and the client socket type are different:
Everything above works unchanged against llama.cpp — swap in a **LlamaCpp Client** and the rest of the graph doesn't know the difference. The one thing llama.cpp needs that Ollama doesn't: **router mode**, a directory of models rather than a single `-m model.gguf`:
```bash
llama-server --models-dir ./models -c 8192
```
In exchange, router mode gives comfydv a richer live status than Ollama can report — `loading` and `downloading`, not just loaded/unloaded — plus the same explicit load/unload primitives Ollama's nodes already use.
### Upgrading a workflow saved before this rename
Nodes were renamed once, to make them backend-generic (`OllamaChatCompletion` → `ChatCompletion`, etc.). If ComfyUI reports old node types as missing when you reopen a saved workflow, reconnect using this table — behavior is unchanged, only the names are:
| Old | New |
|-----|-----|
@@ -169,30 +168,10 @@ If you saved a workflow before this rename, ComfyUI will report the old node typ
| `OllamaUnloadModel` | `LLMUnloadModel` |
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
`OllamaClient` kept its name — delete and re-add any node showing as missing, then rewire it to the same `OllamaClient` node.
---
## llama.cpp
## License
A second backend for the same chat/model-management nodes documented above — **LlamaCpp Client** is the only new node; everything else (Chat Completion, LLM Model Selector, LLM Load Model, LLM Unload Model, structured output, multi-turn history) works unchanged, because they don't know or care which backend they're talking to.
### Prerequisite: router mode
`llama-server` needs to be launched in **router mode** — a directory of models, not a single `-m model.gguf`:
```bash
llama-server --models-dir ./models -c 8192
```
This gives `comfydv` live model status (including `loading`/`downloading`, not just loaded/unloaded — a richer picture than Ollama can report) and explicit load/unload, the same way the Ollama nodes already work.
### LlamaCpp Client node
Configure the server address once (default `http://localhost:8080`); every downstream node inherits it automatically — same pattern as Ollama Client, same `LLM_CLIENT` socket.
### Switching an existing workflow from Ollama to llama.cpp
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed at your running `llama-server`. Nothing else changes — same Chat Completion node, same Load/Unload nodes, same structured-output behavior. That's the entire point of sharing one `LLM_CLIENT` socket type across backends.
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
[AGPL-3.0](LICENSE)
Binary file not shown.

Before

Width:  |  Height:  |  Size: 12 KiB

After

Width:  |  Height:  |  Size: 12 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 32 KiB

After

Width:  |  Height:  |  Size: 33 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 40 KiB

After

Width:  |  Height:  |  Size: 41 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 13 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 37 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 38 KiB

After

Width:  |  Height:  |  Size: 55 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 11 KiB

After

Width:  |  Height:  |  Size: 12 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 50 KiB

After

Width:  |  Height:  |  Size: 55 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 40 KiB

After

Width:  |  Height:  |  Size: 41 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 58 KiB

After

Width:  |  Height:  |  Size: 64 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 19 KiB

After

Width:  |  Height:  |  Size: 18 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 68 KiB

+86 -97
View File
@@ -1,25 +1,37 @@
# comfydv
A collection of workflow efficiency and quality-of-life nodes built out of necessity for personal ComfyUI use.
**Quality-of-life nodes for ComfyUI, built to disappear into your workflow.**
`comfydv` fills the gaps ComfyUI's built-in library leaves on the table: string templates that build their own sockets as you type, seed-controlled randomisation, graceful mid-queue interruption, and a local-LLM integration that doesn't care whether you're running Ollama or llama.cpp. No Python required — install it, drop the nodes on your canvas, wire them up.
![Chat Completion in action](assets/ollama_chat.png)
## What is comfydv?
A small, focused ComfyUI utility pack. It exists because:
- **It reads your intent, not just your syntax.** Format String detects `{variables}` in a template and adds/removes input sockets live, as you type — no manual socket wrangling.
- **One LLM integration, any local backend.** Wire a Chat Completion node once; swap between Ollama and llama.cpp by changing a single upstream client node. Structured output, multi-turn history, and model load/unload work identically on both.
- **It fails politely.** Circuit Breaker halts a queue run cleanly instead of throwing a stack trace at you; a disconnected LLM server gets a specific, actionable error instead of a silent empty dropdown.
- **Small, tested, boring in the best way.** Every node is unit-tested and the local-LLM nodes are verified against real running servers, not just mocks.
## What's inside
| Node | What it does |
|------|-------------|
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the connection through the graph as an `LLM_CLIENT` socket — a generic connection type any backend's client node emits. |
| **LlamaCpp Client** | Configures a connection to a `llama-server` instance running in router mode (default: `http://localhost:8080`). Emits the same `LLM_CLIENT` socket as Ollama Client — every node below works with either. |
| **LLM Model Selector** | Fetches the live model list from the connected server and presents it as a dropdown. Outputs the selected model name. |
| **LLM Load Model** | Loads a model into memory on the connected server. |
| **LLM Unload Model** | Evicts a model from memory on the connected server. |
| **Chat Completion** | Sends a prompt (and optional conversation history) to the connected server. Response and history are shown inline in the node body and available as output sockets. |
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
| **Ollama History Length** | Returns the number of messages in an `OLLAMA_HISTORY` list as an integer. |
|------|---------------|
| **Format String** | Renders a Python f-string or Jinja2 template. Sockets appear and disappear automatically as you type variables. |
| **Random Choice** | Accepts any number of typed inputs and returns one at random, with a seed for reproducibility. |
| **Circuit Breaker** | Halts the current queue run gracefully — no crash, just a clean stop — when a condition isn't met. |
| **Ollama Client** / **LlamaCpp Client** | Configure a connection to a local Ollama or llama.cpp server. Both emit the same `LLM_CLIENT` socket — every node below works with either. |
| **LLM Model Selector** | Live dropdown of models available on the connected server. |
| **LLM Load Model** / **LLM Unload Model** | Explicit VRAM management — pin a model in memory before inference, evict it after. |
| **Chat Completion** | Send a prompt (optionally with history) to the connected server; response shown inline and as an output socket. |
| **Ollama Option — \*** | Seven composable parameter nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into Chat Completion's `options` input. |
| **Ollama Debug History** / **Ollama History Length** | Inspect an `OLLAMA_HISTORY` conversation list — pretty-print it or count its messages. |
## Install
**Via ComfyUI Manager** (recommended): search for `comfydv` and click Install.
**Via ComfyUI Manager** (recommended): search for `comfydv`, click Install.
**Manual:**
@@ -28,122 +40,123 @@ cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/darth-veitcher/comfydv.git
```
Restart ComfyUI. The nodes appear under the **dv/**, **dv/ollama**, and **dv/llamacpp** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) are installed automatically via `requirements.txt`.
Restart ComfyUI. Nodes appear under **dv/**, **dv/ollama**, and **dv/llamacpp** in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) install automatically via `requirements.txt`.
For local LLM nodes, pick one backend (or both):
For local-LLM nodes, bring your own backend:
- **Ollama**: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`).
- **llama.cpp**: [build/install `llama-server`](https://github.com/ggml-org/llama.cpp) and launch it in router mode (`llama-server --models-dir ./models`) — see the [llama.cpp section](#llamacpp) below.
- **Ollama** — [install Ollama](https://ollama.com/download), pull a model (`ollama pull qwen2.5:latest`).
- **llama.cpp** — [build/install `llama-server`](https://github.com/ggml-org/llama.cpp), launch it in [router mode](#llamacpp).
## Quickstart
1. Right-click the canvas → Add Node → **dv/** for Format String, Random Choice, and Circuit Breaker.
2. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
---
## Format String
Formats text from a Python f-string or Jinja2 template. As you type the template, input sockets appear and disappear automatically — one per variable detected.
### Python f-strings
Type `{variable_name}` and a socket appears. Wire it to any string output in your workflow.
Type a template, get sockets. `{variable_name}` in f-string mode, `{{ variable_name }}` in Jinja2 mode — either way, comfydv watches what you type and keeps the node's inputs in sync automatically.
![Format String — f-string mode](assets/fstring.png)
| Output | Content |
|--------|---------|
| `formatted_string` | The rendered result |
| `saved_file_path` | Path written to disk (if `save_path` is set) |
| `<var>` … | Pass-through of each input value, for easy chaining |
| `saved_file_path` | Where it was written, if `save_path` is set |
| `<var>` … | Each input passed through unchanged, for easy chaining |
### Jinja2 templates
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`, …), conditionals (`{% if %}…{% endif %}`), and loops.
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`), conditionals, and loops:
![Format String — Jinja2 mode](assets/jinja2.png)
Variables detected in `{{ }}` expressions become input sockets exactly as in Simple mode. See the [Jinja2 documentation](https://jinja.palletsprojects.com/en/latest/) for the full filter/test reference.
---
## Random Choice
Connect any number of inputs of the same type. Each run picks one at random. Set `seed` for reproducibility.
Wire in any number of same-typed inputs — images, strings, conditioning, anything ComfyUI can carry over a socket — and get one back at random.
![Random Choice](assets/random.png)
- Accepts any ComfyUI type (STRING, IMAGE, CONDITIONING, …)
- Add as many inputs as you like; unused slots are removed automatically when disconnected
- `seed = 0` randomises on every run; any other value locks the selection
`seed = 0` randomises every run; any other value locks the selection. Unused input slots vanish automatically when you disconnect them.
---
## Circuit Breaker
Stops the queue gracefully when a condition isn't met — no crash, no error, just a clean halt.
Stop a queue run cleanly when a condition isn't met, instead of letting a downstream node crash on bad input.
![Circuit Breaker](assets/circuit_breaker.png)
Wire an image (or any trigger) into `trigger` and a boolean into `status`. When `status` is **false** the node raises `InterruptProcessingException`, which tells ComfyUI to stop the current run cleanly. When `status` is **true** the image passes through unchanged.
Wire a trigger (an image, or anything) into `trigger` and a boolean into `status`. `status = false` raises `InterruptProcessingException` — ComfyUI stops the run without an error dialog. `status = true` passes the trigger straight through.
Typical use: skip an expensive upscale step when a quality-check node says the draft is already good enough.
Typical use: skip an expensive upscale pass when an upstream quality-check node says the draft's already good enough.
---
## Ollama
## Local LLMs
Nodes for integrating a local Ollama LLM into your ComfyUI workflow. The host is configured once in **Ollama Client** and threaded through the graph as an `LLM_CLIENT` socket — a generic connection type any future backend's client node can also emit, so the chat/model-management nodes below aren't Ollama-specific.
One set of nodes, two interchangeable backends. Configure a connection once with **Ollama Client** or **LlamaCpp Client** — both output the same `LLM_CLIENT` socket — and every downstream node (model selection, load/unload, chat, structured output, multi-turn history) works exactly the same way regardless of which one you picked. Swapping backends means rewiring one node, not rebuilding your graph.
### Ollama Client node
Configure the server address once; all downstream nodes inherit it automatically.
### Connect and chat
![Ollama Client](assets/ollama_client.png)
### Model lifecycle (load and unload)
On memory-constrained machines and single-GPU setups, explicitly loading and unloading the model before and after inference is critical. **LLM Load Model** pins the model into VRAM (`keep_alive=-1`); **LLM Unload Model** evicts it immediately (`keep_alive=0`), freeing memory for image generation or other models.
![Ollama Load / Unload](assets/ollama_lifecycle.png)
The correct chain is **Load → Chat → Unload**, enforced through data dependencies:
1. Wire `LLMLoadModel.model_name` → `ChatCompletion.model`. This creates the data dependency that guarantees Load runs before Chat and passes the model name into the Chat node's plain-string `model` input.
2. Wire `ChatCompletion.model_name` → `LLMUnloadModel.model`. This guarantees Unload runs after Chat completes.
3. Optionally wire `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged so the rest of your workflow can still consume it.
### Minimal chat workflow
1. **Ollama Client** → set host (default `http://localhost:11434`)
2. **LLM Model Selector** → pick a model from the live dropdown (or type/wire a model name directly into Chat Completion's `model` input)
3. **Chat Completion** → wire client + model + prompt; the response appears inline in the node body and is also available as an output socket
1. **Ollama Client** (default `http://localhost:11434`) or **LlamaCpp Client** (default `http://localhost:8080`) — set the host.
2. **LLM Model Selector** — pick a model from the live dropdown, or wire a model name straight into Chat Completion.
3. **Chat Completion** — wire in client, model, and prompt. The response renders inline in the node and is also available as an output socket.
![Chat Completion](assets/ollama_chat.png)
Wire multiple nodes together for a complete end-to-end workflow:
A complete graph looks like this:
![Ollama Full Workflow](assets/ollama_workflow.png)
![Full LLM workflow](assets/ollama_workflow.png)
### Option nodes
### Manual memory management
Chain any combination of **Ollama Option —** nodes before Chat Completion to override inference parameters:
Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline.
| Option node | Ollama param |
|-------------|-------------|
![Load / Unload lifecycle](assets/ollama_lifecycle.png)
The **Load → Chat → Unload** chain is enforced by data dependencies, not by convention:
1. `LLMLoadModel.model_name` → `ChatCompletion.model` — guarantees Load runs before Chat, and feeds the model name straight in.
2. `ChatCompletion.model_name` → `LLMUnloadModel.model` — guarantees Unload runs after Chat completes.
3. *(Optional)* `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged, so the rest of your workflow can still consume it.
### Tuning generation
Chain any combination of **Ollama Option —** nodes ahead of Chat Completion to override inference parameters:
![Ollama Option nodes](assets/ollama_options.png)
| Option node | Parameter |
|-------------|-----------|
| Temperature | `temperature` |
| Seed | `seed` |
| Max Tokens | `num_predict` |
| Top P | `top_p` |
| Top K | `top_k` |
| Repeat Penalty | `repeat_penalty` |
| Extra Body | arbitrary JSON merged into options |
![Ollama Option Nodes](assets/ollama_options.png)
| Extra Body | arbitrary JSON, merged into `options` |
### Multi-turn conversations
`OLLAMA_HISTORY` flows out of Chat Completion as a list of `{"role", "content"}` dicts. Wire it back into the next Chat Completion for multi-turn conversations, or inspect it with **Ollama Debug History** / **Ollama History Length**.
`OLLAMA_HISTORY` flows out of Chat Completion as a `{"role", "content"}` list. Feed it back into the next Chat Completion call for multi-turn context, or inspect it with **Ollama Debug History** / **Ollama History Length**.
### Upgrading an older workflow
### llama.cpp
If you saved a workflow before this rename, ComfyUI will report the old node types as missing when you reopen it. Reconnect using this mapping, then re-run — behavior is unchanged, only the names and the client socket type are different:
Everything above works unchanged against llama.cpp — swap in a **LlamaCpp Client** and the rest of the graph doesn't know the difference. The one thing llama.cpp needs that Ollama doesn't: **router mode**, a directory of models rather than a single `-m model.gguf`:
```bash
llama-server --models-dir ./models -c 8192
```
In exchange, router mode gives comfydv a richer live status than Ollama can report — `loading` and `downloading`, not just loaded/unloaded — plus the same explicit load/unload primitives Ollama's nodes already use.
### Upgrading a workflow saved before this rename
Nodes were renamed once, to make them backend-generic (`OllamaChatCompletion` → `ChatCompletion`, etc.). If ComfyUI reports old node types as missing when you reopen a saved workflow, reconnect using this table — behavior is unchanged, only the names are:
| Old | New |
|-----|-----|
@@ -153,28 +166,4 @@ If you saved a workflow before this rename, ComfyUI will report the old node typ
| `OllamaUnloadModel` | `LLMUnloadModel` |
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
---
## llama.cpp
A second backend for the same chat/model-management nodes documented above — **LlamaCpp Client** is the only new node; everything else (Chat Completion, LLM Model Selector, LLM Load Model, LLM Unload Model, structured output, multi-turn history) works unchanged, because they don't know or care which backend they're talking to.
### Prerequisite: router mode
`llama-server` needs to be launched in **router mode** — a directory of models, not a single `-m model.gguf`:
```bash
llama-server --models-dir ./models -c 8192
```
This gives `comfydv` live model status (including `loading`/`downloading`, not just loaded/unloaded — a richer picture than Ollama can report) and explicit load/unload, the same way the Ollama nodes already work.
### LlamaCpp Client node
Configure the server address once (default `http://localhost:8080`); every downstream node inherits it automatically — same pattern as Ollama Client, same `LLM_CLIENT` socket.
### Switching an existing workflow from Ollama to llama.cpp
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed at your running `llama-server`. Nothing else changes — same Chat Completion node, same Load/Unload nodes, same structured-output behavior. That's the entire point of sharing one `LLM_CLIENT` socket type across backends.
`OllamaClient` kept its name — delete and re-add any node showing as missing, then rewire it to the same `OllamaClient` node.
+154 -11
View File
@@ -117,7 +117,11 @@ async def _capture(
pad: int = 60,
) -> None:
"""Screenshot the canvas, optionally cropped tightly around the node."""
canvas = page.locator("canvas#graph-canvas, canvas").first
# ComfyUI's frontend added a minimap canvas since this script was last
# verified — it appears before #graph-canvas in DOM order, so a bare
# union selector's `.first` silently grabbed the 250x200 minimap
# instead of the real graph. Target #graph-canvas explicitly.
canvas = page.locator("canvas#graph-canvas")
box = await canvas.bounding_box()
if box is None:
await page.screenshot(path=str(out))
@@ -307,13 +311,13 @@ async def scene_ollama_client(page: Page, out: Path) -> None:
async def scene_ollama_chat(page: Page, out: Path) -> None:
"""OllamaChatCompletion — showing the live model dropdown and prompt widget."""
"""ChatCompletion — showing the live model dropdown and prompt widget."""
await _clear(page)
info = await page.evaluate(
"""
async () => {
const node = LiteGraph.createNode("OllamaChatCompletion");
const node = LiteGraph.createNode("ChatCompletion");
node.pos = [60, 60];
window.app.graph.add(node);
@@ -327,7 +331,7 @@ async def scene_ollama_chat(page: Page, out: Path) -> None:
} else {
// Manually fetch and populate the COMBO
try {
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
if (resp.ok) {
const data = await resp.json();
const models = data.models || [];
@@ -357,8 +361,142 @@ async def scene_ollama_chat(page: Page, out: Path) -> None:
await _capture(page, out, info["pos"], info["size"])
async def scene_structured_output(page: Page, out: Path) -> None:
"""ChatCompletion with structured_output enabled — the live dynamic
output sockets (summary/sentiment/confidence) that appear as soon as
output_schema is edited, no graph run required. Exercises the real
js/ollama.js widget callbacks (structuredWidget.callback /
schemaWidget.callback), the same code path a live user's checkbox click
and schema edit trigger — not a hand-simulated approximation."""
await _clear(page)
schema = (
'{"type": "object", "properties": '
'{"summary": {"type": "string"}, '
'"sentiment": {"type": "string"}, '
'"confidence": {"type": "number"}}, '
'"required": ["summary", "sentiment", "confidence"]}'
)
info = await page.evaluate(
f"""
async () => {{
const node = LiteGraph.createNode("ChatCompletion");
node.pos = [60, 60];
window.app.graph.add(node);
const promptWidget = node.widgets.find(w => w.name === "prompt");
if (promptWidget) promptWidget.value =
"The new render pipeline cut our export time in half and the team is thrilled.";
const structuredWidget = node.widgets.find(w => w.name === "structured_output");
const schemaWidget = node.widgets.find(w => w.name === "output_schema");
if (structuredWidget) structuredWidget.value = true;
if (schemaWidget) schemaWidget.value = {json.dumps(schema)};
// Fire the same callbacks js/ollama.js attaches on node creation —
// real widget-edit code path, not a re-implementation.
if (structuredWidget?.callback) await structuredWidget.callback(true);
if (schemaWidget?.callback) await schemaWidget.callback({json.dumps(schema)});
await new Promise(r => setTimeout(r, 600));
window.app.canvas.setDirty(true, true);
window.app.canvas.draw(true, true);
return {{ pos: [node.pos[0], node.pos[1]], size: [node.size[0], node.size[1]] }};
}}
"""
)
await asyncio.sleep(0.8)
await _redraw(page)
await _frame_node(page, info["pos"], info["size"])
await _capture(page, out, info["pos"], info["size"])
async def scene_llamacpp_client(page: Page, out: Path) -> None:
"""LlamaCppClient — single node showing the router-mode host URL widget."""
await _clear(page)
info = await page.evaluate(
"""
() => {
const node = LiteGraph.createNode("LlamaCppClient");
node.pos = [60, 60];
window.app.graph.add(node);
const hostWidget = node.widgets && node.widgets.find(w => w.name === "host");
if (hostWidget) hostWidget.value = "http://localhost:8080";
window.app.canvas.setDirty(true, true);
window.app.canvas.draw(true, true);
return { pos: [node.pos[0], node.pos[1]], size: [node.size[0], node.size[1]] };
}
"""
)
await _frame_node(page, info["pos"], info["size"])
await _capture(page, out, info["pos"], info["size"])
async def scene_llamacpp_workflow(page: Page, out: Path) -> None:
"""LlamaCppClient → the same ChatCompletion node the Ollama workflow
uses, unmodified — the actual point of the adapter pattern. Attempts a
live model-list refresh via backend=llamacpp against
host.docker.internal:8080; degrades gracefully (same as a real user's
"no server running yet" state) if nothing is listening there."""
await _clear(page)
info = await page.evaluate(
"""
async () => {
const graph = window.app.graph;
const client = LiteGraph.createNode("LlamaCppClient");
client.pos = [40, 60];
graph.add(client);
const hostWidget = client.widgets.find(w => w.name === "host");
if (hostWidget) hostWidget.value = "http://host.docker.internal:8080";
const chat = LiteGraph.createNode("ChatCompletion");
chat.pos = [380, 40];
graph.add(chat);
const promptWidget = chat.widgets.find(w => w.name === "prompt");
if (promptWidget) promptWidget.value = "Write a haiku about ComfyUI.";
client.connect(0, chat, 0);
try {
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:8080&backend=llamacpp");
if (resp.ok) {
const data = await resp.json();
const models = data.models || [];
if (models.length) {
const modelWidget = chat.widgets.find(w => w.name === "model");
if (modelWidget) modelWidget.value = models[0];
}
}
} catch (e) {}
await new Promise(r => setTimeout(r, 800));
window.app.canvas.setDirty(true, true);
window.app.canvas.draw(true, true);
const nodes = [client, chat];
const minX = Math.min(...nodes.map(n => n.pos[0])) - 20;
const minY = Math.min(...nodes.map(n => n.pos[1])) - 20;
const maxX = Math.max(...nodes.map(n => n.pos[0] + n.size[0])) + 20;
const maxY = Math.max(...nodes.map(n => n.pos[1] + n.size[1])) + 20;
return { pos: [minX, minY], size: [maxX - minX, maxY - minY] };
}
"""
)
await asyncio.sleep(1.0)
await _redraw(page)
await _frame_node(page, info["pos"], info["size"], scale=1.0)
await _capture(page, out, info["pos"], info["size"], scale=1.0)
async def scene_ollama_workflow(page: Page, out: Path) -> None:
"""Full mini-workflow: OllamaClient → OllamaChatCompletion + Temperature + Seed options."""
"""Full mini-workflow: OllamaClient → ChatCompletion + Temperature + Seed options."""
await _clear(page)
info = await page.evaluate(
@@ -386,7 +524,7 @@ async def scene_ollama_workflow(page: Page, out: Path) -> None:
if (twSeed) twSeed.value = 42;
// 4. ChatCompletion — right
const chat = LiteGraph.createNode("OllamaChatCompletion");
const chat = LiteGraph.createNode("ChatCompletion");
chat.pos = [380, 100];
graph.add(chat);
const twPrompt = chat.widgets && chat.widgets.find(w => w.name === "prompt");
@@ -409,7 +547,7 @@ async def scene_ollama_workflow(page: Page, out: Path) -> None:
// Refresh model dropdowns for chat node
try {
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
if (resp.ok) {
const data = await resp.json();
const models = data.models || [];
@@ -466,25 +604,25 @@ async def scene_ollama_lifecycle(page: Page, out: Path) -> None:
if (hostWidget) hostWidget.value = "http://localhost:11434";
// OllamaLoadModel — generous gap right of client
const load = LiteGraph.createNode("OllamaLoadModel");
const load = LiteGraph.createNode("LLMLoadModel");
load.pos = [380, 80];
graph.add(load);
// OllamaChatCompletion — wide node, plenty of space to the right of load
const chat = LiteGraph.createNode("OllamaChatCompletion");
const chat = LiteGraph.createNode("ChatCompletion");
chat.pos = [720, 40];
graph.add(chat);
const twPrompt = chat.widgets && chat.widgets.find(w => w.name === "prompt");
if (twPrompt) twPrompt.value = "Describe this image in one sentence.";
// OllamaUnloadModel — far right, vertically offset to match chat's outputs
const unload = LiteGraph.createNode("OllamaUnloadModel");
const unload = LiteGraph.createNode("LLMUnloadModel");
unload.pos = [1200, 280];
graph.add(unload);
// Populate model dropdowns from live Ollama
try {
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
if (resp.ok) {
const data = await resp.json();
const models = data.models || [];
@@ -604,6 +742,11 @@ SCENES = [
("ollama_workflow.png", scene_ollama_workflow),
("ollama_options.png", scene_ollama_options),
("ollama_lifecycle.png", scene_ollama_lifecycle),
# Structured output (ADR-007 / pydantic-ai)
("structured_output.png", scene_structured_output),
# llama.cpp (spec 008)
("llamacpp_client.png", scene_llamacpp_client),
("llamacpp_workflow.png", scene_llamacpp_workflow),
]
+1 -1
View File
@@ -29,7 +29,7 @@ from pydantic_ai.models.openai import OpenAIChatModel
from pydantic_ai.providers.openai import OpenAIProvider
from pydantic_ai.settings import ModelSettings
from comfydv._llm.provider import Message
from .provider import Message
_STRUCTURED_OUTPUT_FAILURE_EXCEPTIONS = (
UnexpectedModelBehavior,
+23 -3
View File
@@ -17,8 +17,8 @@ import logging
from pydantic import BaseModel
from comfydv._llm.ollama_provider import _TTLLRUCache, _cache_key, _get_json, _post_json
from comfydv._llm.provider import Message, ModelInfo, ModelStatus
from .ollama_provider import _TTLLRUCache, _cache_key, _get_json, _post_json
from .provider import Message, ModelInfo, ModelStatus
logger = logging.getLogger(__name__)
@@ -31,6 +31,26 @@ _MODEL_LIST_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=20.0)
_CHAT_RESPONSE_CACHE = _TTLLRUCache(maxsize=64, ttl_seconds=None)
async def _fetch_models(host: str, headers: dict | None = None) -> list[str]:
"""Name-only view for ComfyUI's combo-widget population (the JS refresh
button and node-creation auto-populate) — mirrors
ollama_provider._fetch_models's narrower, gracefully-degrading contract.
Deliberately more forgiving than LlamaCppProvider.list_models(): that
method raises on a non-router-mode server (FR-006, for real workflow
execution, where a silent empty result would be misleading). This
combo-population use case wants the same quiet "just show an empty
dropdown" degradation Ollama's nodes already give for *any* failure —
consistent UX across backends for this specific, lower-stakes path.
"""
try:
models = await LlamaCppProvider(host, headers).list_models()
except Exception as exc:
logger.warning("Could not fetch llama.cpp models from %s: %s", host, exc)
return []
return [m.name for m in models]
class LlamaCppProvider:
"""LLMProvider implementation backed by llama-server's router mode.
@@ -184,7 +204,7 @@ class LlamaCppProvider:
timeout_secs: float = 300.0,
max_retries: int = 2,
) -> BaseModel:
from comfydv._llm.chat import chat_structured as _chat_structured_impl
from .chat import chat_structured as _chat_structured_impl
payload_messages = [m.model_dump() for m in messages]
cache_key = _cache_key(
+2 -2
View File
@@ -18,7 +18,7 @@ import time
from pydantic import BaseModel
from comfydv._llm.provider import Message, ModelInfo, ModelStatus
from .provider import Message, ModelInfo, ModelStatus
logger = logging.getLogger(__name__)
@@ -296,7 +296,7 @@ class OllamaProvider:
timeout_secs: float = 300.0,
max_retries: int = 2,
) -> BaseModel:
from comfydv._llm.chat import chat_structured as _chat_structured_impl
from .chat import chat_structured as _chat_structured_impl
payload_messages = [m.model_dump() for m in messages]
cache_key = _cache_key(
+1 -1
View File
@@ -10,7 +10,7 @@ Deployment prerequisite: llama-server must be launched in router mode
(--models-dir or --models-preset) — see specs/008-llamacpp-integration/quickstart.md.
"""
from comfydv._llm.llamacpp_provider import LlamaCppProvider
from ._llm.llamacpp_provider import LlamaCppProvider
class LlamaCppClient:
+9 -3
View File
@@ -21,8 +21,8 @@ import logging
import os
import sys
from comfydv._llm.ollama_provider import OllamaProvider, _fetch_models, _run_async
from comfydv._llm.provider import Message
from ._llm.ollama_provider import OllamaProvider, _fetch_models, _run_async
from ._llm.provider import Message
logger = logging.getLogger(__name__)
@@ -119,7 +119,13 @@ if "comfy" in sys.modules:
from aiohttp import web
host = request.rel_url.query.get("host", "http://localhost:11434")
models = await _fetch_models(host)
backend = request.rel_url.query.get("backend", "ollama")
if backend == "llamacpp":
from ._llm.llamacpp_provider import _fetch_models as _fetch
models = await _fetch(host)
else:
models = await _fetch_models(host)
if models:
return web.json_response({"models": models})
return web.json_response({"error": f"No models found at {host}"}, status=503)
+42 -33
View File
@@ -1,11 +1,15 @@
/**
* ollama.js — ComfyUI frontend extension for comfydv Ollama nodes.
* ollama.js — ComfyUI frontend extension for comfydv's generic LLM nodes.
*
* Populates the model widget on Ollama nodes from a live call to
* GET /dv/ollama/models?host=<url>.
* Populates the model widget on LLM nodes from a live call to
* GET /dv/ollama/models?host=<url>&backend=<ollama|llamacpp>. Despite the
* file/route name (kept for historical reasons — see MIGRATION_MAP in
* comfydv.ollama), this now serves both backends: which one a given node's
* upstream client is determines the `backend` param (see
* getHostAndBackendFromNode below).
*
* OllamaModelSelector and OllamaLoadModel use a COMBO widget (dropdown).
* OllamaChatCompletion uses a plain STRING widget (accepts wired values).
* LLMModelSelector and LLMLoadModel use a COMBO widget (dropdown).
* ChatCompletion uses a plain STRING widget (accepts wired values).
* The Refresh button works the same way for both: it fetches the live list
* and sets the widget value / updates COMBO options as appropriate.
*/
@@ -13,12 +17,15 @@
import { app } from "../../scripts/app.js";
/** Nodes whose model widget is a COMBO dropdown. */
const OLLAMA_COMBO_NODES = new Set(["OllamaModelSelector", "OllamaLoadModel"]);
const LLM_COMBO_NODES = new Set(["LLMModelSelector", "LLMLoadModel"]);
/** Nodes whose model widget is a plain STRING (accepts wired input). */
const OLLAMA_STRING_MODEL_NODES = new Set(["OllamaChatCompletion"]);
const LLM_STRING_MODEL_NODES = new Set(["ChatCompletion"]);
const OLLAMA_ALL_NODES = new Set([...OLLAMA_COMBO_NODES, ...OLLAMA_STRING_MODEL_NODES]);
const LLM_ALL_NODES = new Set([...LLM_COMBO_NODES, ...LLM_STRING_MODEL_NODES]);
/** Registered client node type -> backend param the /dv/ollama/models route expects. */
const CLIENT_NODE_BACKENDS = { OllamaClient: "ollama", LlamaCppClient: "llamacpp" };
/**
* Fetch model list and update the node's model widget.
@@ -27,9 +34,11 @@ const OLLAMA_ALL_NODES = new Set([...OLLAMA_COMBO_NODES, ...OLLAMA_STRING_MODEL_
* - STRING: sets the value to the first model; keeps existing value if it
* still appears in the live list (user may have typed a valid name).
*/
async function refreshModelWidget(node, host) {
async function refreshModelWidget(node, host, backend) {
try {
const resp = await fetch(`/dv/ollama/models?host=${encodeURIComponent(host)}`);
const resp = await fetch(
`/dv/ollama/models?host=${encodeURIComponent(host)}&backend=${encodeURIComponent(backend)}`
);
if (!resp.ok) return;
const data = await resp.json();
const models = data.models ?? [];
@@ -53,54 +62,53 @@ async function refreshModelWidget(node, host) {
node.setDirtyCanvas(true, false);
} catch (_) {
// Ollama unreachable — leave widget unchanged
// Server unreachable — leave widget unchanged
}
}
/**
* Locate the host string for a node.
* Locate the host and backend for a node's connected LLM client.
*
* First checks the node's own widgets (OllamaClient has a "host" widget).
* Otherwise traverses graph links to find a connected OllamaClient node and
* reads its "host" widget — this is the common case for downstream nodes.
* Traverses graph links to find a connected OllamaClient or LlamaCppClient
* node and reads its "host" widget. Falls back to Ollama's default if
* nothing is wired yet, matching the pre-existing fallback behavior.
*/
function getHostFromNode(node) {
const ownHostWidget = node.widgets?.find(w => w.name === "host");
if (ownHostWidget) return ownHostWidget.value;
function getHostAndBackendFromNode(node) {
for (const input of node.inputs ?? []) {
if (!input.link) continue;
const link = node.graph?.links[input.link];
if (!link) continue;
const sourceNode = node.graph?.getNodeById(link.origin_id);
if (sourceNode?.type === "OllamaClient") {
const backend = sourceNode ? CLIENT_NODE_BACKENDS[sourceNode.type] : undefined;
if (backend) {
const hostWidget = sourceNode.widgets?.find(w => w.name === "host");
if (hostWidget?.value) return hostWidget.value;
if (hostWidget?.value) return { host: hostWidget.value, backend };
}
}
return "http://localhost:11434";
return { host: "http://localhost:11434", backend: "ollama" };
}
app.registerExtension({
name: "comfydv.ollama",
async beforeRegisterNodeDef(nodeType, nodeData) {
if (!OLLAMA_ALL_NODES.has(nodeData.name)) return;
if (!LLM_ALL_NODES.has(nodeData.name)) return;
const onNodeCreated = nodeType.prototype.onNodeCreated;
nodeType.prototype.onNodeCreated = function () {
const result = onNodeCreated?.apply(this, arguments);
const refresh = () => {
const { host, backend } = getHostAndBackendFromNode(this);
refreshModelWidget(this, host, backend);
};
// Add a Refresh button below the model widget
this.addWidget("button", "⟳ Refresh models", null, () => {
const host = getHostFromNode(this);
refreshModelWidget(this, host);
});
this.addWidget("button", "⟳ Refresh models", null, refresh);
// Initial population on node creation
const host = getHostFromNode(this);
refreshModelWidget(this, host);
refresh();
return result;
};
@@ -108,15 +116,16 @@ app.registerExtension({
});
/**
* Live structured-output dynamic sockets for OllamaChatCompletion.
* Live structured-output dynamic sockets for ChatCompletion.
*
* Mirrors FormatString's live dynamic-output pattern (see format_string.js):
* editing structured_output or output_schema posts to a backend route that
* recomputes OllamaChatCompletion.RETURN_TYPES/RETURN_NAMES (the same
* recomputes ChatCompletion.RETURN_TYPES/RETURN_NAMES (the same
* update_outputs() path chat() itself uses at execution time) and returns
* the resulting output list — applied to this node's sockets immediately,
* so you see the extracted fields appear without having to run the graph
* first.
* first. Backend-agnostic: ChatCompletion is the one generic node both
* OllamaProvider and LlamaCppProvider feed.
*/
async function updateStructuredOutputs(node, structuredOutput, outputSchema) {
try {
@@ -148,7 +157,7 @@ app.registerExtension({
name: "comfydv.ollama.structuredOutput",
async beforeRegisterNodeDef(nodeType, nodeData) {
if (nodeData.name !== "OllamaChatCompletion") return;
if (nodeData.name !== "ChatCompletion") return;
const onNodeCreated = nodeType.prototype.onNodeCreated;
nodeType.prototype.onNodeCreated = function () {
+148
View File
@@ -0,0 +1,148 @@
"""Guards against the exact bug found while validating spec 008 against a
real ComfyUI dev harness (docker-compose): every `from comfydv._llm.X
import Y`-style absolute self-import inside src/comfydv/ silently broke the
*entire* plugin (every node, not just LLM ones) as soon as ComfyUI actually
loaded it.
ComfyUI's custom_nodes loader imports the plugin via a *relative* chain —
the repo-root __init__.py does `from .src.comfydv import ...`, nesting
comfydv under whatever top-level name the folder has (never `comfydv`
itself). An absolute `from comfydv...` self-import only resolves if `src/`
has separately been placed on sys.path — which conftest.py does for every
other test file in this suite, masking the bug completely. This file
deliberately does NOT rely on that sys.path insertion: it reproduces
ComfyUI's actual nested-relative-import shape in a subprocess.
Confirmed via git history: this predates spec 008 entirely — it was already
broken immediately after PR #17 merged (spec 007), well before llamacpp.py
existed. No test caught it because none exercised this exact loading shape
until the docker harness was run by hand.
"""
import subprocess
import sys
import textwrap
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).parent.parent
@pytest.fixture(autouse=True)
def _clear_ollama_caches():
"""Shadow conftest.py's autouse fixture of the same name for this module
only. That fixture's own setup does `from comfydv._llm.ollama_provider
import ...` — this file's tests are the exact reproduction of an
environment where `comfydv` resolving correctly can't be assumed (that's
the point of the file), so depending on it for an unrelated cache-reset
would make these tests order-dependent on whichever other test file
happens to import `comfydv` "the normal way" first in the session. These
tests touch no OllamaProvider/ChatCompletion state, so there is nothing
to reset."""
yield
_SUBPROCESS_SCRIPT = textwrap.dedent(
"""
import sys
import types
# Minimal ComfyUI stubs — same shape as conftest.py's pytest_configure,
# but this script intentionally runs outside pytest so it isn't reusing
# (or accidentally validated by) that fixture's sys.path setup.
class _InterruptProcessingException(Exception):
pass
comfy_module = types.ModuleType("comfy")
comfy_module.model_management = types.SimpleNamespace(
InterruptProcessingException=_InterruptProcessingException
)
sys.modules["comfy"] = comfy_module
sys.modules["comfy.model_management"] = comfy_module.model_management
class _Routes:
def post(self, path):
return lambda fn: fn
def get(self, path):
return lambda fn: fn
class _PromptServer:
pass
_PromptServer.instance = _PromptServer()
_PromptServer.instance.routes = _Routes()
server_module = types.ModuleType("server")
server_module.PromptServer = _PromptServer
sys.modules["server"] = server_module
folder_paths_module = types.ModuleType("folder_paths")
folder_paths_module.get_output_directory = lambda: "/tmp/comfydv_test"
sys.modules["folder_paths"] = folder_paths_module
# The critical part: put the repo's *parent* directory on sys.path, so
# `import comfydv` resolves to the repo-root __init__.py — exactly how
# ComfyUI resolves a folder under custom_nodes/ — NOT to src/comfydv
# directly (that's what conftest.py's sys.path.insert(0, ".../src")
# does for the rest of this test suite, and why it never caught this).
sys.path.insert(0, sys.argv[1])
import comfydv
required = {"FormatString", "RandomChoice", "CircuitBreaker",
"OllamaClient", "LlamaCppClient", "ChatCompletion"}
missing = required - set(comfydv.NODE_CLASS_MAPPINGS)
if missing:
print(f"MISSING_NODES:{missing}")
sys.exit(1)
print("OK")
"""
)
def test_package_imports_under_comfyui_style_relative_nesting():
"""Reproduces ComfyUI's real loading shape and fails loudly — with the
actual traceback — if any internal module reverts to an absolute
`from comfydv...` self-import."""
result = subprocess.run(
[sys.executable, "-c", _SUBPROCESS_SCRIPT, str(REPO_ROOT.parent)],
cwd=REPO_ROOT,
capture_output=True,
text=True,
timeout=30,
)
assert result.returncode == 0, (
"comfydv failed to import the way ComfyUI actually loads it "
"(relative nesting, not a top-level `comfydv` on sys.path). "
f"This means every node in the plugin would fail to register.\n"
f"--- stdout ---\n{result.stdout}\n--- stderr ---\n{result.stderr}"
)
assert "OK" in result.stdout
def test_no_absolute_self_imports_in_package():
"""Cheap, fast static guard alongside the dynamic test above: no file
under src/comfydv/ should import itself as `comfydv.X` — internal
imports must be relative (`.X` / `..X`) so they resolve regardless of
what the outer package happens to be named at load time."""
import ast
offenders = []
for path in (REPO_ROOT / "src" / "comfydv").rglob("*.py"):
tree = ast.parse(path.read_text(), filename=str(path))
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom):
if node.module and (
node.module == "comfydv" or node.module.startswith("comfydv.")
):
offenders.append(f"{path.relative_to(REPO_ROOT)}:{node.lineno}")
elif isinstance(node, ast.Import):
for alias in node.names:
if alias.name == "comfydv" or alias.name.startswith("comfydv."):
offenders.append(f"{path.relative_to(REPO_ROOT)}:{node.lineno}")
assert not offenders, (
"Absolute self-imports found — use relative imports instead "
f"(they break under ComfyUI's actual loader): {offenders}"
)
+48 -1
View File
@@ -13,7 +13,7 @@ BDD coverage:
import pytest
import comfydv._llm.llamacpp_provider as provider_mod
from comfydv._llm.llamacpp_provider import LlamaCppProvider
from comfydv._llm.llamacpp_provider import LlamaCppProvider, _fetch_models
from comfydv._llm.ollama_provider import _run_async
from comfydv._llm.provider import Message, ModelStatus
@@ -355,3 +355,50 @@ def test_chat_structured_forwards_options(monkeypatch):
)
assert captured["options"] == {"temperature": 0.0}
# ---------------------------------------------------------------------------
# _fetch_models — name-only view used by ComfyUI's /dv/ollama/models?backend=
# llamacpp route (the JS refresh button / node-creation auto-populate).
# Deliberately more forgiving than list_models(): degrades to [] on any
# failure rather than raising on a non-router-mode server, matching the
# combo-widget UX OllamaProvider's own _fetch_models already gives.
# ---------------------------------------------------------------------------
def test_fetch_models_returns_name_only_list(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
return {
"data": [
{"id": "a", "status": {"value": "loaded"}},
{"id": "b", "status": {"value": "unloaded"}},
]
}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
names = _run_async(_fetch_models("http://localhost:8080"))
assert names == ["a", "b"]
def test_fetch_models_degrades_to_empty_on_non_router_mode(monkeypatch):
"""Unlike list_models() (FR-006), this combo-population view swallows
even the non-router-mode error — a quiet empty dropdown, not a toast."""
async def fake_get(url, *, timeout=5.0, headers=None):
raise RuntimeError("Server returned HTTP 404 for http://x/models: not found")
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
names = _run_async(_fetch_models("http://localhost:8080"))
assert names == []
def test_fetch_models_degrades_to_empty_when_unreachable(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
raise ConnectionError("no route to host")
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
names = _run_async(_fetch_models("http://localhost:19999"))
assert names == []
+86
View File
@@ -824,6 +824,92 @@ class _FakeRequest:
return self._body
class _FakeRelUrl:
def __init__(self, query: dict):
self.query = query
class _FakeGetRequest:
def __init__(self, query: dict):
self.rel_url = _FakeRelUrl(query)
class TestModelsRoute:
"""GET /dv/ollama/models — the JS refresh-button/auto-populate endpoint.
Previously untested (no test survived the atomic cutover rewrite), which
is exactly how a real bug shipped undetected: the `backend` dispatch
added here didn't exist until a live js/ollama.js bug was found (it
matched on pre-rename node names and always spoke Ollama's wire
protocol regardless of which client was actually connected)."""
def _call(self, query: dict):
import comfydv.ollama as ollama_mod
from comfydv._llm.ollama_provider import _run_async
resp = _run_async(ollama_mod._models_endpoint(_FakeGetRequest(query)))
return json.loads(resp.text), resp.status
def test_default_backend_dispatches_to_ollama(self, monkeypatch):
async def fake_fetch(host, headers=None):
assert host == "http://localhost:11434"
return ["a", "b"]
monkeypatch.setattr("comfydv.ollama._fetch_models", fake_fetch)
data, status = self._call({"host": "http://localhost:11434"})
assert status == 200
assert data == {"models": ["a", "b"]}
def test_explicit_ollama_backend_dispatches_to_ollama(self, monkeypatch):
async def fake_fetch(host, headers=None):
return ["m"]
monkeypatch.setattr("comfydv.ollama._fetch_models", fake_fetch)
data, status = self._call(
{"host": "http://localhost:11434", "backend": "ollama"}
)
assert status == 200
assert data == {"models": ["m"]}
def test_llamacpp_backend_dispatches_to_llamacpp_fetch(self, monkeypatch):
"""The bug this regression-tests: before the fix, this endpoint
always called Ollama's _fetch_models regardless of `backend`, so a
llama.cpp host's models never populated the dropdown."""
called = {}
async def fake_llamacpp_fetch(host, headers=None):
called["host"] = host
return ["gemma-3-4b"]
async def fail_ollama_fetch(host, headers=None):
raise AssertionError("must not call Ollama's fetcher for backend=llamacpp")
monkeypatch.setattr(
"comfydv._llm.llamacpp_provider._fetch_models", fake_llamacpp_fetch
)
monkeypatch.setattr("comfydv.ollama._fetch_models", fail_ollama_fetch)
data, status = self._call(
{"host": "http://localhost:8080", "backend": "llamacpp"}
)
assert status == 200
assert data == {"models": ["gemma-3-4b"]}
assert called["host"] == "http://localhost:8080"
def test_no_models_returns_503(self, monkeypatch):
async def fake_fetch(host, headers=None):
return []
monkeypatch.setattr("comfydv.ollama._fetch_models", fake_fetch)
data, status = self._call({"host": "http://localhost:11434"})
assert status == 503
assert "error" in data
class TestUpdateStructuredOutputsRoute:
def teardown_method(self):
# This route mutates ChatCompletion's class-level RETURN_TYPES