Merge pull request #21 from darth-veitcher/fix/comfyui-relative-import-loading
@@ -1,29 +1,37 @@
|
||||
# comfydv
|
||||
|
||||
A collection of workflow efficiency and quality-of-life nodes built out of necessity for personal ComfyUI use.
|
||||
**Quality-of-life nodes for ComfyUI, built to disappear into your workflow.**
|
||||
|
||||
## What is this?
|
||||
`comfydv` fills the gaps ComfyUI's built-in library leaves on the table: string templates that build their own sockets as you type, seed-controlled randomisation, graceful mid-queue interruption, and a local-LLM integration that doesn't care whether you're running Ollama or llama.cpp. No Python required — install it, drop the nodes on your canvas, wire them up.
|
||||
|
||||
`comfydv` fills gaps in ComfyUI's built-in node library: dynamic string formatting, seed-controlled random selection, graceful workflow interruption, and local LLM integration (Ollama and llama.cpp). Install it once and connect the nodes like any other — no Python knowledge required.
|
||||

|
||||
|
||||
## What is comfydv?
|
||||
|
||||
A small, focused ComfyUI utility pack. It exists because:
|
||||
|
||||
- **It reads your intent, not just your syntax.** Format String detects `{variables}` in a template and adds/removes input sockets live, as you type — no manual socket wrangling.
|
||||
- **One LLM integration, any local backend.** Wire a Chat Completion node once; swap between Ollama and llama.cpp by changing a single upstream client node. Structured output, multi-turn history, and model load/unload work identically on both.
|
||||
- **It fails politely.** Circuit Breaker halts a queue run cleanly instead of throwing a stack trace at you; a disconnected LLM server gets a specific, actionable error instead of a silent empty dropdown.
|
||||
- **Small, tested, boring in the best way.** Every node is unit-tested and the local-LLM nodes are verified against real running servers, not just mocks.
|
||||
|
||||
## What's inside
|
||||
|
||||
| Node | What it does |
|
||||
|------|-------------|
|
||||
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
|
||||
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
|
||||
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
|
||||
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the connection through the graph as an `LLM_CLIENT` socket — a generic connection type any backend's client node emits. |
|
||||
| **LlamaCpp Client** | Configures a connection to a `llama-server` instance running in router mode (default: `http://localhost:8080`). Emits the same `LLM_CLIENT` socket as Ollama Client — every node below works with either. |
|
||||
| **LLM Model Selector** | Fetches the live model list from the connected server and presents it as a dropdown. Outputs the selected model name. |
|
||||
| **LLM Load Model** | Loads a model into memory on the connected server. |
|
||||
| **LLM Unload Model** | Evicts a model from memory on the connected server. |
|
||||
| **Chat Completion** | Sends a prompt (and optional conversation history) to the connected server. Response and history are shown inline in the node body and available as output sockets. |
|
||||
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
|
||||
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
|
||||
| **Ollama History Length** | Returns the number of messages in an `OLLAMA_HISTORY` list as an integer. |
|
||||
|------|---------------|
|
||||
| **Format String** | Renders a Python f-string or Jinja2 template. Sockets appear and disappear automatically as you type variables. |
|
||||
| **Random Choice** | Accepts any number of typed inputs and returns one at random, with a seed for reproducibility. |
|
||||
| **Circuit Breaker** | Halts the current queue run gracefully — no crash, just a clean stop — when a condition isn't met. |
|
||||
| **Ollama Client** / **LlamaCpp Client** | Configure a connection to a local Ollama or llama.cpp server. Both emit the same `LLM_CLIENT` socket — every node below works with either. |
|
||||
| **LLM Model Selector** | Live dropdown of models available on the connected server. |
|
||||
| **LLM Load Model** / **LLM Unload Model** | Explicit VRAM management — pin a model in memory before inference, evict it after. |
|
||||
| **Chat Completion** | Send a prompt (optionally with history) to the connected server; response shown inline and as an output socket. |
|
||||
| **Ollama Option — \*** | Seven composable parameter nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into Chat Completion's `options` input. |
|
||||
| **Ollama Debug History** / **Ollama History Length** | Inspect an `OLLAMA_HISTORY` conversation list — pretty-print it or count its messages. |
|
||||
|
||||
## Install
|
||||
|
||||
**Via ComfyUI Manager** (recommended): search for `comfydv` and click Install.
|
||||
**Via ComfyUI Manager** (recommended): search for `comfydv`, click Install.
|
||||
|
||||
**Manual:**
|
||||
|
||||
@@ -32,134 +40,125 @@ cd /path/to/ComfyUI/custom_nodes
|
||||
git clone https://github.com/darth-veitcher/comfydv.git
|
||||
```
|
||||
|
||||
Restart ComfyUI. The nodes appear under the **dv/**, **dv/ollama**, and **dv/llamacpp** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) are installed automatically via `requirements.txt`.
|
||||
Restart ComfyUI. Nodes appear under **dv/**, **dv/ollama**, and **dv/llamacpp** in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) install automatically via `requirements.txt`.
|
||||
|
||||
For local LLM nodes, pick one backend (or both):
|
||||
For local-LLM nodes, bring your own backend:
|
||||
|
||||
- **Ollama**: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`).
|
||||
- **llama.cpp**: [build/install `llama-server`](https://github.com/ggml-org/llama.cpp) and launch it in router mode (`llama-server --models-dir ./models`) — see the [llama.cpp section](#llamacpp) below.
|
||||
- **Ollama** — [install Ollama](https://ollama.com/download), pull a model (`ollama pull qwen2.5:latest`).
|
||||
- **llama.cpp** — [build/install `llama-server`](https://github.com/ggml-org/llama.cpp), launch it in [router mode](#llamacpp).
|
||||
|
||||
## Quickstart
|
||||
|
||||
1. Install via ComfyUI Manager (search `comfydv`) or clone manually into `custom_nodes/`.
|
||||
2. Right-click the canvas → Add Node → **dv/** to find Format String, Random Choice, and Circuit Breaker.
|
||||
3. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
|
||||
1. Right-click the canvas → Add Node → **dv/** for Format String, Random Choice, and Circuit Breaker.
|
||||
2. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
|
||||
|
||||
## Documentation
|
||||
|
||||
Full documentation: [darth-veitcher.github.io/comfydv](https://darth-veitcher.github.io/comfydv/stable/)
|
||||
Full documentation, including every node's inputs/outputs: **[darth-veitcher.github.io/comfydv](https://darth-veitcher.github.io/comfydv/stable/)**
|
||||
|
||||
---
|
||||
|
||||
## Format String
|
||||
|
||||
Formats text from a Python f-string or Jinja2 template. As you type the template, input sockets appear and disappear automatically — one per variable detected.
|
||||
|
||||
### Python f-strings
|
||||
|
||||
Type `{variable_name}` and a socket appears. Wire it to any string output in your workflow.
|
||||
Type a template, get sockets. `{variable_name}` in f-string mode, `{{ variable_name }}` in Jinja2 mode — either way, comfydv watches what you type and keeps the node's inputs in sync automatically.
|
||||
|
||||

|
||||
|
||||
Outputs are always in a stable order:
|
||||
|
||||
| Output | Content |
|
||||
|--------|---------|
|
||||
| `formatted_string` | The rendered result |
|
||||
| `saved_file_path` | Path written to disk (if `save_path` is set) |
|
||||
| `<var>` … | Pass-through of each input value, for easy chaining |
|
||||
| `saved_file_path` | Where it was written, if `save_path` is set |
|
||||
| `<var>` … | Each input passed through unchanged, for easy chaining |
|
||||
|
||||
### Jinja2 templates
|
||||
|
||||
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`, …), conditionals (`{% if %}…{% endif %}`), and loops.
|
||||
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`), conditionals, and loops:
|
||||
|
||||

|
||||
|
||||
Variables detected in `{{ }}` expressions become input sockets exactly as in Simple mode. See the [Jinja2 documentation](https://jinja.palletsprojects.com/en/latest/) for the full filter/test reference.
|
||||
|
||||
---
|
||||
|
||||
## Random Choice
|
||||
|
||||
Connect any number of inputs of the same type. Each run picks one at random. Set `seed` for reproducibility.
|
||||
Wire in any number of same-typed inputs — images, strings, conditioning, anything ComfyUI can carry over a socket — and get one back at random.
|
||||
|
||||

|
||||
|
||||
- Accepts any ComfyUI type (STRING, IMAGE, CONDITIONING, …)
|
||||
- Add as many inputs as you like; unused slots are removed automatically when disconnected
|
||||
- `seed = 0` randomises on every run; any other value locks the selection
|
||||
`seed = 0` randomises every run; any other value locks the selection. Unused input slots vanish automatically when you disconnect them.
|
||||
|
||||
---
|
||||
|
||||
## Circuit Breaker
|
||||
|
||||
Stops the queue gracefully when a condition isn't met — no crash, no error, just a clean halt.
|
||||
Stop a queue run cleanly when a condition isn't met, instead of letting a downstream node crash on bad input.
|
||||
|
||||

|
||||
|
||||
Wire an image (or any trigger) into `trigger` and a boolean into `status`. When `status` is **false** the node raises `InterruptProcessingException`, which tells ComfyUI to stop the current run cleanly. When `status` is **true** the image passes through unchanged.
|
||||
Wire a trigger (an image, or anything) into `trigger` and a boolean into `status`. `status = false` raises `InterruptProcessingException` — ComfyUI stops the run without an error dialog. `status = true` passes the trigger straight through.
|
||||
|
||||
Typical use: skip an expensive upscale step when a quality-check node says the draft is already good enough.
|
||||
Typical use: skip an expensive upscale pass when an upstream quality-check node says the draft's already good enough.
|
||||
|
||||
---
|
||||
|
||||
## Ollama
|
||||
## Local LLMs
|
||||
|
||||
Nodes for integrating a local Ollama LLM into your ComfyUI workflow. The host is configured once in **Ollama Client** and threaded through the graph as an `LLM_CLIENT` socket — a generic connection type any future backend's client node can also emit, so the chat/model-management nodes below aren't Ollama-specific.
|
||||
One set of nodes, two interchangeable backends. Configure a connection once with **Ollama Client** or **LlamaCpp Client** — both output the same `LLM_CLIENT` socket — and every downstream node (model selection, load/unload, chat, structured output, multi-turn history) works exactly the same way regardless of which one you picked. Swapping backends means rewiring one node, not rebuilding your graph.
|
||||
|
||||
### Ollama Client node
|
||||
|
||||
Configure the server address once; all downstream nodes inherit it automatically.
|
||||
### Connect and chat
|
||||
|
||||

|
||||
|
||||
### Model lifecycle (load and unload)
|
||||
|
||||
On memory-constrained machines and single-GPU setups, explicitly loading and unloading the model before and after inference is critical. **LLM Load Model** pins the model into VRAM (`keep_alive=-1`); **LLM Unload Model** evicts it immediately (`keep_alive=0`), freeing memory for image generation or other models.
|
||||
|
||||

|
||||
|
||||
The correct chain is **Load → Chat → Unload**, enforced through data dependencies:
|
||||
|
||||
1. Wire `LLMLoadModel.model_name` → `ChatCompletion.model`. This creates the data dependency that guarantees Load runs before Chat and passes the model name into the Chat node's plain-string `model` input.
|
||||
2. Wire `ChatCompletion.model_name` → `LLMUnloadModel.model`. This guarantees Unload runs after Chat completes.
|
||||
3. Optionally wire `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged so the rest of your workflow can still consume it.
|
||||
|
||||
### Minimal chat workflow
|
||||
|
||||
1. **Ollama Client** → set host (default `http://localhost:11434`)
|
||||
2. **LLM Model Selector** → pick a model from the live dropdown (or type/wire a model name directly into Chat Completion's `model` input)
|
||||
3. **Chat Completion** → wire client + model + prompt; the response appears inline in the node body and is also available as an output socket
|
||||
1. **Ollama Client** (default `http://localhost:11434`) or **LlamaCpp Client** (default `http://localhost:8080`) — set the host.
|
||||
2. **LLM Model Selector** — pick a model from the live dropdown, or wire a model name straight into Chat Completion.
|
||||
3. **Chat Completion** — wire in client, model, and prompt. The response renders inline in the node and is also available as an output socket.
|
||||
|
||||

|
||||
|
||||
Wire multiple nodes together for a complete end-to-end workflow:
|
||||
A complete graph looks like this:
|
||||
|
||||

|
||||

|
||||
|
||||
### Option nodes
|
||||
### Manual memory management
|
||||
|
||||
Chain any combination of **Ollama Option —** nodes before Chat Completion to override inference parameters:
|
||||
Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline.
|
||||
|
||||
| Option node | Ollama param |
|
||||
|-------------|-------------|
|
||||

|
||||
|
||||
The **Load → Chat → Unload** chain is enforced by data dependencies, not by convention:
|
||||
|
||||
1. `LLMLoadModel.model_name` → `ChatCompletion.model` — guarantees Load runs before Chat, and feeds the model name straight in.
|
||||
2. `ChatCompletion.model_name` → `LLMUnloadModel.model` — guarantees Unload runs after Chat completes.
|
||||
3. *(Optional)* `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged, so the rest of your workflow can still consume it.
|
||||
|
||||
### Tuning generation
|
||||
|
||||
Chain any combination of **Ollama Option —** nodes ahead of Chat Completion to override inference parameters:
|
||||
|
||||

|
||||
|
||||
| Option node | Parameter |
|
||||
|-------------|-----------|
|
||||
| Temperature | `temperature` |
|
||||
| Seed | `seed` |
|
||||
| Max Tokens | `num_predict` |
|
||||
| Top P | `top_p` |
|
||||
| Top K | `top_k` |
|
||||
| Repeat Penalty | `repeat_penalty` |
|
||||
| Extra Body | arbitrary JSON merged into options |
|
||||
|
||||

|
||||
| Extra Body | arbitrary JSON, merged into `options` |
|
||||
|
||||
### Multi-turn conversations
|
||||
|
||||
`OLLAMA_HISTORY` flows out of Chat Completion as a list of `{"role", "content"}` dicts. Wire it back into the next Chat Completion for multi-turn conversations, or inspect it with **Ollama Debug History** / **Ollama History Length**.
|
||||
`OLLAMA_HISTORY` flows out of Chat Completion as a `{"role", "content"}` list. Feed it back into the next Chat Completion call for multi-turn context, or inspect it with **Ollama Debug History** / **Ollama History Length**.
|
||||
|
||||
### Upgrading an older workflow
|
||||
### llama.cpp
|
||||
|
||||
If you saved a workflow before this rename, ComfyUI will report the old node types as missing when you reopen it. Reconnect using this mapping, then re-run — behavior is unchanged, only the names and the client socket type are different:
|
||||
Everything above works unchanged against llama.cpp — swap in a **LlamaCpp Client** and the rest of the graph doesn't know the difference. The one thing llama.cpp needs that Ollama doesn't: **router mode**, a directory of models rather than a single `-m model.gguf`:
|
||||
|
||||
```bash
|
||||
llama-server --models-dir ./models -c 8192
|
||||
```
|
||||
|
||||
In exchange, router mode gives comfydv a richer live status than Ollama can report — `loading` and `downloading`, not just loaded/unloaded — plus the same explicit load/unload primitives Ollama's nodes already use.
|
||||
|
||||
### Upgrading a workflow saved before this rename
|
||||
|
||||
Nodes were renamed once, to make them backend-generic (`OllamaChatCompletion` → `ChatCompletion`, etc.). If ComfyUI reports old node types as missing when you reopen a saved workflow, reconnect using this table — behavior is unchanged, only the names are:
|
||||
|
||||
| Old | New |
|
||||
|-----|-----|
|
||||
@@ -169,30 +168,10 @@ If you saved a workflow before this rename, ComfyUI will report the old node typ
|
||||
| `OllamaUnloadModel` | `LLMUnloadModel` |
|
||||
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
|
||||
|
||||
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
|
||||
`OllamaClient` kept its name — delete and re-add any node showing as missing, then rewire it to the same `OllamaClient` node.
|
||||
|
||||
---
|
||||
|
||||
## llama.cpp
|
||||
## License
|
||||
|
||||
A second backend for the same chat/model-management nodes documented above — **LlamaCpp Client** is the only new node; everything else (Chat Completion, LLM Model Selector, LLM Load Model, LLM Unload Model, structured output, multi-turn history) works unchanged, because they don't know or care which backend they're talking to.
|
||||
|
||||
### Prerequisite: router mode
|
||||
|
||||
`llama-server` needs to be launched in **router mode** — a directory of models, not a single `-m model.gguf`:
|
||||
|
||||
```bash
|
||||
llama-server --models-dir ./models -c 8192
|
||||
```
|
||||
|
||||
This gives `comfydv` live model status (including `loading`/`downloading`, not just loaded/unloaded — a richer picture than Ollama can report) and explicit load/unload, the same way the Ollama nodes already work.
|
||||
|
||||
### LlamaCpp Client node
|
||||
|
||||
Configure the server address once (default `http://localhost:8080`); every downstream node inherits it automatically — same pattern as Ollama Client, same `LLM_CLIENT` socket.
|
||||
|
||||
### Switching an existing workflow from Ollama to llama.cpp
|
||||
|
||||
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed at your running `llama-server`. Nothing else changes — same Chat Completion node, same Load/Unload nodes, same structured-output behavior. That's the entire point of sharing one `LLM_CLIENT` socket type across backends.
|
||||
|
||||
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
|
||||
[AGPL-3.0](LICENSE)
|
||||
|
||||
|
Before Width: | Height: | Size: 12 KiB After Width: | Height: | Size: 12 KiB |
|
Before Width: | Height: | Size: 32 KiB After Width: | Height: | Size: 33 KiB |
|
Before Width: | Height: | Size: 40 KiB After Width: | Height: | Size: 41 KiB |
|
After Width: | Height: | Size: 13 KiB |
|
After Width: | Height: | Size: 37 KiB |
|
Before Width: | Height: | Size: 38 KiB After Width: | Height: | Size: 55 KiB |
|
Before Width: | Height: | Size: 11 KiB After Width: | Height: | Size: 12 KiB |
|
Before Width: | Height: | Size: 50 KiB After Width: | Height: | Size: 55 KiB |
|
Before Width: | Height: | Size: 40 KiB After Width: | Height: | Size: 41 KiB |
|
Before Width: | Height: | Size: 58 KiB After Width: | Height: | Size: 64 KiB |
|
Before Width: | Height: | Size: 19 KiB After Width: | Height: | Size: 18 KiB |
|
After Width: | Height: | Size: 68 KiB |
@@ -1,25 +1,37 @@
|
||||
# comfydv
|
||||
|
||||
A collection of workflow efficiency and quality-of-life nodes built out of necessity for personal ComfyUI use.
|
||||
**Quality-of-life nodes for ComfyUI, built to disappear into your workflow.**
|
||||
|
||||
`comfydv` fills the gaps ComfyUI's built-in library leaves on the table: string templates that build their own sockets as you type, seed-controlled randomisation, graceful mid-queue interruption, and a local-LLM integration that doesn't care whether you're running Ollama or llama.cpp. No Python required — install it, drop the nodes on your canvas, wire them up.
|
||||
|
||||

|
||||
|
||||
## What is comfydv?
|
||||
|
||||
A small, focused ComfyUI utility pack. It exists because:
|
||||
|
||||
- **It reads your intent, not just your syntax.** Format String detects `{variables}` in a template and adds/removes input sockets live, as you type — no manual socket wrangling.
|
||||
- **One LLM integration, any local backend.** Wire a Chat Completion node once; swap between Ollama and llama.cpp by changing a single upstream client node. Structured output, multi-turn history, and model load/unload work identically on both.
|
||||
- **It fails politely.** Circuit Breaker halts a queue run cleanly instead of throwing a stack trace at you; a disconnected LLM server gets a specific, actionable error instead of a silent empty dropdown.
|
||||
- **Small, tested, boring in the best way.** Every node is unit-tested and the local-LLM nodes are verified against real running servers, not just mocks.
|
||||
|
||||
## What's inside
|
||||
|
||||
| Node | What it does |
|
||||
|------|-------------|
|
||||
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
|
||||
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
|
||||
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
|
||||
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the connection through the graph as an `LLM_CLIENT` socket — a generic connection type any backend's client node emits. |
|
||||
| **LlamaCpp Client** | Configures a connection to a `llama-server` instance running in router mode (default: `http://localhost:8080`). Emits the same `LLM_CLIENT` socket as Ollama Client — every node below works with either. |
|
||||
| **LLM Model Selector** | Fetches the live model list from the connected server and presents it as a dropdown. Outputs the selected model name. |
|
||||
| **LLM Load Model** | Loads a model into memory on the connected server. |
|
||||
| **LLM Unload Model** | Evicts a model from memory on the connected server. |
|
||||
| **Chat Completion** | Sends a prompt (and optional conversation history) to the connected server. Response and history are shown inline in the node body and available as output sockets. |
|
||||
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
|
||||
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
|
||||
| **Ollama History Length** | Returns the number of messages in an `OLLAMA_HISTORY` list as an integer. |
|
||||
|------|---------------|
|
||||
| **Format String** | Renders a Python f-string or Jinja2 template. Sockets appear and disappear automatically as you type variables. |
|
||||
| **Random Choice** | Accepts any number of typed inputs and returns one at random, with a seed for reproducibility. |
|
||||
| **Circuit Breaker** | Halts the current queue run gracefully — no crash, just a clean stop — when a condition isn't met. |
|
||||
| **Ollama Client** / **LlamaCpp Client** | Configure a connection to a local Ollama or llama.cpp server. Both emit the same `LLM_CLIENT` socket — every node below works with either. |
|
||||
| **LLM Model Selector** | Live dropdown of models available on the connected server. |
|
||||
| **LLM Load Model** / **LLM Unload Model** | Explicit VRAM management — pin a model in memory before inference, evict it after. |
|
||||
| **Chat Completion** | Send a prompt (optionally with history) to the connected server; response shown inline and as an output socket. |
|
||||
| **Ollama Option — \*** | Seven composable parameter nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into Chat Completion's `options` input. |
|
||||
| **Ollama Debug History** / **Ollama History Length** | Inspect an `OLLAMA_HISTORY` conversation list — pretty-print it or count its messages. |
|
||||
|
||||
## Install
|
||||
|
||||
**Via ComfyUI Manager** (recommended): search for `comfydv` and click Install.
|
||||
**Via ComfyUI Manager** (recommended): search for `comfydv`, click Install.
|
||||
|
||||
**Manual:**
|
||||
|
||||
@@ -28,122 +40,123 @@ cd /path/to/ComfyUI/custom_nodes
|
||||
git clone https://github.com/darth-veitcher/comfydv.git
|
||||
```
|
||||
|
||||
Restart ComfyUI. The nodes appear under the **dv/**, **dv/ollama**, and **dv/llamacpp** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) are installed automatically via `requirements.txt`.
|
||||
Restart ComfyUI. Nodes appear under **dv/**, **dv/ollama**, and **dv/llamacpp** in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) install automatically via `requirements.txt`.
|
||||
|
||||
For local LLM nodes, pick one backend (or both):
|
||||
For local-LLM nodes, bring your own backend:
|
||||
|
||||
- **Ollama**: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`).
|
||||
- **llama.cpp**: [build/install `llama-server`](https://github.com/ggml-org/llama.cpp) and launch it in router mode (`llama-server --models-dir ./models`) — see the [llama.cpp section](#llamacpp) below.
|
||||
- **Ollama** — [install Ollama](https://ollama.com/download), pull a model (`ollama pull qwen2.5:latest`).
|
||||
- **llama.cpp** — [build/install `llama-server`](https://github.com/ggml-org/llama.cpp), launch it in [router mode](#llamacpp).
|
||||
|
||||
## Quickstart
|
||||
|
||||
1. Right-click the canvas → Add Node → **dv/** for Format String, Random Choice, and Circuit Breaker.
|
||||
2. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
|
||||
|
||||
---
|
||||
|
||||
## Format String
|
||||
|
||||
Formats text from a Python f-string or Jinja2 template. As you type the template, input sockets appear and disappear automatically — one per variable detected.
|
||||
|
||||
### Python f-strings
|
||||
|
||||
Type `{variable_name}` and a socket appears. Wire it to any string output in your workflow.
|
||||
Type a template, get sockets. `{variable_name}` in f-string mode, `{{ variable_name }}` in Jinja2 mode — either way, comfydv watches what you type and keeps the node's inputs in sync automatically.
|
||||
|
||||

|
||||
|
||||
| Output | Content |
|
||||
|--------|---------|
|
||||
| `formatted_string` | The rendered result |
|
||||
| `saved_file_path` | Path written to disk (if `save_path` is set) |
|
||||
| `<var>` … | Pass-through of each input value, for easy chaining |
|
||||
| `saved_file_path` | Where it was written, if `save_path` is set |
|
||||
| `<var>` … | Each input passed through unchanged, for easy chaining |
|
||||
|
||||
### Jinja2 templates
|
||||
|
||||
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`, …), conditionals (`{% if %}…{% endif %}`), and loops.
|
||||
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`), conditionals, and loops:
|
||||
|
||||

|
||||
|
||||
Variables detected in `{{ }}` expressions become input sockets exactly as in Simple mode. See the [Jinja2 documentation](https://jinja.palletsprojects.com/en/latest/) for the full filter/test reference.
|
||||
|
||||
---
|
||||
|
||||
## Random Choice
|
||||
|
||||
Connect any number of inputs of the same type. Each run picks one at random. Set `seed` for reproducibility.
|
||||
Wire in any number of same-typed inputs — images, strings, conditioning, anything ComfyUI can carry over a socket — and get one back at random.
|
||||
|
||||

|
||||
|
||||
- Accepts any ComfyUI type (STRING, IMAGE, CONDITIONING, …)
|
||||
- Add as many inputs as you like; unused slots are removed automatically when disconnected
|
||||
- `seed = 0` randomises on every run; any other value locks the selection
|
||||
`seed = 0` randomises every run; any other value locks the selection. Unused input slots vanish automatically when you disconnect them.
|
||||
|
||||
---
|
||||
|
||||
## Circuit Breaker
|
||||
|
||||
Stops the queue gracefully when a condition isn't met — no crash, no error, just a clean halt.
|
||||
Stop a queue run cleanly when a condition isn't met, instead of letting a downstream node crash on bad input.
|
||||
|
||||

|
||||
|
||||
Wire an image (or any trigger) into `trigger` and a boolean into `status`. When `status` is **false** the node raises `InterruptProcessingException`, which tells ComfyUI to stop the current run cleanly. When `status` is **true** the image passes through unchanged.
|
||||
Wire a trigger (an image, or anything) into `trigger` and a boolean into `status`. `status = false` raises `InterruptProcessingException` — ComfyUI stops the run without an error dialog. `status = true` passes the trigger straight through.
|
||||
|
||||
Typical use: skip an expensive upscale step when a quality-check node says the draft is already good enough.
|
||||
Typical use: skip an expensive upscale pass when an upstream quality-check node says the draft's already good enough.
|
||||
|
||||
---
|
||||
|
||||
## Ollama
|
||||
## Local LLMs
|
||||
|
||||
Nodes for integrating a local Ollama LLM into your ComfyUI workflow. The host is configured once in **Ollama Client** and threaded through the graph as an `LLM_CLIENT` socket — a generic connection type any future backend's client node can also emit, so the chat/model-management nodes below aren't Ollama-specific.
|
||||
One set of nodes, two interchangeable backends. Configure a connection once with **Ollama Client** or **LlamaCpp Client** — both output the same `LLM_CLIENT` socket — and every downstream node (model selection, load/unload, chat, structured output, multi-turn history) works exactly the same way regardless of which one you picked. Swapping backends means rewiring one node, not rebuilding your graph.
|
||||
|
||||
### Ollama Client node
|
||||
|
||||
Configure the server address once; all downstream nodes inherit it automatically.
|
||||
### Connect and chat
|
||||
|
||||

|
||||
|
||||
### Model lifecycle (load and unload)
|
||||
|
||||
On memory-constrained machines and single-GPU setups, explicitly loading and unloading the model before and after inference is critical. **LLM Load Model** pins the model into VRAM (`keep_alive=-1`); **LLM Unload Model** evicts it immediately (`keep_alive=0`), freeing memory for image generation or other models.
|
||||
|
||||

|
||||
|
||||
The correct chain is **Load → Chat → Unload**, enforced through data dependencies:
|
||||
|
||||
1. Wire `LLMLoadModel.model_name` → `ChatCompletion.model`. This creates the data dependency that guarantees Load runs before Chat and passes the model name into the Chat node's plain-string `model` input.
|
||||
2. Wire `ChatCompletion.model_name` → `LLMUnloadModel.model`. This guarantees Unload runs after Chat completes.
|
||||
3. Optionally wire `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged so the rest of your workflow can still consume it.
|
||||
|
||||
### Minimal chat workflow
|
||||
|
||||
1. **Ollama Client** → set host (default `http://localhost:11434`)
|
||||
2. **LLM Model Selector** → pick a model from the live dropdown (or type/wire a model name directly into Chat Completion's `model` input)
|
||||
3. **Chat Completion** → wire client + model + prompt; the response appears inline in the node body and is also available as an output socket
|
||||
1. **Ollama Client** (default `http://localhost:11434`) or **LlamaCpp Client** (default `http://localhost:8080`) — set the host.
|
||||
2. **LLM Model Selector** — pick a model from the live dropdown, or wire a model name straight into Chat Completion.
|
||||
3. **Chat Completion** — wire in client, model, and prompt. The response renders inline in the node and is also available as an output socket.
|
||||
|
||||

|
||||
|
||||
Wire multiple nodes together for a complete end-to-end workflow:
|
||||
A complete graph looks like this:
|
||||
|
||||

|
||||

|
||||
|
||||
### Option nodes
|
||||
### Manual memory management
|
||||
|
||||
Chain any combination of **Ollama Option —** nodes before Chat Completion to override inference parameters:
|
||||
Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline.
|
||||
|
||||
| Option node | Ollama param |
|
||||
|-------------|-------------|
|
||||

|
||||
|
||||
The **Load → Chat → Unload** chain is enforced by data dependencies, not by convention:
|
||||
|
||||
1. `LLMLoadModel.model_name` → `ChatCompletion.model` — guarantees Load runs before Chat, and feeds the model name straight in.
|
||||
2. `ChatCompletion.model_name` → `LLMUnloadModel.model` — guarantees Unload runs after Chat completes.
|
||||
3. *(Optional)* `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged, so the rest of your workflow can still consume it.
|
||||
|
||||
### Tuning generation
|
||||
|
||||
Chain any combination of **Ollama Option —** nodes ahead of Chat Completion to override inference parameters:
|
||||
|
||||

|
||||
|
||||
| Option node | Parameter |
|
||||
|-------------|-----------|
|
||||
| Temperature | `temperature` |
|
||||
| Seed | `seed` |
|
||||
| Max Tokens | `num_predict` |
|
||||
| Top P | `top_p` |
|
||||
| Top K | `top_k` |
|
||||
| Repeat Penalty | `repeat_penalty` |
|
||||
| Extra Body | arbitrary JSON merged into options |
|
||||
|
||||

|
||||
| Extra Body | arbitrary JSON, merged into `options` |
|
||||
|
||||
### Multi-turn conversations
|
||||
|
||||
`OLLAMA_HISTORY` flows out of Chat Completion as a list of `{"role", "content"}` dicts. Wire it back into the next Chat Completion for multi-turn conversations, or inspect it with **Ollama Debug History** / **Ollama History Length**.
|
||||
`OLLAMA_HISTORY` flows out of Chat Completion as a `{"role", "content"}` list. Feed it back into the next Chat Completion call for multi-turn context, or inspect it with **Ollama Debug History** / **Ollama History Length**.
|
||||
|
||||
### Upgrading an older workflow
|
||||
### llama.cpp
|
||||
|
||||
If you saved a workflow before this rename, ComfyUI will report the old node types as missing when you reopen it. Reconnect using this mapping, then re-run — behavior is unchanged, only the names and the client socket type are different:
|
||||
Everything above works unchanged against llama.cpp — swap in a **LlamaCpp Client** and the rest of the graph doesn't know the difference. The one thing llama.cpp needs that Ollama doesn't: **router mode**, a directory of models rather than a single `-m model.gguf`:
|
||||
|
||||
```bash
|
||||
llama-server --models-dir ./models -c 8192
|
||||
```
|
||||
|
||||
In exchange, router mode gives comfydv a richer live status than Ollama can report — `loading` and `downloading`, not just loaded/unloaded — plus the same explicit load/unload primitives Ollama's nodes already use.
|
||||
|
||||
### Upgrading a workflow saved before this rename
|
||||
|
||||
Nodes were renamed once, to make them backend-generic (`OllamaChatCompletion` → `ChatCompletion`, etc.). If ComfyUI reports old node types as missing when you reopen a saved workflow, reconnect using this table — behavior is unchanged, only the names are:
|
||||
|
||||
| Old | New |
|
||||
|-----|-----|
|
||||
@@ -153,28 +166,4 @@ If you saved a workflow before this rename, ComfyUI will report the old node typ
|
||||
| `OllamaUnloadModel` | `LLMUnloadModel` |
|
||||
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
|
||||
|
||||
`OllamaClient` keeps its name — just delete and re-add any downstream node showing as missing, then rewire it to the same `OllamaClient` node.
|
||||
|
||||
---
|
||||
|
||||
## llama.cpp
|
||||
|
||||
A second backend for the same chat/model-management nodes documented above — **LlamaCpp Client** is the only new node; everything else (Chat Completion, LLM Model Selector, LLM Load Model, LLM Unload Model, structured output, multi-turn history) works unchanged, because they don't know or care which backend they're talking to.
|
||||
|
||||
### Prerequisite: router mode
|
||||
|
||||
`llama-server` needs to be launched in **router mode** — a directory of models, not a single `-m model.gguf`:
|
||||
|
||||
```bash
|
||||
llama-server --models-dir ./models -c 8192
|
||||
```
|
||||
|
||||
This gives `comfydv` live model status (including `loading`/`downloading`, not just loaded/unloaded — a richer picture than Ollama can report) and explicit load/unload, the same way the Ollama nodes already work.
|
||||
|
||||
### LlamaCpp Client node
|
||||
|
||||
Configure the server address once (default `http://localhost:8080`); every downstream node inherits it automatically — same pattern as Ollama Client, same `LLM_CLIENT` socket.
|
||||
|
||||
### Switching an existing workflow from Ollama to llama.cpp
|
||||
|
||||
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed at your running `llama-server`. Nothing else changes — same Chat Completion node, same Load/Unload nodes, same structured-output behavior. That's the entire point of sharing one `LLM_CLIENT` socket type across backends.
|
||||
`OllamaClient` kept its name — delete and re-add any node showing as missing, then rewire it to the same `OllamaClient` node.
|
||||
|
||||
@@ -117,7 +117,11 @@ async def _capture(
|
||||
pad: int = 60,
|
||||
) -> None:
|
||||
"""Screenshot the canvas, optionally cropped tightly around the node."""
|
||||
canvas = page.locator("canvas#graph-canvas, canvas").first
|
||||
# ComfyUI's frontend added a minimap canvas since this script was last
|
||||
# verified — it appears before #graph-canvas in DOM order, so a bare
|
||||
# union selector's `.first` silently grabbed the 250x200 minimap
|
||||
# instead of the real graph. Target #graph-canvas explicitly.
|
||||
canvas = page.locator("canvas#graph-canvas")
|
||||
box = await canvas.bounding_box()
|
||||
if box is None:
|
||||
await page.screenshot(path=str(out))
|
||||
@@ -307,13 +311,13 @@ async def scene_ollama_client(page: Page, out: Path) -> None:
|
||||
|
||||
|
||||
async def scene_ollama_chat(page: Page, out: Path) -> None:
|
||||
"""OllamaChatCompletion — showing the live model dropdown and prompt widget."""
|
||||
"""ChatCompletion — showing the live model dropdown and prompt widget."""
|
||||
await _clear(page)
|
||||
|
||||
info = await page.evaluate(
|
||||
"""
|
||||
async () => {
|
||||
const node = LiteGraph.createNode("OllamaChatCompletion");
|
||||
const node = LiteGraph.createNode("ChatCompletion");
|
||||
node.pos = [60, 60];
|
||||
window.app.graph.add(node);
|
||||
|
||||
@@ -327,7 +331,7 @@ async def scene_ollama_chat(page: Page, out: Path) -> None:
|
||||
} else {
|
||||
// Manually fetch and populate the COMBO
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
|
||||
if (resp.ok) {
|
||||
const data = await resp.json();
|
||||
const models = data.models || [];
|
||||
@@ -357,8 +361,142 @@ async def scene_ollama_chat(page: Page, out: Path) -> None:
|
||||
await _capture(page, out, info["pos"], info["size"])
|
||||
|
||||
|
||||
async def scene_structured_output(page: Page, out: Path) -> None:
|
||||
"""ChatCompletion with structured_output enabled — the live dynamic
|
||||
output sockets (summary/sentiment/confidence) that appear as soon as
|
||||
output_schema is edited, no graph run required. Exercises the real
|
||||
js/ollama.js widget callbacks (structuredWidget.callback /
|
||||
schemaWidget.callback), the same code path a live user's checkbox click
|
||||
and schema edit trigger — not a hand-simulated approximation."""
|
||||
await _clear(page)
|
||||
|
||||
schema = (
|
||||
'{"type": "object", "properties": '
|
||||
'{"summary": {"type": "string"}, '
|
||||
'"sentiment": {"type": "string"}, '
|
||||
'"confidence": {"type": "number"}}, '
|
||||
'"required": ["summary", "sentiment", "confidence"]}'
|
||||
)
|
||||
|
||||
info = await page.evaluate(
|
||||
f"""
|
||||
async () => {{
|
||||
const node = LiteGraph.createNode("ChatCompletion");
|
||||
node.pos = [60, 60];
|
||||
window.app.graph.add(node);
|
||||
|
||||
const promptWidget = node.widgets.find(w => w.name === "prompt");
|
||||
if (promptWidget) promptWidget.value =
|
||||
"The new render pipeline cut our export time in half and the team is thrilled.";
|
||||
|
||||
const structuredWidget = node.widgets.find(w => w.name === "structured_output");
|
||||
const schemaWidget = node.widgets.find(w => w.name === "output_schema");
|
||||
if (structuredWidget) structuredWidget.value = true;
|
||||
if (schemaWidget) schemaWidget.value = {json.dumps(schema)};
|
||||
|
||||
// Fire the same callbacks js/ollama.js attaches on node creation —
|
||||
// real widget-edit code path, not a re-implementation.
|
||||
if (structuredWidget?.callback) await structuredWidget.callback(true);
|
||||
if (schemaWidget?.callback) await schemaWidget.callback({json.dumps(schema)});
|
||||
|
||||
await new Promise(r => setTimeout(r, 600));
|
||||
window.app.canvas.setDirty(true, true);
|
||||
window.app.canvas.draw(true, true);
|
||||
return {{ pos: [node.pos[0], node.pos[1]], size: [node.size[0], node.size[1]] }};
|
||||
}}
|
||||
"""
|
||||
)
|
||||
|
||||
await asyncio.sleep(0.8)
|
||||
await _redraw(page)
|
||||
await _frame_node(page, info["pos"], info["size"])
|
||||
await _capture(page, out, info["pos"], info["size"])
|
||||
|
||||
|
||||
async def scene_llamacpp_client(page: Page, out: Path) -> None:
|
||||
"""LlamaCppClient — single node showing the router-mode host URL widget."""
|
||||
await _clear(page)
|
||||
|
||||
info = await page.evaluate(
|
||||
"""
|
||||
() => {
|
||||
const node = LiteGraph.createNode("LlamaCppClient");
|
||||
node.pos = [60, 60];
|
||||
window.app.graph.add(node);
|
||||
const hostWidget = node.widgets && node.widgets.find(w => w.name === "host");
|
||||
if (hostWidget) hostWidget.value = "http://localhost:8080";
|
||||
window.app.canvas.setDirty(true, true);
|
||||
window.app.canvas.draw(true, true);
|
||||
return { pos: [node.pos[0], node.pos[1]], size: [node.size[0], node.size[1]] };
|
||||
}
|
||||
"""
|
||||
)
|
||||
|
||||
await _frame_node(page, info["pos"], info["size"])
|
||||
await _capture(page, out, info["pos"], info["size"])
|
||||
|
||||
|
||||
async def scene_llamacpp_workflow(page: Page, out: Path) -> None:
|
||||
"""LlamaCppClient → the same ChatCompletion node the Ollama workflow
|
||||
uses, unmodified — the actual point of the adapter pattern. Attempts a
|
||||
live model-list refresh via backend=llamacpp against
|
||||
host.docker.internal:8080; degrades gracefully (same as a real user's
|
||||
"no server running yet" state) if nothing is listening there."""
|
||||
await _clear(page)
|
||||
|
||||
info = await page.evaluate(
|
||||
"""
|
||||
async () => {
|
||||
const graph = window.app.graph;
|
||||
|
||||
const client = LiteGraph.createNode("LlamaCppClient");
|
||||
client.pos = [40, 60];
|
||||
graph.add(client);
|
||||
const hostWidget = client.widgets.find(w => w.name === "host");
|
||||
if (hostWidget) hostWidget.value = "http://host.docker.internal:8080";
|
||||
|
||||
const chat = LiteGraph.createNode("ChatCompletion");
|
||||
chat.pos = [380, 40];
|
||||
graph.add(chat);
|
||||
const promptWidget = chat.widgets.find(w => w.name === "prompt");
|
||||
if (promptWidget) promptWidget.value = "Write a haiku about ComfyUI.";
|
||||
|
||||
client.connect(0, chat, 0);
|
||||
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:8080&backend=llamacpp");
|
||||
if (resp.ok) {
|
||||
const data = await resp.json();
|
||||
const models = data.models || [];
|
||||
if (models.length) {
|
||||
const modelWidget = chat.widgets.find(w => w.name === "model");
|
||||
if (modelWidget) modelWidget.value = models[0];
|
||||
}
|
||||
}
|
||||
} catch (e) {}
|
||||
|
||||
await new Promise(r => setTimeout(r, 800));
|
||||
window.app.canvas.setDirty(true, true);
|
||||
window.app.canvas.draw(true, true);
|
||||
|
||||
const nodes = [client, chat];
|
||||
const minX = Math.min(...nodes.map(n => n.pos[0])) - 20;
|
||||
const minY = Math.min(...nodes.map(n => n.pos[1])) - 20;
|
||||
const maxX = Math.max(...nodes.map(n => n.pos[0] + n.size[0])) + 20;
|
||||
const maxY = Math.max(...nodes.map(n => n.pos[1] + n.size[1])) + 20;
|
||||
return { pos: [minX, minY], size: [maxX - minX, maxY - minY] };
|
||||
}
|
||||
"""
|
||||
)
|
||||
|
||||
await asyncio.sleep(1.0)
|
||||
await _redraw(page)
|
||||
await _frame_node(page, info["pos"], info["size"], scale=1.0)
|
||||
await _capture(page, out, info["pos"], info["size"], scale=1.0)
|
||||
|
||||
|
||||
async def scene_ollama_workflow(page: Page, out: Path) -> None:
|
||||
"""Full mini-workflow: OllamaClient → OllamaChatCompletion + Temperature + Seed options."""
|
||||
"""Full mini-workflow: OllamaClient → ChatCompletion + Temperature + Seed options."""
|
||||
await _clear(page)
|
||||
|
||||
info = await page.evaluate(
|
||||
@@ -386,7 +524,7 @@ async def scene_ollama_workflow(page: Page, out: Path) -> None:
|
||||
if (twSeed) twSeed.value = 42;
|
||||
|
||||
// 4. ChatCompletion — right
|
||||
const chat = LiteGraph.createNode("OllamaChatCompletion");
|
||||
const chat = LiteGraph.createNode("ChatCompletion");
|
||||
chat.pos = [380, 100];
|
||||
graph.add(chat);
|
||||
const twPrompt = chat.widgets && chat.widgets.find(w => w.name === "prompt");
|
||||
@@ -409,7 +547,7 @@ async def scene_ollama_workflow(page: Page, out: Path) -> None:
|
||||
|
||||
// Refresh model dropdowns for chat node
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
|
||||
if (resp.ok) {
|
||||
const data = await resp.json();
|
||||
const models = data.models || [];
|
||||
@@ -466,25 +604,25 @@ async def scene_ollama_lifecycle(page: Page, out: Path) -> None:
|
||||
if (hostWidget) hostWidget.value = "http://localhost:11434";
|
||||
|
||||
// OllamaLoadModel — generous gap right of client
|
||||
const load = LiteGraph.createNode("OllamaLoadModel");
|
||||
const load = LiteGraph.createNode("LLMLoadModel");
|
||||
load.pos = [380, 80];
|
||||
graph.add(load);
|
||||
|
||||
// OllamaChatCompletion — wide node, plenty of space to the right of load
|
||||
const chat = LiteGraph.createNode("OllamaChatCompletion");
|
||||
const chat = LiteGraph.createNode("ChatCompletion");
|
||||
chat.pos = [720, 40];
|
||||
graph.add(chat);
|
||||
const twPrompt = chat.widgets && chat.widgets.find(w => w.name === "prompt");
|
||||
if (twPrompt) twPrompt.value = "Describe this image in one sentence.";
|
||||
|
||||
// OllamaUnloadModel — far right, vertically offset to match chat's outputs
|
||||
const unload = LiteGraph.createNode("OllamaUnloadModel");
|
||||
const unload = LiteGraph.createNode("LLMUnloadModel");
|
||||
unload.pos = [1200, 280];
|
||||
graph.add(unload);
|
||||
|
||||
// Populate model dropdowns from live Ollama
|
||||
try {
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
|
||||
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
|
||||
if (resp.ok) {
|
||||
const data = await resp.json();
|
||||
const models = data.models || [];
|
||||
@@ -604,6 +742,11 @@ SCENES = [
|
||||
("ollama_workflow.png", scene_ollama_workflow),
|
||||
("ollama_options.png", scene_ollama_options),
|
||||
("ollama_lifecycle.png", scene_ollama_lifecycle),
|
||||
# Structured output (ADR-007 / pydantic-ai)
|
||||
("structured_output.png", scene_structured_output),
|
||||
# llama.cpp (spec 008)
|
||||
("llamacpp_client.png", scene_llamacpp_client),
|
||||
("llamacpp_workflow.png", scene_llamacpp_workflow),
|
||||
]
|
||||
|
||||
|
||||
|
||||
@@ -29,7 +29,7 @@ from pydantic_ai.models.openai import OpenAIChatModel
|
||||
from pydantic_ai.providers.openai import OpenAIProvider
|
||||
from pydantic_ai.settings import ModelSettings
|
||||
|
||||
from comfydv._llm.provider import Message
|
||||
from .provider import Message
|
||||
|
||||
_STRUCTURED_OUTPUT_FAILURE_EXCEPTIONS = (
|
||||
UnexpectedModelBehavior,
|
||||
|
||||
@@ -17,8 +17,8 @@ import logging
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
from comfydv._llm.ollama_provider import _TTLLRUCache, _cache_key, _get_json, _post_json
|
||||
from comfydv._llm.provider import Message, ModelInfo, ModelStatus
|
||||
from .ollama_provider import _TTLLRUCache, _cache_key, _get_json, _post_json
|
||||
from .provider import Message, ModelInfo, ModelStatus
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -31,6 +31,26 @@ _MODEL_LIST_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=20.0)
|
||||
_CHAT_RESPONSE_CACHE = _TTLLRUCache(maxsize=64, ttl_seconds=None)
|
||||
|
||||
|
||||
async def _fetch_models(host: str, headers: dict | None = None) -> list[str]:
|
||||
"""Name-only view for ComfyUI's combo-widget population (the JS refresh
|
||||
button and node-creation auto-populate) — mirrors
|
||||
ollama_provider._fetch_models's narrower, gracefully-degrading contract.
|
||||
|
||||
Deliberately more forgiving than LlamaCppProvider.list_models(): that
|
||||
method raises on a non-router-mode server (FR-006, for real workflow
|
||||
execution, where a silent empty result would be misleading). This
|
||||
combo-population use case wants the same quiet "just show an empty
|
||||
dropdown" degradation Ollama's nodes already give for *any* failure —
|
||||
consistent UX across backends for this specific, lower-stakes path.
|
||||
"""
|
||||
try:
|
||||
models = await LlamaCppProvider(host, headers).list_models()
|
||||
except Exception as exc:
|
||||
logger.warning("Could not fetch llama.cpp models from %s: %s", host, exc)
|
||||
return []
|
||||
return [m.name for m in models]
|
||||
|
||||
|
||||
class LlamaCppProvider:
|
||||
"""LLMProvider implementation backed by llama-server's router mode.
|
||||
|
||||
@@ -184,7 +204,7 @@ class LlamaCppProvider:
|
||||
timeout_secs: float = 300.0,
|
||||
max_retries: int = 2,
|
||||
) -> BaseModel:
|
||||
from comfydv._llm.chat import chat_structured as _chat_structured_impl
|
||||
from .chat import chat_structured as _chat_structured_impl
|
||||
|
||||
payload_messages = [m.model_dump() for m in messages]
|
||||
cache_key = _cache_key(
|
||||
|
||||
@@ -18,7 +18,7 @@ import time
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
from comfydv._llm.provider import Message, ModelInfo, ModelStatus
|
||||
from .provider import Message, ModelInfo, ModelStatus
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -296,7 +296,7 @@ class OllamaProvider:
|
||||
timeout_secs: float = 300.0,
|
||||
max_retries: int = 2,
|
||||
) -> BaseModel:
|
||||
from comfydv._llm.chat import chat_structured as _chat_structured_impl
|
||||
from .chat import chat_structured as _chat_structured_impl
|
||||
|
||||
payload_messages = [m.model_dump() for m in messages]
|
||||
cache_key = _cache_key(
|
||||
|
||||
@@ -10,7 +10,7 @@ Deployment prerequisite: llama-server must be launched in router mode
|
||||
(--models-dir or --models-preset) — see specs/008-llamacpp-integration/quickstart.md.
|
||||
"""
|
||||
|
||||
from comfydv._llm.llamacpp_provider import LlamaCppProvider
|
||||
from ._llm.llamacpp_provider import LlamaCppProvider
|
||||
|
||||
|
||||
class LlamaCppClient:
|
||||
|
||||
@@ -21,8 +21,8 @@ import logging
|
||||
import os
|
||||
import sys
|
||||
|
||||
from comfydv._llm.ollama_provider import OllamaProvider, _fetch_models, _run_async
|
||||
from comfydv._llm.provider import Message
|
||||
from ._llm.ollama_provider import OllamaProvider, _fetch_models, _run_async
|
||||
from ._llm.provider import Message
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -119,7 +119,13 @@ if "comfy" in sys.modules:
|
||||
from aiohttp import web
|
||||
|
||||
host = request.rel_url.query.get("host", "http://localhost:11434")
|
||||
models = await _fetch_models(host)
|
||||
backend = request.rel_url.query.get("backend", "ollama")
|
||||
if backend == "llamacpp":
|
||||
from ._llm.llamacpp_provider import _fetch_models as _fetch
|
||||
|
||||
models = await _fetch(host)
|
||||
else:
|
||||
models = await _fetch_models(host)
|
||||
if models:
|
||||
return web.json_response({"models": models})
|
||||
return web.json_response({"error": f"No models found at {host}"}, status=503)
|
||||
|
||||
@@ -1,11 +1,15 @@
|
||||
/**
|
||||
* ollama.js — ComfyUI frontend extension for comfydv Ollama nodes.
|
||||
* ollama.js — ComfyUI frontend extension for comfydv's generic LLM nodes.
|
||||
*
|
||||
* Populates the model widget on Ollama nodes from a live call to
|
||||
* GET /dv/ollama/models?host=<url>.
|
||||
* Populates the model widget on LLM nodes from a live call to
|
||||
* GET /dv/ollama/models?host=<url>&backend=<ollama|llamacpp>. Despite the
|
||||
* file/route name (kept for historical reasons — see MIGRATION_MAP in
|
||||
* comfydv.ollama), this now serves both backends: which one a given node's
|
||||
* upstream client is determines the `backend` param (see
|
||||
* getHostAndBackendFromNode below).
|
||||
*
|
||||
* OllamaModelSelector and OllamaLoadModel use a COMBO widget (dropdown).
|
||||
* OllamaChatCompletion uses a plain STRING widget (accepts wired values).
|
||||
* LLMModelSelector and LLMLoadModel use a COMBO widget (dropdown).
|
||||
* ChatCompletion uses a plain STRING widget (accepts wired values).
|
||||
* The Refresh button works the same way for both: it fetches the live list
|
||||
* and sets the widget value / updates COMBO options as appropriate.
|
||||
*/
|
||||
@@ -13,12 +17,15 @@
|
||||
import { app } from "../../scripts/app.js";
|
||||
|
||||
/** Nodes whose model widget is a COMBO dropdown. */
|
||||
const OLLAMA_COMBO_NODES = new Set(["OllamaModelSelector", "OllamaLoadModel"]);
|
||||
const LLM_COMBO_NODES = new Set(["LLMModelSelector", "LLMLoadModel"]);
|
||||
|
||||
/** Nodes whose model widget is a plain STRING (accepts wired input). */
|
||||
const OLLAMA_STRING_MODEL_NODES = new Set(["OllamaChatCompletion"]);
|
||||
const LLM_STRING_MODEL_NODES = new Set(["ChatCompletion"]);
|
||||
|
||||
const OLLAMA_ALL_NODES = new Set([...OLLAMA_COMBO_NODES, ...OLLAMA_STRING_MODEL_NODES]);
|
||||
const LLM_ALL_NODES = new Set([...LLM_COMBO_NODES, ...LLM_STRING_MODEL_NODES]);
|
||||
|
||||
/** Registered client node type -> backend param the /dv/ollama/models route expects. */
|
||||
const CLIENT_NODE_BACKENDS = { OllamaClient: "ollama", LlamaCppClient: "llamacpp" };
|
||||
|
||||
/**
|
||||
* Fetch model list and update the node's model widget.
|
||||
@@ -27,9 +34,11 @@ const OLLAMA_ALL_NODES = new Set([...OLLAMA_COMBO_NODES, ...OLLAMA_STRING_MODEL_
|
||||
* - STRING: sets the value to the first model; keeps existing value if it
|
||||
* still appears in the live list (user may have typed a valid name).
|
||||
*/
|
||||
async function refreshModelWidget(node, host) {
|
||||
async function refreshModelWidget(node, host, backend) {
|
||||
try {
|
||||
const resp = await fetch(`/dv/ollama/models?host=${encodeURIComponent(host)}`);
|
||||
const resp = await fetch(
|
||||
`/dv/ollama/models?host=${encodeURIComponent(host)}&backend=${encodeURIComponent(backend)}`
|
||||
);
|
||||
if (!resp.ok) return;
|
||||
const data = await resp.json();
|
||||
const models = data.models ?? [];
|
||||
@@ -53,54 +62,53 @@ async function refreshModelWidget(node, host) {
|
||||
|
||||
node.setDirtyCanvas(true, false);
|
||||
} catch (_) {
|
||||
// Ollama unreachable — leave widget unchanged
|
||||
// Server unreachable — leave widget unchanged
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Locate the host string for a node.
|
||||
* Locate the host and backend for a node's connected LLM client.
|
||||
*
|
||||
* First checks the node's own widgets (OllamaClient has a "host" widget).
|
||||
* Otherwise traverses graph links to find a connected OllamaClient node and
|
||||
* reads its "host" widget — this is the common case for downstream nodes.
|
||||
* Traverses graph links to find a connected OllamaClient or LlamaCppClient
|
||||
* node and reads its "host" widget. Falls back to Ollama's default if
|
||||
* nothing is wired yet, matching the pre-existing fallback behavior.
|
||||
*/
|
||||
function getHostFromNode(node) {
|
||||
const ownHostWidget = node.widgets?.find(w => w.name === "host");
|
||||
if (ownHostWidget) return ownHostWidget.value;
|
||||
|
||||
function getHostAndBackendFromNode(node) {
|
||||
for (const input of node.inputs ?? []) {
|
||||
if (!input.link) continue;
|
||||
const link = node.graph?.links[input.link];
|
||||
if (!link) continue;
|
||||
const sourceNode = node.graph?.getNodeById(link.origin_id);
|
||||
if (sourceNode?.type === "OllamaClient") {
|
||||
const backend = sourceNode ? CLIENT_NODE_BACKENDS[sourceNode.type] : undefined;
|
||||
if (backend) {
|
||||
const hostWidget = sourceNode.widgets?.find(w => w.name === "host");
|
||||
if (hostWidget?.value) return hostWidget.value;
|
||||
if (hostWidget?.value) return { host: hostWidget.value, backend };
|
||||
}
|
||||
}
|
||||
|
||||
return "http://localhost:11434";
|
||||
return { host: "http://localhost:11434", backend: "ollama" };
|
||||
}
|
||||
|
||||
app.registerExtension({
|
||||
name: "comfydv.ollama",
|
||||
|
||||
async beforeRegisterNodeDef(nodeType, nodeData) {
|
||||
if (!OLLAMA_ALL_NODES.has(nodeData.name)) return;
|
||||
if (!LLM_ALL_NODES.has(nodeData.name)) return;
|
||||
|
||||
const onNodeCreated = nodeType.prototype.onNodeCreated;
|
||||
nodeType.prototype.onNodeCreated = function () {
|
||||
const result = onNodeCreated?.apply(this, arguments);
|
||||
|
||||
const refresh = () => {
|
||||
const { host, backend } = getHostAndBackendFromNode(this);
|
||||
refreshModelWidget(this, host, backend);
|
||||
};
|
||||
|
||||
// Add a Refresh button below the model widget
|
||||
this.addWidget("button", "⟳ Refresh models", null, () => {
|
||||
const host = getHostFromNode(this);
|
||||
refreshModelWidget(this, host);
|
||||
});
|
||||
this.addWidget("button", "⟳ Refresh models", null, refresh);
|
||||
|
||||
// Initial population on node creation
|
||||
const host = getHostFromNode(this);
|
||||
refreshModelWidget(this, host);
|
||||
refresh();
|
||||
|
||||
return result;
|
||||
};
|
||||
@@ -108,15 +116,16 @@ app.registerExtension({
|
||||
});
|
||||
|
||||
/**
|
||||
* Live structured-output dynamic sockets for OllamaChatCompletion.
|
||||
* Live structured-output dynamic sockets for ChatCompletion.
|
||||
*
|
||||
* Mirrors FormatString's live dynamic-output pattern (see format_string.js):
|
||||
* editing structured_output or output_schema posts to a backend route that
|
||||
* recomputes OllamaChatCompletion.RETURN_TYPES/RETURN_NAMES (the same
|
||||
* recomputes ChatCompletion.RETURN_TYPES/RETURN_NAMES (the same
|
||||
* update_outputs() path chat() itself uses at execution time) and returns
|
||||
* the resulting output list — applied to this node's sockets immediately,
|
||||
* so you see the extracted fields appear without having to run the graph
|
||||
* first.
|
||||
* first. Backend-agnostic: ChatCompletion is the one generic node both
|
||||
* OllamaProvider and LlamaCppProvider feed.
|
||||
*/
|
||||
async function updateStructuredOutputs(node, structuredOutput, outputSchema) {
|
||||
try {
|
||||
@@ -148,7 +157,7 @@ app.registerExtension({
|
||||
name: "comfydv.ollama.structuredOutput",
|
||||
|
||||
async beforeRegisterNodeDef(nodeType, nodeData) {
|
||||
if (nodeData.name !== "OllamaChatCompletion") return;
|
||||
if (nodeData.name !== "ChatCompletion") return;
|
||||
|
||||
const onNodeCreated = nodeType.prototype.onNodeCreated;
|
||||
nodeType.prototype.onNodeCreated = function () {
|
||||
|
||||
@@ -0,0 +1,148 @@
|
||||
"""Guards against the exact bug found while validating spec 008 against a
|
||||
real ComfyUI dev harness (docker-compose): every `from comfydv._llm.X
|
||||
import Y`-style absolute self-import inside src/comfydv/ silently broke the
|
||||
*entire* plugin (every node, not just LLM ones) as soon as ComfyUI actually
|
||||
loaded it.
|
||||
|
||||
ComfyUI's custom_nodes loader imports the plugin via a *relative* chain —
|
||||
the repo-root __init__.py does `from .src.comfydv import ...`, nesting
|
||||
comfydv under whatever top-level name the folder has (never `comfydv`
|
||||
itself). An absolute `from comfydv...` self-import only resolves if `src/`
|
||||
has separately been placed on sys.path — which conftest.py does for every
|
||||
other test file in this suite, masking the bug completely. This file
|
||||
deliberately does NOT rely on that sys.path insertion: it reproduces
|
||||
ComfyUI's actual nested-relative-import shape in a subprocess.
|
||||
|
||||
Confirmed via git history: this predates spec 008 entirely — it was already
|
||||
broken immediately after PR #17 merged (spec 007), well before llamacpp.py
|
||||
existed. No test caught it because none exercised this exact loading shape
|
||||
until the docker harness was run by hand.
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import sys
|
||||
import textwrap
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
REPO_ROOT = Path(__file__).parent.parent
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_ollama_caches():
|
||||
"""Shadow conftest.py's autouse fixture of the same name for this module
|
||||
only. That fixture's own setup does `from comfydv._llm.ollama_provider
|
||||
import ...` — this file's tests are the exact reproduction of an
|
||||
environment where `comfydv` resolving correctly can't be assumed (that's
|
||||
the point of the file), so depending on it for an unrelated cache-reset
|
||||
would make these tests order-dependent on whichever other test file
|
||||
happens to import `comfydv` "the normal way" first in the session. These
|
||||
tests touch no OllamaProvider/ChatCompletion state, so there is nothing
|
||||
to reset."""
|
||||
yield
|
||||
|
||||
|
||||
_SUBPROCESS_SCRIPT = textwrap.dedent(
|
||||
"""
|
||||
import sys
|
||||
import types
|
||||
|
||||
# Minimal ComfyUI stubs — same shape as conftest.py's pytest_configure,
|
||||
# but this script intentionally runs outside pytest so it isn't reusing
|
||||
# (or accidentally validated by) that fixture's sys.path setup.
|
||||
class _InterruptProcessingException(Exception):
|
||||
pass
|
||||
|
||||
comfy_module = types.ModuleType("comfy")
|
||||
comfy_module.model_management = types.SimpleNamespace(
|
||||
InterruptProcessingException=_InterruptProcessingException
|
||||
)
|
||||
sys.modules["comfy"] = comfy_module
|
||||
sys.modules["comfy.model_management"] = comfy_module.model_management
|
||||
|
||||
class _Routes:
|
||||
def post(self, path):
|
||||
return lambda fn: fn
|
||||
|
||||
def get(self, path):
|
||||
return lambda fn: fn
|
||||
|
||||
class _PromptServer:
|
||||
pass
|
||||
|
||||
_PromptServer.instance = _PromptServer()
|
||||
_PromptServer.instance.routes = _Routes()
|
||||
server_module = types.ModuleType("server")
|
||||
server_module.PromptServer = _PromptServer
|
||||
sys.modules["server"] = server_module
|
||||
|
||||
folder_paths_module = types.ModuleType("folder_paths")
|
||||
folder_paths_module.get_output_directory = lambda: "/tmp/comfydv_test"
|
||||
sys.modules["folder_paths"] = folder_paths_module
|
||||
|
||||
# The critical part: put the repo's *parent* directory on sys.path, so
|
||||
# `import comfydv` resolves to the repo-root __init__.py — exactly how
|
||||
# ComfyUI resolves a folder under custom_nodes/ — NOT to src/comfydv
|
||||
# directly (that's what conftest.py's sys.path.insert(0, ".../src")
|
||||
# does for the rest of this test suite, and why it never caught this).
|
||||
sys.path.insert(0, sys.argv[1])
|
||||
|
||||
import comfydv
|
||||
|
||||
required = {"FormatString", "RandomChoice", "CircuitBreaker",
|
||||
"OllamaClient", "LlamaCppClient", "ChatCompletion"}
|
||||
missing = required - set(comfydv.NODE_CLASS_MAPPINGS)
|
||||
if missing:
|
||||
print(f"MISSING_NODES:{missing}")
|
||||
sys.exit(1)
|
||||
print("OK")
|
||||
"""
|
||||
)
|
||||
|
||||
|
||||
def test_package_imports_under_comfyui_style_relative_nesting():
|
||||
"""Reproduces ComfyUI's real loading shape and fails loudly — with the
|
||||
actual traceback — if any internal module reverts to an absolute
|
||||
`from comfydv...` self-import."""
|
||||
result = subprocess.run(
|
||||
[sys.executable, "-c", _SUBPROCESS_SCRIPT, str(REPO_ROOT.parent)],
|
||||
cwd=REPO_ROOT,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=30,
|
||||
)
|
||||
assert result.returncode == 0, (
|
||||
"comfydv failed to import the way ComfyUI actually loads it "
|
||||
"(relative nesting, not a top-level `comfydv` on sys.path). "
|
||||
f"This means every node in the plugin would fail to register.\n"
|
||||
f"--- stdout ---\n{result.stdout}\n--- stderr ---\n{result.stderr}"
|
||||
)
|
||||
assert "OK" in result.stdout
|
||||
|
||||
|
||||
def test_no_absolute_self_imports_in_package():
|
||||
"""Cheap, fast static guard alongside the dynamic test above: no file
|
||||
under src/comfydv/ should import itself as `comfydv.X` — internal
|
||||
imports must be relative (`.X` / `..X`) so they resolve regardless of
|
||||
what the outer package happens to be named at load time."""
|
||||
import ast
|
||||
|
||||
offenders = []
|
||||
for path in (REPO_ROOT / "src" / "comfydv").rglob("*.py"):
|
||||
tree = ast.parse(path.read_text(), filename=str(path))
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom):
|
||||
if node.module and (
|
||||
node.module == "comfydv" or node.module.startswith("comfydv.")
|
||||
):
|
||||
offenders.append(f"{path.relative_to(REPO_ROOT)}:{node.lineno}")
|
||||
elif isinstance(node, ast.Import):
|
||||
for alias in node.names:
|
||||
if alias.name == "comfydv" or alias.name.startswith("comfydv."):
|
||||
offenders.append(f"{path.relative_to(REPO_ROOT)}:{node.lineno}")
|
||||
|
||||
assert not offenders, (
|
||||
"Absolute self-imports found — use relative imports instead "
|
||||
f"(they break under ComfyUI's actual loader): {offenders}"
|
||||
)
|
||||
@@ -13,7 +13,7 @@ BDD coverage:
|
||||
import pytest
|
||||
|
||||
import comfydv._llm.llamacpp_provider as provider_mod
|
||||
from comfydv._llm.llamacpp_provider import LlamaCppProvider
|
||||
from comfydv._llm.llamacpp_provider import LlamaCppProvider, _fetch_models
|
||||
from comfydv._llm.ollama_provider import _run_async
|
||||
from comfydv._llm.provider import Message, ModelStatus
|
||||
|
||||
@@ -355,3 +355,50 @@ def test_chat_structured_forwards_options(monkeypatch):
|
||||
)
|
||||
|
||||
assert captured["options"] == {"temperature": 0.0}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# _fetch_models — name-only view used by ComfyUI's /dv/ollama/models?backend=
|
||||
# llamacpp route (the JS refresh button / node-creation auto-populate).
|
||||
# Deliberately more forgiving than list_models(): degrades to [] on any
|
||||
# failure rather than raising on a non-router-mode server, matching the
|
||||
# combo-widget UX OllamaProvider's own _fetch_models already gives.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_fetch_models_returns_name_only_list(monkeypatch):
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
return {
|
||||
"data": [
|
||||
{"id": "a", "status": {"value": "loaded"}},
|
||||
{"id": "b", "status": {"value": "unloaded"}},
|
||||
]
|
||||
}
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
names = _run_async(_fetch_models("http://localhost:8080"))
|
||||
|
||||
assert names == ["a", "b"]
|
||||
|
||||
|
||||
def test_fetch_models_degrades_to_empty_on_non_router_mode(monkeypatch):
|
||||
"""Unlike list_models() (FR-006), this combo-population view swallows
|
||||
even the non-router-mode error — a quiet empty dropdown, not a toast."""
|
||||
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
raise RuntimeError("Server returned HTTP 404 for http://x/models: not found")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
names = _run_async(_fetch_models("http://localhost:8080"))
|
||||
|
||||
assert names == []
|
||||
|
||||
|
||||
def test_fetch_models_degrades_to_empty_when_unreachable(monkeypatch):
|
||||
async def fake_get(url, *, timeout=5.0, headers=None):
|
||||
raise ConnectionError("no route to host")
|
||||
|
||||
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
|
||||
names = _run_async(_fetch_models("http://localhost:19999"))
|
||||
|
||||
assert names == []
|
||||
|
||||
@@ -824,6 +824,92 @@ class _FakeRequest:
|
||||
return self._body
|
||||
|
||||
|
||||
class _FakeRelUrl:
|
||||
def __init__(self, query: dict):
|
||||
self.query = query
|
||||
|
||||
|
||||
class _FakeGetRequest:
|
||||
def __init__(self, query: dict):
|
||||
self.rel_url = _FakeRelUrl(query)
|
||||
|
||||
|
||||
class TestModelsRoute:
|
||||
"""GET /dv/ollama/models — the JS refresh-button/auto-populate endpoint.
|
||||
|
||||
Previously untested (no test survived the atomic cutover rewrite), which
|
||||
is exactly how a real bug shipped undetected: the `backend` dispatch
|
||||
added here didn't exist until a live js/ollama.js bug was found (it
|
||||
matched on pre-rename node names and always spoke Ollama's wire
|
||||
protocol regardless of which client was actually connected)."""
|
||||
|
||||
def _call(self, query: dict):
|
||||
import comfydv.ollama as ollama_mod
|
||||
from comfydv._llm.ollama_provider import _run_async
|
||||
|
||||
resp = _run_async(ollama_mod._models_endpoint(_FakeGetRequest(query)))
|
||||
return json.loads(resp.text), resp.status
|
||||
|
||||
def test_default_backend_dispatches_to_ollama(self, monkeypatch):
|
||||
async def fake_fetch(host, headers=None):
|
||||
assert host == "http://localhost:11434"
|
||||
return ["a", "b"]
|
||||
|
||||
monkeypatch.setattr("comfydv.ollama._fetch_models", fake_fetch)
|
||||
data, status = self._call({"host": "http://localhost:11434"})
|
||||
|
||||
assert status == 200
|
||||
assert data == {"models": ["a", "b"]}
|
||||
|
||||
def test_explicit_ollama_backend_dispatches_to_ollama(self, monkeypatch):
|
||||
async def fake_fetch(host, headers=None):
|
||||
return ["m"]
|
||||
|
||||
monkeypatch.setattr("comfydv.ollama._fetch_models", fake_fetch)
|
||||
data, status = self._call(
|
||||
{"host": "http://localhost:11434", "backend": "ollama"}
|
||||
)
|
||||
|
||||
assert status == 200
|
||||
assert data == {"models": ["m"]}
|
||||
|
||||
def test_llamacpp_backend_dispatches_to_llamacpp_fetch(self, monkeypatch):
|
||||
"""The bug this regression-tests: before the fix, this endpoint
|
||||
always called Ollama's _fetch_models regardless of `backend`, so a
|
||||
llama.cpp host's models never populated the dropdown."""
|
||||
called = {}
|
||||
|
||||
async def fake_llamacpp_fetch(host, headers=None):
|
||||
called["host"] = host
|
||||
return ["gemma-3-4b"]
|
||||
|
||||
async def fail_ollama_fetch(host, headers=None):
|
||||
raise AssertionError("must not call Ollama's fetcher for backend=llamacpp")
|
||||
|
||||
monkeypatch.setattr(
|
||||
"comfydv._llm.llamacpp_provider._fetch_models", fake_llamacpp_fetch
|
||||
)
|
||||
monkeypatch.setattr("comfydv.ollama._fetch_models", fail_ollama_fetch)
|
||||
|
||||
data, status = self._call(
|
||||
{"host": "http://localhost:8080", "backend": "llamacpp"}
|
||||
)
|
||||
|
||||
assert status == 200
|
||||
assert data == {"models": ["gemma-3-4b"]}
|
||||
assert called["host"] == "http://localhost:8080"
|
||||
|
||||
def test_no_models_returns_503(self, monkeypatch):
|
||||
async def fake_fetch(host, headers=None):
|
||||
return []
|
||||
|
||||
monkeypatch.setattr("comfydv.ollama._fetch_models", fake_fetch)
|
||||
data, status = self._call({"host": "http://localhost:11434"})
|
||||
|
||||
assert status == 503
|
||||
assert "error" in data
|
||||
|
||||
|
||||
class TestUpdateStructuredOutputsRoute:
|
||||
def teardown_method(self):
|
||||
# This route mutates ChatCompletion's class-level RETURN_TYPES
|
||||
|
||||