88 Commits
Author SHA1 Message Date
James VeitchandClaude Sonnet 5 f16da579bc feat(ui): add output-text previews and live retry status to comfydv nodes
Fixes a real gap discovered this session: ChatCompletion/FormatString's
"ui": {"text": [...]} return value was never rendered anywhere — no JS in
this repo ever implemented an onExecuted handler to read it, so the
retry-status banner from the previous commit was invisible in practice.
RandomChoice never returned a ui payload at all (OUTPUT_NODE was False).

- src/js/preview_text.js: read-only STRING preview widget for
  ChatCompletion/FormatString/RandomChoice, mirroring ComfyUI core's own
  PreviewAny node/JS extension (confirmed via the actual bundled frontend
  source, not guessed) rather than inventing a new convention.
- RandomChoice: OUTPUT_NODE=True + a ui.text preview of its arbitrary-typed
  output (str/number passthrough, else JSON, else str() — same fallback
  chain as core's PreviewAny). IS_CHANGED still returns the raw picked
  value, not the new ui-wrapped dict, so change-detection is unaffected.
- Live retry feedback: chat()/chat_structured() gain an on_status
  callback, invoked at each retry boundary with a human-readable status
  (_llm/retry.py's format_retry_status/format_recovered_status).
  ChatCompletion wires this to ComfyUI's own PromptServer.send_progress_text
  (the same first-party mechanism core's gaussian-splat/image nodes use)
  rather than a hand-rolled custom event — the frontend already renders a
  live "progressText" widget for it, so no additional JS was needed for
  the live half of this feature.

Verified live against a real Ollama backend (docker-compose.dev.yml +
Playwright): forced a genuine refusal/retry via custom_phrases, confirmed
both the final preview widget and the live progress widget render
correctly on ChatCompletion, and confirmed FormatString/RandomChoice's
preview widgets render too. 410 tests passing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 23:43:13 +01:00
James VeitchandClaude Sonnet 5 7a40d22230 feat(llm): escalate retry timeouts, expose seed_used, add retry UI feedback
Three follow-ups to the refusal-retry feature:

- Each retry's request timeout now escalates (next_timeout_secs: base *
  attempt) across every chat()/chat_structured() retry loop, instead of
  reusing the same budget that just ran out.
- LLMProvider.chat()/chat_structured() gain an optional attempt_info
  out-param (record_attempt_info), populated with the seed/timeout
  actually used and the attempt/refusal counts — backward compatible,
  no return-type change.
- ChatCompletion exposes a new seed_used output and prepends a status
  line to its existing text preview when a retry/refusal happened, so
  both are visible without any frontend/JS changes.

seed_used is appended as the LAST output (after any dynamic
structured-output fields), not inserted after model_name — dynamic
fields must keep starting at their existing index (3) so already-wired
workflows (workflows/ltx-i2v-pipeline.json, and the still-open PR #35
canvas-format file) that link a FormatString input to a ChatCompletion
field by output index aren't silently repointed at the wrong socket.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 22:19:22 +01:00
James VeitchandClaude Sonnet 5 42c614d5f8 feat(llm): add user-configurable custom_phrases to refusal detection
Ship a comma-separated custom_phrases field on OllamaOptionRefusalRetry
so a new deflection phrasing a specific model uses can be added at
runtime, without waiting on a code change/release. Each phrase is
checked as a free, case-insensitive substring match (no embedding
model required) and, when embedding_model is set, folded in as extra
exemplars for the similarity check too.

Also ship two more default patterns spotted this session: bare
"cannot generate" (no leading "I") and "I am/I'm restricted from" —
both verified against the existing ordinary-content regression cases
to confirm no new false positives.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 17:44:32 +01:00
James VeitchandClaude Sonnet 5 a0f391f212 fix(llm): catch bare "cannot generate" phrasing without a leading "I"
The lexical refusal pattern required an "I" before cannot/can't/won't
so a refusal fragment like "Cannot generate this image due to the
content" (no leading pronoun — the shape it takes as a JSON string
value or a mid-sentence fragment) went undetected. Made the "I"
prefix optional; the cannot/can't/won't-plus-nearby-generate/create/
etc. shape is specific enough that dropping the pronoun requirement
doesn't pick up ordinary content (verified against the existing
"ordinary content" regression cases).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 17:34:41 +01:00
James VeitchandClaude Sonnet 5 3c5136af70 fix(llm): stop gating the embedding refusal check on is_ambiguous
An explicitly configured embedding_model is an opt-in to pay for the
similarity check, but is_refusal still gated it on is_ambiguous — a
length/hedge-keyword heuristic that skips anything long or plainly
worded. For chat_structured() this checks the full serialized JSON
output, whose braces/field-name overhead alone often pushed responses
past the 400-char cutoff, so the embedding call silently never fired
for exactly the subtle, on-topic-looking deflections this feature was
built to catch.

Now embed_fn, once supplied, always runs on any non-blank response that
didn't already match the lexical pass. Also widened is_ambiguous's own
threshold/keyword list, since it remains a standalone tested heuristic
for other callers even though is_refusal no longer wires it in.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 17:30:19 +01:00
James VeitchandClaude Sonnet 5 97249ee281 fix(llm): catch "I am unable to" phrasing in refusal detection
The lexical pattern only matched the "I'm unable to" contraction, so
spelled-out refusals like "I am unable to comply with your instruction"
fell through undetected without an embedding model configured.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 17:16:38 +01:00
James VeitchandClaude Sonnet 5 1f138e5991 feat(llm): detect refusal/deflection responses and retry with a new seed (#36)
Some models (observed with an abliterated Qwen variant) answer a request
they judge "politically sensitive" with hedging refusal language instead of
erroring or returning blank — neither of which the existing blank-response
or schema-validation retry triggers catch, so the refusal just passed
through as a normal result.

Adds a hybrid detector to _llm/retry.py: a free lexical/regex pass catches
blatant refusals ("I cannot generate...") without any network call; a
response that's short and/or hedge-y enough to be ambiguous additionally
gets an embedding-similarity check against canonical refusal exemplars, via
a new optional LLMProvider.embed() (Ollama's /api/embed, llama-server's
/v1/embeddings) — deliberately model-agnostic to the backend, since this is
a model-behavior concern, not a backend one. A detected refusal is treated
exactly like a blank response or validation failure: retried with a bumped
seed (existing next_seed()/RETRY_BACKOFF_SECS machinery), never returned to
the caller.

Wired into all four retry loops: OllamaProvider.chat()/chat_structured(),
LlamaCppProvider.chat(), and the shared _llm/chat.py chat_structured() used
by LlamaCppProvider.chat_structured(). New OllamaOptionRefusalRetry node
(same composable OLLAMA_OPTIONS chain as OllamaOptionDisableThinking)
configures it: enabled toggle, optional embedding_model (blank = lexical-
only), similarity threshold.


Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 17:11:08 +01:00
James VeitchandClaude Sonnet 5 2724f535f5 fix(workflows): disable thinking by default in LTX-2.3 pipeline (#34)
Every ChatCompletion node in ltx-i2v-pipeline.json now reads its options
through a new OllamaOptionDisableThinking node (id 18, disable_thinking=True)
appended to the end of the existing max-tokens/extra-body chain. Without it,
a thinking-capable model routinely burns its whole max_tokens budget on
chain-of-thought and never emits the closing JSON, so chat_structured fails
validation against an empty string — confirmed live against
saracen9/amoral-qwen3.5-9b. README's thinking-headroom section updated to
match: thinking off by default, headroom advice now framed as what to do if
you deliberately re-enable it via node 18.


Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 08:15:49 +01:00
James VeitchandClaude Sonnet 5 a613006f5a fix(ollama): bypass stale combo validation for LLM model widgets (#33)
LLMModelSelector and LLMLoadModel validated the model COMBO against
_DEFAULT_MODELS, a list fetched once at ComfyUI startup. Models pulled
into Ollama afterward showed up in the live-refreshed dropdown but
failed prompt validation with "value ... is not available" until the
server was restarted. VALIDATE_INPUTS now accepts any model string,
deferring to Ollama itself to reject an actually-missing model.


Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-25 23:45:02 +01:00
James Veitch 56c7d4ca56 Merge pull request #32 from darth-veitcher/feat/disable-thinking-toggle
feat(ollama): add disable-thinking toggle for both providers
2026-07-25 10:38:17 +01:00
James VeitchandClaude Sonnet 5 cf7cfeb6aa feat(ollama): add disable-thinking toggle for both providers
Adds OllamaOptionDisableThinking, a composable node chaining into the
same OLLAMA_OPTIONS socket every other OllamaOption* node uses. Unlike
those (Ollama-native sampling params passed through verbatim), the
"think" key it emits is a comfydv-level convention: every LLMProvider
implementation pops it out of options and translates it to its own wire
shape before building a request, since neither backend recognizes a
literal "think" key nested inside a generic options object.

Confirmed live: Ollama's native /api/chat and OpenAI-compatible
/v1/chat/completions both silently ignore "think" nested in options —
it must be a top-level request field, or the model burns its whole
token budget on chain-of-thought reasoning before ever responding
(eval_count: 223 vs 2 in a direct comparison). llama.cpp's translation
(chat_template_kwargs/reasoning_effort) is sourced from llama-server's
documented request-body fields, not live-verified against a running
instance.

This also fixes two existing live integration tests that already passed
options={"think": False} under the mistaken assumption it worked — it
was a silent no-op until now.

See ADR-010 for the full design discussion, including why this ended up
as a composable option node rather than a new ChatCompletion input.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-25 10:37:46 +01:00
James Veitch c9aefef915 Merge pull request #31 from darth-veitcher/fix/ollama-structured-output-native-format
fix(ollama): route structured output through native /api/chat, fix nullable field validation
2026-07-25 08:26:04 +01:00
James VeitchandClaude Sonnet 5 f060698e7f feat(workflows): add LTX-2.3 multi-agent I2V pipeline example
Wires the 6-agent prompt-compiler pipeline from
project-management/Work/planning/ltx.md onto comfydv's existing generic
LLM nodes (OllamaClient, ChatCompletion, FormatString) — no new node code,
purely a wiring exercise. Verified end-to-end against a live local Ollama
server through the actual ComfyUI node graph (all 6 agents plus the final
judge/refiner output), which is what surfaced and validated the fixes in
the preceding commit.

workflows/README.md documents the deliberate adaptations from the
reference design (single-round Judge/Refiner since ComfyUI graphs can't
express the retry loop, JSON-string list handling for FormatString's
STRING-only dynamic sockets, etc.) and how the wiring was verified against
the real node source before ever touching a live server.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-25 08:25:12 +01:00
James VeitchandClaude Sonnet 5 0f4b2234db fix(ollama): route structured output through native /api/chat, fix nullable field validation
Ollama's OpenAI-compatible endpoint silently reloads the model at its
default context size on every call, discarding any options.num_ctx
override — confirmed live even when the same options are included in
that request. OllamaProvider.chat_structured() now hand-rolls a call to
Ollama's native /api/chat + "format" instead of routing through the
shared pydantic-ai path, since that endpoint correctly preserves and
applies context-size options. LlamaCppProvider keeps the shared path
(llama-server's context is fixed at process launch, not per-request),
switched to pydantic-ai's NativeOutput mode so a thinking-capable model's
reasoning no longer competes with the structured response for token
budget.

Also fixes _build_structured_model typing non-required schema fields as
bare `py_type` with a None default, which only tolerates the field being
omitted — an explicit `null` (which models routinely emit) failed
pydantic validation until the field was typed `py_type | None`.

See ADR-009 for the full investigation and rejected alternatives.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-25 08:25:01 +01:00
James Veitch ecba52999d Merge pull request #30 from darth-veitcher/feat/jinja2-fromjson-filter
feat(format-string): add fromjson Jinja2 filter
2026-07-24 22:40:32 +01:00
James VeitchandClaude Sonnet 5 24119e3ab6 feat(format-string): add fromjson Jinja2 filter
ComfyUI has no native list/array socket type — FormatString's dynamic
inputs are always STRING (see update_widget). Passing a list of values
(e.g. extraction hints) into a {% for %} loop therefore requires either
manual comma-splitting or a way to carry structured data through a
STRING socket as JSON text.

Jinja2 ships `tojson` but not its inverse. Register `fromjson` (a thin
wrapper over json.loads) on FormatString.jinja_env so a STRING input
containing a JSON array/object can be parsed back into real Python data:

    {% for hint in extraction_hints | fromjson %}
    - {{ hint }}
    {% endfor %}

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-24 22:40:08 +01:00
James Veitch 1f30c2ff2f Merge pull request #29 from darth-veitcher/fix/jinja2-template-variable-extraction
fix(format-string): correct Jinja2 template variable extraction
2026-07-24 22:00:15 +01:00
James VeitchandClaude Sonnet 5 ecfc7eca59 fix(tests): pin comfydv module resolution before pytest fixture setup
uv run pytest was failing every single test with ModuleNotFoundError:
No module named 'comfydv._llm' — reproduces identically on a clean
checkout, unrelated to any test content.

Root cause: the repo root's own __init__.py (ComfyUI's custom-node
entry point) is also a valid "comfydv" package from a sys.path state
pytest transiently constructs during fixture setup, and it doesn't
expose the _llm submodule. A bare `import comfydv` issued from inside
a fixture body (e.g. _clear_ollama_caches) could resolve to that root
package instead of src/comfydv.

Forcing an explicit `import comfydv` in pytest_configure — while our
sys.path.insert(0, src) is still the definitive answer — caches the
correct module in sys.modules before anything else gets a chance to
resolve it ambiguously.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-24 21:59:38 +01:00
James VeitchandClaude Sonnet 5 e534fac121 fix(format-string): extract Jinja2 template variables via AST, not regex
FormatString._extract_keys used hand-rolled regexes to detect variables
in Jinja2 templates, which broke on anything beyond bare {{ var }} and
{{ var | filter }}:

- {% for hint in extraction_hints %}{{ hint }}{% endfor %} incorrectly
  surfaced the loop-local `hint` as a required input and never detected
  `extraction_hints` itself, since the control-structure regex only
  matched filter syntax (word followed by `|`).
- {% if extraction_hints is defined %} never matched at all.
- Filters called with arguments, e.g. {{ x | tojson(indent=2) }}, broke
  the regex's anchor to the closing `}}` and silently extracted nothing.

Replaced with jinja2.meta.find_undeclared_variables() over the parsed
AST, which handles all of Jinja2's syntax correctly and excludes names
bound within the template (loop targets, {% set %}) by construction.
Since that call returns an unordered set, sort by first textual
occurrence to keep extraction order deterministic for callers that rely
on positional outputs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-24 21:59:23 +01:00
James Veitch 6c3c341913 Merge pull request #24 from darth-veitcher/claude/chatcompletion-image-input-7mm6q5
docs(spec): draft VLM image input for ChatCompletion (009)
2026-07-23 12:32:31 +01:00
James Veitch 8cb9a74462 Merge pull request #28 from darth-veitcher/fix/ollama-incomplete-response-race
fix(ollama): raise a clear error on incomplete done:false stub
2026-07-23 07:59:27 +01:00
James VeitchandClaude Sonnet 5 5e678c2700 fix(ollama): raise a clear error on Ollama's incomplete done:false stub
Fixes #27. Live-observed against a real Ollama server under model-swap
load: /api/chat can answer HTTP 200 with model/created_at/message all
blank and done: false, before the target model has actually finished
loading. chat()'s existing blank-response retry (PR #23) couldn't tell
this apart from a genuine blank generation, so once retries were
exhausted it silently returned "" — success, with a wrong answer.

chat() now tracks whether the last attempt's response was this explicit
done: false stub. If retries are exhausted and the last response was one,
it raises a clear RuntimeError instead of returning blank. A genuine
blank generation (done: true, or no `done` field at all) keeps the
pre-existing silent-return behavior unchanged — only the stub signal
escalates.

3 new tests cover: raises when every attempt is a stub, recovers cleanly
when a later attempt is a real response, and the regression guard that a
genuine done:true blank still returns silently as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-23 07:58:31 +01:00
James VeitchandClaude Sonnet 5 3854b27cba fix(009): surface a clear error for non-vision Ollama models given an image
FR-006/SC-005 require a clear, specific error when an image is wired to
a model that can't process it. Live testing against a real Ollama server
found this unimplemented for the Ollama backend: /api/chat silently
accepts an unsupported `images` field and answers HTTP 200 with a blank
message instead of erroring, which the existing blank-response retry
(PR #23) swallows as an ordinary blank generation rather than surfacing.

OllamaProvider.chat()/chat_structured() now check /api/show's
`capabilities` list before sending a request that carries an image, and
raise a clear ValueError if `vision` is absent. The check is skipped
entirely for image-less requests (FR-003) and fails open on any lookup
problem (older Ollama, network hiccup) so it never blocks a request that
would otherwise have worked.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-23 07:42:57 +01:00
James Veitch e432054b47 Merge pull request #26 from darth-veitcher/fix/dockerfile-test-python-313
fix(tests): update Dockerfile test to expect python:3.13-slim
2026-07-22 21:46:27 +01:00
James VeitchandClaude Sonnet 5 09fe013ff9 fix(tests): update Dockerfile test to expect python:3.13-slim
docker/Dockerfile was bumped from 3.11 to 3.13 in 50076d7 without
updating the assertion, breaking the suite on main. Verified 3.13
compatibility by building the image and booting ComfyUI 0.27.0 in it
(custom node loads, DB migrations run, no errors) before updating the
pinned version the test enforces.

Fixes #20

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
2026-07-22 20:20:33 +01:00
Claude 0f73f3c2cc chore(009): register build bullet, record quality-gate status (T009)
Start the vlm-image-input bullet for this branch (BEACON BUILD tracking);
mark T009 with an honest note on pre-existing gate exceptions (repo-wide ty
diagnostics, a baseline Docker test, 007's deferred tasks) — all out of scope
for spec 009, whose own code is ruff/format/pytest green and doctor-clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:11:50 +00:00
Claude 24917be802 docs(009): document image input on Chat Completion (T008)
README gains a "Describing images (vision)" section (vision-model /
llama.cpp --mmproj prerequisite, batch-as-multi-image, text-only unchanged);
CHANGELOG Unreleased notes the optional image input and the ADR-008 carrier.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:10:10 +00:00
Claude 46c14a872d feat(009): structured-output image input via BinaryContent (T007-I)
_user_prompt_content() lifts a turn's images onto the pydantic-ai user
prompt as BinaryContent PNGs (current turn and history user turns); text-only
turns stay plain-string. One shared helper covers structured image output on
both backends through OpenAIChatModel. Makes T007-T pass. US3 done.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:08:19 +00:00
Claude 0cfe0d22ff test(009): failing tests for structured-output image input (T007-T)
Images on the current turn ride on Agent.run user_prompt as BinaryContent;
images on a history user turn ride on its UserPromptPart; text-only calls
stay plain-string. Red: chat.py passes only .content today. Witnesses both
us3_structured_image.feature scenarios.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:07:39 +00:00
Claude bc7b33922b feat(009): map images to OpenAI content-parts for llama.cpp (T006-I)
_to_openai_message() renders a turn's images as image_url data-URI parts on
/v1/chat/completions; text-only turns keep a plain-string content,
byte-identical to today. Makes T006-T pass. US2: same node/wiring now
describes an image on llama.cpp too (parity with Ollama).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:06:40 +00:00
Claude 1028315d37 test(009): failing tests for llama.cpp image content-parts (T006-T)
llama.cpp's OpenAI endpoint takes images as image_url parts inside content;
text-only messages must keep a plain-string content (byte-identical). Red:
mapping absent and model_dump() leaks images:None. Witnesses both
us2_both_backends.feature scenarios.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:05:56 +00:00
Claude a2f2464fdb feat(009): forward images via Ollama /api/chat, byte-identical text path (T005-I)
model_dump(exclude_none=True) keeps a turn's native flat images array while
omitting the key entirely for text-only turns, so image-less requests are
byte-for-byte unchanged (FR-003/SC-004). Makes T005-T pass. Completes US1
(MVP): describe-an-image end-to-end on Ollama.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:05:20 +00:00
Claude 4a631cc8dc test(009): failing test for Ollama image passthrough (T005-T)
Ollama /api/chat carries images as a flat per-message base64 array; a
text-only request must omit the images key (byte-identical to today,
FR-003/SC-004). Red: model_dump() currently leaks images:None into the
payload. Witnesses us1_describe_image.feature "Describe a wired image".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:05:00 +00:00
Claude fd9054c03f feat(009): optional IMAGE input on ChatCompletion (T004-I)
Adds an optional image input (with a vision-model / --mmproj tooltip) and
attaches encoded base64 image(s) to the current user turn only; history
turns untouched (FR-007). Outputs unchanged (Constitution VI). Makes
T004-T pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:04:06 +00:00
Claude 23d96f1559 test(009): failing tests for ChatCompletion IMAGE input (T004-T)
Optional image input; wired image attaches to the last user turn only
(history untouched, FR-007); un-wired path stays text-only; output
positions unchanged. Red: chat() has no image parameter yet. Witnesses
both us1_describe_image.feature scenarios.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:03:24 +00:00
Claude 9ece5c0efa feat(009): _encode_image_tensor for ChatCompletion image input (T003-I)
Converts a ComfyUI IMAGE [B,H,W,C] float tensor to base64 PNG(s), one per
frame; None/empty -> []. Lazy numpy/PIL import keeps module load clean
outside ComfyUI (Constitution IV). Makes T003-T pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:02:36 +00:00
Claude e225dcaaa3 test(009): failing tests for _encode_image_tensor (T003-T)
ComfyUI IMAGE tensor [B,H,W,C] float 0..1 -> decodable base64 PNG(s);
batch -> one string per frame; None/empty -> []. Red: helper not yet
defined. Witnesses us1_describe_image.feature "Describe a wired image".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:02:06 +00:00
Claude 1c72e0b2ea feat(009): add optional images carrier to Message (T002-I)
Message.images: list[str] | None = None — base64 payloads per turn, the
provider-neutral carrier ADR-008 defines. Default None keeps text-only
turns text-only. Makes T002-T pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:01:00 +00:00
Claude 7e56f2e5bf test(009): failing tests for Message.images carrier (T002-T)
Message gains an optional images field carrying base64 payloads per turn;
text-only messages must dump without an images key (FR-003/SC-004). Red:
current Message has no images attribute.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 19:00:21 +00:00
Claude f4a3390f13 build(deps): add pillow to dev group for image-encode tests (T001)
Enables unit-testing the ChatCompletion tensor→PNG encoder without a live
ComfyUI. Runtime Pillow/numpy remain ComfyUI-provided — no core runtime dep.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 18:59:57 +00:00
Claude a40c1ba373 docs(tasks): task breakdown + BDD witnesses for VLM image input (009)
Phase 2 task decomposition for spec 009 with BEACON test-first discipline.

- tasks.md — interleaved -T/-I TDD pairs mapped to contracts T1–T6:
  * Setup: add pillow to dev deps (no core runtime dep)
  * Foundational: Message.images carrier (blocks all stories)
  * US1 (MVP): node optional IMAGE input + tensor→PNG encode helper +
    Ollama flat-images passthrough; text-only path byte-identical
  * US2: llama.cpp images→OpenAI image_url content-parts (parity)
  * US3: chat_structured images→pydantic-ai BinaryContent (shared, both backends)
  * Polish: docs/CHANGELOG, quality gate; live-backend e2e marked [-] deferred
    (needs a vision model / llama-server --mmproj not available in CI)
- features/ — 3 .feature files, 6 scenarios verbatim from spec.md, referenced
  by name from the -T tasks (spec.md -> .feature -> test traceability)

beacon doctor: spec-bdd-coverage green (6 scenarios witnessed), fail=0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 18:56:26 +00:00
Claude eab86b0324 docs(plan): plan VLM image input (009) — research, data-model, contracts
Phase 0/1 design artifacts for spec 009, resolving the wire shapes ADR-008
deferred for live verification.

- research.md — verified pydantic-ai BinaryContent against the installed
  pydantic-ai-slim 2.9.0 source (structured path is shared via OpenAIChatModel,
  one change covers both backends); Ollama flat images passthrough; llama.cpp
  OpenAI image_url content-parts (requires --mmproj); tensor→PNG via Pillow
  (dev dep) with _llm/ kept torch/numpy/Pillow-free per Constitution IV;
  byte-identical text path by omitting an empty images key
- data-model.md — Message gains optional images: list[str] | None; one carrier,
  three wire shapes; optional IMAGE node input (outputs untouched)
- contracts/image-input-contract.md — behavioural + test contracts T1–T6
- quickstart.md — describe-an-image workflow, structured variant, backend swap
- plan.md — Constitution Check passes (no new class, no new runtime dep,
  outputs unchanged); no Complexity Tracking needed
- CLAUDE.md SpecKit marker repointed to 009 plan

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 18:50:09 +00:00
Claude 45e49a98eb docs(spec): draft VLM image input for ChatCompletion (009)
Add the BEACON DESIGN scaffold for wiring a ComfyUI IMAGE into the existing
generic ChatCompletion node so a vision-capable model can describe/understand
images on either backend.

- specs/009-vlm-image-input/spec.md — WHAT/WHY: optional IMAGE input, same
  node on both backends, structured-output-with-image, graceful degradation
  on non-vision models; text-only path unchanged when no image is wired
- ADR-008 (Proposed) — carry images on an optional Message.images field; each
  provider translates to its own wire shape (Ollama /api/chat images,
  llama.cpp OpenAI image_url parts, shared chat_structured multimodal
  content). Extends ADR-007's adapter pattern to a second input modality
- epics/vlm-image-input.md — new epic with success criteria/non-goals; spec
  backlinked via .beacon.toml and listed under the epic
- Roadmap + ADR index updated

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvS9TFMCFYNHC4MMvzJaZS
2026-07-22 18:38:39 +00:00
James Veitch 87ec79be1b Merge pull request #23 from darth-veitcher/fix/retry-blank-output-with-seed-and-backoff
fix: retry blank/failed LLM responses with a new seed and backoff
2026-07-20 09:03:16 +01:00
James VeitchandClaude Sonnet 5 1dcef77a2b fix: keep chat_structured retry's nested seed in sync with the top-level one
beacon-reviewer caught a real gap: when a caller pins options["seed"],
the retry loop set the new seed on ModelSettings' top-level "seed"
field but left the same request's extra_body.options.seed (the
Ollama-native passthrough) at the original pinned value. A backend
that honors the nested field over the top-level OpenAI one would keep
sending the identical effective seed on every retry, silently
defeating this fix for exactly the pinned-seed case it needs to cover.

Now both fields are kept in sync on every retry attempt, built via
fresh copies rather than mutating the caller's options dict in place
(that dict is shared across every attempt, and possibly across other
calls). Added a regression test asserting the caller's options dict is
untouched after a multi-attempt retry.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-20 08:59:06 +01:00
James VeitchandClaude Sonnet 5 09a173e552 fix: retry blank/failed LLM responses with a new seed and backoff
Reported live on a fresh runpod: structured_output extraction on
ChatCompletion would fail with

  RuntimeError: chat_structured: response failed validation against
  schema after 3 attempt(s) ... Last error: Exceeded maximum output
  retries (0)

...but succeed if a plain (non-structured) chat call with the same
prompt was run first. Root cause: pydantic-ai's Agent is built with
retries=0 (comfydv drives its own outer retry loop, not pydantic-ai's
internal one), so a single bad/empty tool-call response instantly
exhausts one attempt. comfydv's own retry loop then retried with the
*exact same request* every time — if the model is genuinely stuck
(e.g. still warming up on a freshly-started backend) or produces a
degenerate response for a given seed, every attempt fails identically.
Confirmed the seed was unset in the reporting user's workflow (Ollama
picks its own per call) and it still failed 3/3 in a row, which rules
out "stuck on one bad seed" and points at request-level determinism
plus zero backoff between attempts as the real gap.

Plain chat() had zero retry/validation logic at all before this change
— it silently returned whatever content came back, even blank. Given
the underlying failure mode (a fresh model's first response sometimes
being blank) is identical for both chat modes, this fix applies to
both, not just structured output.

Adds src/comfydv/_llm/retry.py: next_seed() (deterministic, starting
from any user-pinned options["seed"], incrementing per retry — attempt
1 is always the caller's original, untouched request) and
RETRY_BACKOFF_SECS (a flat delay between retries, giving a still-
loading model time to finish rather than hammering it with identical
requests back-to-back).

Both OllamaProvider.chat() and LlamaCppProvider.chat() now retry (per
the node's existing max_retries widget) when a response comes back
blank, injecting the incremented seed each retry — Ollama via the
native options.seed field, llama.cpp via the OpenAI spec's top-level
"seed" field (nesting it under "options" like the rest of the
Ollama-native passthrough would silently not work — llama-server's
OpenAI-compatible endpoint doesn't read seed from there, a pre-existing,
documented non-goal for general options passthrough that doesn't apply
to this internal retry mechanism). Neither raises if every retry stays
blank — matches chat()'s existing never-validates contract.

chat_structured() in _llm/chat.py now injects the same incrementing
seed via pydantic-ai's ModelSettings["seed"] (which correctly maps to
the OpenAI API's top-level seed param for both backends) and adds the
same backoff between attempts, on top of its existing validation-retry
loop.

Live-verified against real Ollama: the happy path (model responds
normally first try) still makes exactly one network call for chat()
and completes chat_structured() normally — no regression from the new
retry scaffolding. The specific cold/stuck-model failure itself is
runpod-cold-start-dependent and wasn't independently reproducible in
this session's environment (a second live attempt succeeded on the
very first try) — unit tests directly exercise the retry/seed/backoff
logic instead (test_llm_retry.py, and new cases in
test_ollama_provider.py, test_llamacpp_provider.py, and
test_llm_chat_structured.py), verified against both providers.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-20 08:49:35 +01:00
James Veitch 92da4fc117 Merge pull request #22 from darth-veitcher/fix/run-async-real-comfyui-event-loop
fix: chat_structured() crashes in real ComfyUI's async execution engine
2026-07-12 10:33:50 +01:00
James VeitchandClaude Sonnet 5 1fe3bef663 docs(tests): stop overclaiming what the _run_async regression tests prove
beacon-reviewer caught it: reconstructing the old buggy _run_async and
running it against test_run_async_works_when_called_from_within_a_running_loop
and test_run_async_propagates_exceptions_from_within_a_running_loop shows
both pass against the old code too. asyncio.get_running_loop() never
spuriously raises under plain CPython/pytest, so the old conditional also
takes the safe worker-thread branch there — the actual reported bug is not
reproducible outside real ComfyUI's execution engine, and only the live
end-to-end run against it proved the fix.

Corrects the block comment to say what the tests actually do: lock in the
documented contract so a regression to a no-thread asyncio.run(coro) call
is caught immediately, without claiming they reproduce the specific
environment-dependent failure.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-12 09:39:55 +01:00
James VeitchandClaude Sonnet 5 3bd857f0ab fix: chat_structured() crashes with real asyncio.run() error in live ComfyUI
Found by running a genuine end-to-end test — real docker-compose
ComfyUI harness, real Ollama, real llama-server (router mode),
actual workflow graphs submitted via POST /prompt and polled via
/history, not direct Python calls to provider classes. Every
structured_output=True ChatCompletion call failed with:

  RuntimeError: asyncio.run() cannot be called from a running event loop

Root cause: _run_async() tried asyncio.get_running_loop() first and
only spun up an isolated worker thread if that succeeded; otherwise
it called asyncio.run(coro) directly on the current thread. Under
pytest or a standalone script, get_running_loop() never spuriously
raises, so this worked and the worker-thread path was rarely
exercised for real. Under ComfyUI's actual async execution engine
(Python 3.13, node functions invoked synchronously from inside an
already-running event loop), get_running_loop() sometimes raised
anyway — routing straight into asyncio.run(coro) on the one thread
guaranteed to already have a loop running, reproducing the crash
every time. Confirmed via a live-instrumented debug run against the
real container: the coroutine executed correctly whenever the
worker-thread path was taken (get_running_loop() succeeding), and
crashed only via the direct-call fallback.

Fixed by removing the conditional entirely: always run the coroutine
in a freshly spawned worker thread. A new thread never has an
ambient loop, so asyncio.run() is safe there unconditionally,
regardless of what the calling thread's loop state actually is.

Also traced down a second symptom from the same end-to-end run:
non-structured chat() calls sometimes returned an empty response
with no error. Verified via direct curl calls to Ollama and
llama-server (bypassing comfydv entirely) that this was a real,
transient GPU/Metal memory-pressure issue on the test machine
(`ggml_metal_synchronize: error: Insufficient Memory`, Ollama's own
backend left in a wedged state returning HTTP 200 with an empty
zero-value body) — not a comfydv bug. Restarting the backend
processes resolved it; no code change was needed for that part.

Added tests/test_ollama_provider.py coverage for the specific shape
that broke: _run_async called synchronously from within code that is
itself already executing inside asyncio.run() — the actual pattern
ComfyUI's execution engine uses, which no prior test exercised (every
existing call site invoked _run_async from a plain synchronous pytest
function with no ambient loop).

Re-verified end-to-end after the fix: 5/5 real workflows passing
through actual ComfyUI execution — Ollama basic chat, Ollama full
lifecycle (Load->Chat->Unload), Ollama structured output, llama.cpp
basic chat, llama.cpp structured output — all with real generated
content, not mocks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-12 09:36:21 +01:00
James Veitch 834ae3e9e3 Merge pull request #18 from darth-veitcher/chore/finish-llm-provider-abstraction-epic
chore: finish llm-provider-abstraction epic
2026-07-11 21:55:08 +01:00
copilot-swe-agent[bot] 2df9e87838 chore: resolve merge conflicts with main 2026-07-11 20:47:06 +00:00
James Veitch 27b26ec7a4 Merge pull request #21 from darth-veitcher/fix/comfyui-relative-import-loading 2026-07-11 18:09:56 +01:00
James VeitchandClaude Sonnet 5 16b7e944f0 fix(tests): make the import-compat proof order-independent
beacon-reviewer caught a real gap: tests/test_comfyui_import_compat.py
ERRORed (not passed) when run in isolation or under a file-sharded
runner — the autouse _clear_ollama_caches fixture from conftest.py
does `from comfydv._llm.ollama_provider import ...`, and this file
was the one case where nothing had already bound `comfydv` in
sys.modules correctly first, so it resolved to the wrong package.
Only the full single-process suite happened to be green.

Fixed by shadowing that fixture with a local no-op override scoped
to this module: these tests exercise package import resolution
itself, not OllamaProvider/ChatCompletion state, so there is nothing
for the original fixture to reset here — and depending on it made
the test order-dependent on an unrelated file's import order, which
is exactly the kind of fragility this file exists to guard against.
Verified: passes standalone (`pytest tests/test_comfyui_import_compat.py`)
and as part of the full suite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 18:02:57 +01:00
James VeitchandClaude Sonnet 5 36e2cd123c docs: rebuild README around live screenshots, fix stale doc tooling
Full README/docs/index.md rewrite from a product/UX-first structure
(what-is-this, why, then a deep-dive per node group with screenshots)
instead of an incremental changelog-style read.

All screenshots regenerated from a real, running ComfyUI instance —
not reused. This surfaced and fixed two more bugs in
scripts/take_screenshots.py, the project's own screenshot-automation
tool, which nobody had run since before the ADR-007 rename (the old
committed screenshots still showed pre-rename node names, e.g.
"Ollama Chat Completion" instead of "Chat Completion"):

- Every LiteGraph.createNode() call used the pre-rename node type
  strings (OllamaChatCompletion/OllamaLoadModel/OllamaUnloadModel);
  fixed to the current names.
- The canvas locator ("canvas#graph-canvas, canvas") silently
  grabbed a 250x200 minimap canvas ComfyUI's frontend added since
  this script was last verified, instead of the real graph — every
  screenshot failed with a clipping error. Fixed to target
  #graph-canvas explicitly.

Two new scenes added, addressing the actual gap: existing
screenshots showed connectivity and options, never the two things
this whole epic sequence was about — structured output (the live
dynamic output sockets appearing as output_schema is edited, via the
real js/ollama.js widget callbacks the previous commit fixed) and
llama.cpp (LlamaCppClient wired into the same generic ChatCompletion
node Ollama uses — the actual point of the adapter pattern).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 17:48:37 +01:00
James VeitchandClaude Sonnet 5 1e9a8a3b6a fix: comfydv failed to load in real ComfyUI since PR #17 (critical)
Every internal `from comfydv._llm.X import Y`-style absolute
self-import broke the entire plugin the moment ComfyUI actually
loaded it — every node, not just the LLM ones, since the whole
src/comfydv/__init__.py chain aborted on the first such import.

ComfyUI's custom_nodes loader imports this plugin via a *relative*
chain (repo-root __init__.py does `from .src.comfydv import ...`),
nesting comfydv under whatever top-level name the folder gets —
never `comfydv` itself. An absolute `from comfydv...` self-import
only resolves if `src/` has separately been placed on sys.path,
which conftest.py does for every test — masking this completely.
No test ever exercised the real loading shape.

Confirmed via git bisection against the actual docker-compose dev
harness (built and ran real ComfyUI): this predates spec 008
entirely — checking out the commit right after PR #17 merged,
before llamacpp.py existed, reproduces the identical failure at
ollama.py's own absolute import. Fixed by converting every internal
self-import across src/comfydv/ to a relative import, which resolves
correctly under both loading shapes. Verified fixed by rebuilding
the harness and confirming all 21 nodes register via /object_info.

Added tests/test_comfyui_import_compat.py: a subprocess-based test
reproducing ComfyUI's exact nested-relative-import shape (not
conftest.py's sys.path-patched shape), plus a static AST guard
against any future absolute self-import creeping back in.

Also fixed, found via the same live-harness investigation:
- src/js/ollama.js matched on pre-rename node names
  (OllamaModelSelector/OllamaLoadModel/OllamaChatCompletion), so the
  "Refresh models" button and live structured-output socket preview
  were silently absent from every node in the real UI.
- The /dv/ollama/models route always spoke Ollama's wire protocol
  regardless of which backend was actually connected — pointing it
  at a llama.cpp host could never populate real models. Route now
  takes a `backend` param; getHostAndBackendFromNode() in ollama.js
  determines it from the connected client node's registered type;
  LlamaCppProvider gained a matching _fetch_models() name-only view
  (mirrors OllamaProvider's, same graceful-degradation contract).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 17:48:20 +01:00
James Veitch d59a85dba6 Merge pull request #19 from darth-veitcher/008-llamacpp-integration
feat: llama.cpp support via LLMProvider adapter (issue #15)
2026-07-11 16:48:09 +01:00
James VeitchandClaude Sonnet 5 15bb25efa8 test(llamacpp): live-verify chat_structured(), strengthen idempotency tests
Ran a second live smoke test — load_model() then chat_structured()
with a real pydantic.BaseModel schema — against the same router-mode
llama-server used for the earlier fix. Confirmed pydantic-ai's
Agent/OpenAIProvider mechanism works against llama-server's
OpenAI-compatible endpoint, not just Ollama's (the only one verified
live in the prerequisite epic). No gap found; every LlamaCppProvider
method has now been exercised against a real server, not just mocks.

Also strengthened the two idempotency regression tests added in the
previous commit: they previously asserted only "does not raise", which
would also pass if the method silently no-op'd for an unrelated bug.
Now assert the mocked _post_json was actually called with the expected
request, so the test verifies real behavior, not just absence of a
crash.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 16:24:52 +01:00
James VeitchandClaude Sonnet 5 282198a7df fix(llamacpp): make load_model/unload_model actually idempotent
Ran the live smoke test (T017) that had been wrongly marked deferred
for "no llama-server reachable" — llama-server was installed via
Homebrew the whole time and a router-mode server was launched against
a real local GGUF model to verify.

This caught a real gap the mocked suite couldn't: router mode's
/models/load and /models/unload are NOT idempotent at the wire level.
Calling load on an already-loaded model returns HTTP 400 "model is
already running" (and unload/"model is not running" symmetrically),
not the {"success": true} the conformance contract assumed without
live verification. LlamaCppProvider now absorbs exactly those two
error messages as the desired end-state already reached; any other
error still propagates. Contract doc corrected to describe the real
behavior; regression tests added (mocked, run in CI).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 16:10:51 +01:00
James VeitchandClaude Sonnet 5 ac248e1e99 chore(beacon): regenerate ROADMAP.md rollup
Follow-up to the llamacpp-integration epic status edit — beacon epic
refresh regenerates this file automatically.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 15:49:27 +01:00
James VeitchandClaude Sonnet 5 ea66ceee9d docs(beacon): flip llamacpp-integration epic Planning->Active
Spec 008-llamacpp-integration is complete; epic-finish bookkeeping
(archive) follows in a separate chore PR once this merges, matching
the llm-provider-abstraction epic's pattern.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 15:48:03 +01:00
James VeitchandClaude Sonnet 5 cf3f7fc174 fix(llamacpp): surface a clear error for non-router-mode servers (FR-006)
LlamaCppProvider.list_models() caught every exception and returned [],
indistinguishable from "no models installed". Fixed _get_json to raise
on HTTP error status (matching _post_json's existing behavior — its
docstring already claimed this), and list_models() to distinguish
OSError (genuinely unreachable, degrades to []) from RuntimeError
(server responded with an error — surfaced with a router-mode hint).

Found by beacon-reviewer ahead of PR open.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 15:43:12 +01:00
James VeitchandClaude Sonnet 5 5af502d07e feat(llamacpp): add LlamaCppProvider + LlamaCppClient node (spec 008)
Second LLMProvider (ADR-007), backed by llama-server's router mode. Only
one new ComfyUI node — LlamaCppClient, emitting the same LLM_CLIENT socket
Ollama Client does. ChatCompletion, LLMModelSelector, LLMLoadModel,
LLMUnloadModel work unmodified once wired to it — the actual proof the
provider abstraction generalizes, not just Ollama-shaped in practice.

API shape verified live against ggml-org/llama.cpp's tools/server/README.md
(research.md) rather than assumed: model identifier field is "id" (not
"name"), status is a nested {"value": "..."} object. list_models() needs
no normalization — llama.cpp's status vocabulary is exactly ModelStatus's
full set, unlike Ollama's narrower one.

38 new tests (test_llamacpp_provider.py, test_llamacpp.py), including a
dedicated US4 test proving no generic node branches on provider type —
the same call sequence succeeds against an Ollama-shaped or llama.cpp-
shaped fake provider.

README/docs updated for the new node this time, not left stale.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 15:27:29 +01:00
James VeitchandClaude Sonnet 5 d3ee35bfda docs(beacon): DESIGN artifacts for llama.cpp Model Integration (spec 008)
Full spec/plan/tasks/BDD scaffold for the llamacpp-integration epic, now
that its dependency (llm-provider-abstraction, PR #17) is satisfied.
Researched llama-server's router-mode API shape live (postdates training
data) rather than assuming it — two details that would have been wrong by
assumption: the model identifier field is "id" not "name" (differs from
Ollama's /api/tags), and "status" is a nested object ({"value": "..."}) not
a flat string.

Unlike the prerequisite epic, this decomposition genuinely holds as 4
independent user stories — LlamaCppClient is a new node, not a changed one,
so there's no shared "output type" migration forcing an atomic cutover.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 15:17:54 +01:00
James VeitchandClaude Sonnet 5 27916f4fb0 chore(beacon): regenerate ROADMAP.md gantt/table for llamacpp-integration Active
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 15:11:25 +01:00
James VeitchandClaude Sonnet 5 689aa2a391 chore(beacon): finish llm-provider-abstraction epic
All owned specs shipped (PR #17), so beacon epic finish moves it to
archive/ and flips status to Done. Marks llamacpp-integration Active — its
dependency is now satisfied.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 15:10:57 +01:00
James Veitch 9ca65bc29a Merge pull request #17 from darth-veitcher/007-llm-provider-abstraction
LLM Provider Abstraction: generic adapter pattern for Ollama (ADR-007)
2026-07-11 15:08:14 +01:00
James VeitchandClaude Sonnet 5 c32b6a7e42 docs(beacon): note PR #17 open on the llm-provider-abstraction epic
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 14:42:59 +01:00
James VeitchandClaude Sonnet 5 3df38485d2 docs: update README/docs for the generic node rename, add migration table (FR-009)
README.md and docs/index.md still named OllamaChatCompletion/OllamaModelSelector/
OllamaLoadModel/OllamaUnloadModel and the OLLAMA_CLIENT socket — stale after
the atomic cutover. Updated both (they're content-duplicates) and added the
old->new migration table FR-009 requires, which previously only existed as
a code comment (comfydv.ollama.MIGRATION_MAP), not surfaced to users.

Found via a dedicated docs-freshness pass, not caught by the beacon-reviewer
correctness review (which is scoped to code, not prose).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 14:42:19 +01:00
James VeitchandClaude Sonnet 5 8e227790a9 fix(llm): forward options through to structured output (beacon-reviewer finding)
chat_structured() silently dropped the options dict (temperature, seed,
num_predict, repeat_penalty, etc. — set via the OllamaOption* nodes) in
structured_output mode: OllamaProvider.chat_structured() only used it for
the cache key, and the shared comfydv._llm.chat.chat_structured() helper
had no options parameter at all. On main, the hand-rolled implementation
sent options verbatim in the request body; this migration silently lost it,
contradicting FR-008 (structured-output behavior must be unchanged).

Fixed by forwarding options as pydantic-ai's model_settings.extra_body,
matching the original payload shape exactly rather than lossily remapping
onto ModelSettings' own standardized field names (which don't cover
Ollama-native params like num_predict/repeat_penalty/top_k anyway).

Verified against the live local Ollama server: structured output with
temperature=0.0/seed=42 succeeds on first attempt.

Caught by an independent beacon-reviewer pass on the diff before opening a
PR — exactly what that review step is for.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 14:35:51 +01:00
James VeitchandClaude Sonnet 5 6c2c906c4a docs(beacon): mark tasks.md Phase 7 subsumed by T-CUT-10/12
Same lint/typecheck/doctor/live-validation work, done together with the
cutover rather than as a separate pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 14:29:20 +01:00
James VeitchandClaude Sonnet 5 1879ca2265 test(llm): complete atomic node cutover — test_ollama.py rewrite (T-CUT-08/10/12)
Rewrote tests/test_ollama.py against a _FakeProvider double per
atomic-cutover-plan.md's D4/D5 test-layer split: node-contract/delegation
tests only, no aiohttp/_post_json mocking (that lives in
test_ollama_provider.py now). 98 unit tests, collapsed from the original
~125 largely mechanical monkeypatch-per-scenario tests where the retry/
validation coverage they duplicated already lives in
test_llm_chat_structured.py.

Fixed a real gap the rename surfaced along the way: comfy-manager-entry.json
(ComfyUI Manager registry listing) still had the old Ollama-specific
display names — updated it and its matching test expectation.

Full validation (T-CUT-10): 218 tests green (only the pre-existing,
unrelated Dockerfile-python-version test fails); ruff/ty clean (confirmed
via git stash comparison that the remaining ty diagnostics pre-date this
work); beacon doctor --strict shows only already-disclosed items.

Manual smoke test (T-CUT-12): Ollama was reachable in this environment, so
ran the real integration suite rather than a walkthrough. 6/8 passed
including the critical paths (real load/unload, real structured-output
retry-then-raise, real error handling). 2 failures traced via direct curl
to Ollama's /api/chat (bypassing this codebase) to a pre-existing,
ADR-006-documented unreliability in the specific test model, not a
regression.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 14:28:34 +01:00
James VeitchandClaude Sonnet 5 d90ae98fed feat(llm): atomic node cutover — generic ChatCompletion/LLM* nodes (T-CUT-01/04-07/11)
OllamaClient now emits an OllamaProvider via the generic LLM_CLIENT socket
(was OllamaClientType/OLLAMA_CLIENT). OllamaChatCompletion, OllamaModelSelector,
OllamaLoadModel, OllamaUnloadModel renamed to ChatCompletion, LLMModelSelector,
LLMLoadModel, LLMUnloadModel — all delegate to client.* (LLMProvider protocol)
instead of building Ollama URLs/payloads inline. Structured output now goes
through client.chat_structured() (pydantic-ai), not hand-rolled tool-calling.

ollama.py's own duplicate HTTP/cache infra (_post_json, _TTLLRUCache,
_MODEL_LIST_CACHE, _CHAT_RESPONSE_CACHE) is gone — single source of truth in
comfydv._llm.ollama_provider (plan D1); the combo-widget path
(_load_default_models, /dv/ollama/models) imports _fetch_models/_run_async
from there instead of duplicating them. _client_headers deleted (dead code
once nodes delegate to client.*).

Added MIGRATION_MAP (old->new node/socket names, FR-009). NODE_CLASS_MAPPINGS
keys changed to match (breaking rename, confirmed acceptable per ADR-007).

conftest.py's autouse cache-clearing fixture repointed to the single cache
source and ChatCompletion (plan D8).

tests/test_ollama.py itself is NOT yet updated (T-CUT-08, next commit) — it
will fail to collect until then; this is the expected atomic-cutover
intermediate state, not a regression to fix independently.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 14:19:07 +01:00
James VeitchandClaude Sonnet 5 fb111c0142 feat(llm): implement OllamaProvider method bodies (T-CUT-02/03)
list_models()/load_model()/unload_model()/chat()/chat_structured() —
behavior-preserving ports of the existing inline ollama.py logic, plus one
genuinely new capability: list_models() also queries /api/ps to report
real loaded/unloaded status (the old OllamaModelSelector only ever showed
names via /api/tags, never live status).

15 new tests in tests/test_ollama_provider.py, mocking at the module's own
_post_json/_get_json seam per atomic-cutover-plan.md's D4 test-layer split.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 14:14:51 +01:00
James VeitchandClaude Sonnet 5 22dbddc8eb docs(beacon): fix tasks.md checkbox hygiene, add trackable T-CUT phase
Struck-through historical tasks (- [ ] ~~...~~) still read as open
checkboxes to beacon's tooling — bullet status was picking one as "active"
instead of the real next step. Marked all superseded Phase 3/5/T014/T019-21
entries [-] (deferred, per the project's own convention), and added a real
Phase 8 with T-CUT-01..12 as trackable checkboxes matching
atomic-cutover-plan.md, so tooling and humans see the same source of truth.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 14:12:02 +01:00
James VeitchandClaude Sonnet 5 ef2464aebb docs(beacon): properly spec the atomic node cutover (issue #16)
Full line-by-line inventory of every reference in src/comfydv/ollama.py
and tests/test_ollama.py (1820 lines, read in full — not estimated)
affected by the OllamaClient/OllamaChatCompletion/OllamaModelSelector/
OllamaLoadModel/OllamaUnloadModel rename, plus the design decisions it
surfaced that "just rename it" was hiding:

- Single source of truth for HTTP/cache infra (comfydv._llm.ollama_provider)
- client == "<host string>" equality is removed, not preserved (OllamaProvider
  isn't a str subclass)
- Bare-string client backward compatibility is removed (undocumented side
  effect of the old str-subclass trick, never an intended feature)
- Test-layer split: node-contract/delegation tests stay in test_ollama.py
  using a _FakeProvider double; Ollama-wire-protocol tests move to a new
  test_ollama_provider.py — this is what actually resolves the 35 relocated
  _post_json monkeypatches, rather than patching them 1:1 at a seam that no
  longer exists
- Structured-output retry tests are not ported 1:1 — already covered by
  tests/test_llm_chat_structured.py's existing 6 tests

Produces a 12-step sequenced task list (T-CUT-01..12) in
atomic-cutover-plan.md that supersedes tasks.md's original US1/US3
decomposition (struck through, kept for history). This is now the
authoritative plan for the next BUILD session to execute.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 13:55:39 +01:00
James VeitchandClaude Sonnet 5 736ffd1e77 feat(llm): shared chat_structured() via pydantic-ai (T011-T013)
Provider-layer only — src/comfydv/_llm/chat.py doesn't touch ollama.py, so
it's safe to land ahead of the deferred atomic node cutover (issue #16).
Ports ADR-006's exact retry/validation contract: bounded retries (0-5,
clamped), RuntimeError naming model/attempt-count/truncated-last-response
on exhaustion. Agent's own internal retries disabled so the error contract
is comfydv's, not pydantic-ai's.

Required an API-discovery spike (Agent/OpenAIProvider/OpenAIChatModel
constructors and exception surface aren't reliably knowable from training
data for a fast-moving library) before tests could be written meaningfully
— tests and implementation were validated together rather than strictly
red-first; noted in tasks.md rather than presented as pure TDD.

Also fixes a pre-existing gap in test_packaging.py's requirements.txt/
pyproject.toml consistency check: it stripped `[extras]` on one side of the
comparison but not the other, which this change's pydantic-ai-slim[openai]
entry exposed as the first dependency with an extras marker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 13:37:16 +01:00
James VeitchandClaude Sonnet 5 1b1c7830d5 docs(beacon): correct tasks.md — US1/US3 cutover is atomic, not incremental
Discovered mid-implementation: OllamaClient is a single shared producer for
every downstream Ollama node, so changing its output type per ADR-007
breaks all consumers simultaneously, and the class rename breaks
tests/test_ollama.py's 125 references atomically, not per-story. Confirmed
by independent product + engineering review (agent-trio deliberation,
aligned). Re-scoped the node-layer cutover as its own follow-up session,
tracked in issue #16. ADR-007's decision itself is unaffected — only
delivery sequencing changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 13:31:21 +01:00
James VeitchandClaude Sonnet 5 9d90f39854 feat(llm): scaffold LLMProvider protocol and OllamaProvider (T001-T006)
Adds pydantic-ai-slim[openai] per ADR-007, the Message/ModelStatus/
ModelInfo/LLMProvider protocol, and an OllamaProvider skeleton carrying
host/headers as instance state — ported _post_json/_fetch_models/_run_async/
_TTLLRUCache infra from comfydv.ollama unchanged. Method bodies land per
user story in specs/007-llm-provider-abstraction/tasks.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 13:25:47 +01:00
James VeitchandClaude Sonnet 5 fb0556fff3 docs(beacon): DESIGN artifacts for LLM provider abstraction (ADR-007)
Adopt a shared LLMProvider protocol (list/load/unload/chat/structured-chat)
as the adapter pattern answer to GitHub issue #15 — Ollama and llama.cpp
share the chat/structured-output mechanism via pydantic-ai while model
lifecycle stays backend-specific per provider. Supersedes ADR-006, narrows
ADR-004's scope to non-chat REST calls.

Adds epic 007 (llm-provider-abstraction, prerequisite) and the
llamacpp-integration epic it unblocks, plus the full spec/plan/tasks/BDD
scaffold for the prerequisite epic.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj
2026-07-11 13:25:40 +01:00
James Veitch b3374c88c4 add loop 2026-07-11 13:07:58 +01:00
James Veitch 50076d7c6b add manager to local harness 2026-07-10 16:40:51 +01:00
James Veitch aa7ff8dae3 Merge pull request #14 from darth-veitcher/feature/ollama-structured-output
feat(ollama): opt-in structured output via tool-calling + pydantic validation
2026-07-09 17:51:08 +01:00
James VeitchandClaude Sonnet 5 253569d483 feat(ollama): live dynamic output sockets for structured_output
Adds a POST /dv/ollama/update_structured_outputs route and a matching JS
extension so editing structured_output/output_schema on
OllamaChatCompletion updates the node's output sockets immediately in the
UI — one per schema property — mirroring FormatString's existing live
dynamic-output pattern. No need to run the graph first to see the extracted
fields appear. Invalid/incomplete JSON while typing falls back to the base
3-output shape rather than erroring the request.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-09 17:37:41 +01:00
James VeitchandClaude Sonnet 5 40e851588d feat(ollama): opt-in structured output via tool-calling + pydantic validation
Adds structured_output/output_schema/max_retries inputs to
OllamaChatCompletion, fixing three real reliability problems: models
prepending commentary, wrapping responses in code fences, and occasionally
returning blank output. When enabled, requests go through Ollama's
OpenAI-compatible tool-calling endpoint (a single forced tool call matching
output_schema) instead of native /api/chat, and each response is validated
against a dynamically-built pydantic model — invalid/blank responses retry
with fresh network calls, exhausting into a clear RuntimeError rather than
silently passing bad data downstream. One additional ComfyUI output socket
is exposed per schema property, mirroring FormatString's existing
dynamic-output pattern. Fully backward compatible: structured_output
defaults off and non-structured behavior is untouched.

Ollama's native /api/chat "format" field (JSON-Schema-constrained decoding)
was tried first but proved unreliable against a real local model during
implementation — see ADR-006 for the full investigation and why
tool-calling was chosen over both native format and pydantic-ai (the latter
would reintroduce httpx/openai-sdk, reversing ADR-004).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-09 12:50:54 +01:00
James Veitch 23f8f55493 Merge pull request #13 from darth-veitcher/chore/finish-ollama-integration-epic
chore(beacon): finish ollama-integration epic
2026-07-04 14:09:52 +01:00
James VeitchandClaude Sonnet 5 d38cd5e794 chore(beacon): finish ollama-integration epic
All owned specs shipped, so beacon epic finish moves it to archive/ and
flips status to Done. Clears the epic-spec-lifecycle and epic-gates
warnings beacon doctor --strict was raising.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-04 14:09:15 +01:00
James Veitch e42913d587 Merge pull request #12 from darth-veitcher/chore/dynamic-versioning-from-vcs
chore: derive package version from git tags via hatch-vcs
2026-07-04 14:03:15 +01:00
James VeitchandClaude Sonnet 5 335ff5e31a chore: derive package version from git tags via hatch-vcs
Replaces the static project.version string with dynamic = ["version"],
sourced from the nearest git tag (hatch-vcs). Fixes docs.yml, which read
pyproject.toml's version field directly and would have broken once that
field stopped being a plain string; it now reads the resolved version via
importlib.metadata after `uv sync` installs the package.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-04 13:51:55 +01:00
106 changed files with 14563 additions and 1147 deletions
+19
View File
@@ -0,0 +1,19 @@
/beacon:continue --full-auto keep delivering and - if none exist to complete - use the product and engineering triumvirate and identify the next epics and specs required. Use github issues to communicate with me and to track issues. Don't manufacture marginal work.
Use `gh issue create` to raise topics/questions/blockers async (and `gh issue comment` to post progress/track them); reserve `AskUserQuestion` for genuinely loop-blocking decisions only. Check for new comments, issues, and state changes as part of the loop assessment.
<separation_of_concerns>
Use specialist and targeted subagents for deliver whilst you own the orchestration, planning, and review of their outputs.
</separation_of_concerns>
<the_token_trap>
Avoid the urge to produce content simply because it is rewarded. Perfection is when there is nothing left to remove. Keep in mind that les - invariably - is more.
Challenge yourself to deliver outcomes without flamboyance, excessive verbosity, or unnecessary codebase bloat.
</the_token_trap>
**Note:** At the end of each iteration critically review the state of our README and documentation and then spawn a dedicated subagent to keep things fresh. Documentation and the README should never feel like an incremental read (referring to older versions) and should always read like a fresh "this is the the thing and how it works", not "the thing was this, we've done X, and now it's Y". Write this from a Product and UX perspective so that it is engaging and remember that humans are visual creatures. Screenshots are important both to explain and to engage.
DO NOT MAKE DESIGN DESIGNS. If the plan is underdeliverable raise an issue and abort. If you deviate but can justify and demonstrate it delivers the scope to the specification please document this clearly with rationale and update the necessary designs, decisions, and user facing documentation where required.
THREE (!) SENIOR DEVELOPERS WITH MORE THAN 30 YEARS EXPERIENCE WILL REVIEW YOUR PR. They will not accept shortcuts, monkeypatched tests, fake tests, hacky approaches or deviations from the plan. Produce Senior dev grade production code aligned to the plan at all times.
+1 -1
View File
@@ -29,5 +29,5 @@ jobs:
- name: Deploy docs with mike
run: |
VERSION=$(uv run python -c "import tomllib; d=tomllib.load(open('pyproject.toml','rb')); print(d['project']['version'])")
VERSION=$(uv run python -c "from importlib.metadata import version; print(version('comfydv'))")
uv run mike deploy --push --update-aliases "$VERSION" stable
+1 -1
View File
@@ -1,3 +1,3 @@
{
"feature_directory": "specs/006-ollama-model-integration"
"feature_directory": "specs/009-vlm-image-input"
}
+11
View File
@@ -10,6 +10,17 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
- BEACON framework bootstrap: problem statement, constitution, roadmap, architecture document
- `CHANGELOG.md` (this file)
- README: What-is-this, Install, and Quickstart sections
- `LLMProvider` protocol (`comfydv._llm`) — a shared adapter boundary so ComfyUI LLM nodes work with any backend that implements it, starting with `OllamaProvider`. Structured output now goes through `pydantic-ai` (ADR-007), superseding the hand-rolled Ollama tool-calling approach.
- **Chat Completion** now accepts an optional `image` input for vision-capable models (VLMs): wire a ComfyUI `IMAGE` and the connected model can describe or reason about it. Works identically on both backends (Ollama multimodal models; llama.cpp launched with `--mmproj`), and composes with structured output and multi-turn history. Images are carried on `Message.images` and translated to each backend's native shape (Ollama's flat `images` array, llama.cpp's OpenAI `image_url` parts, pydantic-ai `BinaryContent` on the structured path) — ADR-008, extending ADR-007's adapter pattern to a second input modality. Text-only workflows are unchanged when no image is wired.
- **Ollama Option — Disable Thinking** node: turn off (or explicitly re-enable) a "thinking"-capable model's chain-of-thought reasoning. Chains into the same composable `OLLAMA_OPTIONS` socket every other `OllamaOption*` node uses, but works for both backends — each `LLMProvider` implementation pops the `think` key back out and translates it to its own wire shape (Ollama: a top-level `think` field; llama.cpp: `chat_template_kwargs`/`reasoning_effort` request-body fields, not live-verified — see ADR-010).
### Changed
- **Breaking:** `OllamaChatCompletion` → `ChatCompletion`, `OllamaModelSelector` → `LLMModelSelector`, `OllamaLoadModel` → `LLMLoadModel`, `OllamaUnloadModel` → `LLMUnloadModel`, and the `OLLAMA_CLIENT` socket type → `LLM_CLIENT` — these nodes are now backend-generic. `OllamaClient` is unchanged by name but now outputs an `OllamaProvider` rather than a plain string; existing saved workflows using the old node/socket names need reconnecting (see `comfydv.ollama.MIGRATION_MAP` for the full old→new mapping).
- `ChatCompletion`'s `structured_output=True` path now routes Ollama through Ollama's native `/api/chat` + `"format"` instead of the shared `pydantic-ai` OpenAI-compat path — Ollama's OpenAI-compatible endpoint was found to silently reload the model at its default context size on every call, discarding any `options` (e.g. `num_ctx`) override. `LlamaCppProvider` is unaffected and keeps the shared path, switched to `pydantic-ai`'s `NativeOutput` mode (ADR-009).
### Fixed
- `structured_output=True` requests could fail validation ("token limit exceeded before any response was generated") against "thinking"-capable models, which spent their whole token budget on chain-of-thought reasoning before ever producing the structured response (ADR-009).
- A non-required structured-output schema field rejected an explicit `null` value from the model (only an *omitted* field was tolerated), even though models routinely emit explicit `null` for absent optional fields.
## [0.1.0] — 2026-06-01
+1 -1
View File
@@ -1,5 +1,5 @@
<!-- SPECKIT START -->
For additional context about technologies to be used, project structure,
shell commands, and other important information, read the current plan
at specs/006-ollama-model-integration/plan.md
at specs/009-vlm-image-input/plan.md
<!-- SPECKIT END -->
+104 -72
View File
@@ -1,28 +1,37 @@
# comfydv
A collection of workflow efficiency and quality-of-life nodes built out of necessity for personal ComfyUI use.
**Quality-of-life nodes for ComfyUI, built to disappear into your workflow.**
## What is this?
`comfydv` fills the gaps ComfyUI's built-in library leaves on the table: string templates that build their own sockets as you type, seed-controlled randomisation, graceful mid-queue interruption, and a local-LLM integration that doesn't care whether you're running Ollama or llama.cpp. No Python required — install it, drop the nodes on your canvas, wire them up.
`comfydv` fills gaps in ComfyUI's built-in node library: dynamic string formatting, seed-controlled random selection, graceful workflow interruption, and Ollama LLM integration. Install it once and connect the nodes like any other — no Python knowledge required.
![Chat Completion in action](docs/assets/ollama_chat.png)
## What is comfydv?
A small, focused ComfyUI utility pack. It exists because:
- **It reads your intent, not just your syntax.** Format String detects `{variables}` in a template and adds/removes input sockets live, as you type — no manual socket wrangling.
- **One LLM integration, any local backend.** Wire a Chat Completion node once; swap between Ollama and llama.cpp by changing a single upstream client node. Structured output, multi-turn history, and model load/unload work identically on both.
- **It fails politely.** Circuit Breaker halts a queue run cleanly instead of throwing a stack trace at you; a disconnected LLM server gets a specific, actionable error instead of a silent empty dropdown.
- **Small, tested, boring in the best way.** Every node is unit-tested and the local-LLM nodes are verified against real running servers, not just mocks.
## What's inside
| Node | What it does |
|------|-------------|
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the host URL through the graph as an `OLLAMA_CLIENT` socket. |
| **Ollama Model Selector** | Fetches the live model list from Ollama and presents it as a dropdown. Outputs the selected model name. |
| **Ollama Load Model** | Loads a model into Ollama's memory using `/api/generate` with `keep_alive=-1`. |
| **Ollama Unload Model** | Evicts a model from Ollama's memory using `/api/generate` with `keep_alive=0`. |
| **Ollama Chat Completion** | Sends a prompt (and optional conversation history) to Ollama `/api/chat`. Response and history are shown inline in the node body and available as output sockets. |
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
| **Ollama History Length** | Returns the number of messages in an `OLLAMA_HISTORY` list as an integer. |
|------|---------------|
| **Format String** | Renders a Python f-string or Jinja2 template. Sockets appear and disappear automatically as you type variables. |
| **Random Choice** | Accepts any number of typed inputs and returns one at random, with a seed for reproducibility. |
| **Circuit Breaker** | Halts the current queue run gracefully — no crash, just a clean stop — when a condition isn't met. |
| **Ollama Client** / **LlamaCpp Client** | Configure a connection to a local Ollama or llama.cpp server. Both emit the same `LLM_CLIENT` socket — every node below works with either. |
| **LLM Model Selector** | Live dropdown of models available on the connected server. |
| **LLM Load Model** / **LLM Unload Model** | Explicit VRAM management — pin a model in memory before inference, evict it after. |
| **Chat Completion** | Send a prompt (optionally with history) to the connected server; response shown inline and as an output socket. |
| **Ollama Option — \*** | Seven composable parameter nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into Chat Completion's `options` input. |
| **Ollama Debug History** / **Ollama History Length** | Inspect an `OLLAMA_HISTORY` conversation list — pretty-print it or count its messages. |
## Install
**Via ComfyUI Manager** (recommended): search for `comfydv` and click Install.
**Via ComfyUI Manager** (recommended): search for `comfydv`, click Install.
**Manual:**
@@ -31,124 +40,147 @@ cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/darth-veitcher/comfydv.git
```
Restart ComfyUI. The nodes appear under the **dv/** and **dv/ollama** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`) are installed automatically via `requirements.txt`.
Restart ComfyUI. Nodes appear under **dv/**, **dv/ollama**, and **dv/llamacpp** in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) install automatically via `requirements.txt`.
For Ollama nodes: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`) before using the Ollama nodes.
For local-LLM nodes, bring your own backend:
- **Ollama** — [install Ollama](https://ollama.com/download), pull a model (`ollama pull qwen2.5:latest`).
- **llama.cpp** — [build/install `llama-server`](https://github.com/ggml-org/llama.cpp), launch it in [router mode](#llamacpp).
## Quickstart
1. Install via ComfyUI Manager (search `comfydv`) or clone manually into `custom_nodes/`.
2. Right-click the canvas → Add Node → **dv/** to find Format String, Random Choice, and Circuit Breaker.
3. For Ollama nodes: start Ollama (`ollama serve`), pull a model (`ollama pull qwen2.5:latest`), then add nodes from **dv/ollama/**.
1. Right-click the canvas → Add Node → **dv/** for Format String, Random Choice, and Circuit Breaker.
2. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
## Documentation
Full documentation: [darth-veitcher.github.io/comfydv](https://darth-veitcher.github.io/comfydv/stable/)
Full documentation, including every node's inputs/outputs: **[darth-veitcher.github.io/comfydv](https://darth-veitcher.github.io/comfydv/stable/)**
---
## Format String
Formats text from a Python f-string or Jinja2 template. As you type the template, input sockets appear and disappear automatically — one per variable detected.
### Python f-strings
Type `{variable_name}` and a socket appears. Wire it to any string output in your workflow.
Type a template, get sockets. `{variable_name}` in f-string mode, `{{ variable_name }}` in Jinja2 mode — either way, comfydv watches what you type and keeps the node's inputs in sync automatically.
![Format String — f-string mode](docs/assets/fstring.png)
Outputs are always in a stable order:
| Output | Content |
|--------|---------|
| `formatted_string` | The rendered result |
| `saved_file_path` | Path written to disk (if `save_path` is set) |
| `<var>` … | Pass-through of each input value, for easy chaining |
| `saved_file_path` | Where it was written, if `save_path` is set |
| `<var>` … | Each input passed through unchanged, for easy chaining |
### Jinja2 templates
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`, …), conditionals (`{% if %}…{% endif %}`), and loops.
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`), conditionals, and loops:
![Format String — Jinja2 mode](docs/assets/jinja2.png)
Variables detected in `{{ }}` expressions become input sockets exactly as in Simple mode. See the [Jinja2 documentation](https://jinja.palletsprojects.com/en/latest/) for the full filter/test reference.
---
## Random Choice
Connect any number of inputs of the same type. Each run picks one at random. Set `seed` for reproducibility.
Wire in any number of same-typed inputs — images, strings, conditioning, anything ComfyUI can carry over a socket — and get one back at random.
![Random Choice](docs/assets/random.png)
- Accepts any ComfyUI type (STRING, IMAGE, CONDITIONING, …)
- Add as many inputs as you like; unused slots are removed automatically when disconnected
- `seed = 0` randomises on every run; any other value locks the selection
`seed = 0` randomises every run; any other value locks the selection. Unused input slots vanish automatically when you disconnect them.
---
## Circuit Breaker
Stops the queue gracefully when a condition isn't met — no crash, no error, just a clean halt.
Stop a queue run cleanly when a condition isn't met, instead of letting a downstream node crash on bad input.
![Circuit Breaker](docs/assets/circuit_breaker.png)
Wire an image (or any trigger) into `trigger` and a boolean into `status`. When `status` is **false** the node raises `InterruptProcessingException`, which tells ComfyUI to stop the current run cleanly. When `status` is **true** the image passes through unchanged.
Wire a trigger (an image, or anything) into `trigger` and a boolean into `status`. `status = false` raises `InterruptProcessingException` — ComfyUI stops the run without an error dialog. `status = true` passes the trigger straight through.
Typical use: skip an expensive upscale step when a quality-check node says the draft is already good enough.
Typical use: skip an expensive upscale pass when an upstream quality-check node says the draft's already good enough.
---
## Ollama
## Local LLMs
14 nodes for integrating a local Ollama LLM into your ComfyUI workflow. The host URL is configured once in **Ollama Client** and threaded through the graph — all downstream nodes receive it via the `OLLAMA_CLIENT` socket.
One set of nodes, two interchangeable backends. Configure a connection once with **Ollama Client** or **LlamaCpp Client** — both output the same `LLM_CLIENT` socket — and every downstream node (model selection, load/unload, chat, structured output, multi-turn history) works exactly the same way regardless of which one you picked. Swapping backends means rewiring one node, not rebuilding your graph.
### Ollama Client node
Configure the server address once; all downstream Ollama nodes inherit it automatically.
### Connect and chat
![Ollama Client](docs/assets/ollama_client.png)
### Model lifecycle (load and unload)
1. **Ollama Client** (default `http://localhost:11434`) or **LlamaCpp Client** (default `http://localhost:8080`) — set the host.
2. **LLM Model Selector** — pick a model from the live dropdown, or wire a model name straight into Chat Completion.
3. **Chat Completion** — wire in client, model, and prompt. The response renders inline in the node and is also available as an output socket.
On memory-constrained machines and single-GPU setups, explicitly loading and unloading the model before and after inference is critical. **Ollama Load Model** pins the model into VRAM (`keep_alive=-1`); **Ollama Unload Model** evicts it immediately (`keep_alive=0`), freeing memory for image generation or other models.
![Chat Completion](docs/assets/ollama_chat.png)
![Ollama Load / Unload](docs/assets/ollama_lifecycle.png)
A complete graph looks like this:
The correct chain is **Load → Chat → Unload**, enforced through data dependencies:
![Full LLM workflow](docs/assets/ollama_workflow.png)
1. Wire `OllamaLoadModel.model_name` → `OllamaChatCompletion.model`. This creates the data dependency that guarantees Load runs before Chat and passes the model name into the Chat node's plain-string `model` input.
2. Wire `OllamaChatCompletion.model_name` → `OllamaUnloadModel.model`. This guarantees Unload runs after Chat completes.
3. Optionally wire `OllamaChatCompletion.response` → `OllamaUnloadModel.passthrough` — Unload returns the response unchanged so the rest of your workflow can still consume it.
### Describing images (vision)
### Minimal chat workflow
**Chat Completion** has an optional **image** input. Wire any `IMAGE` into it and, with a vision-capable model loaded, the model can describe or reason about the picture — captioning, visual Q&A, reading text in an image, whatever the model supports.
1. **Ollama Client** → set host (default `http://localhost:11434`)
2. **Ollama Model Selector** → pick a model from the live dropdown (or type/wire a model name directly into Chat Completion's `model` input)
3. **Ollama Chat Completion** → wire client + model + prompt; the response appears inline in the node body and is also available as an output socket
- **Ollama** — use a multimodal model (e.g. a llava-class model).
- **llama.cpp** — launch `llama-server` with a multimodal projector: `--mmproj <projector.gguf>` alongside the model.
![Ollama Chat Completion](docs/assets/ollama_chat.png)
Image input works the same on both backends — same node, same wiring — and composes with everything else: structured output (schema-validated fields pulled straight from the image), multi-turn history, and the option nodes. A batch of images is sent as multiple images on the turn. Leave the input unwired and Chat Completion behaves exactly as before, text only.
Wire multiple nodes together for a complete end-to-end workflow:
### Manual memory management
![Ollama Full Workflow](docs/assets/ollama_workflow.png)
Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline.
### Option nodes
![Load / Unload lifecycle](docs/assets/ollama_lifecycle.png)
Chain any combination of **Ollama Option —** nodes before Chat Completion to override inference parameters:
The **Load → Chat → Unload** chain is enforced by data dependencies, not by convention:
| Option node | Ollama param |
|-------------|-------------|
1. `LLMLoadModel.model_name` → `ChatCompletion.model` — guarantees Load runs before Chat, and feeds the model name straight in.
2. `ChatCompletion.model_name` → `LLMUnloadModel.model` — guarantees Unload runs after Chat completes.
3. *(Optional)* `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged, so the rest of your workflow can still consume it.
### Tuning generation
Chain any combination of **Ollama Option —** nodes ahead of Chat Completion to override inference parameters:
![Ollama Option nodes](docs/assets/ollama_options.png)
| Option node | Parameter |
|-------------|-----------|
| Temperature | `temperature` |
| Seed | `seed` |
| Max Tokens | `num_predict` |
| Top P | `top_p` |
| Top K | `top_k` |
| Repeat Penalty | `repeat_penalty` |
| Extra Body | arbitrary JSON merged into options |
![Ollama Option Nodes](docs/assets/ollama_options.png)
| Extra Body | arbitrary JSON, merged into `options` |
### Multi-turn conversations
`OLLAMA_HISTORY` flows out of Chat Completion as a list of `{"role", "content"}` dicts. Wire it back into the next Chat Completion for multi-turn conversations, or inspect it with **Ollama Debug History** / **Ollama History Length**.
`OLLAMA_HISTORY` flows out of Chat Completion as a `{"role", "content"}` list. Feed it back into the next Chat Completion call for multi-turn context, or inspect it with **Ollama Debug History** / **Ollama History Length**.
### llama.cpp
Everything above works unchanged against llama.cpp — swap in a **LlamaCpp Client** and the rest of the graph doesn't know the difference. The one thing llama.cpp needs that Ollama doesn't: **router mode**, a directory of models rather than a single `-m model.gguf`:
```bash
llama-server --models-dir ./models -c 8192
```
In exchange, router mode gives comfydv a richer live status than Ollama can report — `loading` and `downloading`, not just loaded/unloaded — plus the same explicit load/unload primitives Ollama's nodes already use.
### Upgrading a workflow saved before this rename
Nodes were renamed once, to make them backend-generic (`OllamaChatCompletion` → `ChatCompletion`, etc.). If ComfyUI reports old node types as missing when you reopen a saved workflow, reconnect using this table — behavior is unchanged, only the names are:
| Old | New |
|-----|-----|
| `OllamaChatCompletion` | `ChatCompletion` |
| `OllamaModelSelector` | `LLMModelSelector` |
| `OllamaLoadModel` | `LLMLoadModel` |
| `OllamaUnloadModel` | `LLMUnloadModel` |
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
`OllamaClient` kept its name — delete and re-add any node showing as missing, then rewire it to the same `OllamaClient` node.
---
## License
[AGPL-3.0](LICENSE)
+12 -7
View File
@@ -1,6 +1,6 @@
# comfydv — Roadmap
<!-- generated by beacon roadmap export — 2026-06-28 -->
<!-- generated by beacon roadmap export — 2026-07-11 -->
> comfydv is a small, high-quality ComfyUI utility pack that fills the gaps the core node library leaves: composable string formatting, seed-controlled randomisation, and workflow flow-control. Winning looks like: every node is well-tested, installs in one step, produces no surprises in production workflows, and is documented well enough that a non-programmer ComfyUI user can connect it without reading source code.
@@ -13,14 +13,14 @@ gantt
excludes weekends
section Active
llama.cpp Model Integration :active, llamacpp-integration, 2026-07-11, 7d
ComfyUI UX Polish & Manager Compatibility :active, ux-and-install, 2026-06-28, 21d
section Planned
Ollama Model Integration :ollama-integration, 2026-06-28, 7d
section Done
BEACON Bootstrap :done, beacon-bootstrap, 2026-06-28, 7d
Logging Modernisation :done, logging-modernisation, 2026-06-28, 7d
BEACON Bootstrap :done, beacon-bootstrap, 2026-07-11, 7d
LLM Provider Abstraction :done, llm-provider-abstraction, 2026-07-11, 7d
Logging Modernisation :done, logging-modernisation, 2026-07-11, 7d
Ollama Model Integration :done, ollama-integration, 2026-07-11, 7d
```
@@ -28,10 +28,12 @@ gantt
| Epic | Title | Status | Specs | Fidelity |
|---|---|---|---|---|
| [ollama-integration](project-management/Roadmap/epics/ollama-integration.md) | Ollama Model Integration | Planning | — | S? A? T:- |
| [llamacpp-integration](project-management/Roadmap/epics/llamacpp-integration.md) | llama.cpp Model Integration | Active | 1/1 shipped | S+ A+ T:96% |
| [ux-and-install](project-management/Roadmap/epics/ux-and-install.md) | ComfyUI UX Polish & Manager Compatibility | Active | 1/4 shipped | S+ A+ T:100% |
| [beacon-bootstrap](project-management/Roadmap/epics/archive/beacon-bootstrap.md) | BEACON Bootstrap | Done | — | S? A? T:- |
| [llm-provider-abstraction](project-management/Roadmap/epics/archive/llm-provider-abstraction.md) | LLM Provider Abstraction | Done | 1/1 shipped | S+ A+ T:58% |
| [logging-modernisation](project-management/Roadmap/epics/archive/logging-modernisation.md) | Logging Modernisation | Done | 1/1 shipped | S+ A+ T:100% |
| [ollama-integration](project-management/Roadmap/epics/archive/ollama-integration.md) | Ollama Model Integration | Done | 1/1 shipped | S+ A+ T:100% |
_Fidelity: `S+/S?` = has specs / none · `A+/A?` = has ADRs / none · `T:N%` = task completion_
@@ -46,3 +48,6 @@ _No active bullets._
| [ADR-001](../../../ADRs/ADR-001-stdlib-logging-over-console-libraries.md) | Use stdlib logging instead of colored console output libraries | Accepted |
| [ADR-002](../../../ADRs/ADR-002-nullhandler-pattern-for-library-loggers.md) | NullHandler pattern for the comfydv package root logger | Accepted |
| [ADR-003](../../ADRs/ADR-003-requirements-txt-authoring-policy.md) | Hand-authored requirements.txt as a curated subset of pyproject.toml | Accepted |
| [ADR-004](project-management/ADRs/ADR-004-aiohttp-over-httpx-for-ollama.md) | Use aiohttp for Ollama HTTP communication instead of httpx | Accepted |
| [ADR-005](project-management/ADRs/ADR-005-ollama-host-config-via-client-node.md) | Ollama host configuration via OllamaClient node and OLLAMA_CLIENT socket type | Accepted |
| [ADR-007](project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md) | LLMProvider adapter pattern shared across Ollama and llama.cpp | Accepted |
+4 -4
View File
@@ -10,10 +10,10 @@
"Random Choice",
"Circuit Breaker",
"Ollama Client",
"Ollama Model Selector",
"Ollama Load Model",
"Ollama Unload Model",
"Ollama Chat Completion",
"LLM Model Selector",
"LLM Load Model",
"LLM Unload Model",
"Chat Completion",
"Ollama Option — Temperature",
"Ollama Option — Seed",
"Ollama Option — Max Tokens",
+2
View File
@@ -17,5 +17,7 @@ services:
- sh
- -c
- |
git clone https://github.com/Comfy-Org/ComfyUI-Manager.git /app/ComfyUI/custom_nodes/comfyui-manager
pip install -q -r /app/ComfyUI/custom_nodes/comfyui-manager/requirements.txt
pip install -q -r /app/ComfyUI/custom_nodes/comfydv/requirements.txt
exec python main.py --cpu --listen 0.0.0.0 --port 8188
+2 -2
View File
@@ -1,4 +1,4 @@
FROM python:3.11-slim
FROM python:3.13-slim
WORKDIR /app
@@ -10,7 +10,7 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
&& rm -rf /var/lib/apt/lists/*
# Clone ComfyUI (pinned tag for reproducibility; bump manually)
ARG COMFYUI_VERSION=v0.3.44
ARG COMFYUI_VERSION=v0.27.0
RUN git clone --depth 1 --branch ${COMFYUI_VERSION} \
https://github.com/comfyanonymous/ComfyUI.git /app/ComfyUI
Binary file not shown.

Before

Width:  |  Height:  |  Size: 12 KiB

After

Width:  |  Height:  |  Size: 12 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 32 KiB

After

Width:  |  Height:  |  Size: 33 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 40 KiB

After

Width:  |  Height:  |  Size: 41 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 13 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 37 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 38 KiB

After

Width:  |  Height:  |  Size: 55 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 11 KiB

After

Width:  |  Height:  |  Size: 12 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 50 KiB

After

Width:  |  Height:  |  Size: 55 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 40 KiB

After

Width:  |  Height:  |  Size: 41 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 58 KiB

After

Width:  |  Height:  |  Size: 64 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 19 KiB

After

Width:  |  Height:  |  Size: 18 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 68 KiB

+93 -62
View File
@@ -1,24 +1,37 @@
# comfydv
A collection of workflow efficiency and quality-of-life nodes built out of necessity for personal ComfyUI use.
**Quality-of-life nodes for ComfyUI, built to disappear into your workflow.**
`comfydv` fills the gaps ComfyUI's built-in library leaves on the table: string templates that build their own sockets as you type, seed-controlled randomisation, graceful mid-queue interruption, and a local-LLM integration that doesn't care whether you're running Ollama or llama.cpp. No Python required — install it, drop the nodes on your canvas, wire them up.
![Chat Completion in action](assets/ollama_chat.png)
## What is comfydv?
A small, focused ComfyUI utility pack. It exists because:
- **It reads your intent, not just your syntax.** Format String detects `{variables}` in a template and adds/removes input sockets live, as you type — no manual socket wrangling.
- **One LLM integration, any local backend.** Wire a Chat Completion node once; swap between Ollama and llama.cpp by changing a single upstream client node. Structured output, multi-turn history, and model load/unload work identically on both.
- **It fails politely.** Circuit Breaker halts a queue run cleanly instead of throwing a stack trace at you; a disconnected LLM server gets a specific, actionable error instead of a silent empty dropdown.
- **Small, tested, boring in the best way.** Every node is unit-tested and the local-LLM nodes are verified against real running servers, not just mocks.
## What's inside
| Node | What it does |
|------|-------------|
| **Format String** | Formats a string from a Python f-string or Jinja2 template. Detects variables in the template and automatically adds/removes input sockets. |
| **Random Choice** | Accepts any number of typed inputs and outputs one at random, with a configurable seed for reproducibility. |
| **Circuit Breaker** | Halts the current ComfyUI queue run gracefully without crashing the server. Wire the `status` toggle to a boolean condition to skip the rest of the queue when a condition isn't met. |
| **Ollama Client** | Configures a connection to an Ollama server (default: `http://localhost:11434`). Threads the host URL through the graph as an `OLLAMA_CLIENT` socket. |
| **Ollama Model Selector** | Fetches the live model list from Ollama and presents it as a dropdown. Outputs the selected model name. |
| **Ollama Load Model** | Loads a model into Ollama's memory using `/api/generate` with `keep_alive=-1`. |
| **Ollama Unload Model** | Evicts a model from Ollama's memory using `/api/generate` with `keep_alive=0`. |
| **Ollama Chat Completion** | Sends a prompt (and optional conversation history) to Ollama `/api/chat`. Response and history are shown inline in the node body and available as output sockets. |
| **Ollama Option — \*** | Seven composable option nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into an `OLLAMA_OPTIONS` dict wired into Chat Completion. |
| **Ollama Debug History** | Serialises an `OLLAMA_HISTORY` list to a pretty-printed JSON string for inspection. |
| **Ollama History Length** | Returns the number of messages in an `OLLAMA_HISTORY` list as an integer. |
|------|---------------|
| **Format String** | Renders a Python f-string or Jinja2 template. Sockets appear and disappear automatically as you type variables. |
| **Random Choice** | Accepts any number of typed inputs and returns one at random, with a seed for reproducibility. |
| **Circuit Breaker** | Halts the current queue run gracefully — no crash, just a clean stop — when a condition isn't met. |
| **Ollama Client** / **LlamaCpp Client** | Configure a connection to a local Ollama or llama.cpp server. Both emit the same `LLM_CLIENT` socket — every node below works with either. |
| **LLM Model Selector** | Live dropdown of models available on the connected server. |
| **LLM Load Model** / **LLM Unload Model** | Explicit VRAM management — pin a model in memory before inference, evict it after. |
| **Chat Completion** | Send a prompt (optionally with history) to the connected server; response shown inline and as an output socket. |
| **Ollama Option — \*** | Seven composable parameter nodes (Temperature, Seed, Max Tokens, Top P, Top K, Repeat Penalty, Extra Body) that merge into Chat Completion's `options` input. |
| **Ollama Debug History** / **Ollama History Length** | Inspect an `OLLAMA_HISTORY` conversation list — pretty-print it or count its messages. |
## Install
**Via ComfyUI Manager** (recommended): search for `comfydv` and click Install.
**Via ComfyUI Manager** (recommended): search for `comfydv`, click Install.
**Manual:**
@@ -27,112 +40,130 @@ cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/darth-veitcher/comfydv.git
```
Restart ComfyUI. The nodes appear under the **dv/** and **dv/ollama** categories in the node menu. Runtime dependencies (`jinja2`, `aiohttp`) are installed automatically via `requirements.txt`.
Restart ComfyUI. Nodes appear under **dv/**, **dv/ollama**, and **dv/llamacpp** in the node menu. Runtime dependencies (`jinja2`, `aiohttp`, `pydantic-ai`) install automatically via `requirements.txt`.
For Ollama nodes: [install Ollama](https://ollama.com/download) and pull at least one model (`ollama pull qwen2.5:latest`) before using the Ollama nodes.
For local-LLM nodes, bring your own backend:
- **Ollama** — [install Ollama](https://ollama.com/download), pull a model (`ollama pull qwen2.5:latest`).
- **llama.cpp** — [build/install `llama-server`](https://github.com/ggml-org/llama.cpp), launch it in [router mode](#llamacpp).
## Quickstart
1. Right-click the canvas → Add Node → **dv/** for Format String, Random Choice, and Circuit Breaker.
2. For local LLM nodes: start Ollama (`ollama serve`) or `llama-server` (router mode), then add nodes from **dv/ollama/** or **dv/llamacpp/** — the chat/model-management nodes are shared between both backends.
---
## Format String
Formats text from a Python f-string or Jinja2 template. As you type the template, input sockets appear and disappear automatically — one per variable detected.
### Python f-strings
Type `{variable_name}` and a socket appears. Wire it to any string output in your workflow.
Type a template, get sockets. `{variable_name}` in f-string mode, `{{ variable_name }}` in Jinja2 mode — either way, comfydv watches what you type and keeps the node's inputs in sync automatically.
![Format String — f-string mode](assets/fstring.png)
| Output | Content |
|--------|---------|
| `formatted_string` | The rendered result |
| `saved_file_path` | Path written to disk (if `save_path` is set) |
| `<var>` … | Pass-through of each input value, for easy chaining |
| `saved_file_path` | Where it was written, if `save_path` is set |
| `<var>` … | Each input passed through unchanged, for easy chaining |
### Jinja2 templates
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`, …), conditionals (`{% if %}…{% endif %}`), and loops.
Switch `template_type` to **Jinja2** to unlock filters (`| upper`, `| int`), conditionals, and loops:
![Format String — Jinja2 mode](assets/jinja2.png)
Variables detected in `{{ }}` expressions become input sockets exactly as in Simple mode. See the [Jinja2 documentation](https://jinja.palletsprojects.com/en/latest/) for the full filter/test reference.
---
## Random Choice
Connect any number of inputs of the same type. Each run picks one at random. Set `seed` for reproducibility.
Wire in any number of same-typed inputs — images, strings, conditioning, anything ComfyUI can carry over a socket — and get one back at random.
![Random Choice](assets/random.png)
- Accepts any ComfyUI type (STRING, IMAGE, CONDITIONING, …)
- Add as many inputs as you like; unused slots are removed automatically when disconnected
- `seed = 0` randomises on every run; any other value locks the selection
`seed = 0` randomises every run; any other value locks the selection. Unused input slots vanish automatically when you disconnect them.
---
## Circuit Breaker
Stops the queue gracefully when a condition isn't met — no crash, no error, just a clean halt.
Stop a queue run cleanly when a condition isn't met, instead of letting a downstream node crash on bad input.
![Circuit Breaker](assets/circuit_breaker.png)
Wire an image (or any trigger) into `trigger` and a boolean into `status`. When `status` is **false** the node raises `InterruptProcessingException`, which tells ComfyUI to stop the current run cleanly. When `status` is **true** the image passes through unchanged.
Wire a trigger (an image, or anything) into `trigger` and a boolean into `status`. `status = false` raises `InterruptProcessingException` — ComfyUI stops the run without an error dialog. `status = true` passes the trigger straight through.
Typical use: skip an expensive upscale step when a quality-check node says the draft is already good enough.
Typical use: skip an expensive upscale pass when an upstream quality-check node says the draft's already good enough.
---
## Ollama
## Local LLMs
14 nodes for integrating a local Ollama LLM into your ComfyUI workflow. The host URL is configured once in **Ollama Client** and threaded through the graph — all downstream nodes receive it via the `OLLAMA_CLIENT` socket.
One set of nodes, two interchangeable backends. Configure a connection once with **Ollama Client** or **LlamaCpp Client** — both output the same `LLM_CLIENT` socket — and every downstream node (model selection, load/unload, chat, structured output, multi-turn history) works exactly the same way regardless of which one you picked. Swapping backends means rewiring one node, not rebuilding your graph.
### Ollama Client node
Configure the server address once; all downstream Ollama nodes inherit it automatically.
### Connect and chat
![Ollama Client](assets/ollama_client.png)
### Model lifecycle (load and unload)
1. **Ollama Client** (default `http://localhost:11434`) or **LlamaCpp Client** (default `http://localhost:8080`) — set the host.
2. **LLM Model Selector** — pick a model from the live dropdown, or wire a model name straight into Chat Completion.
3. **Chat Completion** — wire in client, model, and prompt. The response renders inline in the node and is also available as an output socket.
On memory-constrained machines and single-GPU setups, explicitly loading and unloading the model before and after inference is critical. **Ollama Load Model** pins the model into VRAM (`keep_alive=-1`); **Ollama Unload Model** evicts it immediately (`keep_alive=0`), freeing memory for image generation or other models.
![Chat Completion](assets/ollama_chat.png)
![Ollama Load / Unload](assets/ollama_lifecycle.png)
A complete graph looks like this:
The correct chain is **Load → Chat → Unload**, enforced through data dependencies:
![Full LLM workflow](assets/ollama_workflow.png)
1. Wire `OllamaLoadModel.model_name` → `OllamaChatCompletion.model`. This creates the data dependency that guarantees Load runs before Chat and passes the model name into the Chat node's plain-string `model` input.
2. Wire `OllamaChatCompletion.model_name` → `OllamaUnloadModel.model`. This guarantees Unload runs after Chat completes.
3. Optionally wire `OllamaChatCompletion.response` → `OllamaUnloadModel.passthrough` — Unload returns the response unchanged so the rest of your workflow can still consume it.
### Manual memory management
### Minimal chat workflow
Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline.
1. **Ollama Client** → set host (default `http://localhost:11434`)
2. **Ollama Model Selector** → pick a model from the live dropdown (or type/wire a model name directly into Chat Completion's `model` input)
3. **Ollama Chat Completion** → wire client + model + prompt; the response appears inline in the node body and is also available as an output socket
![Load / Unload lifecycle](assets/ollama_lifecycle.png)
![Ollama Chat Completion](assets/ollama_chat.png)
The **Load → Chat → Unload** chain is enforced by data dependencies, not by convention:
Wire multiple nodes together for a complete end-to-end workflow:
1. `LLMLoadModel.model_name` → `ChatCompletion.model` — guarantees Load runs before Chat, and feeds the model name straight in.
2. `ChatCompletion.model_name` → `LLMUnloadModel.model` — guarantees Unload runs after Chat completes.
3. *(Optional)* `ChatCompletion.response` → `LLMUnloadModel.passthrough` — Unload returns the response unchanged, so the rest of your workflow can still consume it.
![Ollama Full Workflow](assets/ollama_workflow.png)
### Tuning generation
### Option nodes
Chain any combination of **Ollama Option —** nodes ahead of Chat Completion to override inference parameters:
Chain any combination of **Ollama Option —** nodes before Chat Completion to override inference parameters:
![Ollama Option nodes](assets/ollama_options.png)
| Option node | Ollama param |
|-------------|-------------|
| Option node | Parameter |
|-------------|-----------|
| Temperature | `temperature` |
| Seed | `seed` |
| Max Tokens | `num_predict` |
| Top P | `top_p` |
| Top K | `top_k` |
| Repeat Penalty | `repeat_penalty` |
| Extra Body | arbitrary JSON merged into options |
![Ollama Option Nodes](assets/ollama_options.png)
| Extra Body | arbitrary JSON, merged into `options` |
### Multi-turn conversations
`OLLAMA_HISTORY` flows out of Chat Completion as a list of `{"role", "content"}` dicts. Wire it back into the next Chat Completion for multi-turn conversations, or inspect it with **Ollama Debug History** / **Ollama History Length**.
`OLLAMA_HISTORY` flows out of Chat Completion as a `{"role", "content"}` list. Feed it back into the next Chat Completion call for multi-turn context, or inspect it with **Ollama Debug History** / **Ollama History Length**.
### llama.cpp
Everything above works unchanged against llama.cpp — swap in a **LlamaCpp Client** and the rest of the graph doesn't know the difference. The one thing llama.cpp needs that Ollama doesn't: **router mode**, a directory of models rather than a single `-m model.gguf`:
```bash
llama-server --models-dir ./models -c 8192
```
In exchange, router mode gives comfydv a richer live status than Ollama can report — `loading` and `downloading`, not just loaded/unloaded — plus the same explicit load/unload primitives Ollama's nodes already use.
### Upgrading a workflow saved before this rename
Nodes were renamed once, to make them backend-generic (`OllamaChatCompletion` → `ChatCompletion`, etc.). If ComfyUI reports old node types as missing when you reopen a saved workflow, reconnect using this table — behavior is unchanged, only the names are:
| Old | New |
|-----|-----|
| `OllamaChatCompletion` | `ChatCompletion` |
| `OllamaModelSelector` | `LLMModelSelector` |
| `OllamaLoadModel` | `LLMLoadModel` |
| `OllamaUnloadModel` | `LLMUnloadModel` |
| `OLLAMA_CLIENT` socket | `LLM_CLIENT` socket |
`OllamaClient` kept its name — delete and re-add any node showing as missing, then rewire it to the same `OllamaClient` node.
+6
View File
@@ -5,3 +5,9 @@
# the bullet survives a fresh clone (`beacon doctor`'s active-bullet check then
# works in CI) and no per-branch file is left behind on the trunk after a merge.
# `git log -p` on this file is the audit trail of who started which bullet when.
[bullets."claude/chatcompletion-image-input-7mm6q5"]
title = "VLM image input for ChatCompletion"
owner = "noreply@anthropic.com"
started = "2026-07-22T19:11:08+00:00"
epic = "vlm-image-input"
@@ -0,0 +1,172 @@
# ADR-006: Structured Ollama output via OpenAI-compatible tool-calling + dynamic pydantic validation, not pydantic-ai
## Status
> Superseded by [ADR-007](ADR-007-llm-provider-adapter-pattern.md)
_Date:_ 2026-07-09
_Deciders:_ darth-veitcher
---
## Context
`OllamaChatCompletion` sends free-text prompts to Ollama's `/api/chat` and
returns `result["message"]["content"]` as-is, with no constraint on what the
model may emit. In practice this is unreliable in three concrete ways:
models prepend commentary ("Here you are:"), wrap responses in ` ``` ` code
fences, or occasionally return blank content — all things prompt wording
alone (e.g. a stricter `system` message) cannot reliably prevent, since it
only *asks* the model to behave, it doesn't constrain what tokens the
decoder is able to produce.
The obvious library to reach for is `pydantic-ai`, which offers structured,
validated LLM output as a first-class feature. Its Ollama support, however,
goes through an OpenAI-compatible client — which pulls in the `openai` SDK,
which depends on `httpx`. This directly reverses
[ADR-004](ADR-004-aiohttp-over-httpx-for-ollama.md), which explicitly
rejected `httpx` as "a new runtime dep for functionality aiohttp already
provides" (ComfyUI's own server is aiohttp-based, so aiohttp is a guaranteed
transitive dependency; httpx is not).
Two Ollama-native mechanisms can force structured output without a new HTTP
client, since both are reachable over plain JSON POST via the existing
`aiohttp`-based `_post_json`:
1. **Native `/api/chat` `"format"` field** — a JSON Schema that constrains
**decoding itself** (grammar-constrained sampling): the model's sampler
is restricted to only emit tokens matching the schema.
2. **OpenAI-compatible `/v1/chat/completions` tool-calling** — a `tools`
array plus `tool_choice` forcing a single named function call, relying on
the model's own trained function-calling behavior rather than a
grammar-to-token mapping.
Both were tried against a real local model
(`lukey03/qwen3.5-9b-abliterated-vision`) during implementation. The native
`format` field was silently ignored — the model returned plain unstructured
text (`"pong"`) despite the schema constraint, reproduced twice. Inspecting
`/api/show` revealed this model's `TEMPLATE` is a degenerate `{{ .Prompt }}`
with no role/message structure — consistent with a community "abliteration"
process having modified the tokenizer/vocab in a way that breaks Ollama's
grammar-to-token mapping, causing it to silently fall back to unconstrained
generation instead of erroring. Tool-calling doesn't depend on that mapping;
it succeeded on its first test against the same model. (A follow-up
tool-calling call with `options` included did also fail — this specific
model appears broadly unreliable, consistent with a degraded fine-tune, so
this evidence is suggestive rather than conclusive. No second generative
model was available locally to get a cleaner signal.)
## Decision
Use Ollama's OpenAI-compatible tool-calling (`/v1/chat/completions`,
`tools`/`tool_choice` forcing a single call) for `structured_output=True`
requests, sent through the existing `_post_json` helper — no new HTTP
client, no `pydantic-ai`, no `openai` SDK. Non-structured requests are
completely unaffected and keep using native `/api/chat`.
Use plain `pydantic` (not `pydantic-ai`) purely as a validation layer:
given the user-supplied JSON Schema (`output_schema` input on
`OllamaChatCompletion`), dynamically build a `pydantic.BaseModel` via
`pydantic.create_model(...)` and validate/parse the tool call's `arguments`
JSON against it. Required *string* fields get `min_length=1` — JSON
Schema's `"required"` only checks presence, so a model could satisfy it
with `""`, silently reintroducing the "blank output" problem. On validation
failure (invalid JSON, missing/empty required field, or the model not
calling the tool at all — observed to happen even with `tool_choice`
forcing it), retry with fresh network calls (bounded by a `max_retries`
input, clamped to 0–5); if every attempt fails, raise a clear
`RuntimeError` naming the model, the attempt count, and a truncated
snippet of the last invalid response — never silently degrade to
unvalidated content.
This is opt-in: a new `structured_output: BOOLEAN` input on the existing
`OllamaChatCompletion` node, default `False`. When off, behavior is
unchanged — no `tools`/`tool_choice` sent, native `/api/chat` used, no
dynamic outputs, `RETURN_TYPES` stays the original fixed 3-tuple. When on,
one additional ComfyUI output socket is exposed per schema property
(mirroring `FormatString`'s existing dynamic-output-socket pattern via
`unique_id`/`RETURN_TYPES` mutation), so downstream nodes can consume
individually typed fields instead of parsing JSON themselves.
## Consequences
**Easier:**
- Fixes the three concrete unreliability problems without depending on a
model/tokenizer-sensitive grammar-constraint mechanism that was observed
to fail silently on at least one real model.
- No new HTTP stack: `pydantic` is a validation-only dependency, not a
client library. ADR-004's aiohttp-only stance is preserved.
- Fully backward compatible — `structured_output` defaults off, and
non-structured requests still use native `/api/chat` exactly as before.
**Harder / constrained:**
- Tool-calling depends on the model having usable trained function-calling
behavior. Models with no tool-calling training may perform worse here
than they would under grammar-constrained `format` decoding — this
repo's only local test model was itself too unreliable to fully confirm
either mechanism's ceiling. If well-behaved-model testing later shows
native `format` is meaningfully more reliable in the common case, this
decision should be revisited rather than treated as permanent.
- Only a flat `properties: {name: {type: ...}}` shape is interpreted into
typed ComfyUI sockets. Complex JSON Schema constructs (`$ref`,
`oneOf`/`anyOf`/`allOf`, `enum`, nested `object`/`array` item schemas) are
still forwarded to Ollama verbatim as the tool's `parameters`, but
comfydv's own type mapping falls back to `STRING` for anything it doesn't
recognize — no nested typed sockets.
- `RETURN_TYPES`/`RETURN_NAMES` are class-level state, shared across every
`OllamaChatCompletion` instance in a graph (same accepted limitation
`FormatString.update_widget` already ships with) — the first execution
after toggling `structured_output` or editing `output_schema` may show
stale downstream socket typing until it runs once.
- Neither mechanism is guaranteed 100% across all versions/models — hence
the retry-then-raise defense-in-depth, rather than trusting either
constraint blindly.
**Debt introduced:**
- None. `pydantic` is a widely-used, low-conflict-risk dependency; ComfyUI
itself is expected to already bundle it for its own API layer, though
this repo lists it explicitly in both `pyproject.toml` and
`requirements.txt` per ADR-003 rather than assume so.
## Considered Alternatives
### Alternative A: `pydantic-ai`
**Why rejected:** Its Ollama support goes through an OpenAI-compatible
client, reintroducing `httpx` + the `openai` SDK — the exact dependency
ADR-004 evaluated and rejected. The reliability benefit it offers is the
same tool-calling/validation mechanism this ADR adopts directly over plain
`aiohttp`, without the added dependency weight.
### Alternative B: Native `/api/chat` `"format"` field (JSON-Schema-constrained decoding)
**Why rejected as primary:** Theoretically the stronger guarantee — a
sampler-level constraint rather than learned behavior — and remains a
reasonable mechanism for well-behaved models. Rejected here because it
failed outright (silently ignored, not even erroring) against the one real
model available for testing, traced to that model's modified tokenizer
breaking Ollama's grammar-to-token mapping. Tool-calling succeeded where it
failed. See "Harder / constrained" above — this may be revisited if
broader testing shows native `format` is more reliable in the common case.
### Alternative C: Prompt-only enforcement (stricter `system` message)
**Why rejected:** `system` already exists as an input and users can already
try this — it's what led to the reported problem in the first place.
Wording can reduce commentary/fences/blank output but cannot guarantee
their absence, since nothing constrains the actual token stream.
### Alternative D: Response-side post-processing (regex-strip fences/preamble)
**Why rejected as the primary fix:** Cheap, but fundamentally reactive —
it can strip a fence wrapper after the fact but can't recover genuinely
blank output, and heuristics for "commentary" are unreliable across models.
Not pursued as a fallback either, to keep this change minimal and avoid two
competing "make output clean" mechanisms with unclear precedence.
---
## Links
- Related ADRs: [ADR-003](ADR-003-requirements-txt-authoring-policy.md), [ADR-004](ADR-004-aiohttp-over-httpx-for-ollama.md), [ADR-005](ADR-005-ollama-host-config-via-client-node.md)
- Originating epic (archived, scope predates this decision): `project-management/Roadmap/epics/archive/ollama-integration.md`
@@ -0,0 +1,198 @@
# ADR-007: LLMProvider adapter pattern shared across Ollama and llama.cpp
## Status
> Accepted
_Date:_ 2026-07-11
_Deciders:_ darth-veitcher
---
## Context
GitHub issue #15 asks comfydv to add ComfyUI nodes for llama.cpp, mirroring
the existing Ollama integration
(`project-management/Roadmap/epics/archive/ollama-integration.md`), and
explicitly poses the design question: separate nodes per backend, or an
adapter pattern that shares code? It's motivated by llama.cpp's new "router
mode" (`llama-server --models-dir <dir>`, via
[llama.cpp PR #18228](https://github.com/ggml-org/llama.cpp/pull/18228),
merged 2025-12-21), which exposes `GET /models`, `POST /models/load`,
`POST /models/unload`, and `--sleep-idle-seconds` auto-unload — giving
llama.cpp the same manual load/unload memory-management primitives comfydv
already relies on for Ollama.
`src/comfydv/ollama.py` (1056 lines, 17 node classes) has no abstraction
layer today: two module-level free functions (`_post_json`, `_fetch_models`)
are called directly by every node, with Ollama endpoint paths hardcoded
inline; socket types (`OLLAMA_CLIENT`, `OLLAMA_OPTIONS`, `OLLAMA_HISTORY`)
and `PromptServer` routes (`/dv/ollama/...`) are Ollama-named throughout.
**Two things needed resolving to answer issue #15's question honestly:**
1. **Does model lifecycle management (list/load/unload) actually converge
between backends, or not?** At the wire-protocol level, no: Ollama uses
`/api/tags` + a per-request `keep_alive` TTL on `/api/generate`;
llama.cpp router mode uses `/models` + explicit `/models/load` /
`/models/unload` + named status states
(`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`). Judged at that
level, a shared interface looks forced. But at the *conceptual* level,
both APIs support exactly the same four operations — list models with
status, load a model, unload a model, and generate/chat against a loaded
model — just with different mechanics. Ollama's `keep_alive`-based
load/unload already *is* `load_model()`/`unload_model()`, mechanically
implemented as a side effect of a `/api/generate` call rather than a
dedicated endpoint. A `Protocol` boundary at the operation level, not the
wire-format level, fits both backends without forcing anything.
2. **Does `pydantic-ai` make sense now that a second backend exists?**
[ADR-006](ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md)
(2026-07-09) rejected `pydantic-ai` for a single backend because its
Ollama support pulls in `httpx` + the `openai` SDK, reversing
[ADR-004](ADR-004-aiohttp-over-httpx-for-ollama.md)'s aiohttp-only
stance. A research pass against current `pydantic-ai` docs/source (this
moves fast and postdates training data, so verified live rather than
assumed) found: `httpx` is a **base dependency of `pydantic-ai-slim`
itself**, not merely pulled in by an OpenAI-specific extra; `openai` SDK
+ `tiktoken` are additionally required to reach any OpenAI-compatible
backend; there is no aiohttp transport option anywhere in pydantic-ai.
This is a **fixed, one-time dependency tax**, not one that grows per
backend. Separately, `pydantic.create_model()`-built `BaseModel`
subclasses (comfydv's existing pattern for validating against a
user-supplied JSON-Schema string at workflow-execution time) work as
pydantic-ai's `output_type` with no special-casing — the dynamic-schema
requirement is not a blocker. And `OpenAIProvider(base_url=...)` is the
exact generic mechanism pydantic-ai's own `OllamaProvider` is built on
internally, so llama.cpp's `/v1/chat/completions` reaches an identical
code path with a different `base_url` — genuinely shared implementation,
not just a shared shape.
Both findings point the same direction: a real `Protocol`-based adapter,
with `pydantic-ai` as the mechanism behind its structured-output method.
## Decision
Define a `LLMProvider` `Protocol` (new internal module, e.g.
`src/comfydv/_llm/provider.py`) with the common surface:
```python
class LLMProvider(Protocol):
async def list_models(self) -> list[ModelInfo]: ...
async def load_model(self, model: str) -> None: ...
async def unload_model(self, model: str) -> None: ...
async def chat(self, model: str, messages: ..., options: ...) -> str: ...
async def chat_structured(self, model: str, messages: ..., schema: type[BaseModel], options: ...) -> BaseModel: ...
```
`OllamaProvider` and `LlamaCppProvider` each implement it, absorbing their
own REST mechanics internally (Ollama: `/api/tags`, `/api/generate` with
`keep_alive`; llama.cpp: `/models`, `/models/load`, `/models/unload`) via
`aiohttp`, unchanged from ADR-004's stance for non-chat calls. Both
implement `chat_structured()` via the same `pydantic-ai` `Agent`/
`output_type` call through `OpenAIProvider(base_url=...)` — one shared
implementation, differing only in `base_url` and model name.
ComfyUI nodes become **generic, not per-backend**: `OllamaClient` and
`LlamaCppClient` both output the same `LLM_CLIENT` socket type (each
internally constructs the matching provider); a single `LLMModelSelector`,
`LLMLoadModel`, `LLMUnloadModel`, and `ChatCompletion` node operate against
`LLM_CLIENT` generically. Swapping providers on the canvas means rewiring
which client node feeds the chat/management nodes, not swapping node
classes — this is the direct answer to issue #15's question: **adapter
pattern**, implemented as a protocol boundary at the operation level.
This **supersedes ADR-006**: `OllamaChatCompletion`'s `structured_output=True`
path moves from hand-rolled tool-calling to `pydantic-ai` via the protocol.
This **narrows ADR-004's scope**: aiohttp remains the transport for every
non-chat REST call inside each provider; `httpx`/`openai` enter the
dependency tree scoped specifically to `chat_structured()`, via
`pydantic-ai`.
**Documented approximation:** `ModelStatus` includes `sleeping` and
`downloading`, states that exist in llama.cpp router mode but not in
Ollama's API. `OllamaProvider.list_models()` normalizes into the same enum
rather than inventing Ollama-specific states — a model that's resident and
idle maps to `loaded` (Ollama has no distinct "kept warm but not serving"
signal via this API), and `downloading` is simply never emitted by
`OllamaProvider` (Ollama's pull/download flow is out of scope per the
original Ollama epic's non-goals). This is an accepted, explicit
approximation, not a silent gap.
**Confirmed:** adopting generic node/socket names (`LLM_CLIENT`,
`ChatCompletion`, etc.) means renaming away from `OLLAMA_CLIENT`,
`OllamaChatCompletion`, and similar — a breaking change for any saved
workflow using the current names. The Ollama integration shipped
2026-07-04, so the blast radius is small. Confirmed 2026-07-11: rename in
place now rather than carry Ollama-prefixed names forward or maintain
deprecated aliases indefinitely.
## Consequences
**Easier:**
- Issue #15's question gets a real answer: one generic node set works with
any backend that implements `LLMProvider`, including future ones (a third
local server, or a hosted OpenAI/Anthropic provider) without new node
classes.
- One implementation of tool-calling/structured-output logic instead of
duplicating it per backend; `pydantic-ai` brings built-in retry/validation
machinery, replacing ADR-006's hand-rolled retry loop.
- The protocol boundary keeps each backend's REST quirks contained inside
its provider — the graph never has to know Ollama uses `keep_alive` while
llama.cpp uses explicit load/unload endpoints.
**Harder / constrained:**
- New dependencies (`pydantic-ai`, `openai`, `tiktoken`, `httpx`) land in a
project that was previously aiohttp-only.
- This is a nontrivial migration of tested, shipped Ollama code (the
`ollama-integration` epic is Done) — not purely additive work. Must be
proven regression-safe before it's trusted as the foundation for
llama.cpp.
- `ModelStatus` is not perfectly symmetric across backends — the
`sleeping`/`downloading` states are llama.cpp-only in practice; documented
above, but still a leak of llama.cpp's richer vocabulary into a
nominally-generic type.
- The node/socket rename is a breaking change for existing saved workflows —
confirmed acceptable given the small blast radius (see above).
**Debt introduced:**
- None deliberately, contingent on the migration preserving existing
Ollama behavior exactly (verified against `tests/test_ollama.py`).
## Considered Alternatives
### Alternative A: Shared `pydantic-ai` chat layer only; separate per-backend management nodes
**Why rejected:** This was the first-pass design — judged convergence at
the REST wire-protocol level (Ollama's `/api/tags`+`keep_alive` vs.
llama.cpp's `/models`+`/models/load`+`/models/unload` don't look alike) and
concluded a shared interface would be forced. That framing was wrong: the
right level to judge convergence is the *operation* (list/load/unload/chat),
not the wire format. Both backends genuinely support the same four
operations; only their REST mechanics differ, and those differences belong
inside each provider implementation, not on the graph.
### Alternative B: Hand-roll llama.cpp's structured output too (duplicate ADR-006's approach)
**Why rejected:** Two independent implementations of the same
OpenAI-compatible tool-calling mechanism is the DRY violation issue #15
raises in the first place, with no offsetting benefit now that a second
backend exists to justify a shared layer.
### Alternative C: Keep `pydantic-ai` rejected; extract a shared aiohttp-based internal helper instead
**Why rejected:** Avoids new dependencies entirely, but forces re-deriving
`pydantic-ai`'s retry/validation machinery by hand for no benefit beyond
dependency-avoidance — and doesn't change the model-management convergence
question at all (that's orthogonal to which HTTP client the chat path
uses). The one-time dependency tax is judged worth paying for the fuller
abstraction, now that two backends exist to amortize it against.
---
## Links
- Related epics: `project-management/Roadmap/epics/llm-provider-abstraction.md`, `project-management/Roadmap/epics/llamacpp-integration.md`
- Related ADRs: [ADR-004](ADR-004-aiohttp-over-httpx-for-ollama.md) (narrowed), [ADR-005](ADR-005-ollama-host-config-via-client-node.md) (client-node pattern generalized to `LLM_CLIENT`), [ADR-006](ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md) (superseded)
- External reference: [llama.cpp PR #18228](https://github.com/ggml-org/llama.cpp/pull/18228) (router mode, merged 2025-12-21), GitHub issue #15
@@ -0,0 +1,158 @@
# ADR-008: Multimodal image input carried on the Message across the LLMProvider boundary
## Status
> Proposed
_Date:_ 2026-07-22
_Deciders:_ darth-veitcher
---
## Context
The `ChatCompletion` node and the `LLMProvider` protocol (ADR-007) are
text-only today. `Message` (`src/comfydv/_llm/provider.py`) carries a single
`content: str`; `ChatCompletion`'s `INPUT_TYPES` (`src/comfydv/ollama.py`)
exposes no `IMAGE` socket. Users want to couple the existing chat node with a
vision-capable model (a VLM) to describe or reason about an image produced
elsewhere in a ComfyUI workflow.
ADR-007 deliberately scoped this out: it defined `Message` as text-only and
recorded that "if llama.cpp's router mode needs a protocol capability that
doesn't exist yet, that is a protocol change scoped as its own follow-up, not
silently special-cased." This ADR is that follow-up — it extends the same
adapter pattern to a second input modality.
The two shipped backends carry images very differently on the wire, and the
node has two distinct code paths (free-text vs structured), so a decision is
needed about **where** an image lives as it crosses the provider boundary and
**who** translates it into each backend's native shape:
- **Ollama free-text** — `OllamaProvider.chat()` posts `[m.model_dump() for m
in messages]` to the native `/api/chat`, which accepts a per-message
`images` field: an array of base64-encoded image data alongside the text
`content`.
- **llama.cpp free-text** — `LlamaCppProvider.chat()` posts to the
OpenAI-compatible `/v1/chat/completions`, where a message's `content` is a
list of typed parts (`{"type": "text", ...}`,
`{"type": "image_url", "image_url": {"url": "data:image/...;base64,..."}}`)
— a flat sibling `images` field is not understood.
- **Structured output (both backends)** — routed through the shared
`chat_structured()` helper (`src/comfydv/_llm/chat.py`) over pydantic-ai,
which represents images as typed multimodal content
(`BinaryContent` / `ImageUrl`) inside the user prompt, not as a raw request
field.
The competing concern is DRY vs. leakage: a single carrier keeps the graph and
the node backend-agnostic (ADR-007's whole point), but the per-backend wire
shapes are irreducibly different and must be translated somewhere.
## Decision
**Carry images as an optional field on `Message`, and make each provider
responsible for translating that field into its own native wire shape** — the
exact same division of responsibility ADR-007 established for text and model
management (operation-level protocol, wire-format quirks contained inside each
provider).
1. **Protocol** — extend `Message` with an optional
`images: list[str] | None = None`, where each entry is a base64-encoded
image. `content` stays required; a text-only message sets `images=None` and
is byte-for-byte unchanged from today (`model_dump()` omits it or emits
`null`), so all existing Ollama/llama.cpp behavior is preserved.
2. **Node** — `ChatCompletion` gains one **optional** `image: ("IMAGE",)`
input. When wired, the node encodes the ComfyUI `IMAGE` tensor to base64
and attaches it to the user `Message` it already constructs. When not
wired, the node builds exactly the message it builds today. The node never
branches on which concrete provider it holds — consistent with ADR-007.
3. **Per-provider translation** (the leakage lives here, deliberately):
- `OllamaProvider.chat()` — the flat `images` field on the dumped message
already matches Ollama's native `/api/chat` schema; it flows through with
no transform.
- `LlamaCppProvider.chat()` — maps a message's `images` into OpenAI-style
`image_url` content parts before POSTing to `/v1/chat/completions`.
- `chat_structured()` (shared) — maps the last user message's `images` into
pydantic-ai multimodal content on the `user_prompt`; both backends inherit
this single implementation, mirroring how they already share the
structured text path.
4. **No new node classes and no new socket types** — image support is an
additional optional input on the *existing* generic node, so a workflow
author gains vision by wiring one socket, not by learning a new node. This
is the direct extension of ADR-007's "generic, not per-backend" node stance.
The base64 string is the neutral interchange form at the boundary because it
is the one representation every target consumes (Ollama's `images` array,
OpenAI's `data:` URI, and pydantic-ai's `BinaryContent` all accept it),
keeping the `Message` carrier itself provider-agnostic.
_Wire specifics (exact Ollama `/api/chat` image field, llama.cpp multimodal
readiness via `mmproj`, and pydantic-ai's multimodal content type) are
verified live in this feature's `research.md`/`plan.md` per project
convention, not assumed from training data._
## Consequences
**Easier:**
- One carrier (`Message.images`) and one node change unlock vision on both
backends at once; a future third provider implements image translation in
its own `chat()` exactly as it implements text, with no protocol churn.
- The graph and the node stay backend-agnostic — swapping providers still
means rewiring one client node, now including the image path.
- Text-only workflows are entirely unaffected (additive optional field +
optional socket).
**Harder / constrained:**
- `Message` is no longer a trivially-uniform text struct; each provider's
`chat()` (and the shared structured helper) must handle the `images` field,
even if only to pass it through. This is accepted leakage, localized to the
provider layer — the same tradeoff ADR-007 already made for `keep_alive` vs
explicit load/unload.
- Vision requires a model actually loaded with multimodal weights (Ollama
multimodal models; llama.cpp launched with an `mmproj` projector). A
text-only model receiving images degrades to a backend error, not a node
crash — surfacing that clearly is a spec requirement, not something this
boundary can prevent.
**Debt introduced:**
- None deliberately, contingent on text-only requests remaining byte-identical
to today (guarded by the existing Ollama/llama.cpp provider tests, which must
stay green).
## Considered Alternatives
### Alternative A: A separate `images` parameter threaded through `chat()`/`chat_structured()` signatures
**Why rejected:** Widens every provider method signature and the protocol for
a value that is conceptually part of a message turn. Images belong to a
specific message (which turn the picture accompanies), and multi-turn vision
histories need per-message association — a single side-channel parameter can't
express that. Putting it on `Message` keeps turn/image association intact and
leaves method signatures unchanged.
### Alternative B: A dedicated multimodal node / socket type separate from `ChatCompletion`
**Why rejected:** Reintroduces exactly the per-capability node proliferation
ADR-007 eliminated. A workflow author would maintain two chat nodes and
relearn one for vision. An optional input on the existing node is strictly
simpler and keeps the "one generic node set" promise.
### Alternative C: Normalize images to OpenAI content-parts at the boundary; make Ollama un-translate
**Why rejected:** Picks OpenAI's shape as the canonical form and forces the
Ollama provider — whose native API wants the simpler flat `images` array — to
convert *away* from it. That inverts the "each provider owns its own wire
format" principle and does more work on the currently-simpler path. A neutral
base64 carrier that every backend adapts *from* is the orthogonal choice.
---
## Links
- Related epic: `project-management/Roadmap/epics/vlm-image-input.md`
- Related spec: `specs/009-vlm-image-input/`
- Related ADRs: [ADR-007](ADR-007-llm-provider-adapter-pattern.md) (extended — same adapter pattern, second input modality), [ADR-005](ADR-005-ollama-host-config-via-client-node.md) (client-node pattern, unchanged)
- External reference: GitHub issue #15 (llama.cpp parity), Ollama multimodal `/api/chat` `images`, OpenAI vision `image_url` content parts
@@ -0,0 +1,82 @@
# ADR-009: Provider-specific structured output — NativeOutput for llama.cpp, hand-rolled native `/api/chat` for Ollama
## Status
> Accepted
_Date:_ 2026-07-25
_Deciders:_ darth-veitcher
---
## Context
`chat_structured()` (`src/comfydv/_llm/chat.py`, ADR-007) builds its `pydantic-ai` `Agent` with a bare `output_type=schema`. Passing a raw `pydantic.BaseModel` subclass this way makes `pydantic-ai` default to **tool-calling** (a synthetic forced function call) for structured output — inherited silently from ADR-007's move to `pydantic-ai`, never a deliberate re-decision. ADR-007 doesn't discuss output-mode choice at all.
Testing a real multi-agent ComfyUI workflow (`workflows/ltx-i2v-pipeline.json`) against a live local Ollama server surfaced this as a real reliability problem: against a "thinking"-capable model (`qwen3.5:9b`), tool-calling failed consistently — the model spent its entire token budget on internal chain-of-thought reasoning and never emitted the tool call, so every attempt failed pydantic validation ("token limit exceeded before any response was generated"), each attempt taking 5-7+ minutes before giving up.
### First fix attempt: `NativeOutput` + a priming call (superseded within this same ADR)
`pydantic-ai` 2.9.0 exposes `NativeOutput`, which makes the `Agent` use `response_format: {"type": "json_schema", ...}` over the OpenAI-compatible endpoint instead of tool-calling. Live-tested directly against Ollama via `curl` before touching code: `/v1/chat/completions` with `response_format: json_schema` returned clean, schema-valid JSON, with the model's reasoning in a separate `message.reasoning` field — fast, and reasoning no longer competed with structured output for token budget. This part of the fix is real and is kept — see Decision §1.
A second, separate problem was also found: Ollama's OpenAI-compatible endpoint doesn't honor a per-request `options` override (e.g. `num_ctx`) — sending the same request to native `/api/chat` reloaded the model at the requested context size; `/v1/chat/completions` silently kept whatever was already loaded. The first fix attempt worked around this with a priming call: hit native `/api/generate` with the desired `options` immediately before the real `/v1/chat/completions` request, on the theory that Ollama would keep the just-loaded context for the next call.
**This did not work, and the failure mode looked exactly like the original bug** — confirmed while re-testing the actual workflow end-to-end (`workflows/ltx-i2v-pipeline.json`, Agent 2 "Scene Grounder": long system prompt + 9-property schema + image), which kept failing with the identical "token limit exceeded" error even after the priming fix landed, tests passed, and `max_tokens`/`num_ctx`/`timeout_secs` were all raised generously. Isolated the exact mechanism with a direct, non-ComfyUI-mediated `curl` sequence:
1. `POST /api/generate` with `options: {num_ctx: 20480}`, `keep_alive: -1` → confirmed via `GET /api/ps`: `context_length: 20480`, loaded "forever".
2. Immediately `POST /v1/chat/completions` for the same model — **even with the identical `options: {num_ctx: 20480}` included in that request's body** → `GET /api/ps` immediately after: `context_length: 4096` (back to default), `expires_at` reset to a normal ~5-minute keep-alive.
So `/v1/chat/completions` doesn't merely *ignore* `options.num_ctx` — every call to it silently **reloads the model at the default context size**, discarding whatever was primed, regardless of what that same call's own `options` field says. A priming call immediately before the real request is structurally incapable of working, because the real request itself is what undoes the priming.
Re-ran the same sequence against native `/api/chat` instead of `/v1/chat/completions`: the primed `context_length: 20480` was preserved through and after the call. Native `/api/chat` also accepts `"format": <json schema>` directly, giving grammar-constrained structured output in the same request that correctly honors `options` — no separate priming call needed at all.
### Re-litigating ADR-006's model concern
This also revisits [ADR-006](ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md), which rejected native `format`-based output — but its rejection was scoped to one specific model, `lukey03/qwen3.5-9b-abliterated-vision`, whose degenerate chat template silently ignored the constraint. ADR-006 explicitly flagged this as revisitable: *"If well-behaved-model testing later shows native `format` is meaningfully more reliable in the common case, this decision should be revisited rather than treated as permanent."* Re-tested that exact model against native structured output: it no longer silently ignores the constraint (ADR-006's specific failure mode) — it returns schema-valid JSON, but the *content* is still garbled (`"ponáp∵49\n"` instead of the requested `"pong"`), consistent with ADR-006's "degenerate tokenizer" diagnosis. That model was also already failing under the tool-calling path (hanging without completing, observed live during this same investigation). So neither part of this decision regresses that model — it was already unusable for structured output either way. `structured_output` has no production users yet, so there is no back-compat concern in making this change.
## Decision
**1. `LlamaCppProvider` keeps the shared `pydantic-ai` path, switched to `NativeOutput`.** `chat.py`'s `_build_agent()` builds `Agent(chat_model, output_type=NativeOutput(schema), retries=0)` instead of a bare `output_type=schema`. llama-server's OpenAI-compatible endpoint is its genuine native structured-output surface (no equivalent context-reload bug found or expected — llama-server's context is fixed at process launch via `--ctx-size`, not a per-request concern, so there's nothing for a request to silently reset), so the shared-implementation architecture from ADR-007 stays intact for this provider.
**2. `OllamaProvider.chat_structured()` no longer uses `chat.py` at all.** It hand-rolls its own call to Ollama's **native** `/api/chat` with a `"format"` JSON schema, mirroring the request-building and retry/validation contract `chat.py` established (bounded retries 0–5, `RuntimeError` naming the model/attempt-count/truncated-response on exhaustion) but without pydantic-ai in the loop for this provider — there is no bare-metal native-JSON-schema mode in pydantic-ai's OpenAI-compatible model class to point at Ollama's native (non-OpenAI-shaped) endpoint, so this is a direct `_post_json` call, parsed with `schema.model_validate_json(...)`, retried on `pydantic.ValidationError` (which pydantic v2 also raises for malformed JSON, not just schema mismatches). `options` (from `OllamaOption*` nodes) is included directly in this same request's `"options"` field and is correctly honored, since it's the native endpoint — no separate priming call, because none is needed: structured output and context sizing now apply atomically in one request.
The retry/validation contract itself (bounded retries, `RuntimeError` naming the model/attempt-count/truncated-response on exhaustion) is unchanged and now implemented twice — once in `chat.py` for llama.cpp, once directly in `ollama_provider.py` for Ollama — rather than shared, which is the real cost of this decision (see Consequences).
## Consequences
**Easier:**
- Structured output is now reliable against "thinking"-capable models on both providers — reasoning and structured content are separate response fields (`message.reasoning`/`message.thinking` vs `message.content`) under both `NativeOutput` and Ollama's native `format`, rather than competing for the same token stream under tool-calling.
- `num_ctx` and other Ollama-native options now actually apply to structured-output requests — genuinely fixed this time, confirmed by re-running the actual failing workflow agent, not just by a passing test suite (the first fix attempt passed every test and still didn't work end-to-end).
- Meaningfully faster in the success case than tool-calling against a thinking model.
- Closes ADR-006's own explicit "revisit later" flag with concrete evidence rather than leaving it open indefinitely.
**Harder / constrained:**
- `OllamaProvider` and `LlamaCppProvider` now have two independent structured-output implementations instead of one shared one — ADR-007's "share one implementation" goal no longer holds for this piece. A future structured-output feature (e.g. plumbing reasoning content back to the caller) needs to land in both places.
- Structured output guarantees schema-*shape* validity, not semantic correctness, on both providers now — a genuinely broken model (degenerate tokenizer, as with the abliterated test model) can still return valid-JSON garbage instead of raising a clear error. Downstream consumers should not treat "returned without error" as "returned correct content" for low-quality/unreliable models.
- Ollama's native `/api/chat` endpoint's `"format"` field is only checked against top-level `type`/`properties`/`required` the same way the OpenAI-compat `response_format` was — no change to `_build_structured_model`'s shallow-schema behavior in `ollama.py`.
**Debt introduced:**
- Two structured-output code paths (per provider) instead of one shared one, as noted above — accepted because the two providers' actual constraints (Ollama's context-reset-per-OpenAI-compat-call bug vs. llama-server's fixed-at-launch context) are genuinely different, not incidentally different.
- Not addressed here (flagged for a future ADR if pursued): a model's reasoning/thinking content is available (`message.reasoning` natively for Ollama, parsed into pydantic-ai's `ThinkingPart` for llama.cpp) but discarded by both `chat_structured()` implementations, which only return the validated schema instance. Plumbing this back to `ChatCompletion` as a node output would need a `LLMProvider.chat_structured()` return-type change — a `Protocol`-level change affecting both providers, out of scope here.
- Not addressed here (pre-existing, unrelated): `LlamaCppProvider.chat()`'s own code comments already note that `OllamaOption*` nodes emit Ollama-native option names llama-server's OpenAI-compatible endpoint doesn't recognize — an accepted gap from the llama.cpp integration epic, unrelated to this decision.
## Considered Alternatives
### Alternative A: `NativeOutput` + priming call for both providers (the first fix attempt)
**Why rejected:** This is what ADR-009 originally shipped as. It passed every test (including a new one added specifically for the priming call) and one successful live single-agent ComfyUI run, but failed to actually fix the real workflow — confirmed by re-running the full pipeline and hitting the identical original failure on a later, more complex agent. Root-caused only after that: `/v1/chat/completions` unconditionally reloads the model at default context on *every* call, so priming immediately before the real call is undone by the real call itself. No amount of retrying, raising `max_tokens`, or raising `timeout_secs` fixes a context-size problem that the request itself keeps resetting.
### Alternative B: Fully switch both providers off `pydantic-ai`, hand-roll native structured output everywhere
**Why rejected:** Unnecessary for `LlamaCppProvider` — no evidence llama-server's OpenAI-compatible endpoint has Ollama's context-reset behavior (its context is fixed at process launch regardless of request), so `NativeOutput` over the existing shared path is strictly simpler there and keeps ADR-007's sharing goal intact for at least one provider.
### Alternative C: Do nothing, document Ollama's context-reset behavior as a known limitation
**Why rejected:** The underlying failure mode (indefinite-looking hangs, or outright failures, against any Ollama model needing more than the default 4096-token context while using structured output) is common enough — any long system prompt plus a non-trivial schema hits it — that documenting around it would leave `structured_output=True` effectively broken for Ollama in exactly the cases where structured output is most useful (complex, multi-field extraction tasks).
---
## Links
- Related ADRs: [ADR-006](ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md) (superseded rationale, not superseded status — ADR-006's tool-calling-vs-native evidence and reasoning stand as historical record; this ADR only revisits its "revisit later" flag), [ADR-007](ADR-007-llm-provider-adapter-pattern.md) (provider abstraction this decision partially steps outside of, for Ollama only)
- Discovered while building/testing: `workflows/ltx-i2v-pipeline.json`
@@ -0,0 +1,64 @@
# ADR-010: `"think"` as an options-carried, per-provider-translated toggle
## Status
> Accepted
_Date:_ 2026-07-25
_Deciders:_ darth-veitcher
---
## Context
ADR-009's investigation into structured-output reliability surfaced, as a side effect, how expensive a "thinking"-capable model's chain-of-thought reasoning is: on a real workflow, a single agent call could spend 5-7+ minutes and its entire token budget on reasoning before ever producing the requested response. Both Ollama and llama-server can turn this off, but neither exposes it through the generic `options` dict `ChatCompletion` already forwards — each has a completely different, incompatible wire shape:
- **Ollama** — live-tested directly: `"think": false` must be a **top-level** field on `/api/chat` (and `/v1/chat/completions`). Nested inside `options` (`{"options": {"think": false}}`) it's silently ignored — confirmed live (`eval_count: 223` reasoning tokens burned vs. `eval_count: 2` with it top-level). One existing test (`test_multi_turn_receives_context`) already carried `options={"think": False}` in its docstring's stated intent; it was a no-op the whole time.
- **llama.cpp** — not live-tested (no router-mode `llama-server` instance available; user explicitly chose doc-based research over waiting for one). Per `tools/server/README.md`: `chat_template_kwargs: {"enable_thinking": false}` (Qwen3-style HF chat-template convention) and/or `reasoning_effort: "none"` (a more model-agnostic OpenAI-style convention llama-server also accepts) — both as request-body fields on `/v1/chat/completions`, not nested in `options` either.
Two shapes were considered for exposing this from comfydv:
1. **A first-class `ChatCompletion` input + `LLMProvider` protocol parameter** — mirroring how `Message.images` crossed the provider boundary (ADR-008). Initially implemented this way.
2. **A composable `OllamaOption*`-style node merging a `"think"` key into the existing `OLLAMA_OPTIONS` chain**, with each provider popping that one key back out and translating it before building its own request — proposed as a simplification once (1) was drafted, since every other tunable knob already flows through this exact composition pattern and a new top-level node parameter would be the only one that doesn't.
## Decision
Went with option 2. `OllamaOptionDisableThinking` (`src/comfydv/ollama.py`) is a new node, identical in shape to `OllamaOptionTemperature`/`OllamaOptionSeed`/etc.: `disable_thinking: BOOLEAN` (default `True`), merges `{"think": not disable_thinking}` into whatever `OLLAMA_OPTIONS` chain it's wired into. No `ChatCompletion` or `LLMProvider` protocol signature change.
Both providers now start `chat()`/`chat_structured()` by popping `"think"` out of the incoming `options` dict (`_pop_think()`, `ollama_provider.py`, shared by both — a pure function, doesn't mutate the caller's dict) and translate it into their own shape before building the request:
- `OllamaProvider`: sets `payload["think"]` at the top level (both `chat()`'s native `/api/chat` call and `chat_structured()`'s, per ADR-009's native-endpoint rewrite).
- `LlamaCppProvider`: sets `chat_template_kwargs`/`reasoning_effort` — directly in its own hand-rolled `chat()` payload, and via `chat.py`'s `extra_body` (the same mechanism `options` itself uses) for `chat_structured()`, which still shares the pydantic-ai path per ADR-009.
Despite living in `ollama.py` and following the `OllamaOption*` naming convention (matching every other option node in that module, all genuinely Ollama-native and untranslated for llama.cpp — see `LlamaCppProvider.chat()`'s own comment), `"think"` is the one key from that chain **both** providers recognize and translate; it isn't itself Ollama's native wire format, it's a comfydv-level convention that happens to reuse Ollama's own field name since Ollama's is the more literal of the two backends' conventions.
## Consequences
**Easier:**
- One node works for both backends, reusing the exact composition pattern (`OLLAMA_OPTIONS` chaining into `ChatCompletion`'s `options` input) every other tunable parameter already uses — no new socket type, no `ChatCompletion.INPUT_TYPES` change, no `LLMProvider` protocol change.
- Fixes an existing test's stated-but-unfulfilled intent for free: `test_multi_turn_receives_context` and `test_structured_output_against_unreliable_model_stays_schema_valid` already passed `options={"think": False}` and now it actually works.
- Meaningfully faster for any thinking-capable model, and directly reduces the token-budget pressure ADR-009 had to fix around.
**Harder / constrained:**
- The llama.cpp translation is not live-verified — sourced from the server's documented request-body fields, not confirmed against a running `llama-server`. Verify against your own deployment before relying on it; a follow-up should close this gap once an instance is available (the user explicitly chose this tradeoff over waiting).
- `"think"` living among genuinely-Ollama-native `OllamaOption*` nodes (which llama.cpp does *not* translate — see that class's own code comment) is a small naming/mental-model inconsistency: one key out of that whole chain is special-cased by both providers. Documented here and in `_pop_think()`'s own docstring so it doesn't read as an oversight later.
**Debt introduced:**
- None. No new dependency, no new socket type.
## Considered Alternatives
### Alternative A: First-class `ChatCompletion` input + protocol parameter (mirroring `Message.images`, ADR-008)
**Why rejected:** Correct in principle (this is a cross-provider concern needing real translation, exactly like images), but heavier than necessary — a new node input plus a `LLMProvider.chat()`/`chat_structured()` signature change plus threading a new parameter through every call site, when the existing `options` dict composition already has a clean seam (`_pop_think`) for a value that needs per-provider translation before hitting the wire. Started implementing this way; reverted once the composable-option alternative was raised.
### Alternative B: Separate provider-specific nodes (`OllamaOptionDisableThinking` / a llama.cpp-only equivalent)
**Why rejected:** Splits one concept into two nodes for no real benefit — both backends' translation lives in code either way, so there's no cost to having one node recognize the same key on both.
---
## Links
- Related ADRs: [ADR-007](ADR-007-llm-provider-adapter-pattern.md) (the `LLMProvider` boundary this operates within), [ADR-008](ADR-008-multimodal-image-input-across-llmprovider-boundary.md) (the pattern this ADR considered and didn't need — cross-provider concerns don't always require a protocol change), [ADR-009](ADR-009-native-structured-output-mode.md) (the investigation that surfaced how expensive unmanaged thinking is)
- llama.cpp server docs (request-body fields, not live-verified): `tools/server/README.md` in `ggml-org/llama.cpp`
+1
View File
@@ -31,3 +31,4 @@ Superseded ADRs keep their file; update their status to `Superseded by ADR-###`.
| ADR | Title | Status | Date |
|-----|-------|--------|------|
| [ADR-000](ADR-000-template.md) | Template | — | — |
| [ADR-008](ADR-008-multimodal-image-input-across-llmprovider-boundary.md) | Multimodal image input across the LLMProvider boundary | Proposed | 2026-07-22 |
+4
View File
@@ -26,6 +26,9 @@
- **BEACON bootstrap** — `epics/archive/beacon-bootstrap.md` — ✅ DONE — BEACON framework wired up; problem statement, constitution, roadmap, and ADR template populated; quality gates clean
- **Logging modernisation** — `epics/logging-modernisation.md` — ✅ DONE — stdlib logging, NullHandler, silent-by-default; colorama/rich/termcolor removed; 11 tests
- **ComfyUI UX Polish & Manager Compatibility** — `epics/ux-and-install.md` — 🔄 ACTIVE — Fix installation, core UX bugs (debounce, connection drops, alert dialogs), correctness bugs (class-level mutation, IS_CHANGED, seed=0), and metadata drift
- **LLM Provider Abstraction** — `epics/archive/llm-provider-abstraction.md` — ✅ DONE — shared `LLMProvider` protocol (list/load/unload/chat/structured-output) and generic ComfyUI nodes; Ollama integration migrated onto it (ADR-007, supersedes ADR-006); merged via PR #17
- **llama.cpp Model Integration** — `epics/llamacpp-integration.md` — 🔄 ACTIVE — Add a `LlamaCppProvider` implementing the shared protocol via llama-server's router mode (GitHub issue #15); dependency on LLM Provider Abstraction now satisfied
- **VLM Image Input for ChatCompletion** — `epics/vlm-image-input.md` — 📋 PLANNING — Wire a ComfyUI IMAGE into the existing generic ChatCompletion node so a vision-capable model can describe/understand images; images carried on the `Message` and translated per-provider (ADR-008 extends ADR-007)
For the live rollup (specs per epic, % tasks complete, last-commit age):
@@ -40,6 +43,7 @@ beacon epic list --detailed
- **BEACON bootstrap** is a prerequisite for all other epics (quality gates need to pass before new work merges)
- **Test hardening** is independent of documentation and can run in parallel
- **Documentation** depends on the final node API (output order, input names) being stable — start after test hardening locks the contracts
- **llama.cpp Model Integration** depends on **LLM Provider Abstraction** landing first — its `LlamaCppProvider` implements the protocol that epic defines, and reuses its generic nodes as-is
---
@@ -0,0 +1,119 @@
# Epic: LLM Provider Abstraction
## Status
Done — completed 2026-07-11
## Why now
GitHub issue #15 asks for llama.cpp support "similar to Ollama," and
explicitly raises the question of separate nodes vs. an adapter pattern.
[ADR-007](../../ADRs/ADR-007-llm-provider-adapter-pattern.md) answers that
with a real adapter: a `LLMProvider` protocol (`list_models`/`load_model`/
`unload_model`/`chat`/`chat_structured`) that any backend implements, backing
a set of generic ComfyUI nodes (`ChatCompletion`, `LLMModelSelector`,
`LLMLoadModel`, `LLMUnloadModel`) that work with whichever provider is wired
in. For that to be real — not just aspirational — the existing Ollama
integration has to migrate onto the protocol first, including moving its
structured-output mechanism from ADR-006's hand-rolled tool-calling onto
`pydantic-ai`. This epic is that migration; it's a prerequisite for
`llamacpp-integration`, not additive scope on top of it — building
llama.cpp against generic nodes that don't exist yet isn't possible.
## Dependencies
_None to start — this epic can begin immediately._ The
`llamacpp-integration` epic depends on this one landing first: its
`LlamaCppProvider` implements the protocol this epic defines, and its
`LlamaCppClient` node emits the same `LLM_CLIENT` socket type this epic
introduces.
## Specs
_Filled by `beacon specify --epic llm-provider-abstraction` /
`/speckit-specify` once this epic is accepted._
- specs/007-llm-provider-abstraction/
## ADRs
- project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md — defines the `LLMProvider` protocol as the adapter boundary; supersedes ADR-006; narrows ADR-004's scope to non-chat REST calls per-provider
## Success criteria
- `LLMProvider` protocol defined (new internal module, e.g. `src/comfydv/_llm/provider.py`): `list_models()`, `load_model(name)`, `unload_model(name)`, `chat(...)`, `chat_structured(..., schema)`
- `ModelStatus` enum defined (`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`) per ADR-007's documented approximation
- `OllamaProvider` implements the protocol, wrapping all existing Ollama REST logic (`/api/tags`, `/api/generate` with `keep_alive`) over `aiohttp` — behavior-preserving port of the current `_post_json`/`_fetch_models` logic, not a rewrite of the underlying calls
- `chat_structured()` implemented via `pydantic-ai`'s `Agent`/`output_type`, called through `OpenAIProvider(base_url=<host>/v1)`, using the same dynamic `pydantic.create_model()`-from-JSON-Schema pattern as ADR-006 — same dynamic-socket UX, same retry/validation contract (bounded `max_retries`, required-string-non-empty check, clear `RuntimeError` on exhaustion)
- `pydantic-ai` and `openai` added to `pyproject.toml`, curated into `requirements.txt` per [ADR-003](../../ADRs/ADR-003-requirements-txt-authoring-policy.md)
- Generic ComfyUI nodes replace the current Ollama-specific ones: `OllamaClient` now outputs `LLM_CLIENT` (constructing an `OllamaProvider` internally); `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`, `ChatCompletion` operate against `LLM_CLIENT` generically
- The non-structured-output chat path (native `/api/chat`) is behavior-unchanged
- All existing `tests/test_ollama.py` coverage passes against the migrated implementation (adjusted for renamed node/socket types, unchanged in behavior otherwise)
- Model-management calls remain entirely on `aiohttp`, inside `OllamaProvider` — untouched transport-wise
- CI smoke test passes
## Non-goals
- No `LlamaCppProvider` or llama.cpp nodes in this epic — that's `llamacpp-integration`
- No behavior change to non-structured-output chat
- No tracing/observability integration (e.g. Logfire), even though `pydantic-ai` supports it
- No multi-turn agentic tool use beyond the existing single structured-output call
- No backward-compat aliases for the old `Ollama`-prefixed node/socket names — confirmed 2026-07-11 to rename in place (see ADR-007)
## Notes
This is the riskiest part of the whole llama.cpp proposal: it changes the
implementation of tested, shipped code from a Done epic
(`archive/ollama-integration.md`), not just adding new code, and it's a
breaking rename (`OLLAMA_CLIENT`→`LLM_CLIENT`,
`OllamaChatCompletion`→`ChatCompletion`, etc.) for anyone with saved
workflows using the current node/socket names. Confirmed 2026-07-11 (per
ADR-007): rename in place now — the Ollama integration only shipped
2026-07-04, so the blast radius is small — rather than carrying
`Ollama`-prefixed generic nodes forward or maintaining deprecated aliases
indefinitely.
Recommend an adversarial pass (`/beacon:review` or `/beacon:engineering`)
before merging, specifically checking that the retry/validation contract
from ADR-006's `## Decision` section is preserved exactly by the
`pydantic-ai` reimplementation, and that `OllamaProvider`'s REST calls are a
faithful port of the current `_post_json`/`_fetch_models` logic.
**2026-07-11 — mid-build correction, tracked in [issue #16](https://github.com/darth-veitcher/comfydv/issues/16):**
the Foundational layer (`LLMProvider` protocol + `OllamaProvider` skeleton)
shipped safely, but the planned per-user-story incremental cutover doesn't
hold — `OllamaClient` is a single shared producer for every downstream
Ollama node, so the node-layer rename/cutover (`tasks.md`'s US1 + US3) must
land as one atomic change, not four independent ones. Confirmed by
independent product + engineering review. Re-scoped as its own dedicated
follow-up BUILD session — see `specs/007-llm-provider-abstraction/tasks.md`'s
correction note for full detail. Open question for the next session: does
this take priority over `ux-and-install` (active, 1/4 specs shipped), since
llama.cpp (issue #15) has no deadline.
**2026-07-11 — properly specced:** the deferred cutover is now fully
inventoried and planned in
`specs/007-llm-provider-abstraction/atomic-cutover-plan.md` — a full
line-by-line read of the ~125 affected references (not an estimate), the
design decisions it surfaced (cache-singleton duplication, `client ==
"<string>"` equality breaking, bare-string-client backward compat removal,
and a test-layer split so the 35 `_post_json` monkeypatches land at the
right architectural seam), and a 12-step sequenced task list (T-CUT-01 …
T-CUT-12). This is now the authoritative implementation plan for the
cutover — the next BUILD session executes it directly rather than
re-deriving the approach.
`pydantic-ai`'s `StructuredDict` (raw-JSON-Schema output, no Python class)
was considered as a lighter-weight alternative to `create_model()` during
research and rejected: it performs no pydantic validation at all, which
would silently drop the "reject blank required strings" safeguard ADR-006
introduced. Stick with `create_model()`-built `BaseModel` subclasses.
**2026-07-11 — cutover executed, PR open:** T-CUT-01 through T-CUT-12
complete (`specs/007-llm-provider-abstraction/tasks.md` — every task `[x]`
or explicitly `[-]` superseded/deferred with a reason). `beacon epic
refresh` reports 1/1 owned specs complete. An independent `beacon-reviewer`
pass caught one real regression (`options` silently dropped in
structured-output mode) before merge — fixed and re-verified clear. README
and docs/index.md updated for the rename, including the FR-009 migration
table. **PR: [#17](https://github.com/darth-veitcher/comfydv/pull/17)** —
not yet merged; `beacon epic finish` waits for that, per
`beacon epic refresh`'s own guidance.
@@ -1,7 +1,7 @@
# Epic: Ollama Model Integration
## Status
Planning — started 2026-06-28
Done — completed 2026-07-04
## Why now
@@ -0,0 +1,69 @@
# Epic: llama.cpp Model Integration
## Status
Active — spec 008-llamacpp-integration complete, ready to finish once merged
## Why now
GitHub issue #15 requests llama.cpp support "similar to Ollama." This is now
practical because llama.cpp's `llama-server` gained a "router mode" via
[llama.cpp PR #18228](https://github.com/ggml-org/llama.cpp/pull/18228)
(merged 2025-12-21): launched with `--models-dir <dir>` (or
`--models-preset <file>.ini`) instead of `-m`, it exposes `GET /models`
(with live status: `unloaded`/`loading`/`loaded`/`sleeping`/`downloading`),
`POST /models/load`, `POST /models/unload`, and `--sleep-idle-seconds`
auto-unload — giving llama.cpp the same manual load/unload
memory-management primitives comfydv already relies on for Ollama. With the
`llm-provider-abstraction` epic in place, adding llama.cpp is now a matter
of implementing one more `LLMProvider`, not building a parallel set of
ComfyUI nodes.
## Dependencies
Depended on the `llm-provider-abstraction` epic landing first — **satisfied**
2026-07-11, merged via [PR #17](https://github.com/darth-veitcher/comfydv/pull/17)
(`project-management/Roadmap/epics/archive/llm-provider-abstraction.md`).
This epic's `LlamaCppProvider` implements the `LLMProvider` protocol that
epic defined, and its `LlamaCppClient` node emits the same `LLM_CLIENT`
socket type the generic `ChatCompletion`/`LLMModelSelector`/`LLMLoadModel`/
`LLMUnloadModel` nodes already consume — none of those node classes are
touched by this epic.
## Specs
_Filled by `beacon specify --epic llamacpp-integration` / `/speckit-specify`
once this epic is accepted._
- specs/008-llamacpp-integration/
## ADRs
- project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md — decided during the prerequisite epic; this epic implements the second `LLMProvider` the ADR anticipated
## Success criteria
- `LlamaCppProvider` implements the `LLMProvider` protocol from the prerequisite epic:
- `list_models()` via `GET /models`, surfacing native status (`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`) directly — no normalization needed, since llama.cpp's vocabulary is the `ModelStatus` enum's superset
- `load_model()` / `unload_model()` via `POST /models/load` / `POST /models/unload`
- `chat_structured()` via the same shared `pydantic-ai` mechanism as `OllamaProvider`, `OpenAIProvider(base_url=<llama-server host>/v1)` — no new structured-output code, just a different `base_url`
- `LlamaCppClient` config node (reuses the [ADR-005](../../ADRs/ADR-005-ollama-host-config-via-client-node.md) config-node pattern), outputs the same `LLM_CLIENT` socket type `OllamaClient` does
- No new node classes for model selection, load/unload, or chat — the generic `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`, and `ChatCompletion` nodes from the prerequisite epic work unchanged once a `LlamaCppClient` is wired in
- `LlamaCppClient` registered in `NODE_CLASS_MAPPINGS` / `NODE_DISPLAY_NAME_MAPPINGS`
- Test coverage for `LlamaCppProvider` mirrors the `OllamaProvider` test conventions established in the prerequisite epic
- No new runtime dependencies beyond what the prerequisite epic already introduced (`aiohttp` for model management, `pydantic-ai`/`openai` for chat)
- CI smoke test passes
## Non-goals
- No support for llama-server's non-router single-model launch mode (`-m`) — router mode only, since that's what gives load/unload parity with Ollama
- No GPU inference optimisation or quantisation tuning — CPU-first dev harness, consistent with the Ollama epic's own non-goal
- No auth/TLS/remote-serving hardening — localhost/configurable host via client node only, consistent with the Ollama epic
- No ComfyUI Manager registry listing in this epic
- No changes to the generic nodes or `LLMProvider` protocol themselves — if llama.cpp's router mode needs something the protocol doesn't support, that's a protocol change scoped back into the prerequisite epic's follow-up, not silently special-cased here
## Notes
Router mode is a deployment prerequisite, not something comfydv configures:
the user must launch `llama-server` with `--models-dir`/`--models-preset`
themselves. Document this clearly in the eventual spec/node tooltips.
Reference: [llama.cpp PR #18228](https://github.com/ggml-org/llama.cpp/pull/18228), GitHub issue #15.
@@ -0,0 +1,61 @@
# Epic: VLM Image Input for ChatCompletion
## Status
Planning — started 2026-07-22
## Why now
The generic `ChatCompletion` node and the `LLMProvider` protocol landed
text-only (ADR-007), which explicitly deferred any protocol change for a new
capability as "its own follow-up." Both shipped backends can already serve
vision models — Ollama multimodal models via `/api/chat`'s per-message
`images`, and llama.cpp via a multimodal projector (`mmproj`) on the same
OpenAI-compatible `/v1/chat/completions` the provider already calls — so the
gap is entirely on comfydv's client side, not the servers'. Coupling the chat
node with a VLM to describe or reason about images produced elsewhere in a
workflow is a frequently-wanted next step, and the adapter pattern makes it a
small, symmetric addition rather than a new node family.
## Specs
_SpecKit specs that contribute to this epic._
- specs/009-vlm-image-input/ — wire a ComfyUI IMAGE into the existing ChatCompletion node; images carried on the Message and translated per-provider
## ADRs
_Cross-cutting decisions this epic required._
- project-management/ADRs/ADR-008-multimodal-image-input-across-llmprovider-boundary.md — carry images as an optional `Message.images` field; each provider translates to its own wire shape (extends ADR-007's adapter pattern to a second input modality)
- project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md — the adapter pattern this epic extends; the generic node/protocol it adds an image path to
## Success criteria
- `Message` carries an optional `images` field; text-only requests remain byte-for-byte unchanged (existing Ollama + llama.cpp provider tests stay green)
- `ChatCompletion` gains one **optional** `IMAGE` input — no new node classes, no new socket types; a workflow author gains vision by wiring one socket
- A wired image reaches a vision model and produces a description/answer on **both** backends:
- Ollama: flat per-message `images` passes through `/api/chat` untransformed
- llama.cpp: mapped to OpenAI `image_url` content parts on `/v1/chat/completions`
- Structured output with an image works via the shared `chat_structured()` (pydantic-ai multimodal content) — one implementation, both backends
- A text-only model that receives an image degrades to a clear backend error surfaced by the node, not a crash
- Test coverage mirrors the `OllamaProvider`/`LlamaCppProvider` conventions (mock at the provider's own transport seam); CI smoke test passes
- No new runtime dependencies beyond what ADR-007 already introduced
## Non-goals
- No image **output** or image generation — input-to-VLM only
- No new node classes or socket types — additive optional input on the existing generic node
- No changes to the client/config nodes (`OllamaClient`, `LlamaCppClient`) or model-management nodes
- No auto-provisioning of vision models — the user must have a multimodal model loaded (Ollama multimodal model; llama.cpp launched with an `mmproj` projector); this epic does not install or configure it
- No video, audio, or document/PDF modalities — still images only
- No image preprocessing beyond what's needed to hand a ComfyUI IMAGE tensor to a backend (no resizing policy, tiling, or OCR of our own)
- No `OllamaOption*` parameter translation work — inherited unchanged from ADR-007's scope
## Notes
Multimodal readiness is a deployment prerequisite, not something comfydv
configures: document in node tooltips that the wired model must be
vision-capable, and that llama.cpp needs `--mmproj`. The exact wire shapes
(Ollama `/api/chat` `images`, OpenAI `image_url`, pydantic-ai `BinaryContent`)
are verified live in the spec's `research.md`/`plan.md`, consistent with how
the llama.cpp epic verified router-mode endpoints.
Reference: ADR-008, ADR-007, GitHub issue #15.
File diff suppressed because it is too large Load Diff
+8 -2
View File
@@ -1,6 +1,6 @@
[project]
name = "comfydv"
version = "0.1.0"
dynamic = ["version"]
requires-python = ">=3.11"
description = "Quality of life ComfyUI nodes: dynamic string formatting, seed-controlled random selection, and conditional queue interruption."
readme = "README.md"
@@ -10,6 +10,8 @@ authors = [
dependencies = [
"aiohttp>=3.9.0",
"jinja2>=3.1.6",
"pydantic>=2.0",
"pydantic-ai-slim[openai]>=2.9.0",
]
[project.urls]
@@ -17,11 +19,15 @@ Repository = "https://github.com/darth-veitcher/comfydv"
Documentation = "https://darth-veitcher.github.io/comfydv/stable/"
[build-system]
requires = ["hatchling"]
requires = ["hatchling", "hatch-vcs"]
build-backend = "hatchling.build"
[tool.hatch.version]
source = "vcs"
[dependency-groups]
dev = [
"pillow>=10.0.0",
"playwright>=1.60.0",
"pytest>=8.4.2",
"pytest-cov>=6.0.0",
+2
View File
@@ -2,3 +2,5 @@
# Lists only deps NOT already provided by ComfyUI's own environment.
# See ADR-003: never auto-generate this file with `uv export`.
jinja2>=3.1.6
pydantic>=2.0
pydantic-ai-slim[openai]>=2.9.0
+154 -11
View File
@@ -117,7 +117,11 @@ async def _capture(
pad: int = 60,
) -> None:
"""Screenshot the canvas, optionally cropped tightly around the node."""
canvas = page.locator("canvas#graph-canvas, canvas").first
# ComfyUI's frontend added a minimap canvas since this script was last
# verified — it appears before #graph-canvas in DOM order, so a bare
# union selector's `.first` silently grabbed the 250x200 minimap
# instead of the real graph. Target #graph-canvas explicitly.
canvas = page.locator("canvas#graph-canvas")
box = await canvas.bounding_box()
if box is None:
await page.screenshot(path=str(out))
@@ -307,13 +311,13 @@ async def scene_ollama_client(page: Page, out: Path) -> None:
async def scene_ollama_chat(page: Page, out: Path) -> None:
"""OllamaChatCompletion — showing the live model dropdown and prompt widget."""
"""ChatCompletion — showing the live model dropdown and prompt widget."""
await _clear(page)
info = await page.evaluate(
"""
async () => {
const node = LiteGraph.createNode("OllamaChatCompletion");
const node = LiteGraph.createNode("ChatCompletion");
node.pos = [60, 60];
window.app.graph.add(node);
@@ -327,7 +331,7 @@ async def scene_ollama_chat(page: Page, out: Path) -> None:
} else {
// Manually fetch and populate the COMBO
try {
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
if (resp.ok) {
const data = await resp.json();
const models = data.models || [];
@@ -357,8 +361,142 @@ async def scene_ollama_chat(page: Page, out: Path) -> None:
await _capture(page, out, info["pos"], info["size"])
async def scene_structured_output(page: Page, out: Path) -> None:
"""ChatCompletion with structured_output enabled — the live dynamic
output sockets (summary/sentiment/confidence) that appear as soon as
output_schema is edited, no graph run required. Exercises the real
js/ollama.js widget callbacks (structuredWidget.callback /
schemaWidget.callback), the same code path a live user's checkbox click
and schema edit trigger — not a hand-simulated approximation."""
await _clear(page)
schema = (
'{"type": "object", "properties": '
'{"summary": {"type": "string"}, '
'"sentiment": {"type": "string"}, '
'"confidence": {"type": "number"}}, '
'"required": ["summary", "sentiment", "confidence"]}'
)
info = await page.evaluate(
f"""
async () => {{
const node = LiteGraph.createNode("ChatCompletion");
node.pos = [60, 60];
window.app.graph.add(node);
const promptWidget = node.widgets.find(w => w.name === "prompt");
if (promptWidget) promptWidget.value =
"The new render pipeline cut our export time in half and the team is thrilled.";
const structuredWidget = node.widgets.find(w => w.name === "structured_output");
const schemaWidget = node.widgets.find(w => w.name === "output_schema");
if (structuredWidget) structuredWidget.value = true;
if (schemaWidget) schemaWidget.value = {json.dumps(schema)};
// Fire the same callbacks js/ollama.js attaches on node creation —
// real widget-edit code path, not a re-implementation.
if (structuredWidget?.callback) await structuredWidget.callback(true);
if (schemaWidget?.callback) await schemaWidget.callback({json.dumps(schema)});
await new Promise(r => setTimeout(r, 600));
window.app.canvas.setDirty(true, true);
window.app.canvas.draw(true, true);
return {{ pos: [node.pos[0], node.pos[1]], size: [node.size[0], node.size[1]] }};
}}
"""
)
await asyncio.sleep(0.8)
await _redraw(page)
await _frame_node(page, info["pos"], info["size"])
await _capture(page, out, info["pos"], info["size"])
async def scene_llamacpp_client(page: Page, out: Path) -> None:
"""LlamaCppClient — single node showing the router-mode host URL widget."""
await _clear(page)
info = await page.evaluate(
"""
() => {
const node = LiteGraph.createNode("LlamaCppClient");
node.pos = [60, 60];
window.app.graph.add(node);
const hostWidget = node.widgets && node.widgets.find(w => w.name === "host");
if (hostWidget) hostWidget.value = "http://localhost:8080";
window.app.canvas.setDirty(true, true);
window.app.canvas.draw(true, true);
return { pos: [node.pos[0], node.pos[1]], size: [node.size[0], node.size[1]] };
}
"""
)
await _frame_node(page, info["pos"], info["size"])
await _capture(page, out, info["pos"], info["size"])
async def scene_llamacpp_workflow(page: Page, out: Path) -> None:
"""LlamaCppClient → the same ChatCompletion node the Ollama workflow
uses, unmodified — the actual point of the adapter pattern. Attempts a
live model-list refresh via backend=llamacpp against
host.docker.internal:8080; degrades gracefully (same as a real user's
"no server running yet" state) if nothing is listening there."""
await _clear(page)
info = await page.evaluate(
"""
async () => {
const graph = window.app.graph;
const client = LiteGraph.createNode("LlamaCppClient");
client.pos = [40, 60];
graph.add(client);
const hostWidget = client.widgets.find(w => w.name === "host");
if (hostWidget) hostWidget.value = "http://host.docker.internal:8080";
const chat = LiteGraph.createNode("ChatCompletion");
chat.pos = [380, 40];
graph.add(chat);
const promptWidget = chat.widgets.find(w => w.name === "prompt");
if (promptWidget) promptWidget.value = "Write a haiku about ComfyUI.";
client.connect(0, chat, 0);
try {
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:8080&backend=llamacpp");
if (resp.ok) {
const data = await resp.json();
const models = data.models || [];
if (models.length) {
const modelWidget = chat.widgets.find(w => w.name === "model");
if (modelWidget) modelWidget.value = models[0];
}
}
} catch (e) {}
await new Promise(r => setTimeout(r, 800));
window.app.canvas.setDirty(true, true);
window.app.canvas.draw(true, true);
const nodes = [client, chat];
const minX = Math.min(...nodes.map(n => n.pos[0])) - 20;
const minY = Math.min(...nodes.map(n => n.pos[1])) - 20;
const maxX = Math.max(...nodes.map(n => n.pos[0] + n.size[0])) + 20;
const maxY = Math.max(...nodes.map(n => n.pos[1] + n.size[1])) + 20;
return { pos: [minX, minY], size: [maxX - minX, maxY - minY] };
}
"""
)
await asyncio.sleep(1.0)
await _redraw(page)
await _frame_node(page, info["pos"], info["size"], scale=1.0)
await _capture(page, out, info["pos"], info["size"], scale=1.0)
async def scene_ollama_workflow(page: Page, out: Path) -> None:
"""Full mini-workflow: OllamaClient → OllamaChatCompletion + Temperature + Seed options."""
"""Full mini-workflow: OllamaClient → ChatCompletion + Temperature + Seed options."""
await _clear(page)
info = await page.evaluate(
@@ -386,7 +524,7 @@ async def scene_ollama_workflow(page: Page, out: Path) -> None:
if (twSeed) twSeed.value = 42;
// 4. ChatCompletion — right
const chat = LiteGraph.createNode("OllamaChatCompletion");
const chat = LiteGraph.createNode("ChatCompletion");
chat.pos = [380, 100];
graph.add(chat);
const twPrompt = chat.widgets && chat.widgets.find(w => w.name === "prompt");
@@ -409,7 +547,7 @@ async def scene_ollama_workflow(page: Page, out: Path) -> None:
// Refresh model dropdowns for chat node
try {
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
if (resp.ok) {
const data = await resp.json();
const models = data.models || [];
@@ -466,25 +604,25 @@ async def scene_ollama_lifecycle(page: Page, out: Path) -> None:
if (hostWidget) hostWidget.value = "http://localhost:11434";
// OllamaLoadModel — generous gap right of client
const load = LiteGraph.createNode("OllamaLoadModel");
const load = LiteGraph.createNode("LLMLoadModel");
load.pos = [380, 80];
graph.add(load);
// OllamaChatCompletion — wide node, plenty of space to the right of load
const chat = LiteGraph.createNode("OllamaChatCompletion");
const chat = LiteGraph.createNode("ChatCompletion");
chat.pos = [720, 40];
graph.add(chat);
const twPrompt = chat.widgets && chat.widgets.find(w => w.name === "prompt");
if (twPrompt) twPrompt.value = "Describe this image in one sentence.";
// OllamaUnloadModel — far right, vertically offset to match chat's outputs
const unload = LiteGraph.createNode("OllamaUnloadModel");
const unload = LiteGraph.createNode("LLMUnloadModel");
unload.pos = [1200, 280];
graph.add(unload);
// Populate model dropdowns from live Ollama
try {
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434");
const resp = await fetch("/dv/ollama/models?host=http://host.docker.internal:11434&backend=ollama");
if (resp.ok) {
const data = await resp.json();
const models = data.models || [];
@@ -604,6 +742,11 @@ SCENES = [
("ollama_workflow.png", scene_ollama_workflow),
("ollama_options.png", scene_ollama_options),
("ollama_lifecycle.png", scene_ollama_lifecycle),
# Structured output (ADR-007 / pydantic-ai)
("structured_output.png", scene_structured_output),
# llama.cpp (spec 008)
("llamacpp_client.png", scene_llamacpp_client),
("llamacpp_workflow.png", scene_llamacpp_workflow),
]
@@ -0,0 +1 @@
epic = "llm-provider-abstraction"
@@ -0,0 +1,178 @@
# Atomic Cutover Plan: Ollama Node Rename → Generic LLM Nodes
Companion to `tasks.md`'s "⚠️ Correction (2026-07-11)" section and
[issue #16](https://github.com/darth-veitcher/comfydv/issues/16). This is
the properly-specced version of that deferred work — based on a full,
line-by-line inventory of every one of the ~125 affected references in
`src/comfydv/ollama.py` and `tests/test_ollama.py` (1820 lines, read in
full), not an estimate.
The inventory surfaced that this isn't one mechanical find-and-replace —
several genuine design decisions were implicit in "rename it" and needed
resolving before any code changes. Those decisions are below, followed by
the sequenced task list that implements them.
## Resolved decisions
### D1 — Single source of truth for HTTP/cache infra
`comfydv._llm.ollama_provider` owns `_post_json`, `_run_async`,
`_fetch_models`, `_TTLLRUCache`, `_cache_key`, `_MODEL_LIST_CACHE` (already
ported there — see `ollama_provider.py`). `ollama.py` stops defining its
own copies of these. The one remaining non-node use case —
`_load_default_models()` (combo-widget population at import time, before
any `OllamaClient` node exists) and the `/dv/ollama/models` refresh route —
imports `_fetch_models`/`_run_async` from `comfydv._llm.ollama_provider`
instead of duplicating them. `_CHAT_RESPONSE_CACHE` and `_post_json`
disappear from `ollama.py` entirely — nothing there needs them once
load/unload/chat delegate to `client.*`.
**Why not keep two copies:** they were already flagged as byte-identical by
the inventory (§2.13) — the only reason `ollama.py` still has them is that
nothing has repointed the imports yet. Keeping a second copy "just in case"
is exactly the kind of duplication ADR-007 exists to eliminate.
### D2 — `client == "<host string>"` equality is removed, not preserved
`OllamaProvider` is a plain object, not a `str` subclass — this is a
deliberate consequence of the adapter boundary (ADR-007), not an oversight
to work around. Two tests assert string equality on `client`
(`test_client_outputs_ollama_client_type:208-209`,
`test_client_carries_headers:1233-1234`) — both get rewritten to assert
`client.host == "..."` and `isinstance(client, OllamaProvider)`.
`OllamaClientType` (the `str`-subclass, `ollama.py:97-109`) is left in
place but becomes unused by `OllamaClient.create_client()` — not deleted in
this cutover (no test depends on deleting it, and removing a class nothing
references is a separate, lower-risk cleanup, not part of this bullet's
scope).
### D3 — Bare-string `client` backward compatibility is removed
`test_plain_string_client_has_no_headers` (`test_ollama.py:1328-1344`)
documents and tests that wiring a plain `STRING` node directly into
`client` (skipping `OllamaClient` entirely) silently works, because
`f"{client}/..."` succeeds on any string. This was never a documented,
intended feature — it's a side effect of `OllamaClientType` being a `str`
subclass, not mentioned in ADR-005 or the original Ollama epic's spec. Once
node methods call `client.chat(...)`/`client.load_model(...)`, a bare
string raises `AttributeError`. **This test is deleted**, not rewritten —
its premise (bare-string clients are supported) is being intentionally
removed, and asserting the new failure mode would just be testing that
Python raises `AttributeError` on missing methods, which isn't
comfydv-specific behavior worth a test.
The same bare-string pattern appears incidentally in ~8 other tests
(`ollama_host` fixture returns a plain string, used as `client=ollama_host`
in several integration tests — `test_ollama.py:220, 386, 407, 588, ...`).
These need `client=OllamaClient().create_client(ollama_host)[0]` instead of
`client=ollama_host` — a required edit, not optional, since they'll raise
`AttributeError` otherwise. See task T-CUT-08 below.
### D4 — Test layer split (resolves all 35 relocated `_post_json` monkeypatches)
This is the biggest structural decision. Today, `test_ollama.py` tests
ComfyUI node behavior by mocking `aiohttp` at the `ollama_mod._post_json`
seam and asserting on the exact Ollama wire payload (`keep_alive`,
`/api/generate`, tool-calling JSON shape, retry counts) *through* the node.
That seam moves — nodes no longer call `_post_json` directly, they call
`client.chat(...)` etc. Two options: (a) keep patching at whatever the new
seam is, 1:1 per test, or (b) recognize this is an architectural boundary
and split coverage accordingly. Going with **(b)**:
- **`tests/test_ollama.py`** — ComfyUI node **contract + delegation** only.
A new `_FakeProvider` test double (implements `list_models`/
`load_model`/`unload_model`/`chat`/`chat_structured`, records calls made
to it) stands in for `client`. Tests assert: right method called, right
arguments passed, return value flows through to the node's output tuple
correctly. **No `aiohttp`/`_post_json` mocking at this layer anymore.**
This directly matches the protocol contract's own rule ("generic nodes
MUST NOT branch on which concrete provider type they received") — if the
node tests don't need to know Ollama's wire format, they shouldn't mock
it either.
- **`tests/test_ollama_provider.py`** (new file) — `OllamaProvider`'s
actual Ollama-wire-protocol behavior: `/api/generate`+`keep_alive` int
shape, `/api/tags` parsing into `ModelInfo`, header/timeout forwarding,
response caching, cache-key composition. This is where the *substance* of
today's 35 `_post_json` monkeypatches lands — not 1:1, since several
collapse or move (see D5).
- **`tests/test_llm_chat_structured.py`** (exists, unchanged) — already
covers the shared retry/validation/error-contract mechanism.
### D5 — Structured-output retry tests are not ported 1:1
`TestStructuredOutput` has 15 tests monkeypatching `_post_json` to assert
exact retry-count behavior (`test_retries_on_invalid_json_then_succeeds`,
`test_exhausts_retries_raises_runtime_error`, `test_max_retries_clamped_*`,
etc.). Once `OllamaProvider.chat_structured()` delegates to the already-
tested shared `chat_structured()` helper (`src/comfydv/_llm/chat.py`,
covered by `tests/test_llm_chat_structured.py`'s 6 tests), re-asserting
retry counts at the Ollama-node layer duplicates that coverage without
adding confidence. Replaced with:
- A handful of `test_ollama.py` delegation tests: `ChatCompletion` with
`structured_output=True` calls `client.chat_structured(model, messages,
schema, ...)` with the right schema/model/messages.
- One `test_ollama_provider.py` test: `OllamaProvider.chat_structured()`
builds `base_url=f"{self.host}/v1"` and forwards to the shared helper
with the right arguments.
- `test_structured_output_true_sends_tool_call_payload` (asserts the exact
`tools`/`tool_choice` JSON shape) is **deleted** — that's `pydantic-ai`'s
internal tool-calling mechanism now, not comfydv's; asserting on a
third-party library's internals isn't a test worth keeping.
- Pure schema-parsing tests that don't touch HTTP at all (`_parse_output_schema`
fail-fast checks, `_coerce_structured_value`, dynamic-socket
`RETURN_TYPES` mutation) are unaffected — they test code that stays in
`ollama.py` unchanged, only need the class-name rename.
### D6 — `_client_headers()` is deleted
Dead code once `OllamaLoadModel`/`OllamaUnloadModel`/`OllamaChatCompletion`
delegate to `client.*` (headers become internal to `OllamaProvider`,
captured once at construction). No test calls it directly.
### D7 — Contract doc gets a small fix
`contracts/llm_provider_protocol.md`'s illustrative code sample is missing
`timeout_secs` on `chat()`/`chat_structured()` — the actual `provider.py`
(built after the doc) has it. Fix the doc to match the real protocol; docs
follow code here, not the reverse.
### D8 — `conftest.py`'s `_clear_ollama_caches` fixture repoints
Per D1, there's now one cache source (`comfydv._llm.ollama_provider`).
Fixture imports `_CHAT_RESPONSE_CACHE`/`_MODEL_LIST_CACHE` from there
instead of `comfydv.ollama`, and `ChatCompletion` instead of
`OllamaChatCompletion`. This is the single highest-priority fixture change
— every `TestResponseCache` test (12) and every `structured_output`-
toggling test depends on it for isolation.
## Sequenced task list
Replaces `tasks.md`'s Phase 3 (US1) + Phase 5 (US3) + the deferred T014.
One coordinated PR/session, ordered so the codebase stays important at each
step even though it can't be split across separate merges (per the
2026-07-11 correction — this is genuinely atomic).
1. **T-CUT-01** — `ollama.py`: add `from comfydv._llm.ollama_provider import OllamaProvider, _fetch_models, _run_async` (drop the local `_post_json`, `_TTLLRUCache`, `_cache_key`, `_MODEL_LIST_CACHE`, `_CHAT_RESPONSE_CACHE`, `_run_async`, `_fetch_models`, `_post_json` definitions — lines 42-194 collapse to the import). Repoint `_load_default_models()` and the `/dv/ollama/models` route to the imported `_fetch_models`. (D1)
2. **T-CUT-02** — `ollama_provider.py`: implement `OllamaProvider.list_models()` (port `_fetch_models`'s `/api/tags` logic, map to `ModelInfo`/`ModelStatus.UNLOADED`/`LOADED` — Ollama never emits `SLEEPING`/`DOWNLOADING`, per ADR-007's documented approximation), `load_model()` (port `/api/generate` + `keep_alive: -1`), `unload_model()` (port `/api/generate` + `keep_alive: 0`), `chat()` (port native `/api/chat` non-structured path), `chat_structured()` (build `base_url=f"{self.host}/v1"`, delegate to `comfydv._llm.chat.chat_structured()`).
3. **T-CUT-03** — `tests/test_ollama_provider.py` (new): tests for T-CUT-02's method bodies, mocking at `ollama_provider_mod._post_json`/`aiohttp.ClientSession` — ports the *substance* of the 35 relocated monkeypatches per D4/D5 (not 1:1 — collapses redundant retry-count tests per D5).
4. **T-CUT-04** — `ollama.py`: `OllamaClient.RETURN_TYPES = ("LLM_CLIENT",)`, `create_client()` returns `OllamaProvider(host, headers)`. (D2)
5. **T-CUT-05** — `ollama.py`: rename `OllamaModelSelector`→`LLMModelSelector`, `OllamaLoadModel`→`LLMLoadModel`, `OllamaUnloadModel`→`LLMUnloadModel`, `OllamaChatCompletion`→`ChatCompletion`; every `"OLLAMA_CLIENT"` input socket → `"LLM_CLIENT"`; rewrite the 3 method bodies (`load_model`, `unload_model`, `chat`) to delegate to `client.*` instead of `_post_json`/f-string URLs; delete `_client_headers` (D6); update the 3 `OllamaChatCompletion.*` references in the `/dv/ollama/update_structured_outputs` route body.
6. **T-CUT-06** — `src/comfydv/__init__.py`: update imports and `NODE_CLASS_MAPPINGS`/`NODE_DISPLAY_NAME_MAPPINGS` for the 4 renamed classes.
7. **T-CUT-07** — `tests/conftest.py`: repoint `_clear_ollama_caches` (D8) and `first_generative_model`'s `_fetch_models` import (D1).
8. **T-CUT-08** — `tests/test_ollama.py`: update the import block (4 class renames); add `_FakeProvider` test double; convert every `_post_json`-monkeypatched test to use `_FakeProvider` as `client` instead (D4); replace bare-string `client=ollama_host`/`client="http://..."` usages with a constructed provider (D3); rewrite the 2 `client == "<string>"` assertions (D2); delete `test_plain_string_client_has_no_headers` (D3) and `test_structured_output_true_sends_tool_call_payload` (D5); collapse the 15 `TestStructuredOutput` retry-count tests per D5; update `TestNodeContracts`'s `NODE_CLASSES` list (4 renames).
9. **T-CUT-09** — `contracts/llm_provider_protocol.md`: add missing `timeout_secs` params (D7).
10. **T-CUT-10** — Full suite green (`uv run pytest -m "not integration and not system"`), `ruff check --fix && ruff format`, `ty check`, `beacon doctor --strict`.
11. **T-CUT-11** — `tasks.md`: mark T007-T010/T015-T018/T014 done, referencing this plan; migration mapping (old→new names, FR-009) as a module-level constant/docstring in `ollama.py`.
12. **T-CUT-12** — Manual smoke test against a live local Ollama server per `quickstart.md`.
## What stays exactly as originally scoped
`OllamaHeader*`, `OllamaOption*`, `OllamaDebugHistory`, `OllamaHistoryLength`
classes and the `OLLAMA_HEADERS`/`OLLAMA_OPTIONS`/`OLLAMA_HISTORY` socket
types are **out of scope** — confirmed zero test dependencies force a
change, and ADR-007 never proposed touching them (only the
model-management/chat surface generalizes). `_parse_output_schema`,
`_comfy_types_for_schema`, `_build_structured_model`,
`_coerce_structured_value` stay in `ollama.py` unchanged — pure/local
schema logic with no network dependency, still needed by the live-preview
route.
@@ -0,0 +1,38 @@
# Specification Quality Checklist: LLM Provider Abstraction
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-07-11
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
All items pass on first pass — no [NEEDS CLARIFICATION] markers were needed;
scope boundaries (llama.cpp out of scope, no automatic workflow migration,
no new tracing capability) came directly from the parent epic's Non-goals
(`project-management/Roadmap/epics/llm-provider-abstraction.md`) and
ADR-007, so no ambiguity required flagging back to the user.
@@ -0,0 +1,81 @@
# Contract: `LLMProvider` protocol
This is the interface the follow-on `llamacpp-integration` epic implements
against (`LlamaCppProvider`) — it's the actual deliverable that makes ADR-007's
adapter pattern real, not internal implementation detail. Treat changes to
this contract as requiring epic-level sign-off (per ADR-007's own scope),
not a routine refactor.
```python
class ModelStatus(str, Enum):
UNLOADED = "unloaded"
LOADING = "loading"
LOADED = "loaded"
SLEEPING = "sleeping" # not all providers emit this
DOWNLOADING = "downloading" # not all providers emit this
class ModelInfo(BaseModel):
name: str
status: ModelStatus
size: int | None = None
class Message(BaseModel):
role: Literal["system", "user", "assistant"]
content: str
class LLMProvider(Protocol):
async def list_models(self) -> list[ModelInfo]: ...
async def load_model(self, model: str) -> None: ...
async def unload_model(self, model: str) -> None: ...
async def chat(
self, model: str, messages: list[Message], options: dict | None = None,
timeout_secs: float = 300.0,
) -> str: ...
async def chat_structured(
self, model: str, messages: list[Message], schema: type[BaseModel],
options: dict | None = None, timeout_secs: float = 300.0, max_retries: int = 2,
) -> BaseModel: ...
```
## Behavioral requirements (every implementation MUST satisfy)
- `load_model`/`unload_model` are **idempotent** — calling either on a model
already in that state is not an error.
- `chat_structured` **MUST NOT** return a `BaseModel` instance with a blank
required `str` field — validate and retry (bounded, provider-internal)
rather than pass through invalid data. On exhausted retries, raise
`RuntimeError` naming the model, attempt count, and a truncated snippet of
the last invalid response (FR-004 in `../spec.md`).
- `list_models` MUST return every model the server currently knows about,
including ones not currently loaded — this is a status listing, not a
"loaded models only" filter.
- A provider that cannot represent a given `ModelStatus` value (e.g. Ollama
has no `sleeping`/`downloading` concept) MUST normalize to the closest
applicable status rather than omit the model or invent a new status value
outside this enum.
- Connection state (host, auth headers, or equivalent) is captured once at
provider-construction time; no method takes connection details as a
parameter.
## Non-requirements (explicitly not part of this contract)
- No requirement that every provider support every `ModelStatus` value —
see `data-model.md`'s per-provider emission notes.
- No streaming contract — `chat`/`chat_structured` return a complete result,
not a stream. (Not requested by the parent spec; a future contract change
if ever needed.)
- No multi-turn agent/tool-use contract beyond a single structured-output
call — out of scope per the parent epic's Non-goals.
## ComfyUI-facing contract: `LLM_CLIENT` socket
An `LLMProvider`-implementing instance is the value carried by ComfyUI's
`LLM_CLIENT` custom socket type. Any node that outputs `LLM_CLIENT` (e.g.
`OllamaClient`, and later `LlamaCppClient`) is committing to have constructed
a fully-configured provider instance — no partial/lazy construction that
defers connection details to the consuming node.
Generic nodes (`LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`,
`ChatCompletion`) accept `LLM_CLIENT` as their only connection-related input
and MUST NOT branch on which concrete provider type they received — doing so
would defeat the point of the protocol boundary (ADR-007).
@@ -0,0 +1,69 @@
# Data Model: LLM Provider Abstraction
## `ModelStatus` (enum)
Residency status of a model on a provider's server.
| Value | Meaning | Emitted by |
|---|---|---|
| `unloaded` | Known to the server, not resident in memory | all providers |
| `loading` | Transitioning into memory | all providers |
| `loaded` | Resident and ready to serve requests | all providers |
| `sleeping` | Resident but idle-parked | llama.cpp only; Ollama has no distinct signal for this via its API and normalizes resident-and-idle to `loaded` (documented approximation, ADR-007) |
| `downloading` | Server is fetching model weights | llama.cpp only; `OllamaProvider` never emits this (Ollama's pull/download flow is out of scope, per the original Ollama epic's non-goals) |
## `ModelInfo`
One entry returned by `list_models()`.
| Field | Type | Notes |
|---|---|---|
| `name` | `str` | Model identifier as the provider's server knows it |
| `status` | `ModelStatus` | See above |
| `size` | `int \| None` | Bytes, if the provider reports it; `None` otherwise |
## `LLMProvider` (Protocol)
The adapter boundary. Every backend (`OllamaProvider` now, `LlamaCppProvider`
in the follow-on epic) implements this shape; ComfyUI nodes depend only on
the protocol, never on a concrete provider class.
| Method | Signature | Notes |
|---|---|---|
| `list_models` | `async def list_models(self) -> list[ModelInfo]` | |
| `load_model` | `async def load_model(self, model: str) -> None` | Idempotent: loading an already-loaded model is not an error |
| `unload_model` | `async def unload_model(self, model: str) -> None` | Idempotent: unloading an already-unloaded model is not an error |
| `chat` | `async def chat(self, model: str, messages: list[Message], options: dict) -> str` | Free-text response |
| `chat_structured` | `async def chat_structured(self, model: str, messages: list[Message], schema: type[BaseModel], options: dict) -> BaseModel` | Validated response; raises on exhausted retries (see FR-004) |
A concrete provider instance is constructed once per ComfyUI client node with
its connection's host/headers as instance state (Constitution Principle V
justification — see `research.md`), and that instance is the value carried
by the `LLM_CLIENT` ComfyUI socket type.
## `Message`
One turn in a chat request, matching the existing shape already sent to
Ollama's `/api/chat`/`/v1/chat/completions` (`role` + `content`); unchanged
by this feature, carried forward as-is.
| Field | Type | Notes |
|---|---|---|
| `role` | `Literal["system", "user", "assistant"]` | |
| `content` | `str` | |
## Relationships
```
ProviderConnection (ComfyUI client node)
└─ produces → LLM_CLIENT socket value (an LLMProvider instance)
└─ consumed by → LLMModelSelector, LLMLoadModel, LLMUnloadModel, ChatCompletion (ComfyUI nodes)
├─ list_models() → ModelInfo[]
├─ load_model()/unload_model() → mutates server-side residency, no return value
└─ chat()/chat_structured() → str | BaseModel
```
No new persistent storage is introduced — every entity above is
constructed per-request or per-node-execution from the connected server's
live state; the only caching is the existing in-memory TTL cache for model
listing (`_TTLLRUCache`, unchanged, reused inside `OllamaProvider`).
@@ -0,0 +1,16 @@
Feature: US1 — Connect to a local inference server and get chat responses
Scenario: Client node feeds a chat node
Given a running local inference server and a workflow with a client node wired into a chat node
When the workflow executes
Then the chat node returns the model's text response
Scenario: Unreachable server surfaces a clear error
Given a client node configured with an unreachable server address
When the workflow executes
Then the chat node reports a clear connection error rather than hanging indefinitely or crashing the workflow
Scenario: One client node configures multiple chat nodes
Given two chat nodes in the same workflow wired to the same client node
When the host address is changed on the client node
Then both chat nodes use the new address without being edited individually
@@ -0,0 +1,16 @@
Feature: US2 — Get structured, validated output instead of parsing raw text
Scenario: Valid structured response exposes typed fields
Given a chat node with structured output enabled and a valid schema
When the workflow executes and the model responds correctly
Then each schema field is available as its own typed output, and no required field is blank
Scenario: Invalid response triggers automatic retry
Given a model that returns invalid, incomplete, or empty-required-field output
When the workflow executes
Then the node automatically retries the request up to a configured limit
Scenario: Exhausted retries fail clearly instead of passing through bad data
Given a model that continues to return invalid output after all retries are exhausted
When the workflow executes
Then the node fails with a clear, specific error rather than silently passing through invalid or partial data
@@ -0,0 +1,16 @@
Feature: US3 — Manage which models are resident in memory
Scenario: List models with current status
Given a running local server with at least one available model
When a workflow author uses the model-listing node
Then they see each available model along with its current status
Scenario: Load a model into memory
Given a model that is not currently loaded
When a workflow author runs the load-model node against it
Then the model becomes loaded and is then usable by the chat node
Scenario: Unload a model from memory
Given a model that is loaded and idle
When a workflow author runs the unload-model node against it
Then the model is freed from memory and its reported status updates accordingly
@@ -0,0 +1,12 @@
Feature: US4 — Reconnect an existing workflow after upgrading
Scenario: Renamed nodes are reported with a documented replacement
Given a saved workflow using the current Ollama-specific node and connection-socket names
When it is opened after upgrading
Then ComfyUI reports the now-missing node types
And documentation identifies the replacement node for each one
Scenario: Reconnected workflow produces equivalent output
Given a workflow that has been reconnected to the new generic nodes
When it executes with the same inputs and model as before the upgrade
Then it produces equivalent output
+121
View File
@@ -0,0 +1,121 @@
# Implementation Plan: LLM Provider Abstraction
**Branch**: `007-llm-provider-abstraction` | **Date**: 2026-07-11 | **Spec**: [spec.md](./spec.md)
**Input**: Feature specification from `/specs/007-llm-provider-abstraction/spec.md`
**Note**: This template is filled in by the `/speckit-plan` command. See `.specify/templates/plan-template.md` for the execution workflow.
## Summary
Define a shared `LLMProvider` protocol (list/load/unload/chat/structured-chat)
and generic ComfyUI nodes so workflow authors can connect any supported local
inference backend the same way. Migrate the existing Ollama integration onto
it — `OllamaProvider` becomes the first (and, in this feature, only)
implementation — including moving structured-output from ADR-006's
hand-rolled tool-calling onto `pydantic-ai`, per ADR-007. This is a
behavior-preserving mechanism swap for existing capability, plus the new
protocol boundary that the follow-on `llamacpp-integration` epic builds a
second provider against.
## Technical Context
**Language/Version**: Python ≥3.11 (per `pyproject.toml`)
**Primary Dependencies**: `aiohttp` (existing, unchanged — model-management
REST calls), `pydantic` (existing, unchanged — validation), `pydantic-ai` +
`openai` (new, per ADR-007 — powers `chat_structured()` only)
**Storage**: N/A — no persistent storage; the existing in-memory
`_TTLLRUCache` for model listing is reused unchanged inside `OllamaProvider`
**Testing**: `pytest` via `uv run pytest`, following `tests/test_ollama.py`'s
existing conventions (mocked `aiohttp`/`pydantic-ai` calls, no live server
required for unit tests; the existing `integration` pytest marker — "requiring
live Ollama at localhost:11434" — is reused for tests that exercise a real
server)
**Target Platform**: ComfyUI custom-node runtime, cross-platform wherever
ComfyUI runs; CPU-only dev harness per the project's stated vision
**Project Type**: Library / ComfyUI custom-node pack (single project,
existing `src/comfydv/` layout — no new top-level project)
**Performance Goals**: No new numeric target; must not add latency beyond the
existing bounded retry loop already in ADR-006 (`max_retries`, 0–5)
**Constraints**: Behavior-preserving for existing Ollama structured/
non-structured chat and all model-management calls (FR-007, FR-008); no new
dependency beyond what ADR-007 already accepted (`pydantic-ai`, `openai`,
their transitive `httpx`/`tiktoken`); model-management stays on `aiohttp`
**Scale/Scope**: One new internal package (`src/comfydv/_llm/`), migration of
the existing 1056-line `ollama.py` node/HTTP logic to consume it, five
ComfyUI node classes renamed to generic names — no change to the project's
single-repo, single-package scope
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Principle | Verdict | Notes |
|---|---|---|
| I. ComfyUI Contract First | PASS | Generic nodes (`ChatCompletion`, `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`, `OllamaClient`) still expose `INPUT_TYPES`/`RETURN_TYPES`/`RETURN_NAMES`/`FUNCTION`/`CATEGORY`; `NODE_CLASS_MAPPINGS` in `__init__.py` remains the only install-time interface. `LLMProvider` is internal, not a ComfyUI-facing contract change beyond the node/socket rename. |
| II. Sandbox All User-Supplied Code | N/A | No template/expression evaluation in this feature — structured-output schemas are parsed as JSON Schema by `pydantic`, never `eval`/`exec`. |
| III. Test-First | PASS (binding on tasks/implement phases) | `tests/test_ollama.py`'s existing assertions are the regression oracle (see `research.md`); new `_llm` package gets tests written before implementation, red→green→refactor. |
| IV. Graceful Degradation Outside ComfyUI | PASS (binding on implementation) | `src/comfydv/_llm/` must not import `comfy`/`server` at module scope, matching `ollama.py`'s existing guarded-import pattern. |
| V. Simplicity — Function Before Class | **Justified exception — see Complexity Tracking** | `LLMProvider` is a `Protocol` implemented by stateful provider classes, not module-level functions. |
| VI. Fixed Output Positions | PASS (binding on implementation) | `ChatCompletion`'s (renamed from `OllamaChatCompletion`) `RETURN_TYPES`/`RETURN_NAMES` positions 0/1 carry forward unchanged — only the class/node name and internal mechanism change. |
Re-checked post-Phase 1 design (data-model.md, contracts/): unchanged — the
`Protocol`-based design in `contracts/llm_provider_protocol.md` is exactly
what was justified below, no new gate violations introduced by the detailed
design.
## Project Structure
### Documentation (this feature)
```text
specs/[###-feature]/
├── plan.md # This file (/speckit-plan command output)
├── research.md # Phase 0 output (/speckit-plan command)
├── data-model.md # Phase 1 output (/speckit-plan command)
├── quickstart.md # Phase 1 output (/speckit-plan command)
├── contracts/ # Phase 1 output (/speckit-plan command)
└── tasks.md # Phase 2 output (/speckit-tasks command - NOT created by /speckit-plan)
```
### Source Code (repository root)
```text
src/comfydv/
├── ollama.py # existing — node classes renamed to generic names,
│ # delegates HTTP/chat logic to _llm/ internally
├── _llm/ # new internal package (not a ComfyUI node module)
│ ├── __init__.py
│ ├── provider.py # LLMProvider Protocol, ModelStatus, ModelInfo, Message
│ ├── ollama_provider.py # OllamaProvider — wraps existing aiohttp REST logic
│ └── chat.py # shared chat_structured() pydantic-ai helper
└── __init__.py # NODE_CLASS_MAPPINGS updated for renamed nodes
tests/
├── test_ollama.py # existing — updated for renamed nodes; behavior-
│ # preserving assertions carried forward unchanged
└── test_llm_provider.py # new — protocol conformance + OllamaProvider unit tests
```
**Structure Decision**: Single project (existing `src/comfydv/` layout, no new
top-level project). New internal package `src/comfydv/_llm/` (underscore
prefix marks it as internal, consistent with existing internal helpers like
`_TTLLRUCache` that already live inside `ollama.py`) hosts the protocol and
shared chat logic; `ollama.py` keeps the actual ComfyUI-registered node
classes and becomes a thin caller into `_llm`.
## Complexity Tracking
> **Fill ONLY if Constitution Check has violations that must be justified**
| Violation | Why Needed | Simpler Alternative Rejected Because |
|-----------|------------|-------------------------------------|
| `LLMProvider` as a `Protocol` implemented by stateful classes (Principle V: Function Before Class) | Every one of the five protocol methods (`list_models`/`load_model`/`unload_model`/`chat`/`chat_structured`) needs the same connection state (host, auth headers) — genuine shared state, the exact condition under which the constitution allows a class. A `Protocol` also lets ComfyUI's `LLM_CLIENT` socket carry one opaque object satisfying the shape, which is what makes the adapter pattern (ADR-007) work on the canvas. | Module-level functions taking host/headers as explicit parameters on every call were considered and rejected: they'd reintroduce the exact per-call-site repetition [ADR-005](../../project-management/ADRs/ADR-005-ollama-host-config-via-client-node.md)'s config-node pattern was built to eliminate, and a bare function can't be the typed payload of a ComfyUI socket the way an object implementing a `Protocol` can. |
@@ -0,0 +1,49 @@
# Quickstart: LLM Provider Abstraction
A minimal ComfyUI workflow using the generic nodes this feature introduces.
## 1. Connect to a local server
Add an **Ollama Client** node. Set its host widget (default
`http://localhost:11434`). This is the only node that knows it's talking to
Ollama specifically — everything downstream just sees `LLM_CLIENT`.
## 2. Chat
Add a **Chat Completion** node. Wire the client node's `LLM_CLIENT` output
into it. Set a model name (or feed one from a model-selector node — see
below) and a prompt. Run the workflow: the node returns the model's text
response.
## 3. Get structured output instead of free text
On the same **Chat Completion** node, enable `structured_output` and supply
a JSON Schema (e.g. `{"type": "object", "properties": {"summary": {"type": "string"}, "score": {"type": "number"}}, "required": ["summary", "score"]}`).
Re-run: the node now exposes one typed output socket per schema property
(`summary`, `score`) instead of a single text blob, and guarantees neither
is blank.
## 4. Manage what's loaded in memory
Add an **LLM Model Selector** node wired to the same client, to see every
model the server knows about and its current status (`unloaded` /
`loading` / `loaded` / …). Add **LLM Load Model** / **LLM Unload Model**
nodes, wired to the same client, to explicitly control residency before a
chat node needs a model.
## 5. (Follow-on epic) Swap backends without touching downstream nodes
Once `llamacpp-integration` ships a **Llama.cpp Client** node, replacing the
**Ollama Client** node in step 1 with it — and nothing else — is the whole
migration: it emits the same `LLM_CLIENT` socket type, so every node from
steps 2–4 keeps working unmodified. That's the point of this feature.
## Migrating an existing pre-upgrade workflow
If you have a saved workflow using the old node names (`OllamaClient`,
`OllamaChatCompletion`, `OllamaModelSelector`, `OllamaLoadModel`,
`OllamaUnloadModel`), ComfyUI will report those node types as missing on
load. Replace each with its generic equivalent from the list above and
reconnect — behavior is unchanged, only the node names and the
`LLM_CLIENT` socket type (replacing `OLLAMA_CLIENT`) are different. See
`spec.md`'s User Story 4 and Edge Cases for the full detail.
@@ -0,0 +1,94 @@
# Research: LLM Provider Abstraction
All unknowns below were already resolved during DESIGN-phase work on
[ADR-007](../../project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md);
this file consolidates that research for the plan gate rather than re-deriving it.
## Decision: `pydantic-ai` for `chat_structured()`, not hand-rolled tool-calling
**Decision**: Both `OllamaProvider` and the future `LlamaCppProvider` implement
`chat_structured()` via `pydantic-ai`'s `Agent`/`output_type`, called through
`OpenAIProvider(base_url=<host>/v1)`.
**Rationale**: A live research pass against current `pydantic-ai` docs/source
found `httpx` is a base dependency of `pydantic-ai-slim` itself (not merely
pulled in by an OpenAI extra), and `openai`+`tiktoken` are required for any
OpenAI-compatible provider — a fixed, one-time dependency tax rather than a
per-backend one. `pydantic.create_model()`-built `BaseModel` subclasses
(comfydv's existing dynamic-schema pattern) work as `output_type` with no
special-casing. `OpenAIProvider(base_url=...)` is one generic code path both
Ollama's and llama.cpp's OpenAI-compatible `/v1/chat/completions` reach
identically.
**Alternatives considered**: hand-roll llama.cpp's structured output too
(duplicates [ADR-006](../../project-management/ADRs/ADR-006-structured-ollama-output-tool-calling-not-pydantic-ai.md)'s
mechanism — rejected, defeats the DRY goal); extract a shared aiohttp-based
helper with no new dependencies (rejected — forces re-deriving pydantic-ai's
retry/validation machinery by hand for no benefit now that two backends exist
to amortize the dependency cost against).
## Decision: aiohttp stays authoritative for model-management REST calls
**Decision**: `list_models()` / `load_model()` / `unload_model()` on every
provider use `aiohttp` — no dependency change from the existing Ollama
integration for this surface.
**Rationale**: [ADR-004](../../project-management/ADRs/ADR-004-aiohttp-over-httpx-for-ollama.md)'s
reasoning (ComfyUI's own server is aiohttp-based; httpx was an unjustified
addition) still applies fully to REST calls that don't need pydantic-ai's
machinery. ADR-007 narrows ADR-004's scope to exactly this surface, rather
than superseding it.
**Alternatives considered**: route everything (including model management)
through `pydantic-ai`/httpx for consistency — rejected, pydantic-ai has no
model-lifecycle-management concept (it's a chat/agent framework, not a
generic REST client) and would add no value over plain aiohttp calls that
already exist and work.
## Decision: `LLMProvider` as a `Protocol` implemented by stateful provider classes
**Decision**: `list_models`/`load_model`/`unload_model`/`chat`/`chat_structured`
are defined as a `typing.Protocol`, implemented by `OllamaProvider` (and later
`LlamaCppProvider`) classes, each constructed once per ComfyUI client node
with the connection's host/headers as instance state.
**Rationale**: This is a Constitution Principle V ("Function Before Class")
gate — classes are only justified when there's shared state a group of
functions would otherwise have to thread through every call. Here there is:
every one of the five protocol methods needs the same host/headers, exactly
the connection config [ADR-005](../../project-management/ADRs/ADR-005-ollama-host-config-via-client-node.md)'s
config-node pattern centralizes. A `Protocol` (structural typing, no
inheritance required) keeps this lightweight — `LlamaCppProvider` doesn't
need to import or subclass `OllamaProvider`, it only needs to match the
method shapes.
**Alternatives considered**: module-level functions taking host/headers as
explicit parameters on every call — rejected, this reintroduces the exact
per-call-site repetition ADR-005 eliminated, and loses the ability for a
ComfyUI `LLM_CLIENT` socket to carry one opaque object implementing the
protocol (functions can't be typed as a socket payload the way an object
implementing a `Protocol` can).
## Decision: behavior-preserving migration, verified against existing tests
**Decision**: `tests/test_ollama.py`'s existing assertions (retry bounds,
required-string validation, error messages) are the acceptance bar for the
migrated `OllamaProvider.chat_structured()` — this is a mechanism swap, not a
new capability, per the parent epic's Non-goals and spec FR-008.
**Rationale**: Constitution Principle III (Test-First) and the epic's
explicit framing of this as the riskiest change in the whole llama.cpp
proposal (touches a Done, shipped epic's code) both point the same way: the
existing test suite is the regression oracle, not a new one written from
scratch.
## Testing approach
Per Constitution Principle IV (Graceful Degradation Outside ComfyUI), the new
`src/comfydv/_llm/` package must not import `comfy`/`server` at module scope,
matching `ollama.py`'s existing runtime-guarded pattern. Unit tests mock
`aiohttp`/`pydantic-ai` calls (no live server required, matching
`tests/test_ollama.py`'s existing convention); the `integration` pytest
marker (already defined in `pyproject.toml`, "requiring live Ollama at
localhost:11434") is reused, not redefined, for tests that exercise a real
local server.
+119
View File
@@ -0,0 +1,119 @@
# Feature Specification: LLM Provider Abstraction
**Feature Branch**: `007-llm-provider-abstraction`
**Created**: 2026-07-11
**Status**: Draft
**Input**: User description: "Introduce a shared LLMProvider protocol for comfydv's LLM backend nodes so ComfyUI workflow authors can swap between local inference servers (starting with Ollama, with llama.cpp planned next) without changing their chat/model-management nodes. Migrate the existing Ollama integration onto generic nodes (client config, model list, load, unload, chat completion with optional structured/validated output) backed by this shared interface, per ADR-007."
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Connect to a local inference server and get chat responses (Priority: P1)
As a ComfyUI workflow author, I want to point a single configuration node at my local LLM server and get chat responses through a generic chat node, so I can generate text without hardcoding a server address into every node that needs one.
**Why this priority**: This is the minimum viable path — without a working connection and a basic chat response, nothing else in this feature has value. It also directly replaces the most-used capability of the existing Ollama integration, so it carries the highest regression risk.
**Independent Test**: Wire a client configuration node into a chat node, run the workflow against a running local server, and confirm the chat node returns the model's text response.
**Acceptance Scenarios**:
1. **Given** a running local inference server and a workflow with a client node wired into a chat node, **When** the workflow executes, **Then** the chat node returns the model's text response.
2. **Given** a client node configured with an unreachable server address, **When** the workflow executes, **Then** the chat node reports a clear connection error rather than hanging indefinitely or crashing the workflow.
3. **Given** two chat nodes in the same workflow wired to the same client node, **When** the host address is changed on the client node, **Then** both chat nodes use the new address without being edited individually.
---
### User Story 2 - Get structured, validated output instead of parsing raw text (Priority: P1)
As a workflow author, I want to describe the shape of data I need and turn on structured output for a chat node, so downstream nodes receive individually typed fields I can trust are present and non-empty, instead of me parsing free text myself.
**Why this priority**: This is an existing, relied-upon capability of the current Ollama integration (structured output with retry-on-invalid-response). Preserving it exactly is required for this migration to be considered safe, so it's equal priority to basic chat.
**Independent Test**: Enable structured output on a chat node with a schema describing two or three fields, run the workflow against a model, and confirm each schema field is exposed as its own typed output socket with a valid value.
**Acceptance Scenarios**:
1. **Given** a chat node with structured output enabled and a valid schema, **When** the workflow executes and the model responds correctly, **Then** each schema field is available as its own typed output, and no required field is blank.
2. **Given** a model that returns invalid, incomplete, or empty-required-field output, **When** the workflow executes, **Then** the node automatically retries the request up to a configured limit.
3. **Given** a model that continues to return invalid output after all retries are exhausted, **When** the workflow executes, **Then** the node fails with a clear, specific error rather than silently passing through invalid or partial data.
---
### User Story 3 - Manage which models are resident in memory (Priority: P2)
As a workflow author running models locally, I want to see which models are currently loaded, loading, or unloaded, and explicitly load or unload a model, so I can control memory usage on my machine without leaving ComfyUI or using a separate terminal.
**Why this priority**: Valuable and already present in the current Ollama integration, but a workflow can still generate output without ever calling load/unload explicitly (servers can auto-load on first use) — so this is lower risk to defer than basic chat.
**Independent Test**: Use a model-listing node against a running server, confirm it shows each available model with a current status; use load/unload nodes against one model and confirm its reported status changes accordingly.
**Acceptance Scenarios**:
1. **Given** a running local server with at least one available model, **When** a workflow author uses the model-listing node, **Then** they see each available model along with its current status.
2. **Given** a model that is not currently loaded, **When** a workflow author runs the load-model node against it, **Then** the model becomes loaded and is then usable by the chat node.
3. **Given** a model that is loaded and idle, **When** a workflow author runs the unload-model node against it, **Then** the model is freed from memory and its reported status updates accordingly.
---
### User Story 4 - Reconnect an existing workflow after upgrading (Priority: P3)
As an existing user of the current Ollama nodes, when I open a workflow I saved before this change, I want it to be clear which new node replaces each renamed one, so I can reconnect my workflow with minimal effort and get the same results as before.
**Why this priority**: This is migration friction, not new capability — it matters for a good upgrade experience but doesn't block anyone building a new workflow from scratch, so it's the lowest priority of the four.
**Independent Test**: Open a workflow saved against the current Ollama-specific node names, follow the provided migration guidance to reconnect it to the new generic nodes, and confirm it produces the same output as before, given the same inputs and model.
**Acceptance Scenarios**:
1. **Given** a saved workflow using the current Ollama-specific node and connection-socket names, **When** it is opened after upgrading, **Then** ComfyUI reports the now-missing node types (standard ComfyUI behavior for renamed nodes), and documentation identifies the replacement node for each one.
2. **Given** a workflow that has been reconnected to the new generic nodes, **When** it executes with the same inputs and model as before the upgrade, **Then** it produces equivalent output.
---
### Edge Cases
- What happens when the configured server address is unreachable at the moment a model-listing, load, or unload node runs (not just the chat node)?
- What happens when a workflow author supplies an invalid or malformed schema to structured output, rather than an invalid model response?
- What happens when the connected server does not support structured/validated output at all?
- What happens to an in-flight chat request if the model it depends on is unloaded by another node in the same workflow run?
- What happens when a workflow author tries to wire a pre-upgrade Ollama-specific node's output into a new generic node, or vice versa? (Expected: ComfyUI's own type-checking refuses the connection, since the socket types differ — this is the intended, safe failure mode, not a bug to work around.)
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: The system MUST allow a workflow author to configure a connection to a local inference server once and reuse that single configuration across multiple nodes in the same workflow.
- **FR-002**: The system MUST allow a workflow author to request either free-text or schema-validated structured output from the same chat node, choosing per request.
- **FR-003**: When structured output is requested, the system MUST validate the response against the supplied schema and MUST NOT deliver output to downstream nodes where a required field is missing or empty.
- **FR-004**: When validation fails, the system MUST retry the request automatically up to a configurable limit before reporting a clear, actionable error that identifies the model, the number of attempts made, and a snippet of the last invalid response.
- **FR-005**: The system MUST allow a workflow author to list available models on a connected server along with each model's current residency status.
- **FR-006**: The system MUST allow a workflow author to explicitly load a model into memory and explicitly unload a model from memory.
- **FR-007**: The system's chat and model-management nodes MUST behave identically regardless of which supported local inference server is connected, given equivalent inputs.
- **FR-008**: The existing chat and structured-output behavior for the currently-supported local inference server (Ollama) MUST be unchanged in outcome after this migration — same retry limits, same validation rules, same error conditions — since this feature changes the underlying mechanism, not the capability.
- **FR-009**: The system MUST document, for each node type renamed or removed by this change, which new node replaces it.
### Key Entities *(include if feature involves data)*
- **Provider connection**: A configured connection to one local inference server (address and any authentication), created once and reused by every model-management and chat node that needs it.
- **Model**: An inference model known to a provider connection, identified by name, with a current residency status (e.g., unloaded, loading, loaded, and — on servers that support it — sleeping or downloading).
- **Chat request/response**: A request for a model's output, optionally carrying a schema describing the required shape of a structured response, and the corresponding validated or free-text result.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: A workflow author can go from no nodes to a working chat response using no more than two nodes (one connection node, one chat node).
- **SC-002**: Structured-output workflows never deliver a blank or missing required field to a downstream node — every request either produces fully valid data or a clear error, with zero silent partial results.
- **SC-003**: Existing example/reference workflows built against the current Ollama nodes remain reproducible on the new nodes with equivalent output, after a workflow author reconnects the renamed nodes.
- **SC-004**: Adding support for a second local inference server (planned as a follow-on feature) requires no visible change to chat or model-management node behavior — only a new connection node is needed.
## Assumptions
- Workflow authors run their own local inference server (e.g., Ollama) reachable over HTTP from the machine running ComfyUI; this feature does not host, install, or manage that server.
- Users with workflows saved against the current Ollama-specific node and socket names will need to manually reconnect them after upgrading. This is an accepted, intentional breaking change (confirmed 2026-07-11), not a defect — see FR-009 for the mitigation (documented replacement mapping), not automatic migration.
- Support for a second local inference server (llama.cpp) is planned as a separate, follow-on feature and is out of scope here — this feature only needs to prove the shared design works end-to-end for one real backend (Ollama).
- Structured-output schemas remain limited to flat object shapes with typed properties, consistent with what the current Ollama integration already supports — deeper nested schemas are unaffected by (neither improved nor degraded by) this change.
- No new observability/tracing capability is introduced for workflow authors as part of this feature, even though the underlying mechanism change makes it feasible to add later.
+267
View File
@@ -0,0 +1,267 @@
# Tasks: LLM Provider Abstraction
**Input**: Design documents from `/specs/007-llm-provider-abstraction/`
**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/llm_provider_protocol.md
**Tests**: First-class (spec carries Acceptance Scenarios) — every implementation task has a paired failing-test task (`-T`/`-I` suffix) per BEACON's test-first discipline.
**Organization**: Tasks are grouped by user story (spec.md priorities P1/P1/P2/P3) to enable independent implementation and testing of each.
## Format: `[ID] [P?] [Story] Description`
- **[P]**: Can run in parallel (different files, no dependencies)
- **[Story]**: Which user story this task belongs to (US1–US4)
- **-T / -I**: paired test (red) / implementation (green) — a `-I` task is never parallel with its own `-T`
## Path Conventions
Single project: `src/comfydv/`, `tests/` at repository root (per plan.md's Project Structure).
---
## Phase 1: Setup
- [x] T001 Add `pydantic-ai` and `openai` to `pyproject.toml` dependencies; curate the addition into `requirements.txt` per [ADR-003](../../project-management/ADRs/ADR-003-requirements-txt-authoring-policy.md)
- [x] T002 [P] Create `src/comfydv/_llm/__init__.py` (empty package init)
---
## Phase 2: Foundational (Blocking Prerequisites)
**⚠️ CRITICAL**: No user story work can begin until this phase is complete.
- [x] T003 [P] Define `Message`, `ModelStatus`, `ModelInfo` in `src/comfydv/_llm/provider.py` per `data-model.md`
- [x] T004 Define the `LLMProvider` `Protocol` in `src/comfydv/_llm/provider.py` per `contracts/llm_provider_protocol.md` (depends on T003)
- [x] T005 [P] `LLM_CLIENT` is introduced as part of T007-I (`OllamaClient`'s output type) rather than as a standalone constant — ComfyUI socket types are plain string literals, not declared objects; folded in, not skipped
- [x] T006 Scaffold `OllamaProvider.__init__(self, host, headers)` in `src/comfydv/_llm/ollama_provider.py`, porting the existing module-level `_post_json`/`_fetch_models`/`_run_async`/`_TTLLRUCache` helpers from `ollama.py` into it — behavior-preserving port, not a rewrite (depends on T004)
**Checkpoint**: protocol + provider skeleton exist; user story work can begin.
---
## Phase 3: User Story 1 — Connect to a local server and get chat responses (Priority: P1) 🎯 MVP
**Goal**: A workflow author wires a client node into a generic chat node and gets a text response.
**Independent Test**: Wire `OllamaClient` → `ChatCompletion`, run against a live server, confirm text output.
**Superseded 2026-07-11 — see `atomic-cutover-plan.md`.** T007–T010 as
written below assumed US1 was independently deliverable; it isn't (see the
correction at the bottom of this file). The actual work is now
`atomic-cutover-plan.md`'s **T-CUT-04, T-CUT-05, T-CUT-06, T-CUT-08**
(`OllamaClient` output-type change, `ChatCompletion` rename+delegation,
`__init__.py` registration, and the corresponding `test_ollama.py`
rewrite using a `_FakeProvider` double). Original text kept below for
history, not as the active task list:
- [-] T007-T [US1] ~~Write FAILING test: `OllamaClient` node constructs and outputs an `OllamaProvider` via the `LLM_CLIENT` socket, in `tests/test_ollama.py`~~ _Superseded, see Phase 8 T-CUT-08._
- [-] T007-I [US1] ~~Update `OllamaClient` in `src/comfydv/ollama.py` to construct and output an `OllamaProvider` via `LLM_CLIENT`~~ _Superseded, see Phase 8 T-CUT-04._
- [-] T008-T [US1] ~~Write FAILING test: `OllamaProvider.chat()` returns model text via the existing `/api/chat` aiohttp call~~ _Superseded, see Phase 8 T-CUT-03._
- [-] T008-I [US1] ~~Implement `OllamaProvider.chat()` in `src/comfydv/_llm/ollama_provider.py`~~ _Superseded, see Phase 8 T-CUT-02._
- [-] T009-T [US1] ~~Write FAILING test: generic `ChatCompletion` node (non-structured) calls `provider.chat()` and surfaces a clear error on an unreachable host~~ _Superseded, see Phase 8 T-CUT-08._
- [-] T009-I [US1] ~~Rename `OllamaChatCompletion` → `ChatCompletion`, delegate the non-structured path to `LLMProvider.chat()`, update `NODE_CLASS_MAPPINGS`~~ _Superseded, see Phase 8 T-CUT-05/T-CUT-06._
- [-] T010-T [US1] ~~Write FAILING test: two `ChatCompletion` nodes sharing one `OllamaClient` both pick up a host change~~ _Superseded, see Phase 8 T-CUT-08._
- [-] T010-I [US1] ~~Verify/adjust that `OllamaClient` → `OllamaProvider` construction happens per node execution~~ _Superseded, folded into Phase 8 T-CUT-04 (construction is already per-call in `create_client()`)._
**Checkpoint**: superseded — see `atomic-cutover-plan.md`'s checkpoint (T-CUT-10, full suite green).
---
## Phase 4: User Story 2 — Structured, validated output (Priority: P1)
**Goal**: The chat node's `structured_output` toggle returns validated typed fields via the shared `pydantic-ai` mechanism.
**Independent Test**: Enable `structured_output` with a schema, run against a model, confirm typed sockets are populated and never blank.
- [x] T011-T [P] [US2] Write test: `chat_structured()` returns a validated schema instance on success, in `tests/test_llm_chat_structured.py` (witnesses `features/us2_structured_output.feature` scenario "Valid structured response exposes typed fields") — mocked at the `_build_agent` seam after an API-discovery spike into pydantic-ai's exact `Agent`/`OpenAIProvider`/`OpenAIChatModel` constructor and exception surface (not guessable from training data alone — verified live against the installed package); tests and implementation validated together rather than strictly red-first, noted honestly rather than presented as pure TDD
- [x] T011-I [US2] Implement the shared `chat_structured()` helper in `src/comfydv/_llm/chat.py` using `pydantic-ai`'s `Agent`/`output_type` through `OpenAIProvider(base_url=<host>/v1)` + `OpenAIChatModel` — makes T011-T pass (depends on T004)
- [x] T012-T [US2] Write test: invalid/failed-validation responses trigger automatic retry up to `max_retries` (clamped 0–5), in `tests/test_llm_chat_structured.py` (witnesses `features/us2_structured_output.feature` scenario "Invalid response triggers automatic retry")
- [x] T012-I [US2] Implement the bounded retry loop (0–5, matching ADR-006's existing contract) around the `pydantic-ai` call in `src/comfydv/_llm/chat.py`, with the Agent's own internal retries disabled (`retries=0`) so the error contract is comfydv's — makes T012-T pass (depends on T011-I)
- [x] T013-T [US2] Write test: exhausted retries raise `RuntimeError` naming the model, attempt count, and a truncated last-response snippet, in `tests/test_llm_chat_structured.py` (witnesses `features/us2_structured_output.feature` scenario "Exhausted retries fail clearly instead of passing through bad data")
- [x] T013-I [US2] Implement the exhausted-retry error path in `src/comfydv/_llm/chat.py`, matching ADR-006's existing error message contract (model, attempt count, truncated last response) — makes T013-T pass (depends on T012-I)
- [-] T014-T [US2] _Superseded — see `atomic-cutover-plan.md` D5 and T-CUT-08. Original: write FAILING test that `ChatCompletion`'s `structured_output=True` path wires a schema through `chat_structured()` to per-field dynamic ComfyUI output sockets. D5 replaces the originally-planned retry-count-style test with a delegation test against a `_FakeProvider`, since retry behavior is already covered by `tests/test_llm_chat_structured.py`._
- [-] T014-I [US2] _Superseded — see `atomic-cutover-plan.md` T-CUT-05. Original: wire `ChatCompletion`'s `structured_output`/`output_schema` inputs to `LLMProvider.chat_structured()`, preserving the dynamic-socket UX from ADR-006 — still the right implementation shape, just executed as part of the coordinated T-CUT-05 rename, not standalone._
**Checkpoint**: US1 + US2 both independently functional — matches today's Ollama capability, now on the shared mechanism.
---
## Phase 5: User Story 3 — Manage model residency (Priority: P2)
**Goal**: List/load/unload models through generic nodes against any connected provider.
**Independent Test**: List models via `LLMModelSelector`; load/unload one via `LLMLoadModel`/`LLMUnloadModel`; confirm status changes.
**Superseded 2026-07-11 — see `atomic-cutover-plan.md`.** Maps to
**T-CUT-02** (`OllamaProvider.list_models`/`load_model`/`unload_model`
method bodies), **T-CUT-03** (`tests/test_ollama_provider.py`, new file),
and **T-CUT-05/T-CUT-06/T-CUT-08** (the node renames + delegation +
registration + test rewrite). Original text kept for history:
- [-] T015-T [US3] ~~Write FAILING test: `OllamaProvider.list_models()` returns `ModelInfo` entries with status normalized into `ModelStatus`~~ _Superseded, see Phase 8 T-CUT-03._
- [-] T015-I [US3] ~~Implement `OllamaProvider.list_models()`~~ _Superseded, see Phase 8 T-CUT-02._
- [-] T016-T [US3] ~~Write FAILING test: `OllamaProvider.load_model()`/`unload_model()` are idempotent~~ _Superseded, see Phase 8 T-CUT-03._
- [-] T016-I [US3] ~~Implement `OllamaProvider.load_model()`/`unload_model()`~~ _Superseded, see Phase 8 T-CUT-02._
- [-] T017-T [US3] ~~Write FAILING test: generic `LLMModelSelector` node returns model+status pairs~~ _Superseded, see Phase 8 T-CUT-08._
- [-] T017-I [US3] ~~Rename `OllamaModelSelector` → `LLMModelSelector`, delegate to `LLMProvider.list_models()`~~ _Superseded, see Phase 8 T-CUT-05/T-CUT-06._
- [-] T018-T [US3] ~~Write FAILING test: generic `LLMLoadModel`/`LLMUnloadModel` nodes call the protocol~~ _Superseded, see Phase 8 T-CUT-08._
- [-] T018-I [US3] ~~Rename `OllamaLoadModel`/`OllamaUnloadModel` → `LLMLoadModel`/`LLMUnloadModel`~~ _Superseded, see Phase 8 T-CUT-05/T-CUT-06._
**Checkpoint**: superseded — see `atomic-cutover-plan.md`.
---
## Phase 6: User Story 4 — Reconnect an existing workflow after upgrading (Priority: P3)
**Goal**: A clear old→new node mapping exists, and migrated workflows are output-equivalent.
**Independent Test**: Follow the mapping to reconnect a pre-upgrade workflow; confirm equivalent output.
- [-] T019 [US4] ~~Add a migration mapping constant~~ _Superseded, see Phase 8 T-CUT-11._
- [-] T020 [US4] ~~Full suite green, SC-003 equivalence~~ _Superseded, see Phase 8 T-CUT-10._
- [-] T021 [US4] ~~Update quickstart.md migration section~~ _Superseded, see Phase 8 T-CUT-12._
**Checkpoint**: all four user stories independently functional; migration path documented.
---
## Phase 8: Atomic Node Cutover (supersedes Phases 3, 5, and T014/T019-T021)
**Goal**: execute `atomic-cutover-plan.md`'s 12-step sequenced plan as one
coordinated change — this is the actual current work; Phases 3/5's `[-]`
entries above are historical only.
**Not TDD-paired** the way earlier phases are — per the correction above,
this genuinely can't be decomposed into independent red/green pairs (a
class rename fails test *collection* for the whole file at once). Each
T-CUT step is still verified incrementally during implementation; the
suite only needs to be green as a whole at T-CUT-10, not after every step.
- [x] T-CUT-01 [P] `ollama.py`: import HTTP/cache infra from `comfydv._llm.ollama_provider` instead of duplicating it; repoint `_load_default_models()`/`/dv/ollama/models` route (plan D1)
- [x] T-CUT-02 `ollama_provider.py`: implement `OllamaProvider.list_models()`/`load_model()`/`unload_model()`/`chat()`/`chat_structured()` method bodies (ports existing inline logic; `chat_structured()` delegates to `_llm/chat.py`; `list_models()` also queries `/api/ps` to distinguish loaded/unloaded, a genuinely new capability the old `OllamaModelSelector` never had)
- [x] T-CUT-03 `tests/test_ollama_provider.py` (new file): tests for T-CUT-02, mocking at the `ollama_provider` seam (plan D4/D5)
- [x] T-CUT-04 `ollama.py`: `OllamaClient.RETURN_TYPES` → `("LLM_CLIENT",)`, `create_client()` returns `OllamaProvider(host, headers)` (plan D2)
- [x] T-CUT-05 `ollama.py`: rename the 4 classes, `"OLLAMA_CLIENT"`→`"LLM_CLIENT"` on every consumer, rewrite the 3 delegating method bodies, delete `_client_headers` (plan D6)
- [x] T-CUT-06 `src/comfydv/__init__.py`: update imports and `NODE_CLASS_MAPPINGS`/`NODE_DISPLAY_NAME_MAPPINGS`
- [x] T-CUT-07 `tests/conftest.py`: repoint `_clear_ollama_caches` and `first_generative_model`'s `_fetch_models` import (plan D8) — `first_generative_model`'s import needed no change (still re-exported from `comfydv.ollama`)
- [x] T-CUT-08 `tests/test_ollama.py`: rewrote against a `_FakeProvider` double per plan D4/D5 — 98 unit tests, all passing. Also fixed a real gap the rename surfaced: `comfy-manager-entry.json`'s `nodename` list (and its matching test expectation) still had the old display names — updated both.
- [x] T-CUT-09 [P] `contracts/llm_provider_protocol.md`: `timeout_secs` fix (commit `ef2464a`)
- [x] T-CUT-10 Full suite green (218 passed, only the pre-existing unrelated Dockerfile-python-version test fails), `ruff check --fix && ruff format` clean, `ty check` clean (confirmed the `create_model`/`RandomChoice` diagnostics pre-date this cutover via `git stash` comparison), `beacon doctor --strict` shows only pre-existing/disclosed items (`tdd-commit-discipline` — already documented as an intentional deviation; `epic-gates` — `llamacpp-integration` correctly has no specs yet)
- [x] T-CUT-11 [P] `tasks.md`/`ollama.py`: migration mapping constant (FR-009) — `MIGRATION_MAP` dict, `ollama.py`
- [x] T-CUT-12 [P] Ollama was reachable in this environment — ran the real `@pytest.mark.integration` suite (not just a manual walkthrough). 6/8 passed, including the critical ones: unreachable-host error handling, real load/unload against the live server, structured-output retry-then-raise against the live server, temperature-determinism. 2 failures (`test_single_turn_returns_non_empty_response`, `test_multi_turn_receives_context`) — confirmed via direct `curl` to `/api/chat` (bypassing this codebase entirely) that the test model (`lukey03/qwen3.5-9b-abliterated-vision`) itself returns a degenerate empty response server-side; this is the exact pre-existing model unreliability ADR-006 already documented, not a cutover regression.
**Checkpoint**: T-CUT-10 green = all four user stories functional on the generic nodes; T-CUT-11/12 close out US4.
---
## Phase 7: Polish & Cross-Cutting Concerns
_Subsumed by T-CUT-10/T-CUT-12 (Phase 8) — same work, done together with the
cutover rather than as a separate pass, since ruff/ty/doctor need to run
against the final state anyway:_
- [x] T022 [P] `ruff check --fix && ruff format` — clean (T-CUT-10)
- [x] T023 [P] `ty check` — clean, pre-existing diagnostics confirmed unrelated via `git stash` comparison (T-CUT-10)
- [x] T024 Constitution Principle IV confirmed: `src/comfydv/_llm/*.py` import no `comfy`/`server`/`folder_paths` at module scope (verified via grep)
- [x] T025 `beacon doctor --strict`: only `tdd-commit-discipline` (documented deviation, see Phase 4's US2 notes) and `epic-gates` (`llamacpp-integration` correctly has no specs yet) — both pre-disclosed, not new findings (T-CUT-10)
- [x] T026 Live-server validation done via the real `@pytest.mark.integration` suite rather than a separate manual walkthrough — Ollama was reachable in this environment (T-CUT-12); a literal `quickstart.md` click-through in ComfyUI itself is still worth doing whenever this branch is reviewed in a real ComfyUI install, but the underlying behavior is now proven against a live server
---
## Dependencies & Execution Order
### Phase Dependencies
- **Setup (Phase 1)**: no dependencies
- **Foundational (Phase 2)**: depends on Setup — BLOCKS all user stories
- **User Stories (Phase 3–6)**: all depend on Foundational; US1 has no dependency on US2/US3/US4; US2's `ChatCompletion` wiring (T014) depends on US1's node rename (T009-I); US3 is independent of US1/US2 except for sharing `OllamaProvider`'s constructor (T006); US4 depends on the node renames done in US1/US3 (T009-I, T017-I, T018-I) since it documents them
- **Polish (Phase 7)**: depends on all four user stories
### Parallel Opportunities
- T002 (package init) can run alongside T001 (dependency addition)
- T003 and T005 can run in parallel (different files/concerns) within Foundational
- T008-T (provider-level test) can run in parallel with T007-T (node-level test) — different files
- T011-T, T015-T, T016-T can each start as soon as Foundational is done, in parallel with US1 — different files, no shared dependency beyond T004/T006
- T022/T023 (lint/type-check) can run in parallel in Polish
---
## Implementation Strategy
### MVP First
1. Phase 1 (Setup) → Phase 2 (Foundational) → Phase 3 (US1) → **STOP and validate US1 independently** against a live local server.
### Incremental Delivery
1. Setup + Foundational → foundation ready.
2. US1 → validate → this alone restores basic chat parity with today's Ollama integration, on the new mechanism.
3. US2 → validate → restores structured-output parity (the ADR-006→ADR-007 migration is now complete in behavior).
4. US3 → validate → restores model-management parity.
5. US4 → validate → migration guidance ships; full regression pass (T020) confirms SC-003.
6. Polish.
Each story adds value without breaking the previous one — this mirrors the epic's own framing: US1+US2 together are the risky "prove the migration is behavior-preserving" core; US3 and US4 round out parity and upgrade experience.
---
## ⚠️ Correction (2026-07-11) — US1/US3 independence claim was wrong
**Discovered mid-implementation, confirmed by independent product + engineering
review (agent-trio deliberation, aligned verdicts):** the "US1 has no
dependency on other stories" and "US3 is independent of US1/US2" claims above
are **false**. `OllamaClient` is a single shared producer node — every
downstream node (`OllamaModelSelector`, `OllamaLoadModel`,
`OllamaUnloadModel`, `OllamaChatCompletion`) consumes its output via
`f"{client}/api/..."` string interpolation (10 call sites in `ollama.py`).
Changing `OllamaClient` to emit an `OllamaProvider` object instead of the
current string-like `OllamaClientType` breaks **all four** consumers
simultaneously — there is no way to migrate just `ChatCompletion` (US1)
while leaving `OllamaModelSelector`/`OllamaLoadModel`/`OllamaUnloadModel`
(US3) on the old string-based access pattern. Separately, renaming these
classes breaks `tests/test_ollama.py`'s imports atomically (125 references
across the file) — a class rename fails test *collection* for the whole
file at once, not test-by-test.
**Rejected fix:** making `OllamaProvider` also subclass `str` (mirroring
`OllamaClientType`'s trick) to preserve incremental per-node migration.
Both reviewers rejected this — it reintroduces the exact hack ADR-007
exists to eliminate into the new clean boundary, and would silently mask an
incomplete cutover (un-migrated consumers keep working via the string
trick, so T020's regression pass would go green for the wrong reason).
**Decision:** T007–T010 (US1) and T015–T018 (US3)'s *node-layer* work
(everything that touches `OllamaClient`'s output type or renames a node
class) must land as **one atomic cutover** — one coordinated change across
`ollama.py` and `tests/test_ollama.py`, verified green as a whole, not as
separable per-story TDD pairs. This is sized beyond a single 2–4h tracer
bullet and is explicitly re-scoped as its own dedicated BUILD session
(tracked in GitHub issue — see epic Notes for the link once filed), not
attempted in the same session as the Foundational layer (T001–T006, already
shipped safely — see git log). The *provider-layer* work that doesn't touch
`ollama.py` (e.g. `OllamaProvider.chat()`/`list_models()`/`load_model()`/
`unload_model()` method bodies, and the `pydantic-ai`-backed
`chat_structured()` helper) remains genuinely independent and safe to build
ahead of the cutover — only the ComfyUI node-layer rename is atomic.
ADR-007's own decision (breaking rename, no deprecated aliases) is
**unaffected** — that call was about user-facing blast radius (small,
Ollama integration shipped 2026-07-04), which this finding doesn't change.
What's re-scoped is delivery sequencing, not the design decision.
---
## ✅ Properly specced (2026-07-11) — see `atomic-cutover-plan.md`
Full line-by-line inventory of every affected reference in `ollama.py` and
`tests/test_ollama.py` (1820 lines, read in full), the design decisions it
surfaced (cache-singleton duplication, `client == "<string>"` equality
breaking, bare-string-client backward compat removal, and — the big one —
a test-layer split so the 35 relocated `_post_json` monkeypatches land at
the right architectural seam instead of being patched 1:1), and a
12-step sequenced task list (T-CUT-01 … T-CUT-12) that supersedes the
struck-through tasks above. That file is now the authoritative task list
for this remaining work; this file's Phase 3/5/6 entries are kept only for
history.
@@ -0,0 +1 @@
epic = "llamacpp-integration"
@@ -0,0 +1,39 @@
# Specification Quality Checklist: llama.cpp Model Integration
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-07-11
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
No [NEEDS CLARIFICATION] markers needed — scope boundaries (router-mode-only,
no GPU tuning, no auth/TLS, no Manager listing) came directly from the
parent epic's Non-goals (`project-management/Roadmap/epics/llamacpp-integration.md`)
and ADR-007. User Story 4 (swap backends without touching downstream nodes)
is the adapter pattern's central promise made concrete and testable, not
padding.
@@ -0,0 +1,44 @@
# Contract: `LlamaCppProvider` conforms to `LLMProvider`
This is the concrete proof of ADR-007's adapter pattern — the same protocol
contract documented in
`specs/007-llm-provider-abstraction/contracts/llm_provider_protocol.md`,
now with a second implementation. Nothing in that contract changes; this
file only documents `LlamaCppProvider`'s specific wire-format bindings.
```python
class LlamaCppProvider:
def __init__(self, host: str, headers: dict | None = None): ...
async def list_models(self) -> list[ModelInfo]:
"""GET {host}/models → data[] → ModelInfo(name=m["id"], status=ModelStatus(m["status"]["value"]), size=None)"""
async def load_model(self, model: str) -> None:
"""POST {host}/models/load {"model": model}"""
async def unload_model(self, model: str) -> None:
"""POST {host}/models/unload {"model": model}"""
async def chat(self, model, messages, options=None, timeout_secs=300.0) -> str:
"""POST {host}/v1/chat/completions → choices[0].message.content"""
async def chat_structured(self, model, messages, schema, options=None, timeout_secs=300.0, max_retries=2) -> BaseModel:
"""Delegates to comfydv._llm.chat.chat_structured(base_url=f"{host}/v1", ...) — identical call OllamaProvider makes"""
```
## Behavioral requirements (inherited from the protocol contract, restated for this implementation)
- `load_model`/`unload_model` MUST be idempotent. **Live-verified against a
real router-mode server**: router mode's own endpoints are *not*
idempotent — `/models/load` on an already-loaded model returns HTTP 400
`"model is already running"`, and `/models/unload` on an already-unloaded
model returns HTTP 400 `"model is not running"`, instead of `{"success": true}`.
`LlamaCppProvider` absorbs this itself: these two specific error messages
are treated as the desired end-state already reached, not a failure; any
other error still propagates.
- `list_models()` MUST NOT normalize away llama.cpp's `sleeping`/`downloading`
states (unlike `OllamaProvider`, which has no choice but to normalize —
see `research.md`).
- A `llama-server` not running in router mode (missing endpoints) MUST
surface a clear, specific error (spec.md FR-006) — not a generic
connection failure indistinguishable from "server not running at all."
@@ -0,0 +1,41 @@
# Data Model: llama.cpp Model Integration
No new types — this feature is a second implementation of the existing
`LLMProvider` protocol, `ModelStatus`, `ModelInfo`, and `Message` types
(`src/comfydv/_llm/provider.py`, unchanged). This file documents
`LlamaCppProvider`'s field mapping from llama-server's router-mode JSON onto
those existing types (see `research.md` for the verified API shapes).
## `LlamaCppProvider.list_models()` → `ModelInfo` mapping
| `ModelInfo` field | Source (`GET /models` response, per model in `data[]`) |
|---|---|
| `name` | `id` — **not** `name` (llama.cpp's field name differs from Ollama's) |
| `status` | `status.value` — nested object, not a flat string |
| `size` | Not provided by this endpoint; `None` |
`status.value` maps directly onto `ModelStatus`'s five values
(`unloaded`/`loading`/`loaded`/`sleeping`/`downloading`) — llama.cpp's
vocabulary is exactly `ModelStatus`'s full set, so unlike `OllamaProvider`
(which normalizes into a narrower subset), `LlamaCppProvider` needs no
approximation. A `"failed": true` state exists outside this vocabulary
(model process crashed) — out of scope per spec.md's edge cases; treated as
whatever `status.value` reports rather than added as a sixth enum value.
## `LlamaCppProvider.load_model()` / `unload_model()`
Both `POST /models/load` and `POST /models/unload` take `{"model": <id>}` —
the same `id` string `list_models()` returns as `ModelInfo.name`. No mapping
ambiguity here (unlike Ollama, where load/unload uses `/api/generate`'s
`keep_alive` side effect rather than a dedicated endpoint).
## `LlamaCppProvider.chat()` / `chat_structured()`
Both reach `llama-server`'s OpenAI-compatible `/v1/chat/completions` —
`chat_structured()` calls the existing shared `comfydv._llm.chat.chat_structured()`
helper unchanged (`base_url=f"{self.host}/v1"`, matching `OllamaProvider`'s
own call exactly). `chat()` parses the response as
`choices[0].message.content` (OpenAI shape), not Ollama's native
`message.content` — the two providers' non-structured paths differ here
because llama-server doesn't have an Ollama-style native `/api/chat`
endpoint to prefer instead.
@@ -0,0 +1,11 @@
Feature: US1 — Connect to a local llama.cpp server and get chat responses
Scenario: llama.cpp connection node feeds the existing chat node
Given a running local llama-server (router mode) and a workflow with a llama.cpp connection node wired into the existing chat node
When the workflow executes
Then the chat node returns the model's text response
Scenario: Unreachable llama.cpp server surfaces a clear error
Given the llama.cpp connection node configured with an unreachable server address
When the workflow executes
Then the chat node reports a clear connection error
@@ -0,0 +1,11 @@
Feature: US2 — Get structured, validated output from llama.cpp
Scenario: Valid structured response exposes typed fields, same as Ollama
Given a chat node connected to llama.cpp with structured output enabled and a valid schema
When the workflow executes and the model responds correctly
Then each schema field is available as its own typed output, and no required field is blank
Scenario: Invalid response retries then fails clearly, same as Ollama
Given a llama.cpp-hosted model that returns invalid or incomplete structured output
When the workflow executes
Then the node retries automatically and, if still unsuccessful, fails with a clear error
@@ -0,0 +1,16 @@
Feature: US3 — See and control which models are loaded on llama.cpp
Scenario: List models with full status vocabulary
Given a running local llama-server with at least one available model
When a workflow author uses the model-listing node
Then they see each available model along with its current status, drawn from llama.cpp's full status vocabulary
Scenario: Load a model into memory
Given a model that is not currently loaded
When a workflow author runs the load-model node against it
Then the model becomes loaded and is then usable by the chat node
Scenario: Unload a model from memory
Given a model that is loaded and idle
When a workflow author runs the unload-model node against it
Then the model is freed from memory and its reported status updates accordingly
@@ -0,0 +1,6 @@
Feature: US4 — Swap from Ollama to llama.cpp without touching the rest of the workflow
Scenario: Replacing only the connection node preserves the workflow
Given a workflow with chat/model-management nodes wired to an Ollama connection node
When a workflow author replaces only the connection node with a llama.cpp one, pointed at a running llama-server
Then the workflow runs successfully with no changes to any other node
+117
View File
@@ -0,0 +1,117 @@
# Implementation Plan: llama.cpp Model Integration
**Branch**: `008-llamacpp-integration` | **Date**: 2026-07-11 | **Spec**: [spec.md](./spec.md)
**Input**: Feature specification from `/specs/008-llamacpp-integration/spec.md`
**Note**: This template is filled in by the `/speckit-plan` command. See `.specify/templates/plan-template.md` for the execution workflow.
## Summary
Implement `LlamaCppProvider` as the second `LLMProvider` (ADR-007), backed by
`llama-server`'s router mode (`GET /models`, `POST /models/load`,
`POST /models/unload`, `/v1/chat/completions`). Add one new ComfyUI node
(`LlamaCppClient`) emitting the existing `LLM_CLIENT` socket type — no other
node classes change. This is the concrete proof the provider abstraction
(prerequisite epic, PR #17) actually generalizes: a second backend, zero
changes to `ChatCompletion`/`LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel`.
## Technical Context
**Language/Version**: Python ≥3.11 (unchanged, per `pyproject.toml`)
**Primary Dependencies**: `aiohttp` (existing — model-management REST calls),
`pydantic-ai`/`openai` (existing, from the prerequisite epic — `chat_structured()`
reuses the shared helper unchanged, zero new structured-output code)
**Storage**: N/A — no persistent storage; reuses the existing
`_MODEL_LIST_CACHE`/`_CHAT_RESPONSE_CACHE` infra pattern from `OllamaProvider`
**Testing**: `pytest` via `uv run pytest`, following `tests/test_ollama_provider.py`'s
established convention (mock at the provider's own `_post_json`/`_get_json`
seam, no live server required for unit tests)
**Target Platform**: ComfyUI custom-node runtime, same as the existing Ollama
integration
**Project Type**: Library / ComfyUI custom-node pack (single project, adds to
existing `src/comfydv/` layout)
**Performance Goals**: No new numeric target; must not add latency beyond
what `OllamaProvider`'s equivalent methods already accept
**Constraints**: Router-mode-only (spec.md Assumptions — a `llama-server`
without `--models-dir`/`--models-preset` doesn't expose these endpoints at
all, FR-006); model identifier field is `id` (llama.cpp) vs `name` (Ollama) —
`LlamaCppProvider.list_models()` must map this correctly (see `research.md`);
`status` is a nested object (`{"value": "..."}`), not a flat string
**Scale/Scope**: One new class (`LlamaCppProvider`, mirrors `OllamaProvider`'s
shape), one new ComfyUI node (`LlamaCppClient`), one new test file — no
changes to `ollama.py`, `_llm/provider.py`, `_llm/chat.py`, or any existing
node class
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Principle | Verdict | Notes |
|---|---|---|
| I. ComfyUI Contract First | PASS | `LlamaCppClient` exposes the standard `INPUT_TYPES`/`RETURN_TYPES`/`FUNCTION`/`CATEGORY`; registered in `NODE_CLASS_MAPPINGS` like every other node. |
| II. Sandbox All User-Supplied Code | N/A | No template/expression evaluation in this feature. |
| III. Test-First | PASS (binding) | `tests/test_llamacpp_provider.py` written test-first, mirroring `test_ollama_provider.py`'s TDD-pair structure. |
| IV. Graceful Degradation Outside ComfyUI | PASS (binding) | `LlamaCppProvider` lives in `src/comfydv/_llm/`, which already has no `comfy`/`server` imports at module scope (verified for the prerequisite epic; this feature adds no new module-scope imports of either). |
| V. Simplicity — Function Before Class | PASS, same justification as `OllamaProvider` | `LlamaCppProvider` carries connection state (host, headers) across 5 methods — the same shared-state condition that already justified `OllamaProvider` as a class (research.md, prerequisite epic). No new gate — same precedent applies. |
| VI. Fixed Output Positions | N/A | `LlamaCppClient`'s single output (`client`) isn't a multi-output node; no positional contract to preserve. |
Re-checked post-Phase 1 design (data-model.md): unchanged — no new gate
violations. No Complexity Tracking entries needed (unlike the prerequisite
epic, this feature introduces no new pattern, just a second instance of an
already-justified one).
## Project Structure
### Documentation (this feature)
```text
specs/008-llamacpp-integration/
├── plan.md # This file
├── research.md # Phase 0 — router-mode API shape, verified live
├── data-model.md # Phase 1 — LlamaCppProvider field mapping
├── quickstart.md # Phase 1 — minimal workflow walkthrough
├── contracts/ # Phase 1 — LlamaCppProvider's protocol conformance
└── tasks.md # Phase 2 (/speckit-tasks)
```
### Source Code (repository root)
```text
src/comfydv/
├── ollama.py # unchanged — add LlamaCppClient node only via a new module
├── llamacpp.py # new — LlamaCppClient node (mirrors OllamaClient's shape)
├── _llm/
│ ├── provider.py # unchanged — LLMProvider/ModelStatus/ModelInfo/Message
│ ├── ollama_provider.py # unchanged
│ ├── llamacpp_provider.py # new — LlamaCppProvider (mirrors ollama_provider.py's shape)
│ └── chat.py # unchanged — chat_structured() reused as-is
└── __init__.py # add LlamaCppClient import + NODE_CLASS_MAPPINGS entry
tests/
├── test_ollama_provider.py # unchanged
├── test_llamacpp_provider.py # new — mirrors test_ollama_provider.py's structure
└── test_llamacpp.py # new — LlamaCppClient node contract test (small; mirrors
# the OllamaClient-specific slice of test_ollama.py)
```
**Structure Decision**: New `src/comfydv/llamacpp.py` module (not added into
`ollama.py`) for the `LlamaCppClient` node, and a new `src/comfydv/_llm/llamacpp_provider.py`
for `LlamaCppProvider` — mirroring the existing `ollama.py`/`ollama_provider.py`
split exactly, so the two backends read as parallel, symmetric implementations
rather than one growing to accommodate the other. No existing file is
modified except `__init__.py`'s registration block.
## Complexity Tracking
> **Fill ONLY if Constitution Check has violations that must be justified**
None — see Constitution Check above.
@@ -0,0 +1,26 @@
# Quickstart: llama.cpp Model Integration
## Prerequisite
Launch `llama-server` in router mode:
```bash
llama-server --models-dir ./models -c 8192
```
## Minimal workflow
1. Add an **LlamaCpp Client** node. Set its host widget (default
`http://localhost:8080`, llama-server's default port).
2. Wire it into a **Chat Completion** node — the exact same node used for
Ollama. Set a model and prompt, run.
3. Structured output, model listing, and load/unload all work exactly as
documented for Ollama in the main README/quickstart — swap the client
node, nothing else changes.
## Swapping an existing Ollama workflow to llama.cpp
Replace the **Ollama Client** node with an **LlamaCpp Client** node, pointed
at your running `llama-server`. Every downstream node (Chat Completion, LLM
Model Selector, LLM Load Model, LLM Unload Model) keeps working unmodified —
this is the whole point of the provider abstraction (ADR-007).
@@ -0,0 +1,78 @@
# Research: llama.cpp Model Integration
## Decision: exact router-mode API shape (verified against `ggml-org/llama.cpp`'s live `tools/server/README.md`, not assumed)
llama.cpp's router mode postdates this session's training data — verified live
against the authoritative source rather than guessed, since getting field
names wrong here would silently produce broken code (wrong key = `KeyError`
or silent `None`, not an obvious failure).
**`GET /models`** response:
```json
{
"data": [
{
"id": "ggml-org/gemma-3-4b-it-GGUF:Q4_K_M",
"path": "/Users/.../gemma-3-4b-it-Q4_K_M.gguf",
"status": {
"value": "loaded",
"args": ["llama-server", "-ctx", "4096"]
},
"architecture": {
"input_modalities": ["text", "image"],
"output_modalities": ["text"]
}
}
]
}
```
**Two details that would have been wrong by assumption:**
1. The model identifier field is **`id`**, not `name` — different from Ollama's
`/api/tags`, which uses `name`. `OllamaProvider.list_models()` maps
`m["name"]`; `LlamaCppProvider.list_models()` must map `m["id"]` instead.
2. **`status` is a nested object** (`{"value": "loaded", ...}`), not a flat
string field. `LlamaCppProvider.list_models()` must read
`m["status"]["value"]`, not `m["status"]` directly. A `"failed"` state
also exists (`{"failed": true, "exit_code": ...}`) outside the five
`ModelStatus` values the protocol defines — not handled by this feature
(see Non-goals/edge cases in `spec.md`); a failed model is reported as
whatever `status.value` degrades to rather than added as a sixth enum
value, keeping `ModelStatus` unchanged across both providers.
**`POST /models/load`** and **`POST /models/unload`**: identical request
shape, `{"model": "<id>"}` (using the same `id` string from `GET /models`,
despite the request field being named `model` not `id`). Response:
`{"success": true}`.
**CLI**: `--models-dir <path>` or `--models-preset <path>.ini` — a deployment
prerequisite (spec.md Assumptions), not something comfydv configures.
## Decision: `chat_structured()` needs zero new code
`llama-server`'s `/v1/chat/completions` is OpenAI-compatible (the same
assumption ADR-007 made when adopting `pydantic-ai`). `LlamaCppProvider.chat_structured()`
calls the exact same `comfydv._llm.chat.chat_structured()` helper
`OllamaProvider` already uses, with `base_url=f"{self.host}/v1"` — the only
per-provider difference. This is the concrete proof the shared mechanism
generalizes (spec.md User Story 2/FR-004), not just an assumption.
## Decision: `chat()` (non-structured) also reuses the OpenAI-compatible endpoint
Unlike Ollama (which has both a native `/api/chat` and an OpenAI-compat
`/v1/chat/completions`), llama-server's primary chat endpoint is the
OpenAI-compatible one. `LlamaCppProvider.chat()` POSTs to
`{host}/v1/chat/completions` (via the existing `_post_json` helper, no new
HTTP client) rather than mirroring Ollama's native-endpoint choice — the
response shape (`choices[0].message.content`) differs from Ollama's native
`message.content` and must be parsed accordingly.
## Decision: no protocol changes needed
`LLMProvider`'s five methods (`list_models`/`load_model`/`unload_model`/
`chat`/`chat_structured`) already cover everything router mode needs — this
was the actual point of designing the protocol at the operation level in
ADR-007, and this research confirms it held up against llama.cpp's real API,
not just Ollama's.
+110
View File
@@ -0,0 +1,110 @@
# Feature Specification: llama.cpp Model Integration
**Feature Branch**: `008-llamacpp-integration`
**Created**: 2026-07-11
**Status**: Draft
**Input**: User description: "Add ComfyUI nodes for llama.cpp local inference via llama-server's router mode, implementing the LlamaCppProvider as the second LLMProvider (ADR-007) alongside the existing OllamaProvider. Router mode exposes GET /models (with live status), POST /models/load, POST /models/unload, giving llama.cpp the same manual load/unload memory-management primitives as Ollama. No new ComfyUI node classes needed for chat/model-selection/load/unload — only a new LlamaCppClient config node; the existing generic ChatCompletion/LLMModelSelector/LLMLoadModel/LLMUnloadModel nodes work unchanged once wired to it."
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Connect to a local llama.cpp server and get chat responses (Priority: P1) 🎯 MVP
As a ComfyUI workflow author running `llama-server` locally, I want a connection node for it — just like the one I already use for Ollama — so I can get chat responses from a llama.cpp-hosted model using the same chat node I already know.
**Why this priority**: This is the entire point of the feature and the proof that the provider abstraction (shipped in the prerequisite epic) actually works: a second backend, zero changes to the chat node.
**Independent Test**: Wire a new llama.cpp connection node into the existing chat node, run against a local `llama-server` (router mode), confirm a text response.
**Acceptance Scenarios**:
1. **Given** a running local `llama-server` (router mode) and a workflow with a llama.cpp connection node wired into the existing chat node, **When** the workflow executes, **Then** the chat node returns the model's text response — using the exact same chat node a workflow author already uses for Ollama.
2. **Given** the llama.cpp connection node configured with an unreachable server address, **When** the workflow executes, **Then** the chat node reports a clear connection error, matching the behavior workflow authors already know from the Ollama connection.
---
### User Story 2 - Get structured, validated output from llama.cpp (Priority: P1)
As a workflow author, I want structured output (a schema-validated response instead of free text) to work identically regardless of whether I'm connected to Ollama or llama.cpp, so I don't have to relearn or rebuild anything when switching backends.
**Why this priority**: Structured output is a core existing capability (already proven for Ollama); this story proves the shared mechanism genuinely generalizes rather than being Ollama-specific in practice, not just in name.
**Independent Test**: Enable structured output on the chat node with a schema, run against a llama.cpp-hosted model, confirm each schema field is populated and never blank — using the same steps as the equivalent Ollama test.
**Acceptance Scenarios**:
1. **Given** a chat node connected to llama.cpp with structured output enabled and a valid schema, **When** the workflow executes and the model responds correctly, **Then** each schema field is available as its own typed output, and no required field is blank.
2. **Given** a llama.cpp-hosted model that returns invalid or incomplete structured output, **When** the workflow executes, **Then** the node retries automatically and, if still unsuccessful, fails with a clear error — identical behavior to the Ollama path.
---
### User Story 3 - See and control which models are loaded on llama.cpp (Priority: P2)
As a workflow author running models locally, I want to see live model status (including whether a model is currently loading or being downloaded, not just loaded/unloaded) and explicitly load or unload a model on my llama.cpp server, so I can manage memory the same way I already do for Ollama — with more visibility, since llama.cpp's router mode reports richer status than Ollama does.
**Why this priority**: Valuable and proves the model-management path generalizes too, but a workflow can still run chat completions without ever calling load/unload explicitly (the server can load on first use), so it's lower risk to defer than basic chat.
**Independent Test**: Use the existing model-listing node against a running `llama-server`, confirm it shows each available model with its current status (including `loading`/`downloading` if applicable); use the existing load/unload nodes against one model and confirm its status changes.
**Acceptance Scenarios**:
1. **Given** a running local `llama-server` with at least one available model, **When** a workflow author uses the model-listing node, **Then** they see each available model along with its current status, drawn from llama.cpp's full status vocabulary (not just loaded/unloaded).
2. **Given** a model that is not currently loaded, **When** a workflow author runs the load-model node against it, **Then** the model becomes loaded and is then usable by the chat node.
3. **Given** a model that is loaded and idle, **When** a workflow author runs the unload-model node against it, **Then** the model is freed from memory and its reported status updates accordingly.
---
### User Story 4 - Swap from Ollama to llama.cpp without touching the rest of the workflow (Priority: P3)
As a workflow author with an existing Ollama-based workflow, I want to switch it to llama.cpp by changing only the connection node, so I don't have to rebuild my chat/model-management logic for a second backend.
**Why this priority**: This is the adapter pattern's actual promise made concrete for a user, but it's a validation/demonstration story rather than new capability — everything it depends on is already covered by User Stories 1–3.
**Independent Test**: Take a workflow using the Ollama connection node, replace it with the llama.cpp connection node (same downstream nodes, no other changes), run it, confirm it still works.
**Acceptance Scenarios**:
1. **Given** a workflow with chat/model-management nodes wired to an Ollama connection node, **When** a workflow author replaces only the connection node with a llama.cpp one (pointed at a running `llama-server`), **Then** the workflow runs successfully with no changes to any other node.
---
### Edge Cases
- What happens when `llama-server` is running but was launched without router mode (i.e. with `-m` instead of `--models-dir`)? The router-mode-only endpoints this feature depends on won't exist — the connection/model-management nodes should fail with a clear error, not hang or silently return empty results.
- What happens when the configured server address is unreachable at the moment a model-listing, load, or unload node runs (not just the chat node)?
- What happens when llama.cpp reports a model status this feature doesn't expect (a router-mode API change)? Should degrade gracefully (surface the status if recognized, don't crash on an unrecognized one), not silently misreport.
- What happens to an in-flight chat request if the model it depends on is unloaded by another node in the same workflow run? (Same question already answered for Ollama — behavior should be consistent.)
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: The system MUST allow a workflow author to configure a connection to a local `llama-server` (router mode) the same way they already configure a connection to Ollama — a dedicated connection node, reusable across multiple nodes in a workflow.
- **FR-002**: The system MUST NOT require any new or different node classes for chat, structured output, model listing, or load/unload when using llama.cpp — the existing generic nodes MUST work unchanged once connected to a llama.cpp connection node.
- **FR-003**: The system MUST report each model's status using llama.cpp's full status vocabulary (unloaded, loading, loaded, sleeping, downloading) when connected to llama.cpp — not degraded to the narrower Ollama-compatible set.
- **FR-004**: The system's chat and structured-output behavior MUST be identical between Ollama and llama.cpp connections, given equivalent inputs — same retry limits, same validation rules, same error conditions (this is the direct continuation of the prerequisite epic's own FR-007/FR-008).
- **FR-005**: The system MUST allow a workflow author to explicitly load a model into memory and explicitly unload a model from memory on a connected llama.cpp server.
- **FR-006**: The system MUST surface a clear, specific error when connected to a `llama-server` instance that isn't running in router mode (the endpoints this feature needs don't exist), rather than an unhelpful generic failure.
### Key Entities *(include if feature involves data)*
- **llama.cpp connection**: A configured connection to a local `llama-server` instance running in router mode (host + any authentication), implementing the same connection concept already established for Ollama.
- **Model status**: Reuses the existing status concept from the prerequisite feature, now populated with llama.cpp's full vocabulary rather than a narrowed subset.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: A workflow author can connect to a llama.cpp server and get a chat response using the same node count and shape as connecting to Ollama (one connection node, one chat node) — no new nodes to learn for the chat path.
- **SC-002**: An existing workflow can be repointed from Ollama to llama.cpp by changing exactly one node (the connection node) — zero edits to any chat or model-management node.
- **SC-003**: Structured-output workflows behave identically (same validation guarantees, zero blank-required-field results) regardless of which backend is connected.
- **SC-004**: Model status reporting for llama.cpp surfaces all five status values where applicable — a strictly richer view than what Ollama can report through the same interface.
## Assumptions
- Workflow authors run their own local `llama-server` instance, launched in router mode (`--models-dir` or `--models-preset`), reachable over HTTP from the machine running ComfyUI; this feature does not install, configure, or launch that server.
- Non-router-mode `llama-server` usage (a single model launched with `-m`) is out of scope — router mode is required for the load/unload/status parity with Ollama that is this feature's whole point.
- GPU inference optimisation, quantisation tuning, authentication/TLS, and ComfyUI Manager registry listing are out of scope, consistent with the prerequisite Ollama epic's own non-goals.
- The `LLMProvider` protocol and generic nodes (`ChatCompletion`, `LLMModelSelector`, `LLMLoadModel`, `LLMUnloadModel`) already exist and are not modified by this feature — if llama.cpp's router mode needs a protocol capability that doesn't exist yet, that is a protocol change scoped as its own follow-up, not silently special-cased here.
+224
View File
@@ -0,0 +1,224 @@
# Tasks: llama.cpp Model Integration
**Input**: Design documents from `/specs/008-llamacpp-integration/`
**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/llamacpp_provider_conformance.md
**Tests**: First-class — every implementation task has a paired failing-test task (`-T`/`-I` suffix).
**Organization**: Grouped by user story (spec.md priorities P1/P1/P2/P3).
## Format: `[ID] [P?] [Story] Description`
- **[P]**: Can run in parallel (different files, no dependencies)
- **[Story]**: US1–US4
- **-T / -I**: paired test (red) / implementation (green)
## Path Conventions
Single project: `src/comfydv/`, `tests/` at repository root, mirroring the
`ollama.py`/`_llm/ollama_provider.py` split exactly (plan.md Structure
Decision).
---
## Phase 1: Setup
- [x] T001 No new dependencies — `aiohttp`/`pydantic-ai` already present from the prerequisite epic (verified in `pyproject.toml`)
- [x] T002 [P] Create `src/comfydv/_llm/llamacpp_provider.py` and `src/comfydv/llamacpp.py` (empty modules with docstrings, mirroring `ollama_provider.py`/`ollama.py`'s module docstring style)
---
## Phase 2: Foundational
None — `LLMProvider`, `ModelStatus`, `ModelInfo`, `Message`, and the shared
`chat_structured()` helper already exist from the prerequisite epic and are
unmodified by this feature (plan.md Constitution Check, research.md).
**Checkpoint**: nothing blocks user story work — it can start immediately.
---
## Phase 3: User Story 1 — Connect to a local llama.cpp server and get chat responses (Priority: P1) 🎯 MVP
**Goal**: A workflow author wires an `LlamaCppClient` node into the existing `ChatCompletion` node and gets a text response.
**Independent Test**: Wire `LlamaCppClient` → `ChatCompletion`, run against a live `llama-server` (router mode), confirm text output.
- [x] T003-T [US1] Write FAILING test: `LlamaCppProvider.chat()` POSTs to `{host}/v1/chat/completions` and parses `choices[0].message.content`, in `tests/test_llamacpp_provider.py` (witnesses `features/us1_connect_and_chat.feature` scenario "llama.cpp connection node feeds the existing chat node")
- [x] T003-I [US1] Implement `LlamaCppProvider.__init__`/`.chat()` in `src/comfydv/_llm/llamacpp_provider.py` (data-model.md — OpenAI-shape response parsing, not Ollama's native shape) — makes T003-T pass
- [x] T004-T [US1] Write FAILING test: `LlamaCppClient` node's `INPUT_TYPES`/`RETURN_TYPES` match `OllamaClient`'s shape (`LLM_CLIENT` output), and `create_client()` constructs a `LlamaCppProvider`, in `tests/test_llamacpp.py`
- [x] T004-I [US1] Implement `LlamaCppClient` node in `src/comfydv/llamacpp.py` (mirrors `OllamaClient` exactly, default host `http://localhost:8080` per llama-server's default port) — makes T004-T pass (depends on T003-I)
- [x] T005-T [US1] Write FAILING test: `LlamaCppClient` registered in `NODE_CLASS_MAPPINGS`/`NODE_DISPLAY_NAME_MAPPINGS`, in `tests/test_llamacpp.py`
- [x] T005-I [US1] Register `LlamaCppClient` in `src/comfydv/__init__.py` — makes T005-T pass (depends on T004-I)
- [x] T006-T [US1] Write FAILING test: `LlamaCppProvider` connection error surfaces a clear message (mirrors `OllamaProvider`'s `_post_json` connection-error contract), in `tests/test_llamacpp_provider.py` (witnesses `features/us1_connect_and_chat.feature` scenario "Unreachable llama.cpp server surfaces a clear error")
- [x] T006-I [US1] Confirm `LlamaCppProvider.chat()` reuses the shared `_post_json` connection-error handling unchanged (likely no code change needed — verify, don't assume) — makes T006-T pass
**Checkpoint**: US1 fully functional and independently testable (MVP) — proves the adapter pattern for the chat path.
---
## Phase 4: User Story 2 — Get structured, validated output from llama.cpp (Priority: P1)
**Goal**: `structured_output=True` on `ChatCompletion` works identically against llama.cpp.
**Independent Test**: Enable `structured_output` with a schema, run against a llama.cpp-hosted model, confirm typed sockets populate and are never blank.
- [x] T007-T [US2] Write FAILING test: `LlamaCppProvider.chat_structured()` builds `base_url=f"{host}/v1"` and delegates to the shared `comfydv._llm.chat.chat_structured()` helper unchanged, in `tests/test_llamacpp_provider.py` (witnesses `features/us2_structured_output.feature` scenario "Valid structured response exposes typed fields, same as Ollama")
- [x] T007-I [US2] Implement `LlamaCppProvider.chat_structured()` in `src/comfydv/_llm/llamacpp_provider.py` — zero new structured-output logic, same call shape `OllamaProvider.chat_structured()` already makes — makes T007-T pass
- [x] T008 [US2] No new test needed for the retry-then-fail path (witnesses `features/us2_structured_output.feature` scenario "Invalid response retries then fails clearly, same as Ollama") — already fully covered by `tests/test_llm_chat_structured.py`'s existing suite, since `LlamaCppProvider.chat_structured()` calls the identical shared helper `OllamaProvider` does; re-testing it here would duplicate coverage without adding confidence (same reasoning as the prerequisite epic's D5)
**Checkpoint**: US1 + US2 both independently functional — the chat surface is now backend-agnostic in practice, not just in name.
---
## Phase 5: User Story 3 — See and control which models are loaded on llama.cpp (Priority: P2)
**Goal**: `LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel` work against llama.cpp via `LlamaCppProvider`.
**Independent Test**: List models via `LLMModelSelector` wired to `LlamaCppClient`; load/unload one; confirm status changes, including `loading`/`downloading` states if triggered.
- [x] T009-T [P] [US3] Write FAILING test: `LlamaCppProvider.list_models()` maps `GET /models`'s `data[].id`→`ModelInfo.name` and `data[].status.value`→`ModelInfo.status`, surfacing all five `ModelStatus` values without normalization (data-model.md), in `tests/test_llamacpp_provider.py` (witnesses `features/us3_model_lifecycle.feature` scenario "List models with full status vocabulary")
- [x] T009-I [US3] Implement `LlamaCppProvider.list_models()` in `src/comfydv/_llm/llamacpp_provider.py` — makes T009-T pass
- [x] T010-T [P] [US3] Write FAILING test: `LlamaCppProvider.load_model()`/`unload_model()` POST `{"model": id}` to `/models/load`/`/models/unload` and are idempotent, in `tests/test_llamacpp_provider.py` (witnesses `features/us3_model_lifecycle.feature` scenarios "Load a model into memory" and "Unload a model from memory")
- [x] T010-I [US3] Implement `LlamaCppProvider.load_model()`/`unload_model()` in `src/comfydv/_llm/llamacpp_provider.py` — makes T010-T pass
- [x] T011 [US3] No new node-layer tests needed — `LLMModelSelector`/`LLMLoadModel`/`LLMUnloadModel` are untouched by this epic (plan.md Structure Decision) and already have delegation-test coverage against a generic `_FakeProvider` in `tests/test_ollama.py`; that coverage is provider-agnostic by construction (FR-002), so it already proves these nodes work with `LlamaCppProvider` too, not just `OllamaProvider`
**Checkpoint**: US1 + US2 + US3 independently functional.
---
## Phase 6: User Story 4 — Swap from Ollama to llama.cpp without touching the rest of the workflow (Priority: P3)
**Goal**: Demonstrate/prove the adapter pattern's actual promise end-to-end.
**Independent Test**: Same workflow, only the connection node changes.
- [x] T012-T [US4] Write FAILING test: a workflow-shaped test (client → `ChatCompletion` → `LLMModelSelector` → `LLMLoadModel` → `LLMUnloadModel`) runs identically whether `client` is an `OllamaProvider`-double or a `LlamaCppProvider`-double — i.e. no node branches on provider type, in `tests/test_llamacpp.py` (witnesses `features/us4_swap_backends.feature` scenario "Replacing only the connection node preserves the workflow")
- [x] T012-I [US4] No implementation expected — this test should already pass given T003-T011 (it's a regression/integration proof, not new functionality); if it fails, that reveals a node secretly branching on provider type, which would be a real bug to fix, not a feature to add
**Checkpoint**: all four user stories independently functional; the adapter pattern is proven end-to-end, not just asserted.
---
## Phase 7: Polish & Cross-Cutting Concerns
- [x] T013 [P] `ruff check --fix && ruff format` — clean
- [x] T014 [P] `ty check` — clean, same pre-existing diagnostics as the prerequisite epic (unrelated to this feature — `comfy`/`server`/`folder_paths` unresolved-import, `format_string.py`'s dynamic RETURN_TYPES, `create_model`/`RandomChoice` — none touch the new files)
- [x] T015 Confirmed via grep: `llamacpp_provider.py`/`llamacpp.py` import no `comfy`/`server`/`folder_paths` at module scope
- [x] T016 `beacon doctor --strict`: only the pre-existing `llm-provider-abstraction: all specs [complete]` epic-gates item (PR #18, the archive-bookkeeping PR for the *prerequisite* epic, not yet merged — unrelated to this feature) and `tdd-commit-discipline` (disclosed pattern, same reasoning as the prerequisite epic)
- [x] T017 Live smoke test — run against a real router-mode `llama-server` (Homebrew-installed, already present in the dev environment; a prior pass wrongly assumed no server was reachable without actually checking). Full lifecycle exercised against a real 5.6GB local GGUF model: `list_models()` → `load_model()` → `chat()` → `unload_model()`, plus explicit idempotency checks (calling `load_model()`/`unload_model()` again in the already-satisfied state). Found and fixed a real gap not caught by the mocked suite — see the finding below.
---
## Dependencies & Execution Order
### Phase Dependencies
- **Setup (Phase 1)**: no dependencies
- **Foundational (Phase 2)**: none — nothing blocks user story work
- **US1**: no dependency on other stories — genuinely the MVP
- **US2**: depends on US1's `LlamaCppProvider` skeleton existing (T003-I), but its own logic (T007) has no dependency on US1's chat() specifically
- **US3**: independent of US1/US2 except sharing `LlamaCppProvider`'s constructor (T003-I) — unlike the prerequisite epic's atomic cutover, there is no shared "client output type" migration risk here, since `LlamaCppClient` is a brand-new node, not a changed one
- **US4**: depends on US1–US3 all being done (it's a proof, not new functionality)
- **Polish**: depends on all four user stories
### Parallel Opportunities
- T002 can start immediately
- T009-T and T010-T can run in parallel (different methods, same file, no shared state)
- T013/T014 can run in parallel in Polish
---
## Implementation Strategy
### MVP First
1. Phase 1 (Setup, trivial) → Phase 3 (US1) → **STOP and validate US1 independently** against a live `llama-server`.
### Incremental Delivery
1. US1 → validate → basic chat parity with Ollama, on a second backend.
2. US2 → validate → structured-output parity — the shared mechanism holds.
3. US3 → validate → model-management parity, with richer status than Ollama can offer.
4. US4 → validate → the adapter pattern is proven, not just asserted.
5. Polish.
Unlike the prerequisite epic, **this decomposition genuinely holds** —
there is no shared "output type" migration forcing an atomic cutover, because
`LlamaCppClient` is new, not a change to an existing node. Each phase really
can land independently.
---
## Post-implementation review finding (fixed)
A `beacon-reviewer` pass ahead of PR open found `LlamaCppProvider.list_models()`
caught *every* exception and returned `[]`, silently indistinguishable from
"no models installed" — violating FR-006 and `contracts/llamacpp_provider_conformance.md`'s
explicit requirement that a non-router-mode `llama-server` (unreachable
endpoints → HTTP error on `GET /models`) surface a clear, specific error.
Fixed: `_get_json` (shared with `OllamaProvider`, in `ollama_provider.py`) now
raises `RuntimeError` on an HTTP error status, matching `_post_json`'s
existing behavior — its docstring already claimed this, it just didn't do it.
`LlamaCppProvider.list_models()` now distinguishes `OSError` (genuinely
unreachable — connection refused, DNS failure, timeout; all aiohttp
connection-level exceptions are `OSError` subclasses) from `RuntimeError`
(server responded, but with an error): the former still degrades gracefully
to `[]` (consistent with `OllamaProvider`'s existing UX), the latter is
re-raised naming router mode as the likely cause. Regression test added:
`test_list_models_non_router_mode_raises_clear_error` in
`tests/test_llamacpp_provider.py`. `OllamaProvider`'s own `list_models()`/
`_fetch_models()` still catch broadly and degrade to `[]` unchanged — no
spec requirement asks Ollama to make this distinction, and this fix doesn't
force it to.
---
## Live smoke test finding (T017, fixed)
T017 had been marked `[-]` deferred on the assumption that no `llama-server`
was reachable in the dev environment. That assumption was never actually
checked — `llama-server` was installed via Homebrew the whole time, and a
router-mode server was launched against a real local GGUF model
(`--models-dir` pointed at a symlinked model file) to run the smoke test for
real.
This caught a genuine gap the mocked suite couldn't: `contracts/llamacpp_provider_conformance.md`
claimed router mode's `/models/load`/`/models/unload` return `{"success": true}`
on an already-loaded/unloaded model, satisfying the `LLMProvider` protocol's
idempotency requirement "without extra handling." That claim was never
live-verified — live testing showed the opposite: both endpoints return HTTP
400 (`"model is already running"` / `"model is not running"`) instead.
`LlamaCppProvider.load_model()`/`unload_model()` now absorb exactly those two
error messages as the desired end-state already reached (any other error
still propagates); the contract doc is corrected to describe the real
behavior. Regression tests added (mocked, so they run in CI):
`test_load_model_already_running_is_idempotent`,
`test_unload_model_not_running_is_idempotent`, and their
`_other_http_error_still_raises` counterparts confirming non-idempotency
errors aren't over-broadly swallowed.
Also observed live (informational, no code change needed): `load_model()`/
`unload_model()` return once the request is *accepted*, not once the state
transition completes — a 5.6GB model reported `LOADING` for several seconds
before `LOADED`. This matches `ModelStatus`'s documented vocabulary (`loading`
is a real, intended state) and how a real UI would behave — fire the request,
poll `list_models()` for the transition. No protocol change; noted here so
it's not mistaken for a future bug report.
**`chat_structured()` live-verified separately** (T007/T008's mocked coverage
only ever exercised the call-shape, never the real network path): ran a
second live smoke test — `load_model()` → `chat_structured()` with a real
`pydantic.BaseModel` schema — against the same router-mode server. Result
validated correctly (`Color(name='Red', hex_code='#FF0000')`), confirming
pydantic-ai's `Agent`/`OpenAIProvider(base_url=f"{host}/v1")` mechanism
genuinely works against llama-server's OpenAI-compatible endpoint, not just
Ollama's (which was the only one live-verified in the prerequisite epic).
No gap found here — recorded as verification evidence, not a fix.
**Net result**: every `LlamaCppProvider` method (`list_models`, `load_model`,
`unload_model`, `chat`, `chat_structured`) has now been exercised against a
real router-mode `llama-server`, not just mocks. T017 is genuinely done.
+1
View File
@@ -0,0 +1 @@
epic = "vlm-image-input"
@@ -0,0 +1,41 @@
# Specification Quality Checklist: VLM Image Input for ChatCompletion
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-07-22
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
- The cross-provider "where does the image live / who translates it" decision
is intentionally kept out of the spec (WHAT/WHY) and recorded in
ADR-008 (HOW). The spec references it via the epic, not inline.
- Multi-image-per-turn is documented as an out-of-MVP extension in Assumptions,
not a functional requirement — keeps scope bounded.
- Items are validated by review, not by an automated gate (`beacon` CLI is not
installed in this environment; placeholder/ADR-reference checks were run
manually — see the session's validation step).
@@ -0,0 +1,64 @@
# Contract: Image Input across the LLMProvider boundary
**Spec**: [spec.md](../spec.md) · **Data model**: [data-model.md](../data-model.md) · **ADR**: [ADR-008](../../../project-management/ADRs/ADR-008-multimodal-image-input-across-llmprovider-boundary.md)
This feature adds no new protocol methods and no new socket types. The contract
below is the **behavioural conformance** every `LLMProvider` must satisfy for
the new `Message.images` field, plus the node's input contract.
---
## C1 — `Message.images` carrier
- `Message.images: list[str] | None = None`, base64 strings (no `data:` prefix).
- `images=None` or `[]` ⇒ the turn is text-only and MUST produce a request
**byte-for-byte identical** to the pre-feature behaviour.
## C2 — `LLMProvider.chat()` conformance (both providers)
Given `messages` where the last user turn carries `images`:
1. The provider MUST transmit those images with that turn to its backend using
its native shape (Ollama flat `images`; llama.cpp OpenAI `image_url` parts).
2. The provider MUST NOT transmit an `images` field for turns that have none
(empty key omitted).
3. All existing behaviour is preserved: blank-retry-with-new-seed loop, response
caching, timeout, and error surfacing are unchanged by the presence of images.
4. A backend that cannot process images (non-vision model / no `--mmproj`) MUST
have its error surfaced to the caller, not swallowed (FR-006).
## C3 — `chat_structured()` conformance (shared helper, both providers)
1. Images on the last user turn MUST be attached as `BinaryContent` on the
`Agent.run` `user_prompt`; images on history user turns MUST be attached to
their `UserPromptPart`.
2. All existing structured guarantees hold unchanged: bounded retries (0–5),
`RuntimeError` on exhaustion naming model/attempts/snippet, never returns a
value that failed schema validation.
3. A text-only structured call MUST be indistinguishable from today's.
## C4 — `ChatCompletion` node input contract
1. Adds exactly one **optional** `image: ("IMAGE",)` input. No required input
added; `RETURN_TYPES`/`RETURN_NAMES` positions unchanged (Constitution VI).
2. Un-wired ⇒ behaviour, request, and outputs identical to today.
3. Wired ⇒ image(s) attached to the current user turn only; `history` turns
unchanged (FR-007).
4. Works with `structured_output=True` (C3) and free-text (C2) alike, on either
backend, with no per-backend wiring difference (FR-004).
---
## Test contracts (test-first — Constitution III)
| ID | Level | Asserts |
|---|---|---|
| T1 | `Message` unit | `images` defaults `None`; round-trips base64 list; text-only dump omits the key |
| T2 | `OllamaProvider.chat` | image turn → payload message has flat `images:[...]`; text-only payload byte-identical to today (regression) |
| T3 | `LlamaCppProvider.chat` | image turn → `content` becomes text+`image_url` parts; text-only `content` stays a plain string (regression) |
| T4 | `chat_structured` | image turn builds `BinaryContent` on the prompt; text-only path unchanged; retry/validation contract intact |
| T5 | node encode helper | synthetic `[1,H,W,3]` tensor → decodable base64 PNG; `None`/empty → `[]` |
| T6 | node contract | optional `image` in `INPUT_TYPES`; un-wired run == today; wired run attaches to last turn only |
All tests run without a live ComfyUI or a live backend (mock at each provider's
own `_post_json`/`Agent.run` seam, per the `test_ollama_provider.py`
convention). T5 uses a synthetic tensor + Pillow (dev dep), no ComfyUI.
+83
View File
@@ -0,0 +1,83 @@
# Phase 1 Data Model: VLM Image Input for ChatCompletion
**Spec**: [spec.md](./spec.md) · **Research**: [research.md](./research.md)
The feature adds **one optional field** to an existing model and defines how it
maps into each backend's wire shape. No new entities, no new socket types.
---
## Modified entity — `Message` (`src/comfydv/_llm/provider.py`)
```python
class Message(BaseModel):
role: Literal["system", "user", "assistant"]
content: str
images: list[str] | None = None # NEW — base64-encoded images (no data: prefix)
```
**Field: `images`**
- **Type**: `list[str] | None`, default `None`.
- **Meaning**: base64-encoded image payloads associated with this turn. `None`
(or empty) means a text-only turn — **byte-for-byte identical to today**.
- **Carrier form**: raw base64 string, no `data:` URI prefix. Chosen because
every target adapts *from* it (Ollama `images` array, OpenAI data-URI,
pydantic-ai `BinaryContent`) — ADR-008.
- **Validation**: no format validation at the model layer (the model stays a
dumb carrier); malformed data surfaces as a backend error (FR-006). A turn
may carry ≥1 image; MVP exercises exactly one.
- **Serialization invariant**: provider payload construction MUST omit the
`images` key when `None`/empty so existing text-only requests are unchanged
(research.md Decision 2; FR-003, SC-004).
---
## Mapping table — one carrier, three wire shapes
| Path | Code site | Transform |
|---|---|---|
| Ollama free-text | `ollama_provider.py::chat` | none — `model_dump()`'s flat `images` array already matches `/api/chat`; only drop the key when empty |
| llama.cpp free-text | `llamacpp_provider.py::chat` | rebuild `content` as OpenAI parts: `[{"type":"text",...},{"type":"image_url","image_url":{"url":"data:image/png;base64,<b64>"}}]` |
| Structured (both) | `chat.py::chat_structured` | build `BinaryContent(data=b64decode(img), media_type="image/png")`; attach to `user_prompt` (last turn) / `UserPromptPart` (history turns) as `[text, *images]` |
---
## Node input — `ChatCompletion` (`src/comfydv/ollama.py`)
Add to `INPUT_TYPES["optional"]`:
```python
"image": ("IMAGE",),
```
- **Optional** — un-wired ⇒ `image=None` ⇒ the node builds exactly today's
text-only user message. No new required input; no output/socket change
(Constitution VI untouched — `RETURN_TYPES` positions 0/1 unchanged).
- When wired: encode the tensor to base64 PNG(s) (research.md Decision 4) and
set them on the appended `Message(role="user", ...)`. History turns are not
modified (FR-007).
### Encode helper (node-local, `comfy`/Pillow lazy)
```
_encode_image_tensor(image) -> list[str]:
# image: ComfyUI IMAGE, torch float tensor [B, H, W, C] in 0..1
# → for each frame: *255 → uint8 → PIL.Image.fromarray → PNG bytes → base64
# returns [] for None/empty so callers treat it as "no image"
```
Lives in `ollama.py` (node module, already `comfy`-guarded). `src/comfydv/_llm/`
never imports torch/numpy/Pillow — it deals only in the base64 strings this
helper produces.
---
## State & relationships
- No persistent state; no new caching entity. Existing `_CHAT_RESPONSE_CACHE`
keys already include the dumped messages, so an added `images` value
participates in the cache key automatically (same image + prompt ⇒ cache
hit), and a text-only turn's key is unchanged since the empty key is omitted.
- Relationship: `images` belongs to exactly one `Message` (one turn) — this is
why the carrier is a message field, not a side-channel parameter (ADR-008
Alternative A rejected).
@@ -0,0 +1,11 @@
Feature: US1 — Describe an image with a chat node
Scenario: Describe a wired image
Given a chat node connected to a backend with a vision-capable model loaded and an image wired into the node's image input
When the workflow executes with a prompt like "describe this image"
Then the node returns a text response that reflects the actual content of the wired image
Scenario: No image wired behaves exactly as today
Given the same chat node with no image wired
When the workflow executes
Then the node behaves exactly as it does today — text-only chat, identical response for identical text input — with no new required inputs and no change in output
@@ -0,0 +1,11 @@
Feature: US2 — Same image input on either backend
Scenario: Swap Ollama for llama.cpp and the image path still works
Given a workflow that describes an image via the chat node wired to Ollama
When the connection node is swapped to a llama.cpp one (pointed at a server with a multimodal model) with no other change
Then the workflow still returns a description of the same image
Scenario: Both backends produce an image-grounded response
Given equivalent image + prompt inputs on both backends
When each workflow executes
Then both produce a coherent image-grounded text response — no backend requires a different node, input shape, or wiring for the image
@@ -0,0 +1,11 @@
Feature: US3 — Structured output about an image
Scenario: Structured fields populated from the image
Given the chat node with an image wired and structured output enabled with a valid schema
When the workflow executes against a vision-capable model
Then each schema field is available as its own typed output, populated from the image, with no required field blank
Scenario: Invalid structured output retries then fails clearly
Given the same setup where the model first returns invalid or incomplete structured output
When the workflow executes
Then the node retries and, if still unsuccessful, fails with a clear error — the same retry/validation behaviour the text-only structured path already guarantees
+125
View File
@@ -0,0 +1,125 @@
# Implementation Plan: VLM Image Input for ChatCompletion
**Branch**: `009-vlm-image-input` | **Date**: 2026-07-22 | **Spec**: [spec.md](./spec.md)
**Input**: Feature specification from `/specs/009-vlm-image-input/spec.md`
## Summary
Let a workflow author wire a ComfyUI `IMAGE` into the existing generic
`ChatCompletion` node so a vision-capable model can describe or reason about it,
on **either** backend. Per ADR-008 (extending ADR-007's adapter pattern to a
second input modality), images ride on an optional `Message.images` carrier and
each provider translates that carrier into its own wire shape: Ollama's flat
`/api/chat` `images` array (passes through untouched), llama.cpp's OpenAI
`image_url` content-parts, and — for structured output — pydantic-ai
`BinaryContent`, shared by both backends through `OpenAIChatModel`. The node
converts its `IMAGE` tensor to base64 PNG; everything below the node deals only
in base64 strings. Text-only behaviour is byte-for-byte unchanged when no image
is wired.
## Technical Context
**Language/Version**: Python ≥3.11 (unchanged, per `pyproject.toml`)
**Primary Dependencies**: existing only for runtime — `aiohttp` (Ollama/llama.cpp
REST), `pydantic-ai-slim[openai]>=2.9.0` (structured path; its `BinaryContent`
multimodal type was verified against the installed 2.9.0 source, see
`research.md`). **No new core runtime dependency.** `pillow` is added to the
**dev** group so the node's tensor→PNG encoder is unit-testable without a live
ComfyUI; at runtime Pillow/numpy are ComfyUI-provided (same stance the repo
already takes for torch).
**Storage**: N/A — no persistent storage; reuses the existing
`_CHAT_RESPONSE_CACHE` (an `images` value participates in the cache key
automatically).
**Testing**: `pytest` via `uv run pytest`, following the
`tests/test_ollama_provider.py` convention (mock at each provider's own
`_post_json` / `Agent.run` seam, no live server or ComfyUI required). Test-first
per Constitution III; the tensor-encode test uses a synthetic tensor + Pillow.
**Target Platform**: ComfyUI custom-node runtime, same as the existing LLM nodes.
**Project Type**: Library / ComfyUI custom-node pack (single project).
**Performance Goals**: No new numeric target; image encoding is a one-shot
per-execution PNG encode, negligible against inference latency.
**Constraints**: Text-only requests MUST stay byte-identical (FR-003/SC-004) —
providers omit an empty `images` key. `src/comfydv/_llm/` must not import
torch/numpy/Pillow (Constitution IV) — tensor handling stays in the node.
llama.cpp image support requires a server launched with `--mmproj` (deployment
prerequisite, surfaced as a clear error when absent, not configured by comfydv).
**Scale/Scope**: One new `Message` field; a per-provider mapping in each
`chat()` plus the shared `chat_structured()`; one optional node input + a
node-local encode helper. No new node classes, no new socket types, no protocol
method changes.
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Principle | Verdict | Notes |
|---|---|---|
| I. ComfyUI Contract First | PASS | `ChatCompletion` keeps its `INPUT_TYPES`/`RETURN_TYPES`/`FUNCTION`/`CATEGORY`; only an optional input is added. No new registration, no ComfyUI changes. |
| II. Sandbox All User-Supplied Code | N/A | No template/expression evaluation in this feature. |
| III. Test-First | PASS (binding) | Test contracts T1–T6 (`contracts/image-input-contract.md`) written test-first, mirroring `test_ollama_provider.py`. Each runs without a live ComfyUI/backend. |
| IV. Graceful Degradation Outside ComfyUI | PASS (binding) | `_llm/` stays torch/numpy/Pillow-free — pure base64 carrier + mapping, unit-testable. Tensor→PNG lives in `ollama.py` (already `comfy`-guarded) with lazy Pillow/numpy import, so module import outside ComfyUI is unaffected. |
| V. Simplicity — Function Before Class | PASS | No new class. New logic is a `Message` field, two small per-provider transforms, one shared helper edit, and one module-level encode function. |
| VI. Fixed Output Positions | PASS | Outputs are untouched — only an optional **input** is added; `RETURN_TYPES`/`RETURN_NAMES` positions 0/1 and the structured extra-outputs contract are unchanged. |
Re-checked post-Phase 1 design (data-model.md, contracts/): unchanged — no new
gate violations. No Complexity Tracking entries needed.
## Project Structure
### Documentation (this feature)
```text
specs/009-vlm-image-input/
├── plan.md # This file
├── research.md # Phase 0 — wire shapes verified against installed deps
├── data-model.md # Phase 1 — Message.images + per-provider mapping
├── quickstart.md # Phase 1 — minimal describe-an-image workflow
├── contracts/
│ └── image-input-contract.md # Phase 1 — behavioural + test contracts (T1–T6)
├── checklists/requirements.md # Spec quality checklist (from /speckit-specify)
└── tasks.md # Phase 2 (/speckit-tasks) — not created here
```
### Source Code (repository root)
```text
src/comfydv/
├── ollama.py # MODIFIED — ChatCompletion: optional `image` input +
│ # node-local _encode_image_tensor() (lazy Pillow/numpy)
├── _llm/
│ ├── provider.py # MODIFIED — Message gains `images: list[str] | None = None`
│ ├── ollama_provider.py # MODIFIED — chat(): pass flat images through; omit empty key
│ ├── llamacpp_provider.py # MODIFIED — chat(): map images → OpenAI image_url parts
│ └── chat.py # MODIFIED — chat_structured(): images → BinaryContent on prompt
└── __init__.py # unchanged — no new node class or mapping
tests/
├── test_provider.py (or test_ollama_provider.py) # T1 Message carrier + regression
├── test_ollama_provider.py # MODIFIED — T2 Ollama image mapping + text regression
├── test_llamacpp_provider.py # MODIFIED — T3 llama.cpp content-parts + text regression
├── test_llm_chat.py / chat tests # T4 chat_structured multimodal + regression
└── test_ollama.py # MODIFIED — T5 encode helper, T6 node input contract
pyproject.toml # MODIFIED — add `pillow` to [dependency-groups].dev only
```
**Structure Decision**: Purely additive edits to the four existing `_llm`/node
files that ADR-007 established — no new module, because there is no new class or
node (contrast 008, which added a provider + node). The image path threads
through the exact seams the text path already uses, which is the whole point of
ADR-008: a second modality on the same adapter, not a parallel structure.
## Complexity Tracking
> **Fill ONLY if Constitution Check has violations that must be justified**
None — see Constitution Check above.
+48
View File
@@ -0,0 +1,48 @@
# Quickstart: Describe an image with ChatCompletion
**Spec**: [spec.md](./spec.md)
Minimal end-to-end walkthrough of the feature once shipped.
## Prerequisites
- A running backend with a **vision-capable** model:
- **Ollama** — a multimodal model pulled and available (e.g. a llava-class model), or
- **llama.cpp** — `llama-server` launched in router mode **with a multimodal projector**: `--mmproj <projector.gguf>` alongside the model.
- comfydv installed in ComfyUI.
## Steps
1. Add an image source to the canvas (e.g. **Load Image**) → gives an `IMAGE`.
2. Add a client node (**OllamaClient** or **LlamaCppClient**) → gives `LLM_CLIENT`.
3. Add **ChatCompletion**. Wire:
- `client` ← the client node
- `model` ← a vision-capable model name (typed or wired)
- `prompt` ← `"Describe this image in one sentence."`
- `image` ← the `IMAGE` from step 1 ← **the only new wire**
4. Queue the prompt. The `response` output is a text description of the image.
## Structured variant (optional)
- On **ChatCompletion**, set `structured_output = True` and provide a schema, e.g.:
```json
{"type":"object","properties":{"caption":{"type":"string"},"has_text":{"type":"boolean"}},"required":["caption","has_text"]}
```
- Run: each field (`caption`, `has_text`) appears as its own typed output,
populated from the image, with no required field blank.
## Swap backends (proves FR-004 / SC-002)
- Replace **OllamaClient** with **LlamaCppClient** (pointed at an `--mmproj`
server) — **change nothing else**. Re-queue: same image description path.
## What stays the same
- Leave `image` un-wired and ChatCompletion behaves exactly as before — text-only,
identical results. No existing workflow changes.
## Expected failure (proves FR-006 / SC-005)
- Wire an image but select a **non-vision** model (or a `llama-server` started
without `--mmproj`): the node reports a clear error that the model/server
can't process images — it does not silently answer as if no image was sent.
+155
View File
@@ -0,0 +1,155 @@
# Phase 0 Research: VLM Image Input for ChatCompletion
**Spec**: [spec.md](./spec.md) · **Plan**: [plan.md](./plan.md) · **ADR**: [ADR-008](../../project-management/ADRs/ADR-008-multimodal-image-input-across-llmprovider-boundary.md)
ADR-008 recorded the boundary decision (images on `Message.images`, translated
per-provider) but deferred the exact wire shapes for live verification. This
file resolves them against the **installed** dependency versions and the
current provider code, not from memory.
---
## Decision 1 — pydantic-ai multimodal vehicle (structured-output path)
**Decision**: In the shared `chat_structured()` helper (`src/comfydv/_llm/chat.py`),
attach images as `pydantic_ai.messages.BinaryContent(data=<png bytes>,
media_type="image/png")` inside a `Sequence[UserContent]`. The current-turn
image rides on `Agent.run(user_prompt=[text, BinaryContent(...)])`; a
history turn's image rides on `UserPromptPart(content=[text, BinaryContent(...)])`.
**Rationale / verified**: Read directly from the pinned
`pydantic_ai_slim==2.9.0` source in this environment:
- `BinaryContent` (`messages.py:521`) — `__init__(self, data: bytes, *,
media_type: ..., identifier=None, ...)`; exposes a `.base64` helper. `data`
is **bytes**, so `chat.py` must `base64.b64decode()` the `Message.images`
string into bytes when building it.
- `UserPromptPart.content: str | Sequence[UserContent]` (`messages.py:1022`)
and `user_prompt` on `Agent.run` accept the same. `UserContent = str |
TextContent | MultiModalContent | CachePoint` (`messages.py:899`), and
`MultiModalContent` includes `BinaryContent`/`ImageUrl` — so a `[text,
image]` list is the supported shape.
- `OpenAIChatModel` renders `BinaryContent` for images as an OpenAI
`image_url` data-URI, and reads `BinaryContent.vendor_metadata['detail']`
for the `detail` setting (documented in the field's own docstring).
**Consequence**: the structured path is provider-agnostic *for free* — both
Ollama and llama.cpp reach `/v1/chat/completions` through the same
`OpenAIChatModel`, so one change in `chat.py` covers structured output on both
backends. No per-provider structured code.
**Alternatives considered**: `ImageUrl(url="data:image/png;base64,...")` — also
supported, but requires assembling a data URI string; `BinaryContent` from raw
bytes + media type is the more direct representation of what we hold and lets
pydantic-ai own the data-URI formatting.
---
## Decision 2 — Ollama free-text path (`/api/chat`)
**Decision**: `OllamaProvider.chat()` sends each message's images as a flat
`images` array of **base64 strings** (no `data:` prefix) alongside `content`,
which is exactly Ollama's native `/api/chat` message schema. Because
`Message.images` already holds base64 strings, `Message.model_dump()` produces
the correct shape with **no transform** — the field flows straight through.
**Rationale / verified**: `OllamaProvider.chat()`
(`src/comfydv/_llm/ollama_provider.py:280`) already builds
`payload_messages = [m.model_dump() for m in messages]` and POSTs to
`/api/chat`. Ollama's documented `/api/chat` message object is
`{"role", "content", "images": [<base64>, ...]}` — the flat sibling field this
carrier maps onto directly. This is the reason base64 is the neutral carrier
form (ADR-008).
**Constraint discovered — byte-identical text path (FR-003/SC-004)**: adding
`images: list[str] | None = None` to `Message` means a text-only message would
dump as `{"role","content","images":null}`, changing today's request body.
Providers MUST drop a `None`/empty `images` before sending. Resolution:
serialize provider payload messages with the images key omitted when empty
(e.g. `model_dump(exclude_none=True)`, or drop the key explicitly). Guarded by
the existing Ollama provider tests, which assert the exact payload.
---
## Decision 3 — llama.cpp free-text path (`/v1/chat/completions`)
**Decision**: `LlamaCppProvider.chat()` maps a message carrying images into
OpenAI-style multimodal `content` **parts** before POSTing:
`content: [{"type":"text","text":<content>},
{"type":"image_url","image_url":{"url":"data:image/png;base64,<b64>"}}]`.
Messages with no images keep the plain-string `content` unchanged.
**Rationale / verified**: `LlamaCppProvider.chat()`
(`src/comfydv/_llm/llamacpp_provider.py:163`) builds
`payload_messages = [m.model_dump() for m in messages]` and POSTs to
`/v1/chat/completions`. Unlike Ollama, a flat `images` sibling is **not**
understood there — OpenAI's vision schema requires images inside `content` as
typed parts. `llama-server` implements this OpenAI-compatible multimodal
format **only when launched with a multimodal projector (`--mmproj`)**; without
it, image parts yield a server error (surfaced per FR-006, not crashed on).
This is the single point where the two providers genuinely diverge — exactly
the leakage ADR-008 localizes inside each provider.
**Alternatives considered**: normalizing Ollama *up* to content-parts too (one
shared mapper) — rejected in ADR-008 Alternative C: it forces the
currently-simpler Ollama path to do extra work and inverts "each provider owns
its wire format."
---
## Decision 4 — ComfyUI IMAGE tensor → base64 PNG (node layer)
**Decision**: The `ChatCompletion` node converts its optional `IMAGE` input to
base64 PNG(s) via Pillow: ComfyUI IMAGE is a float tensor `[B, H, W, C]` in
`0..1`; scale to `uint8`, `PIL.Image.fromarray(...)`, save PNG to an in-memory
buffer, base64-encode. A batch of `B` frames becomes `B` base64 strings in the
turn's `images` list (natural multi-image; MVP exercises `B=1`). The import of
Pillow/numpy is **lazy** (inside the encode function), so the module still
imports cleanly outside ComfyUI (Constitution IV).
**Rationale**: Pillow is the ComfyUI-ecosystem standard for IMAGE tensor ↔
file and is present in every ComfyUI install; numpy comes with torch. Neither
is added to comfydv's **core** runtime deps — they are ComfyUI-provided, the
same stance the repo already takes for torch (dev-only in `pyproject.toml`).
To keep the encoder **test-first** (Constitution III) without a live ComfyUI,
add `pillow` to the **dev** dependency group so a unit test can feed a
synthetic `numpy`/`torch` tensor through the pure encode function and assert a
decodable PNG.
**Boundary kept clean**: only the node (`src/comfydv/ollama.py`, already
`comfy`-guarded) touches tensors/Pillow. Everything in `src/comfydv/_llm/`
deals purely in base64 strings and stays unit-testable with hand-crafted
strings — no torch, numpy, or Pillow import there.
**Edge cases (FR-006, Edge Cases)**: an un-wired optional input arrives as
`None` → node builds today's exact text-only message. A zero-size / empty batch
tensor → treated as "no image". A non-vision model or non-`mmproj` server
returns a backend error → surfaced with a clear message, never a silent
image-less answer.
---
## Decision 5 — where the image attaches on the turn (FR-007)
**Decision**: The node attaches images to the **current user turn only** — the
`Message(role="user", content=prompt, images=[...])` it already appends. Prior
`history` turns are untouched. The structured helper likewise only lifts images
onto the final user turn (and any history turn that already carried them),
matching its existing "last message is the prompt" contract
(`chat.py:106`, which requires `messages[-1].role == "user"`).
---
## Summary of resolved unknowns
| Unknown (from ADR-008) | Resolved to |
|---|---|
| pydantic-ai multimodal type | `BinaryContent(data=bytes, media_type="image/png")` — verified in installed 2.9.0 |
| Structured path per-provider? | No — shared via `OpenAIChatModel`; one change in `chat.py` |
| Ollama wire shape | flat `images: [base64]` on the message; passes through `model_dump()` |
| llama.cpp wire shape | OpenAI `image_url` content-parts; requires `--mmproj` |
| Text-path byte-identity | drop empty `images` key in provider payloads (guarded by existing tests) |
| Tensor → base64 | Pillow, lazy import in node; `pillow` added to dev deps for testability |
| No new runtime deps | Confirmed — Pillow/numpy are ComfyUI-provided, dev-only here |
No `NEEDS CLARIFICATION` remain.
+148
View File
@@ -0,0 +1,148 @@
# Feature Specification: VLM Image Input for ChatCompletion
**Feature Branch**: `009-vlm-image-input`
**Created**: 2026-07-22
**Status**: Draft
**Input**: User description: "Let a workflow author wire a ComfyUI IMAGE into the existing generic ChatCompletion node so a vision-capable model (VLM) on either backend (Ollama multimodal models, llama.cpp multimodal via mmproj) can describe or understand the image. Provider-agnostic per ADR-007/ADR-008: the node attaches the image to the user message; each provider maps it to its own wire format. The Message carrier gains an optional image field; text-only behaviour is unchanged when no image is wired."
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Describe an image with a chat node (Priority: P1) 🎯 MVP
As a ComfyUI workflow author with a vision-capable model available, I want to
wire an image into the chat node I already use and get back a text description
or answer about that image, so I can add image understanding to a workflow
without learning a new node.
**Why this priority**: This is the entire point of the feature — a picture in,
a text understanding out — and the proof that image input works through the
existing generic node on at least one backend.
**Independent Test**: Wire any image source into the chat node's image input,
point the node at a loaded vision-capable model, run the workflow, and confirm
the response text describes the wired image.
**Acceptance Scenarios**:
1. **Given** a chat node connected to a backend with a vision-capable model loaded and an image wired into the node's image input, **When** the workflow executes with a prompt like "describe this image", **Then** the node returns a text response that reflects the actual content of the wired image.
2. **Given** the same chat node with **no** image wired, **When** the workflow executes, **Then** the node behaves exactly as it does today — text-only chat, identical response for identical text input — with no new required inputs and no change in output.
---
### User Story 2 - Same image input on either backend (Priority: P1)
As a workflow author, I want image input to work the same way whether my chat
node is connected to Ollama or to llama.cpp, so I don't have to rebuild or
relearn the image path when I switch backends — exactly as text and structured
output already behave identically across the two.
**Why this priority**: The generic-node promise (ADR-007) is the reason this
feature is small; this story is what proves the image path honours it rather
than quietly becoming backend-specific.
**Independent Test**: Run User Story 1 unchanged against an Ollama connection
and against a llama.cpp connection (each with a vision-capable model), and
confirm both return a description of the wired image using the identical node
setup.
**Acceptance Scenarios**:
1. **Given** a workflow that describes an image via the chat node wired to Ollama, **When** the connection node is swapped to a llama.cpp one (pointed at a server with a multimodal model) with no other change, **Then** the workflow still returns a description of the same image.
2. **Given** equivalent image + prompt inputs on both backends, **When** each workflow executes, **Then** both produce a coherent image-grounded text response — no backend requires a different node, input shape, or wiring for the image.
---
### User Story 3 - Structured output about an image (Priority: P2)
As a workflow author, I want to combine image input with the node's existing
structured-output mode, so a VLM can return schema-validated fields extracted
from an image (for example a caption, a list of detected objects, or a
yes/no), not just free text.
**Why this priority**: Structured output is an existing, valued capability;
making it work with images turns "describe this" into usable, wired,
downstream-typed data. It builds on User Story 1 and is lower risk to defer
than getting basic image chat working at all.
**Independent Test**: Enable structured output on the chat node with a schema,
wire an image, run against a vision-capable model, and confirm each schema
field is populated from the image and no required field is blank.
**Acceptance Scenarios**:
1. **Given** the chat node with an image wired and structured output enabled with a valid schema, **When** the workflow executes against a vision-capable model, **Then** each schema field is available as its own typed output, populated from the image, with no required field blank.
2. **Given** the same setup where the model first returns invalid or incomplete structured output, **When** the workflow executes, **Then** the node retries and, if still unsuccessful, fails with a clear error — the same retry/validation behaviour the text-only structured path already guarantees.
---
### Edge Cases
- What happens when an image is wired but the selected model is **not**
vision-capable? The node must surface a clear error attributable to the model
lacking image support, not crash and not silently drop the image and answer
as if none was sent.
- What happens when the backend server is reachable but was not started with
multimodal support (e.g. a llama.cpp server launched without an `mmproj`
projector)? The node should report a clear, specific error rather than an
unhelpful generic failure.
- What happens with an empty or zero-size image input, or an image input that
is wired but carries no actual image data? The node should treat it as "no
image" or report a clear error — never send a malformed request.
- What happens when both an image and a multi-turn history are present? The
image must be associated with the current user turn, and prior turns must
remain unaffected.
- What happens to the node's text-only path for a model/backend that does not
understand images at all — does an un-wired image input leave the request
byte-for-byte identical to today's? (It must.)
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: The system MUST let a workflow author provide an image to the existing chat node through a single, **optional** image input — no new node and no new required input.
- **FR-002**: When an image is provided, the system MUST include it with the current user turn sent to the connected model, so a vision-capable model can ground its response in that image.
- **FR-003**: When **no** image is provided, the system MUST send exactly the request it sends today — text-only behaviour, inputs, and outputs unchanged, with no regression for existing workflows.
- **FR-004**: Image input MUST work identically across both supported backends from the workflow author's perspective — same node, same wiring, same input shape — with each backend's differing native image format handled internally, not exposed on the graph.
- **FR-005**: Image input MUST be compatible with the node's existing structured-output mode: an image-grounded response can be schema-validated with the same retry and validation guarantees as the text-only structured path.
- **FR-006**: The system MUST surface a clear, specific error when an image is provided but the target model or backend cannot process images (non-vision model, or a server without multimodal support), rather than crashing or silently discarding the image.
- **FR-007**: The system MUST associate a provided image with the current user turn only, leaving any prior conversation history unchanged.
### Key Entities *(include if feature involves data)*
- **Chat message**: The existing per-turn unit of a chat request. Extended so a
turn can optionally carry one or more images in addition to its text; a
turn with no image is unchanged from today.
- **Image input**: An image supplied on the workflow canvas (the standard
ComfyUI image type) and attached to the current user turn; provider-neutral
at the boundary, translated to each backend's native shape internally.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: A workflow author can make an existing chat node describe a wired image by adding exactly one connection (the image), with no new node and no other node changes.
- **SC-002**: The same image-describing workflow runs unchanged when repointed from one backend to the other — zero edits beyond swapping the connection node.
- **SC-003**: Structured-output workflows with an image populate every required schema field from the image content, with zero blank-required-field results, matching the text-only structured guarantee.
- **SC-004**: Every existing text-only workflow produces identical results after this feature ships — no observable change when no image is wired (existing backend behaviour tests remain green).
- **SC-005**: Providing an image to a non-vision model or a non-multimodal server yields a clear, specific error in 100% of such cases — never a crash and never a silently image-less answer presented as if the image was seen.
## Assumptions
- Workflow authors run their own backend (Ollama or llama.cpp) with a
vision-capable model available and loaded; for llama.cpp this means the
server was launched with a multimodal projector (`mmproj`). This feature does
not install, configure, download, or launch vision models.
- Scope is still **images only** — no video, audio, or document modalities; and
image **input** only — no image generation or output.
- A single image per turn is the primary target; carrying more than one image
per turn is a natural extension of the same carrier but is not a required
acceptance criterion of the MVP.
- The generic `ChatCompletion` node, the `LLMProvider` protocol, and both
providers already exist (ADR-007) and are extended, not replaced; the
cross-provider image-carrier decision is recorded in ADR-008.
- The standard ComfyUI image type is the input; converting it to the neutral
form each backend consumes is an internal concern of this feature, not
something the workflow author sees.
+146
View File
@@ -0,0 +1,146 @@
# Tasks: VLM Image Input for ChatCompletion
**Input**: Design documents from `/specs/009-vlm-image-input/`
**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/image-input-contract.md
**Tests**: First-class — every implementation task has a paired failing-test task (`-T`/`-I` suffix). Contracts T1–T6 in `contracts/image-input-contract.md` map to the pairs below.
**Organization**: Grouped by user story (spec.md priorities: US1 P1 🎯 MVP, US2 P1, US3 P2).
## Format: `[ID] [P?] [Story] Description`
- **[P]**: Can run in parallel (different files, no dependencies)
- **[Story]**: US1–US3
- **-T / -I**: paired test (red) / implementation (green) — the `-T` is committed failing before its `-I` partner (no test + impl in one commit)
## Path Conventions
Single project: `src/comfydv/`, `tests/` at repo root. Purely additive edits to
the existing `_llm`/node files (plan.md Structure Decision) — no new module.
`src/comfydv/_llm/` stays torch/numpy/Pillow-free (Constitution IV); tensor
handling lives only in the `comfy`-guarded `ollama.py`.
---
## Phase 1: Setup
- [x] T001 Add `pillow` to `[dependency-groups].dev` in `pyproject.toml` — lets the node's tensor→PNG encoder be unit-tested without a live ComfyUI; runtime Pillow/numpy are ComfyUI-provided, so **no core runtime dependency is added** (research.md Decision 4)
---
## Phase 2: Foundational (Blocking Prerequisites)
**Purpose**: the image carrier every path depends on. **⚠️ No user story work can begin until this is complete.**
- [x] T002-T Write FAILING test: `Message.images` defaults to `None`, round-trips a base64 list, and a text-only message's transport dump **omits** the `images` key (byte-identical to today), in `tests/test_llm_provider.py` (contract T1)
- [x] T002-I Add `images: list[str] | None = None` to `Message` in `src/comfydv/_llm/provider.py` — makes T002-T pass
**Checkpoint**: carrier ready — user stories can begin.
---
## Phase 3: User Story 1 — Describe an image with a chat node (Priority: P1) 🎯 MVP
**Goal**: A workflow author wires a ComfyUI `IMAGE` into the existing `ChatCompletion` node and gets back a text description via an Ollama vision model; the text-only path is unchanged when no image is wired.
**Independent Test**: Wire an image → `ChatCompletion` → Ollama (vision model), confirm the response describes the image; un-wire the image and confirm behaviour/output identical to today.
- [x] T003-T [P] [US1] Write FAILING test: node-local `_encode_image_tensor()` converts a synthetic `[1,H,W,3]` float tensor (0..1) into a **decodable** base64 PNG, encodes a `B>1` batch to a list of that length, and returns `[]` for `None`/empty, in `tests/test_ollama.py` (contract T5; witnesses `features/us1_describe_image.feature` scenario "Describe a wired image")
- [x] T003-I [US1] Implement `_encode_image_tensor()` in `src/comfydv/ollama.py` — lazy `PIL`/`numpy` import so module import stays clean outside ComfyUI (Constitution IV); batch → one base64 string per frame — makes T003-T pass
- [x] T004-T [US1] Write FAILING test: `ChatCompletion.INPUT_TYPES` exposes an **optional** `image: ("IMAGE",)`; `RETURN_TYPES`/`RETURN_NAMES` positions are unchanged; an un-wired run builds the same text-only messages as today; a wired run attaches images to the **last user turn only** (history untouched), in `tests/test_ollama.py` (contract T6; witnesses both `features/us1_describe_image.feature` scenarios)
- [x] T004-I [US1] Add the optional `image` input (with a tooltip noting a vision-capable model is required; llama.cpp needs `--mmproj`) and attach encoded images to the appended user `Message` in `ChatCompletion.chat()` in `src/comfydv/ollama.py` — makes T004-T pass (depends on T003-I, T002-I)
- [x] T005-T [P] [US1] Write FAILING test: `OllamaProvider.chat()` forwards a message's images as a flat `images:[...]` array to `/api/chat`, and a text-only call's payload is **byte-identical to today** (regression), in `tests/test_ollama_provider.py` (contract T2; witnesses `features/us1_describe_image.feature` scenario "Describe a wired image")
- [x] T005-I [US1] Ensure `OllamaProvider.chat()` passes images through and omits the empty `images` key (e.g. `model_dump(exclude_none=True)`) in `src/comfydv/_llm/ollama_provider.py` — makes T005-T pass (depends on T002-I)
**Checkpoint**: describe-an-image works end-to-end on Ollama (MVP); every existing text-only test stays green.
---
## Phase 4: User Story 2 — Same image input on either backend (Priority: P1)
**Goal**: The same node and wiring drive image input on llama.cpp too, via its OpenAI-compatible content-parts shape — proving the generic-node promise (ADR-007/008) holds for the image path.
**Independent Test**: Run the US1 workflow unchanged against a llama.cpp server (launched with `--mmproj`); swap the Ollama client node for the llama.cpp one with no other change and confirm the image is still described.
- [x] T006-T [P] [US2] Write FAILING test: `LlamaCppProvider.chat()` maps a message's images into OpenAI `content` parts (`{"type":"text",...}` + `{"type":"image_url","image_url":{"url":"data:image/png;base64,..."}}`) for `/v1/chat/completions`, and a text-only message keeps a **plain-string** `content` (regression), in `tests/test_llamacpp_provider.py` (contract T3; witnesses both `features/us2_both_backends.feature` scenarios)
- [x] T006-I [US2] Implement the images→content-parts mapping in `LlamaCppProvider.chat()` in `src/comfydv/_llm/llamacpp_provider.py`; leave text-only messages untouched — makes T006-T pass (depends on T002-I)
**Checkpoint**: parity proven — the identical node/wiring describes an image on both backends; swapping the client node is the only change.
---
## Phase 5: User Story 3 — Structured output about an image (Priority: P2)
**Goal**: Image input works with the node's existing structured-output mode, via the shared `chat_structured()` helper (pydantic-ai `BinaryContent`) — one implementation covering both backends through `OpenAIChatModel`.
**Independent Test**: Enable structured output with a schema, wire an image, run against a vision model, confirm each field is populated from the image with no required field blank; a first-invalid response retries then fails clearly.
- [x] T007-T [US3] Write FAILING test: `chat_structured()` attaches a message's images as `BinaryContent(data=b64decode(img), media_type="image/png")` onto the run's `user_prompt` (last turn) and onto history `UserPromptPart`s, a text-only structured call is unchanged, and the retry/validation contract is intact, in `tests/test_llm_chat_structured.py` (contract T4; witnesses both `features/us3_structured_image.feature` scenarios) — mock at the `Agent.run`/`_build_agent` seam per the established convention
- [x] T007-I [US3] Implement image→`BinaryContent` handling in `chat_structured()` and `_history_to_messages()` in `src/comfydv/_llm/chat.py` — makes T007-T pass (depends on T002-I)
**Checkpoint**: structured image output works on both backends via the one shared helper; all prior stories remain green.
---
## Phase 6: Polish & Cross-Cutting Concerns
- [x] T008 [P] Document image input on `ChatCompletion` in `README.md` and add a `CHANGELOG.md` Unreleased entry — note the vision-model / llama.cpp `--mmproj` prerequisite (quickstart.md)
- [x] T009 Run the full quality gate: `ruff check` ✓, `ruff format` ✓, `pytest` ✓ (289 passed, +21 new; all spec-009 code green), `beacon doctor --strict` ✓ for this spec (bullet + BDD + backlinks pass). _Pre-existing, out of scope: `ty check` has 36 diagnostics repo-wide (0 from spec-009 code — verified), one Docker packaging test (`test_dockerfile_uses_python_311_base`) fails at baseline, and `spec-task-alignment` flags 007's deferred tasks under --strict._
- [-] T010 End-to-end `quickstart.md` validation against a live vision backend (Ollama multimodal model and `llama-server --mmproj`) _Deferred — requires a live vision-capable backend not available in CI/this environment; validate manually before release._
---
## Dependencies & Execution Order
- **Setup (T001)** → no dependencies; start immediately.
- **Foundational (T002-T/I)** → depends on nothing; **blocks all user stories** (every path reads `Message.images`).
- **US1 (T003–T005)** → after T002-I. `T004-I` depends on `T003-I`; `T005-I` depends on `T002-I`. MVP.
- **US2 (T006)** → after T002-I. Independent of US1's files; independently testable.
- **US3 (T007)** → after T002-I. Independent of US1/US2's files; independently testable.
- **Polish (T008–T010)** → after the stories you intend to ship.
### Within each story
- The `-T` task is written and committed **failing** before its `-I` partner (`tdd-commit-discipline`).
- `-I` is never `[P]` with its own `-T`.
### Parallel opportunities
- US1: `T003-T` (`tests/test_ollama.py`) and `T005-T` (`tests/test_ollama_provider.py`) are different files → `[P]`.
- Across stories: US1, US2, US3 touch different provider/helper files and can proceed in parallel once T002-I lands.
---
## Parallel Example: User Story 1
```bash
# Different test files, no shared deps — write both failing tests together:
Task: "T003-T encode-helper test in tests/test_ollama.py"
Task: "T005-T Ollama image-passthrough test in tests/test_ollama_provider.py"
```
---
## Implementation Strategy
### MVP first (US1 only)
1. T001 Setup → T002 carrier → T003–T005 US1.
2. **STOP and VALIDATE**: an Ollama vision model describes a wired image; every text-only test stays green.
3. Demoable as-is.
### Incremental delivery
1. Foundation + US1 → describe-an-image on Ollama (MVP).
2. + US2 → same node works on llama.cpp (parity).
3. + US3 → structured output about an image (both backends).
4. Polish → docs, quality gate, manual live validation (T010).
---
## Notes
- `[-]` (T010) is a **known-deferred** follow-up — `beacon bullet finish` skips it rather than flipping to `[x]`; `beacon doctor` reports it as deferred, held under `--strict`.
- `beacon doctor` runs two gates against this discipline: `spec-bdd-coverage` (every acceptance scenario has a `.feature` witness — 6 scenarios across 3 features here) and `tdd-commit-discipline` (no test + implementation in the same commit). Both FAIL under `--strict`.
- Commit after each task or `-T`/`-I` pair; keep existing Ollama/llama.cpp/text tests green throughout (FR-003/SC-004 regression guard).
+24 -14
View File
@@ -2,24 +2,27 @@ import logging
from .circuit_breaker import CircuitBreaker
from .format_string import FormatString
from .llamacpp import LlamaCppClient
from .ollama import (
ChatCompletion,
LLMLoadModel,
LLMModelSelector,
LLMUnloadModel,
OllamaClient,
OllamaChatCompletion,
OllamaDebugHistory,
OllamaHeaderBasicAuth,
OllamaHeaderBearerToken,
OllamaHeaderCustom,
OllamaHistoryLength,
OllamaLoadModel,
OllamaModelSelector,
OllamaOptionDisableThinking,
OllamaOptionExtraBody,
OllamaOptionMaxTokens,
OllamaOptionRefusalRetry,
OllamaOptionRepeatPenalty,
OllamaOptionSeed,
OllamaOptionTemperature,
OllamaOptionTopK,
OllamaOptionTopP,
OllamaUnloadModel,
)
from .random_choice import RandomChoice
@@ -31,18 +34,22 @@ NODE_CLASS_MAPPINGS = {
"RandomChoice": RandomChoice,
"CircuitBreaker": CircuitBreaker,
"FormatString": FormatString,
# Ollama nodes
# LLM nodes (generic, ADR-007) — see comfydv.ollama.MIGRATION_MAP for
# the pre-cutover Ollama-specific names these replace
"OllamaClient": OllamaClient,
"OllamaModelSelector": OllamaModelSelector,
"OllamaLoadModel": OllamaLoadModel,
"OllamaUnloadModel": OllamaUnloadModel,
"OllamaChatCompletion": OllamaChatCompletion,
"LlamaCppClient": LlamaCppClient,
"LLMModelSelector": LLMModelSelector,
"LLMLoadModel": LLMLoadModel,
"LLMUnloadModel": LLMUnloadModel,
"ChatCompletion": ChatCompletion,
"OllamaOptionTemperature": OllamaOptionTemperature,
"OllamaOptionSeed": OllamaOptionSeed,
"OllamaOptionMaxTokens": OllamaOptionMaxTokens,
"OllamaOptionTopP": OllamaOptionTopP,
"OllamaOptionTopK": OllamaOptionTopK,
"OllamaOptionRepeatPenalty": OllamaOptionRepeatPenalty,
"OllamaOptionDisableThinking": OllamaOptionDisableThinking,
"OllamaOptionRefusalRetry": OllamaOptionRefusalRetry,
"OllamaOptionExtraBody": OllamaOptionExtraBody,
"OllamaDebugHistory": OllamaDebugHistory,
"OllamaHistoryLength": OllamaHistoryLength,
@@ -56,18 +63,21 @@ NODE_DISPLAY_NAME_MAPPINGS = {
"RandomChoice": "Random Choice",
"CircuitBreaker": "Circuit Breaker",
"FormatString": "Format String (Python f-strings)",
# Ollama nodes
# LLM nodes (generic, ADR-007)
"OllamaClient": "Ollama Client",
"OllamaModelSelector": "Ollama Model Selector",
"OllamaLoadModel": "Ollama Load Model",
"OllamaUnloadModel": "Ollama Unload Model",
"OllamaChatCompletion": "Ollama Chat Completion",
"LlamaCppClient": "LlamaCpp Client",
"LLMModelSelector": "LLM Model Selector",
"LLMLoadModel": "LLM Load Model",
"LLMUnloadModel": "LLM Unload Model",
"ChatCompletion": "Chat Completion",
"OllamaOptionTemperature": "Ollama Option — Temperature",
"OllamaOptionSeed": "Ollama Option — Seed",
"OllamaOptionMaxTokens": "Ollama Option — Max Tokens",
"OllamaOptionTopP": "Ollama Option — Top P",
"OllamaOptionTopK": "Ollama Option — Top K",
"OllamaOptionRepeatPenalty": "Ollama Option — Repeat Penalty",
"OllamaOptionDisableThinking": "Ollama Option — Disable Thinking",
"OllamaOptionRefusalRetry": "Ollama Option — Refusal Retry",
"OllamaOptionExtraBody": "Ollama Option — Extra Body",
"OllamaDebugHistory": "Ollama Debug History",
"OllamaHistoryLength": "Ollama History Length",
+6
View File
@@ -0,0 +1,6 @@
"""Internal package: shared LLM provider abstraction (ADR-007).
Not a ComfyUI node module — nothing here is registered in
``NODE_CLASS_MAPPINGS``. ``comfydv.ollama`` (and, in a follow-on epic,
``comfydv.llamacpp``) import from here.
"""
+335
View File
@@ -0,0 +1,335 @@
"""Shared chat_structured() helper — pydantic-ai backed structured output.
Used by ``LlamaCppProvider.chat_structured()`` (ADR-007) over llama-server's
OpenAI-compatible ``/v1/chat/completions``. ``OllamaProvider`` no longer uses
this module (ADR-009): Ollama's OpenAI-compatible endpoint was found to
silently reload the model at its default context size on every call,
discarding any ``options.num_ctx`` override even when included in that same
request — a behavior specific to Ollama's compat layer, not llama-server's.
``OllamaProvider.chat_structured()`` now hand-rolls its own structured-output
call over Ollama's *native* ``/api/chat`` + ``"format"``, which doesn't have
that problem.
ADR-009: the Agent uses ``NativeOutput`` (``response_format``/JSON-schema
constrained decoding), not pydantic-ai's default tool-calling. Live-tested
against a "thinking"-capable model: tool-calling let the model spend its
whole token budget on chain-of-thought reasoning and never emit the tool
call; native output keeps reasoning in a separate response field and the
constrained ``content`` always comes back as schema-valid JSON. This benefit
still applies to llama.cpp, which is why this module (and its NativeOutput
choice) is kept for that provider.
Ports ADR-006's retry/validation contract exactly: bounded retries (0-5,
clamped), and a ``RuntimeError`` naming the model, attempt count, and a
truncated snippet of the last invalid response on exhausted retries. The
Agent's own internal retries are disabled (``retries=0``) — this helper
drives its own retry loop so the error contract is comfydv's, not
pydantic-ai's internal one.
"""
import asyncio
from collections.abc import Callable
from typing import cast
from pydantic import BaseModel, ValidationError
from pydantic_ai import Agent, NativeOutput
from pydantic_ai.exceptions import ModelRetry, UnexpectedModelBehavior
from pydantic_ai.messages import (
BinaryContent,
ModelRequest,
ModelResponse,
SystemPromptPart,
TextPart,
UserPromptPart,
)
from pydantic_ai.models.openai import OpenAIChatModel
from pydantic_ai.providers.openai import OpenAIProvider
from pydantic_ai.settings import ModelSettings
from .provider import Message
from .retry import (
RETRY_BACKOFF_SECS,
EmbedFn,
format_recovered_status,
format_retry_status,
is_refusal,
next_seed,
next_timeout_secs,
record_attempt_info,
)
_STRUCTURED_OUTPUT_FAILURE_EXCEPTIONS = (
UnexpectedModelBehavior,
ModelRetry,
ValidationError,
)
def _build_agent(
*,
base_url: str,
model: str,
schema: type[BaseModel],
headers: dict | None,
timeout_secs: float,
) -> Agent:
import httpx
http_client = httpx.AsyncClient(
headers=headers or None, timeout=httpx.Timeout(timeout_secs)
)
provider = OpenAIProvider(
base_url=base_url, api_key="not-needed", http_client=http_client
)
chat_model = OpenAIChatModel(model, provider=provider)
return Agent(chat_model, output_type=NativeOutput(schema), retries=0)
def _user_prompt_content(msg: Message):
"""Render a user turn as pydantic-ai user-prompt content.
Text-only ``msg`` → the plain ``content`` string, byte-identical to the
pre-009 path (FR-003). A turn carrying images → ``[content, *images]``
where each image is a ``BinaryContent`` PNG (ADR-008 / research.md
Decision 1); ``OpenAIChatModel`` renders these as OpenAI ``image_url``
parts, so both backends reach the same multimodal request through one
shared code path.
"""
if not msg.images:
return msg.content
import base64
content: list = [msg.content]
for image in msg.images:
content.append(
BinaryContent(data=base64.b64decode(image), media_type="image/png")
)
return content
def _history_to_messages(messages: list[Message]) -> list:
"""Convert all but the last message into pydantic-ai's typed history.
The last message (the current turn) is passed separately as
``Agent.run()``'s ``user_prompt`` — see ``chat_structured()``.
"""
history: list = []
for msg in messages[:-1]:
if msg.role == "assistant":
history.append(ModelResponse(parts=[TextPart(msg.content)]))
elif msg.role == "system":
history.append(ModelRequest(parts=[SystemPromptPart(msg.content)]))
else:
history.append(
ModelRequest(parts=[UserPromptPart(_user_prompt_content(msg))])
)
return history
async def chat_structured(
*,
base_url: str,
model: str,
messages: list[Message],
schema: type[BaseModel],
headers: dict | None = None,
options: dict | None = None,
max_retries: int = 2,
timeout_secs: float = 300.0,
embed_fn: EmbedFn | None = None,
attempt_info: dict | None = None,
on_status: Callable[[str], None] | None = None,
) -> BaseModel:
"""Call ``model`` at ``base_url`` (an OpenAI-compatible ``/v1`` root) and
return a validated instance of ``schema``.
``options`` is forwarded verbatim as a top-level ``"options"`` field in
the request body via pydantic-ai's ``extra_body`` — the same shape the
pre-ADR-007 hand-rolled implementation sent, so provider-native sampling
params (Ollama's ``num_predict``/``repeat_penalty``/etc., set via the
``OllamaOption*`` nodes) keep working unchanged rather than being
lossily remapped onto pydantic-ai's own standardized ``ModelSettings``
fields.
ADR-010: ``options`` may also carry a ``"think"`` key (bool), popped out
here rather than forwarded inside the nested ``options`` object —
llama-server's OpenAI-compatible endpoint doesn't recognize a literal
``"think"`` key there. Translated to its own two documented
request-body toggles instead: ``chat_template_kwargs:
{"enable_thinking": ...}`` (Qwen3-style models) and, when disabling,
``reasoning_effort: "none"`` (the more model-agnostic OpenAI convention
llama-server also honors) — both via ``extra_body`` the same way
``options`` is. This provider only serves ``LlamaCppProvider`` — see
``OllamaProvider``'s own hand-rolled ``chat_structured`` for why Ollama
needed a different mechanism entirely. Sourced from llama.cpp's server
docs, not live-verified against a running llama-server (no instance
available at implementation time) — verify against your own deployment.
Retries up to ``max_retries`` times (clamped 0-5) on validation failure
before raising ``RuntimeError``. Never returns a value that failed
validation against ``schema``.
``options`` may also carry a ``"refusal_retry"`` config dict (same
comfydv-level convention as ``"think"``, emitted by
``OllamaOptionRefusalRetry``) — a detected refusal/deflection (see
``_llm/retry.py``) is treated exactly like a validation failure: retried
with a bumped seed rather than returned to the caller. ``embed_fn`` is
``LlamaCppProvider``'s own ``embed()``, bound to whatever embedding
model the config names — passed in rather than looked up here since
this module has no provider instance of its own to call.
"""
if not messages or messages[-1].role != "user":
raise ValueError(
"chat_structured requires the last message to have role='user'"
)
history = _history_to_messages(messages)
prompt = _user_prompt_content(messages[-1])
think = None
if options and "think" in options:
options = dict(options)
think = options.pop("think")
options = options or None
refusal_cfg = None
if options and "refusal_retry" in options:
options = dict(options)
refusal_cfg = options.pop("refusal_retry")
options = options or None
extra_body: dict = {}
if options:
extra_body["options"] = options
if think is not None:
extra_body["chat_template_kwargs"] = {"enable_thinking": think}
if not think:
extra_body["reasoning_effort"] = "none"
model_settings: ModelSettings | None = (
{"extra_body": extra_body} if extra_body else None
)
total_attempts = max(0, min(int(max_retries), 5)) + 1
last_error: Exception | None = None
last_invalid_text = ""
refusal_count = 0
attempt_seed = (options or {}).get("seed", 0) if isinstance(options, dict) else 0
attempt_timeout = timeout_secs
def _emit_retry_status(reason: str, attempt: int) -> None:
if on_status is None or attempt >= total_attempts:
return
upcoming_seed = next_seed(options, attempt + 1)
upcoming_timeout = next_timeout_secs(timeout_secs, attempt + 1)
on_status(
format_retry_status(
reason, attempt, total_attempts, upcoming_seed, upcoming_timeout
)
)
for attempt in range(1, total_attempts + 1):
attempt_timeout = next_timeout_secs(timeout_secs, attempt)
# Rebuilt each attempt so the escalated timeout actually takes
# effect — httpx.AsyncClient's timeout is fixed at construction,
# not mutable per-request.
agent = _build_agent(
base_url=base_url,
model=model,
schema=schema,
headers=headers,
timeout_secs=attempt_timeout,
)
attempt_settings = dict(model_settings) if model_settings else {}
if attempt > 1:
# Confirmed live: a freshly-loaded model's first structured-output
# attempt can fail outright (no valid tool call at all) and then
# behave normally on the very next call. Retrying with the exact
# same request reproduces the same failure if the model is
# genuinely stuck rather than just unlucky, so force a new seed
# (pydantic-ai maps ModelSettings["seed"] to the OpenAI API's
# top-level "seed" param, which works against both Ollama's and
# llama-server's OpenAI-compatible endpoints) and give it a beat
# via RETRY_BACKOFF_SECS in case it's still finishing loading.
seed = next_seed(options, attempt)
attempt_seed = seed
attempt_settings["seed"] = seed
if "extra_body" in attempt_settings:
# beacon-reviewer caught this: if a caller pinned options["seed"],
# it's also sitting in extra_body.options.seed (the Ollama-native
# passthrough). Left untouched, a backend that honors that nested
# field over the top-level OpenAI "seed" above would keep sending
# the same old seed on every retry — silently defeating this fix
# for exactly the pinned-seed case. Copy rather than mutate in
# place: extra_body/options here are the caller's own dicts,
# shared across every attempt (and possibly other calls).
# ModelSettings declares extra_body as `object` (it's an
# opaque passthrough field), so a plain dict() call on it
# doesn't type-check — cast first, this module always builds
# it as a dict (see model_settings above).
extra_body = dict(cast(dict, attempt_settings["extra_body"]))
nested_options = dict(extra_body.get("options") or {})
nested_options["seed"] = seed
extra_body["options"] = nested_options
attempt_settings["extra_body"] = extra_body
try:
result = await agent.run(
prompt,
message_history=history,
model_settings=cast(ModelSettings, attempt_settings)
if attempt_settings
else None,
)
# agent's output_type is the caller's `schema` (a runtime value,
# not a static type parameter), so the checker can't narrow
# result.output past Agent's default `str` — cast to the
# function's declared return type, which schema is a subtype of.
output = cast(BaseModel, result.output)
if refusal_cfg and refusal_cfg.get("enabled"):
# Re-serialized, not the original wire text — pydantic-ai's
# NativeOutput doesn't expose that separately, and the
# regex/embedding check works the same either way (same
# textual content, just re-encoded).
content = output.model_dump_json()
refused = await is_refusal(
content,
embed_fn=embed_fn,
embed_cache_key=refusal_cfg.get("embedding_model", ""),
threshold=refusal_cfg.get("threshold", 0.82),
custom_phrases=tuple(refusal_cfg.get("custom_phrases") or ()),
)
if refused:
refusal_count += 1
last_error = RuntimeError("refusal/deflection detected")
last_invalid_text = content
_emit_retry_status("Refusal/deflection detected", attempt)
if attempt < total_attempts:
await asyncio.sleep(RETRY_BACKOFF_SECS)
continue
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=attempt,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
if on_status is not None and attempt > 1:
on_status(
format_recovered_status(attempt, total_attempts, attempt_seed)
)
return output
except _STRUCTURED_OUTPUT_FAILURE_EXCEPTIONS as exc:
last_error = exc
last_invalid_text = str(exc)
_emit_retry_status("Structured output failed", attempt)
if attempt < total_attempts:
await asyncio.sleep(RETRY_BACKOFF_SECS)
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=total_attempts,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
raise RuntimeError(
f"chat_structured: response failed validation against schema after "
f"{total_attempts} attempt(s) (model={model!r}). Last error: "
f"{last_error}. Last response (truncated): {last_invalid_text[:300]!r}"
)
+440
View File
@@ -0,0 +1,440 @@
"""LlamaCppProvider — LLMProvider implementation backed by llama-server's
router mode.
Mirrors comfydv._llm.ollama_provider's structure exactly (ADR-007's parallel-
implementation pattern). Router-mode API shape verified live against
ggml-org/llama.cpp's tools/server/README.md (postdates training data) — see
specs/008-llamacpp-integration/research.md. Two details differ from Ollama:
the model identifier field is "id" (not "name"), and "status" is a nested
object ({"value": "..."}), not a flat string.
Deployment prerequisite: llama-server must be launched with --models-dir or
--models-preset (router mode) — the endpoints this provider calls don't
exist otherwise (spec.md FR-006).
"""
import asyncio
import logging
from collections.abc import Callable
from pydantic import BaseModel
from .ollama_provider import (
_TTLLRUCache,
_cache_key,
_get_json,
_pop_refusal_retry,
_pop_think,
_post_json,
)
from .provider import Message, ModelInfo, ModelStatus
from .retry import (
RETRY_BACKOFF_SECS,
format_recovered_status,
format_retry_status,
is_refusal,
next_seed,
next_timeout_secs,
record_attempt_info,
)
logger = logging.getLogger(__name__)
# Own cache pool, not shared with OllamaProvider's — see plan.md's Structure
# Decision (parallel, symmetric, independent implementations). ChatCompletion
# is OUTPUT_NODE=True and re-executes every queue run regardless of which
# provider is wired in, so caching parity matters for llama.cpp too, not
# just Ollama.
_MODEL_LIST_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=20.0)
_CHAT_RESPONSE_CACHE = _TTLLRUCache(maxsize=64, ttl_seconds=None)
async def _fetch_models(host: str, headers: dict | None = None) -> list[str]:
"""Name-only view for ComfyUI's combo-widget population (the JS refresh
button and node-creation auto-populate) — mirrors
ollama_provider._fetch_models's narrower, gracefully-degrading contract.
Deliberately more forgiving than LlamaCppProvider.list_models(): that
method raises on a non-router-mode server (FR-006, for real workflow
execution, where a silent empty result would be misleading). This
combo-population use case wants the same quiet "just show an empty
dropdown" degradation Ollama's nodes already give for *any* failure —
consistent UX across backends for this specific, lower-stakes path.
"""
try:
models = await LlamaCppProvider(host, headers).list_models()
except Exception as exc:
logger.warning("Could not fetch llama.cpp models from %s: %s", host, exc)
return []
return [m.name for m in models]
def _to_openai_message(message: Message) -> dict:
"""Render a ``Message`` in llama.cpp's OpenAI-compatible shape.
A text-only turn stays ``{"role", "content": <str>}`` — byte-identical to
the pre-009 payload (FR-003). A turn carrying images becomes OpenAI
multimodal ``content`` parts: the text followed by one ``image_url`` part
per base64 image, as a ``data:`` URI (ADR-008). ``llama-server`` only
honours these parts when launched with a multimodal projector
(``--mmproj``); without it the server errors, surfaced to the caller
rather than crashed on (FR-006).
"""
if not message.images:
return {"role": message.role, "content": message.content}
parts: list[dict] = [{"type": "text", "text": message.content}]
for image in message.images:
parts.append(
{
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{image}"},
}
)
return {"role": message.role, "content": parts}
class LlamaCppProvider:
"""LLMProvider implementation backed by llama-server's router mode.
Host and headers are captured once at construction — every method
reuses them, the same pattern OllamaProvider already established.
"""
def __init__(self, host: str, headers: dict | None = None):
self.host = host
self.headers = dict(headers) if headers else None
async def list_models(self) -> list[ModelInfo]:
"""GET {host}/models — every model llama-server's router knows
about, with its live status. Unlike OllamaProvider, no
normalization is needed: llama.cpp's status vocabulary is exactly
ModelStatus's full set.
"""
cache_key = _cache_key("llamacpp_list_models", self.host, self.headers or {})
cached, hit = _MODEL_LIST_CACHE.get(cache_key)
if hit:
return cached
try:
data = await _get_json(f"{self.host}/models", headers=self.headers)
except OSError as exc:
# Genuinely unreachable (connection refused, DNS failure, timed
# out — aiohttp's connection-level exceptions are all OSError
# subclasses) — degrade gracefully like OllamaProvider does, so
# a not-yet-started server just shows an empty dropdown rather
# than a hard error.
logger.warning(
"Could not fetch llama.cpp models from %s: %s", self.host, exc
)
return []
except RuntimeError as exc:
# The server answered but with an HTTP error status — GET
# /models only exists in router mode, so this is almost always
# a llama-server launched without --models-dir/--models-preset.
# Surfacing this distinctly (FR-006) matters: silently returning
# [] here would be indistinguishable from "no models installed".
raise RuntimeError(
f"llama-server at {self.host} did not return a model list from "
f"GET {self.host}/models — is it running in router mode "
f"(--models-dir or --models-preset)? Underlying error: {exc}"
) from exc
models = []
for m in data.get("data", []):
status_value = m.get("status", {}).get("value")
try:
status = ModelStatus(status_value)
except ValueError:
logger.warning(
"llama.cpp reported an unrecognized model status %r for %r — "
"skipping status normalization, this model will be omitted",
status_value,
m.get("id"),
)
continue
models.append(ModelInfo(name=m["id"], status=status, size=None))
if models:
_MODEL_LIST_CACHE.set(cache_key, models)
return models
async def load_model(self, model: str) -> None:
if not model.strip():
raise ValueError("model name cannot be empty")
try:
await _post_json(
f"{self.host}/models/load",
{"model": model},
headers=self.headers,
)
except RuntimeError as exc:
# Confirmed live: router mode's own /models/load is NOT
# idempotent — it 400s "model is already running" rather than
# the {"success": true} the contract assumed. The LLMProvider
# protocol requires load_model() to be idempotent, so this
# error is the desired end-state, not a failure — absorb it
# here rather than leaking the wire-level quirk to callers.
if "model is already running" not in str(exc):
raise
async def unload_model(self, model: str) -> None:
if not model.strip():
raise ValueError("model name cannot be empty")
try:
await _post_json(
f"{self.host}/models/unload",
{"model": model},
headers=self.headers,
)
except RuntimeError as exc:
# Mirror of load_model()'s non-idempotency above, confirmed live:
# /models/unload 400s "model is not running" on an already-
# unloaded model instead of {"success": true}.
if "model is not running" not in str(exc):
raise
async def chat(
self,
model: str,
messages: list[Message],
options: dict | None = None,
timeout_secs: float = 300.0,
max_retries: int = 2,
attempt_info: dict | None = None,
on_status: Callable[[str], None] | None = None,
) -> str:
payload_messages = [_to_openai_message(m) for m in messages]
options, think = _pop_think(options)
options, refusal_cfg = _pop_refusal_retry(options)
embed_fn = None
custom_phrases: tuple[str, ...] = ()
if refusal_cfg and refusal_cfg.get("enabled"):
custom_phrases = tuple(refusal_cfg.get("custom_phrases") or ())
if refusal_cfg.get("embedding_model"):
embedding_model = refusal_cfg["embedding_model"]
embed_fn = lambda t: self.embed(embedding_model, t) # noqa: E731
total_attempts = max(0, min(int(max_retries), 5)) + 1
response_text = ""
refusal_count = 0
attempt_seed = 0
attempt_timeout = timeout_secs
for attempt in range(1, total_attempts + 1):
attempt_timeout = next_timeout_secs(timeout_secs, attempt)
payload: dict = {
"model": model,
"messages": payload_messages,
"stream": False,
}
if options:
# Passed through verbatim, same nesting OllamaProvider.chat()
# uses (payload["options"] = options) — the OllamaOption*
# nodes emit Ollama-native parameter names (num_predict,
# repeat_penalty, ...), which llama-server's OpenAI-compatible
# endpoint won't recognize either way; translating them is
# out of scope for this epic (plan.md Non-goals — no changes
# to the generic nodes). This keeps the two providers'
# handling consistent rather than silently special-casing
# one of them.
payload["options"] = options
if think is not None:
# ADR-010: llama-server's two documented reasoning toggles —
# sourced from server docs, not live-verified (no instance
# available at implementation time).
payload["chat_template_kwargs"] = {"enable_thinking": think}
if not think:
payload["reasoning_effort"] = "none"
if attempt > 1:
# Unlike the options-passthrough above, this IS the OpenAI
# spec's actual top-level "seed" field, so it takes effect
# against llama-server's /v1/chat/completions.
payload["seed"] = next_seed(options, attempt)
attempt_seed = payload["seed"]
else:
# attempt 1 never sets the top-level "seed" field above (only
# retries do) — fall back to whatever the caller pinned in
# options, so attempt_info/seed_used reports the real seed in
# play even on a first-attempt success, not a stale 0.
attempt_seed = (options or {}).get("seed", 0)
cache_key = _cache_key(
"llamacpp_chat",
self.host,
self.headers or {},
model,
payload_messages,
options or {},
think,
payload.get("seed"),
)
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
if hit:
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=attempt,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
return cached
result = await _post_json(
f"{self.host}/v1/chat/completions",
payload,
timeout=attempt_timeout,
headers=self.headers,
)
choices = result.get("choices") or []
response_text = (
choices[0].get("message", {}).get("content", "") or ""
if choices
else ""
)
retry_reason: str | None = None
if response_text.strip():
refused = False
if refusal_cfg and refusal_cfg.get("enabled"):
refused = await is_refusal(
response_text,
embed_fn=embed_fn,
embed_cache_key=refusal_cfg.get("embedding_model", ""),
threshold=refusal_cfg.get("threshold", 0.82),
custom_phrases=custom_phrases,
)
if not refused:
_CHAT_RESPONSE_CACHE.set(cache_key, response_text)
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=attempt,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
if on_status is not None and attempt > 1:
on_status(
format_recovered_status(
attempt, total_attempts, attempt_seed
)
)
return response_text
refusal_count += 1
retry_reason = "Refusal/deflection detected"
else:
retry_reason = "Blank response"
if attempt < total_attempts:
if on_status is not None:
upcoming_seed = next_seed(options, attempt + 1)
upcoming_timeout = next_timeout_secs(timeout_secs, attempt + 1)
on_status(
format_retry_status(
retry_reason,
attempt,
total_attempts,
upcoming_seed,
upcoming_timeout,
)
)
await asyncio.sleep(RETRY_BACKOFF_SECS)
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=total_attempts,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
# Every attempt came back blank — never raises here (chat() has
# never validated its output, unlike chat_structured()); return the
# last (blank) attempt uncached so the next queue run tries fresh.
return response_text
async def chat_structured(
self,
model: str,
messages: list[Message],
schema: type[BaseModel],
options: dict | None = None,
timeout_secs: float = 300.0,
max_retries: int = 2,
attempt_info: dict | None = None,
on_status: Callable[[str], None] | None = None,
) -> BaseModel:
from .chat import chat_structured as _chat_structured_impl
payload_messages = [m.model_dump() for m in messages]
cache_key = _cache_key(
"llamacpp_chat_structured",
self.host,
self.headers or {},
model,
payload_messages,
options or {},
schema.model_json_schema(),
)
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
if hit:
record_attempt_info(
attempt_info,
seed=(options or {}).get("seed", 0),
attempts=1,
timeout_secs=timeout_secs,
refusals=0,
)
return schema.model_validate(cached)
embed_fn = None
refusal_cfg = (options or {}).get("refusal_retry")
if (
refusal_cfg
and refusal_cfg.get("enabled")
and refusal_cfg.get("embedding_model")
):
embedding_model = refusal_cfg["embedding_model"]
embed_fn = lambda t: self.embed(embedding_model, t) # noqa: E731
result = await _chat_structured_impl(
base_url=f"{self.host}/v1",
model=model,
messages=messages,
schema=schema,
headers=self.headers,
options=options,
max_retries=max_retries,
timeout_secs=timeout_secs,
embed_fn=embed_fn,
attempt_info=attempt_info,
on_status=on_status,
)
_CHAT_RESPONSE_CACHE.set(cache_key, result.model_dump())
return result
async def embed(self, model: str, text: str) -> list[float] | None:
"""POST {host}/v1/embeddings — llama-server's OpenAI-compatible
embeddings endpoint.
Requires an embedding-capable model to be loaded in the router
(typically a *different* model from whatever's answering chat
requests) — not live-verified against a running llama-server (no
instance available at implementation time), mirroring this
provider's other sourced-from-docs-not-verified caveats. Returns
``None`` rather than raising on any failure, same contract as
``OllamaProvider.embed()``.
"""
if not model.strip() or not text.strip():
return None
try:
result = await _post_json(
f"{self.host}/v1/embeddings",
{"model": model, "input": text},
timeout=30.0,
headers=self.headers,
)
except Exception:
return None
data = result.get("data")
if not isinstance(data, list) or not data:
return None
vec = data[0].get("embedding")
if not isinstance(vec, list) or not vec:
return None
return vec
+731
View File
@@ -0,0 +1,731 @@
"""OllamaProvider — LLMProvider implementation backed by Ollama's REST API.
Ported from comfydv.ollama's original module-level HTTP/cache helpers
(_post_json, _fetch_models, _run_async, _TTLLRUCache) — behavior-preserving,
not a rewrite. See ADR-007 and
specs/007-llm-provider-abstraction/research.md.
OllamaProvider implements the LLMProvider Protocol structurally (no explicit
inheritance — that's the point of typing.Protocol); conformance is checked
by ``ty check``, not the runtime.
"""
import asyncio
import json
import logging
import threading
import time
from collections.abc import Callable
from pydantic import BaseModel, ValidationError
from .provider import Message, ModelInfo, ModelStatus
from .retry import (
RETRY_BACKOFF_SECS,
format_recovered_status,
format_retry_status,
is_refusal,
next_seed,
next_timeout_secs,
record_attempt_info,
)
logger = logging.getLogger(__name__)
# ---------------------------------------------------------------------------
# Local response cache (ported from comfydv.ollama)
# ---------------------------------------------------------------------------
#
# See comfydv.ollama's original module docstring for why this exists:
# OUTPUT_NODE=True chat nodes re-execute every queue run even when inputs
# are unchanged; this cache absorbs the redundant round-trips.
class _TTLLRUCache:
"""Bounded cache, LRU-evicted, with an optional per-entry TTL."""
def __init__(self, maxsize: int, ttl_seconds: float | None = None):
self.maxsize = maxsize
self.ttl_seconds = ttl_seconds
self._data: dict = {}
self._lock = threading.Lock()
def get(self, key):
with self._lock:
entry = self._data.get(key)
if entry is None:
return None, False
expires_at, value = entry
if self.ttl_seconds is not None and time.monotonic() > expires_at:
del self._data[key]
return None, False
# Re-insert to mark as most-recently-used (dicts preserve insertion order).
del self._data[key]
self._data[key] = (expires_at, value)
return value, True
def set(self, key, value):
with self._lock:
expires_at = (
time.monotonic() + self.ttl_seconds
if self.ttl_seconds is not None
else float("inf")
)
self._data.pop(key, None)
self._data[key] = (expires_at, value)
while len(self._data) > self.maxsize:
oldest_key = next(iter(self._data))
del self._data[oldest_key]
def clear(self):
with self._lock:
self._data.clear()
def _cache_key(*parts) -> str:
"""Deterministic, hashable key from arbitrary JSON-serializable parts."""
return json.dumps(parts, sort_keys=True, default=str)
def _pop_think(options: dict | None) -> tuple[dict | None, bool | None]:
"""Split a ``"think"`` toggle out of a generic ``options`` dict.
ADR-010: ``OllamaOptionDisableThinking`` merges a ``"think": bool`` key
into the same composable ``OLLAMA_OPTIONS`` chain every other
``OllamaOption*`` node feeds into ``ChatCompletion``'s ``options``
input — but unlike those (Ollama-native sampling params, passed through
verbatim), ``"think"`` needs real per-provider translation: neither
Ollama's native ``/api/chat`` nor llama-server's OpenAI-compatible
endpoint recognizes a literal ``"think"`` key nested inside their own
``options``/sampling-params object, so every provider pops it out here
(or in ``LlamaCppProvider``'s own copy) before building its request.
Returns ``options`` with ``"think"`` removed (unchanged if absent, so a
falsy/empty result stays falsy) and the popped value, or ``None`` if the
caller didn't set it — never touches the caller's own dict in place.
"""
if not options or "think" not in options:
return options, None
remaining = dict(options)
think = remaining.pop("think")
return (remaining or None), think
def _pop_refusal_retry(options: dict | None) -> tuple[dict | None, dict | None]:
"""Split a ``"refusal_retry"`` config dict out of a generic ``options``
dict — same convention as ``_pop_think``: ``OllamaOptionRefusalRetry``
merges ``{"refusal_retry": {"enabled", "embedding_model", "threshold"}}``
into the same composable ``OLLAMA_OPTIONS`` chain every other
``OllamaOption*`` node feeds into ``ChatCompletion``'s ``options``
input, and neither Ollama's nor llama.cpp's own API recognizes this key,
so every provider pops it out here before building its request.
"""
if not options or "refusal_retry" not in options:
return options, None
remaining = dict(options)
cfg = remaining.pop("refusal_retry")
return (remaining or None), cfg
_MODEL_LIST_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=20.0)
_CHAT_RESPONSE_CACHE = _TTLLRUCache(maxsize=64, ttl_seconds=None)
_CAPABILITY_CACHE = _TTLLRUCache(maxsize=32, ttl_seconds=300.0)
# ---------------------------------------------------------------------------
# Async infrastructure (ported from comfydv.ollama)
# ---------------------------------------------------------------------------
def _run_async(coro):
"""Run an async coroutine synchronously in an isolated worker thread.
Always spins up a fresh thread rather than conditionally checking
asyncio.get_running_loop() first: live-verified against a real running
ComfyUI instance (its execution engine runs its own event loop, in
Python 3.13, on the same process) that the conditional version — try
get_running_loop(), spin up a thread only if it succeeds, otherwise
call asyncio.run(coro) directly — is unreliable there. Under real
ComfyUI, get_running_loop() sometimes raised inside that try block
(unlike under pytest or a standalone script, where it never does),
which routed straight into `asyncio.run(coro)` on the *current* thread
— the one thread guaranteed to already have ComfyUI's own loop running
— reproducing exactly the "asyncio.run() cannot be called from a
running event loop" crash this function exists to prevent. Always
using a dedicated thread sidesteps the detection entirely: a freshly
spawned thread never has an ambient loop, so asyncio.run() is safe
there unconditionally, regardless of what the calling thread's loop
state actually is.
"""
import concurrent.futures
with concurrent.futures.ThreadPoolExecutor(max_workers=1) as pool:
return pool.submit(asyncio.run, coro).result()
async def _post_json(
url: str,
payload: dict,
*,
timeout: float = 120.0,
headers: dict | None = None,
) -> dict:
"""POST JSON to url, return parsed response dict."""
import aiohttp
try:
async with aiohttp.ClientSession() as session:
async with session.post(
url,
json=payload,
headers=headers or None,
timeout=aiohttp.ClientTimeout(total=timeout),
) as resp:
if resp.status >= 400:
body = await resp.text()
raise RuntimeError(
f"Ollama returned HTTP {resp.status} for {url}: {body[:300]}"
)
return await resp.json()
except aiohttp.ClientConnectionError as exc:
raise RuntimeError(f"Cannot reach Ollama at {url}: {exc}") from exc
async def _get_json(
url: str, *, timeout: float = 5.0, headers: dict | None = None
) -> dict:
"""GET url, return parsed response dict.
Raises RuntimeError on an HTTP error status (distinct message, so callers
can tell "server responded with an error" from "couldn't reach it at
all" — aiohttp connection/timeout errors propagate unwrapped for that
reason). Message is generic, not backend-branded: this helper is shared
by every LLMProvider implementation.
"""
import aiohttp
async with aiohttp.ClientSession() as session:
async with session.get(
url,
headers=headers or None,
timeout=aiohttp.ClientTimeout(total=timeout),
) as resp:
if resp.status >= 400:
body = await resp.text()
raise RuntimeError(
f"Server returned HTTP {resp.status} for {url}: {body[:300]}"
)
return await resp.json()
async def _fetch_models(host: str, headers: dict | None = None) -> list[str]:
"""GET {host}/api/tags — return list of model name strings.
Used by comfydv.ollama's combo-widget population (_load_default_models,
the /dv/ollama/models route) — a narrower, name-only view than
OllamaProvider.list_models(), which returns full ModelInfo with status.
Cached for _MODEL_LIST_CACHE.ttl_seconds per (host, headers) pair.
"""
cache_key = _cache_key("models", host, headers or {})
cached, hit = _MODEL_LIST_CACHE.get(cache_key)
if hit:
return cached
try:
data = await _get_json(f"{host}/api/tags", headers=headers)
models = [m["name"] for m in data.get("models", [])]
except Exception as exc:
logger.warning("Could not fetch Ollama models from %s: %s", host, exc)
return []
if models:
_MODEL_LIST_CACHE.set(cache_key, models)
return models
async def _require_vision_capability(
host: str, model: str, headers: dict | None
) -> None:
"""Raise a clear error if ``model`` lacks Ollama's ``vision`` capability.
Only called when a request carries at least one image (spec 009 FR-006):
Ollama's /api/chat silently accepts an unsupported ``images`` field and
answers with a blank/malformed HTTP 200 instead of an error — which
would otherwise be indistinguishable from an ordinary blank generation
and get swallowed by chat()'s existing blank-response retry. /api/show's
``capabilities`` list is the only place Ollama states support explicitly,
so a request carrying an image is checked against it up front.
Fails open on any lookup problem (older Ollama without ``capabilities``,
unreachable host, unexpected shape) — a lookup failure must not block a
request that would otherwise have worked; the real request surfaces its
own clear error if the host is genuinely unreachable.
"""
cache_key = _cache_key("capabilities", host, headers or {}, model)
cached, hit = _CAPABILITY_CACHE.get(cache_key)
if hit:
capabilities = cached
else:
try:
data = await _post_json(
f"{host}/api/show", {"model": model}, timeout=10.0, headers=headers
)
except Exception:
return
capabilities = data.get("capabilities")
if capabilities is None:
return
_CAPABILITY_CACHE.set(cache_key, capabilities)
if "vision" not in capabilities:
raise ValueError(
f"Model '{model}' does not support image input — Ollama reports "
f"capabilities {capabilities!r} for it, no 'vision'. Wire a "
"vision-capable model, or disconnect the image input for "
"text-only chat."
)
class OllamaProvider:
"""LLMProvider implementation backed by Ollama's REST API.
Host and headers are captured once at construction — every method
reuses them, matching the ADR-005 config-node pattern (one
``OllamaClient`` node's output is one ``OllamaProvider`` instance).
"""
def __init__(self, host: str, headers: dict | None = None):
self.host = host
self.headers = dict(headers) if headers else None
async def list_models(self) -> list[ModelInfo]:
"""Every installed model, with live loaded/unloaded status.
`/api/tags` lists installed models; `/api/ps` lists currently-loaded
ones. Ollama has no `sleeping`/`downloading` concept via this API —
never emitted here (ADR-007's documented approximation).
"""
cache_key = _cache_key("list_models", self.host, self.headers or {})
cached, hit = _MODEL_LIST_CACHE.get(cache_key)
if hit:
return cached
try:
tags = await _get_json(f"{self.host}/api/tags", headers=self.headers)
except Exception as exc:
logger.warning("Could not fetch Ollama models from %s: %s", self.host, exc)
return []
loaded_names: set[str] = set()
try:
ps = await _get_json(f"{self.host}/api/ps", headers=self.headers)
loaded_names = {m["name"] for m in ps.get("models", [])}
except Exception as exc:
logger.warning(
"Could not fetch Ollama running models from %s: %s", self.host, exc
)
models = [
ModelInfo(
name=m["name"],
status=(
ModelStatus.LOADED
if m["name"] in loaded_names
else ModelStatus.UNLOADED
),
size=m.get("size"),
)
for m in tags.get("models", [])
]
if models:
_MODEL_LIST_CACHE.set(cache_key, models)
return models
async def load_model(self, model: str) -> None:
if not model.strip():
raise ValueError("model name cannot be empty")
await _post_json(
f"{self.host}/api/generate",
{"model": model, "keep_alive": -1, "stream": False},
timeout=300.0,
headers=self.headers,
)
async def unload_model(self, model: str) -> None:
if not model.strip():
raise ValueError("model name cannot be empty")
await _post_json(
f"{self.host}/api/generate",
{"model": model, "keep_alive": 0, "stream": False},
timeout=30.0,
headers=self.headers,
)
async def chat(
self,
model: str,
messages: list[Message],
options: dict | None = None,
timeout_secs: float = 300.0,
max_retries: int = 2,
attempt_info: dict | None = None,
on_status: Callable[[str], None] | None = None,
) -> str:
if any(m.images for m in messages):
await _require_vision_capability(self.host, model, self.headers)
# exclude_none drops the images key for text-only turns so an
# image-less request is byte-identical to the pre-009 payload
# (FR-003); a turn with images keeps Ollama's native flat images
# array (ADR-008 — no transform needed for /api/chat).
payload_messages = [m.model_dump(exclude_none=True) for m in messages]
options, think = _pop_think(options)
options, refusal_cfg = _pop_refusal_retry(options)
embed_fn = None
custom_phrases: tuple[str, ...] = ()
if refusal_cfg and refusal_cfg.get("enabled"):
custom_phrases = tuple(refusal_cfg.get("custom_phrases") or ())
if refusal_cfg.get("embedding_model"):
embedding_model = refusal_cfg["embedding_model"]
embed_fn = lambda t: self.embed(embedding_model, t) # noqa: E731
total_attempts = max(0, min(int(max_retries), 5)) + 1
response_text = ""
incomplete = False
refusal_count = 0
attempt_seed = 0
attempt_timeout = timeout_secs
for attempt in range(1, total_attempts + 1):
attempt_options = dict(options) if options else {}
if attempt > 1:
attempt_options["seed"] = next_seed(options, attempt)
attempt_seed = attempt_options.get("seed", 0)
attempt_timeout = next_timeout_secs(timeout_secs, attempt)
payload: dict = {
"model": model,
"messages": payload_messages,
"stream": False,
}
if attempt_options:
payload["options"] = attempt_options
if think is not None:
# ADR-010: confirmed live this must be a top-level field —
# Ollama silently ignores "think" nested inside "options".
payload["think"] = think
cache_key = _cache_key(
"chat",
self.host,
self.headers or {},
model,
payload_messages,
attempt_options,
think,
)
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
if hit:
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=attempt,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
return cached
result = await _post_json(
f"{self.host}/api/chat",
payload,
timeout=attempt_timeout,
headers=self.headers,
)
response_text = result.get("message", {}).get("content", "")
retry_reason: str | None = None
if response_text.strip():
refused = False
if refusal_cfg and refusal_cfg.get("enabled"):
refused = await is_refusal(
response_text,
embed_fn=embed_fn,
embed_cache_key=refusal_cfg.get("embedding_model", ""),
threshold=refusal_cfg.get("threshold", 0.82),
custom_phrases=custom_phrases,
)
if not refused:
_CHAT_RESPONSE_CACHE.set(cache_key, response_text)
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=attempt,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
if on_status is not None and attempt > 1:
on_status(
format_recovered_status(
attempt, total_attempts, attempt_seed
)
)
return response_text
# A detected refusal is handled exactly like a blank
# response below: fall through to the backoff/retry with a
# bumped seed (next_seed), rather than returning the refusal
# text to the caller.
refusal_count += 1
retry_reason = "Refusal/deflection detected"
# done: false alongside blank content is a distinct signal from
# an ordinary blank generation — it's Ollama answering before
# the model has actually finished loading/swapping in, observed
# live under model-swap load (issue #27), not the model having
# genuinely generated nothing. Tracked separately so it can be
# raised on below instead of silently returned like a real
# blank generation would be.
incomplete = result.get("done") is False
if retry_reason is None and not response_text.strip():
retry_reason = (
"Model still loading/swapping" if incomplete else "Blank response"
)
if attempt < total_attempts:
if on_status is not None and retry_reason is not None:
upcoming_seed = next_seed(options, attempt + 1)
upcoming_timeout = next_timeout_secs(timeout_secs, attempt + 1)
on_status(
format_retry_status(
retry_reason,
attempt,
total_attempts,
upcoming_seed,
upcoming_timeout,
)
)
await asyncio.sleep(RETRY_BACKOFF_SECS)
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=total_attempts,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
if incomplete:
raise RuntimeError(
f"Ollama returned an incomplete response after "
f"{total_attempts} attempt(s) for model '{model}' — it may "
"still be loading or swapping in memory. Try again in a "
"few seconds."
)
# Every attempt came back blank (and complete) — never raises here
# (chat() has never validated its output, unlike chat_structured());
# return the last (blank) attempt uncached so the next queue run
# tries fresh.
return response_text
async def chat_structured(
self,
model: str,
messages: list[Message],
schema: type[BaseModel],
options: dict | None = None,
timeout_secs: float = 300.0,
max_retries: int = 2,
attempt_info: dict | None = None,
on_status: Callable[[str], None] | None = None,
) -> BaseModel:
"""Native ``/api/chat`` + ``"format"`` (grammar-constrained JSON
decoding), not the shared pydantic-ai ``chat.py`` helper.
ADR-009 originally routed this through the OpenAI-compatible
``/v1/chat/completions`` endpoint via pydantic-ai's ``NativeOutput``.
Confirmed live that endpoint silently *reloads the model at its
default context size on every call*, discarding any prior
``options.num_ctx`` — even when the same ``options`` are included in
that very request. Priming with a separate native call first
(the original fix) didn't help: the very next OpenAI-compat call
undid it immediately. The native ``/api/chat`` endpoint doesn't
have this problem — confirmed live it preserves an already-primed
context, and it supports structured output directly via
``"format"``, so ``options`` and structured output now apply
atomically in one request. ``LlamaCppProvider`` is unaffected — it
keeps using the shared pydantic-ai path, since llama-server's
context is fixed at process launch, not a per-request concern.
"""
if any(m.images for m in messages):
await _require_vision_capability(self.host, model, self.headers)
payload_messages = [m.model_dump(exclude_none=True) for m in messages]
json_schema = schema.model_json_schema()
options, think = _pop_think(options)
options, refusal_cfg = _pop_refusal_retry(options)
embed_fn = None
custom_phrases: tuple[str, ...] = ()
if refusal_cfg and refusal_cfg.get("enabled"):
custom_phrases = tuple(refusal_cfg.get("custom_phrases") or ())
if refusal_cfg.get("embedding_model"):
embedding_model = refusal_cfg["embedding_model"]
embed_fn = lambda t: self.embed(embedding_model, t) # noqa: E731
cache_key = _cache_key(
"chat_structured",
self.host,
self.headers or {},
model,
payload_messages,
options or {},
json_schema,
think,
)
cached, hit = _CHAT_RESPONSE_CACHE.get(cache_key)
if hit:
record_attempt_info(
attempt_info,
seed=(options or {}).get("seed", 0),
attempts=1,
timeout_secs=timeout_secs,
refusals=0,
)
return schema.model_validate(cached)
total_attempts = max(0, min(int(max_retries), 5)) + 1
last_error: Exception | None = None
last_invalid_text = ""
refusal_count = 0
attempt_seed = 0
attempt_timeout = timeout_secs
def _emit_retry_status(reason: str, attempt: int) -> None:
if on_status is None or attempt >= total_attempts:
return
upcoming_seed = next_seed(options, attempt + 1)
upcoming_timeout = next_timeout_secs(timeout_secs, attempt + 1)
on_status(
format_retry_status(
reason, attempt, total_attempts, upcoming_seed, upcoming_timeout
)
)
for attempt in range(1, total_attempts + 1):
attempt_options = dict(options) if options else {}
if attempt > 1:
attempt_options["seed"] = next_seed(options, attempt)
attempt_seed = attempt_options.get("seed", 0)
attempt_timeout = next_timeout_secs(timeout_secs, attempt)
payload: dict = {
"model": model,
"messages": payload_messages,
"format": json_schema,
"stream": False,
}
if attempt_options:
payload["options"] = attempt_options
if think is not None:
# ADR-010: confirmed live this must be a top-level field —
# Ollama silently ignores "think" nested inside "options".
payload["think"] = think
try:
result = await _post_json(
f"{self.host}/api/chat",
payload,
timeout=attempt_timeout,
headers=self.headers,
)
except RuntimeError as exc:
last_error = exc
last_invalid_text = str(exc)
_emit_retry_status("Request failed", attempt)
if attempt < total_attempts:
await asyncio.sleep(RETRY_BACKOFF_SECS)
continue
content = result.get("message", {}).get("content", "")
try:
parsed = schema.model_validate_json(content)
except ValidationError as exc:
last_error = exc
last_invalid_text = content
_emit_retry_status("Schema validation failed", attempt)
if attempt < total_attempts:
await asyncio.sleep(RETRY_BACKOFF_SECS)
continue
if refusal_cfg and refusal_cfg.get("enabled"):
# Checked against the raw JSON text, not a specific parsed
# field: ChatCompletion's schema is caller-defined and this
# provider has no idea which field would carry refusal
# language — the regex/embedding check still matches text
# sitting inside a JSON string value either way.
refused = await is_refusal(
content,
embed_fn=embed_fn,
embed_cache_key=refusal_cfg.get("embedding_model", ""),
threshold=refusal_cfg.get("threshold", 0.82),
custom_phrases=custom_phrases,
)
if refused:
refusal_count += 1
last_error = RuntimeError("refusal/deflection detected")
last_invalid_text = content
_emit_retry_status("Refusal/deflection detected", attempt)
if attempt < total_attempts:
await asyncio.sleep(RETRY_BACKOFF_SECS)
continue
_CHAT_RESPONSE_CACHE.set(cache_key, parsed.model_dump())
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=attempt,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
if on_status is not None and attempt > 1:
on_status(
format_recovered_status(attempt, total_attempts, attempt_seed)
)
return parsed
record_attempt_info(
attempt_info,
seed=attempt_seed,
attempts=total_attempts,
timeout_secs=attempt_timeout,
refusals=refusal_count,
)
raise RuntimeError(
f"chat_structured: response failed validation against schema after "
f"{total_attempts} attempt(s) (model={model!r}). Last error: "
f"{last_error}. Last response (truncated): {last_invalid_text[:300]!r}"
)
async def embed(self, model: str, text: str) -> list[float] | None:
"""POST {host}/api/embed — Ollama's native embeddings endpoint.
Returns ``None`` rather than raising on any failure (wrong/missing
embedding model, unreachable server, malformed response) — this is
a best-effort capability per the ``LLMProvider`` protocol, and its
one current caller (refusal-retry detection) already treats
``None`` as "skip the embedding check", not an error.
"""
if not model.strip() or not text.strip():
return None
try:
result = await _post_json(
f"{self.host}/api/embed",
{"model": model, "input": text},
timeout=30.0,
headers=self.headers,
)
except Exception:
return None
embeddings = result.get("embeddings")
if not isinstance(embeddings, list) or not embeddings:
return None
vec = embeddings[0]
if not isinstance(vec, list) or not vec:
return None
return vec
+161
View File
@@ -0,0 +1,161 @@
"""LLMProvider protocol — the adapter boundary between ComfyUI nodes and
specific local inference backends.
ADR-007: every backend (OllamaProvider now, LlamaCppProvider in a follow-on
epic) implements this shape; ComfyUI nodes depend only on the protocol,
never on a concrete provider class. See
project-management/ADRs/ADR-007-llm-provider-adapter-pattern.md and
specs/007-llm-provider-abstraction/contracts/llm_provider_protocol.md.
"""
from collections.abc import Callable
from enum import Enum
from typing import Literal, Protocol
from pydantic import BaseModel
class ModelStatus(str, Enum):
"""Residency status of a model on a provider's server.
Not every provider emits every value — e.g. Ollama has no distinct
signal for SLEEPING or DOWNLOADING via its API and normalizes to the
closest applicable status rather than omitting the model (documented
approximation, ADR-007).
"""
UNLOADED = "unloaded"
LOADING = "loading"
LOADED = "loaded"
SLEEPING = "sleeping"
DOWNLOADING = "downloading"
class ModelInfo(BaseModel):
"""One entry returned by ``LLMProvider.list_models()``."""
name: str
status: ModelStatus
size: int | None = None
class Message(BaseModel):
"""One turn in a chat request.
``images`` carries optional base64-encoded image payloads (no ``data:``
prefix) associated with this turn, for vision-capable models. ``None``
(the default) means a text-only turn that serializes byte-for-byte as
before — providers dump with ``exclude_none=True`` so no ``images`` key
reaches the wire for image-less turns. Each provider translates this
neutral carrier into its own native shape (ADR-008): Ollama's flat
per-message ``images`` array, llama.cpp's OpenAI ``image_url`` content
parts, and pydantic-ai ``BinaryContent`` on the structured path.
"""
role: Literal["system", "user", "assistant"]
content: str
images: list[str] | None = None
class LLMProvider(Protocol):
"""Adapter boundary every backend implements.
Connection state (host, auth headers, or equivalent) is captured once
at provider-construction time — no method takes connection details as
a parameter.
"""
async def list_models(self) -> list[ModelInfo]:
"""Every model the server currently knows about, loaded or not."""
...
async def load_model(self, model: str) -> None:
"""Load a model into memory. Idempotent — already-loaded is not an error."""
...
async def unload_model(self, model: str) -> None:
"""Unload a model from memory. Idempotent — already-unloaded is not an error."""
...
async def chat(
self,
model: str,
messages: list[Message],
options: dict | None = None,
timeout_secs: float = 300.0,
max_retries: int = 2,
attempt_info: dict | None = None,
on_status: Callable[[str], None] | None = None,
) -> str:
"""Free-text chat response.
Retries up to ``max_retries`` times (clamped 0-5) with a new seed if
the response comes back blank — confirmed live on a freshly-loaded
model, whose first response is sometimes empty before it settles
into normal behavior. Still returns the (possibly blank) last
attempt's text rather than raising if every retry comes back blank —
this method has never validated its output, unlike
``chat_structured()``. Each retry's request timeout also escalates
(``_llm/retry.py``'s ``next_timeout_secs``) rather than reusing the
same budget that just ran out.
ADR-010: ``options`` may carry a ``"think"`` key (bool) to disable a
"thinking"-capable model's chain-of-thought reasoning — every
implementation pops it out of ``options`` and translates it to its
own wire shape (Ollama: a top-level ``think`` field; llama.cpp:
``chat_template_kwargs``/``reasoning_effort`` in the request body),
since neither backend recognizes a literal ``"think"`` key nested
inside a generic options object.
``attempt_info``, if given, is populated in place with the retry
loop's final outcome (seed/timeout used, attempt count, refusal
count) via ``_llm/retry.py``'s ``record_attempt_info`` — an optional
out-param, not a return-type change, so existing callers that don't
pass it see no behavior change.
``on_status``, if given, is called synchronously at each retry
boundary with a one-line human-readable status (see
``_llm/retry.py``'s ``format_retry_status``/``format_recovered_status``)
— a live counterpart to ``attempt_info``, which only reports the
final outcome after the call returns.
"""
...
async def chat_structured(
self,
model: str,
messages: list[Message],
schema: type[BaseModel],
options: dict | None = None,
timeout_secs: float = 300.0,
max_retries: int = 2,
attempt_info: dict | None = None,
on_status: Callable[[str], None] | None = None,
) -> BaseModel:
"""Schema-validated chat response.
Raises ``RuntimeError`` (naming the model, attempt count, and a
truncated snippet of the last invalid response) if every retry is
exhausted — never returns a value with a missing/blank required
field.
ADR-010: see ``chat()`` — same ``options["think"]`` convention,
same per-provider translation, same escalating per-attempt timeout,
and the same ``attempt_info``/``on_status`` conventions.
"""
...
async def embed(self, model: str, text: str) -> list[float] | None:
"""Embedding vector for ``text``, or ``None`` if unavailable.
Best-effort, not a core capability every deployment has configured:
``model`` must itself be embedding-capable, which is typically a
*different* model from whatever's answering chat requests (e.g.
``nomic-embed-text``, not the model passed to ``chat()``). Returns
``None`` rather than raising when embeddings aren't usable right now
(wrong/missing model, unreachable server) — the one current caller,
refusal-retry detection (see ``_llm/retry.py``), degrades gracefully
to lexical-only detection when this returns ``None``, so a provider
with no embedding model configured is never a hard failure.
"""
...
+333
View File
@@ -0,0 +1,333 @@
"""Shared retry-on-empty-output helpers for chat()/chat_structured().
Both providers' chat() calls (ADR-007) and the shared chat_structured()
helper (_llm/chat.py) hit the same class of failure, confirmed live against
a freshly-started Ollama instance on a fresh runpod: the model's first
response after loading is sometimes blank or fails structured-output
validation outright, then behaves normally on the very next call. Centralized
here so both providers and both chat modes retry the same way rather than
each re-deriving the policy.
Only blank/whitespace-only responses trigger a retry for plain chat() —
not merely "short" ones — because a fixed length threshold would misfire on
legitimately short, valid answers (single-word replies, labels, "yes"/"no").
Refusal/deflection detection (below) is a separate, opt-in trigger for the
same retry-with-a-new-seed mechanism: some models (observed with an
abliterated Qwen variant) answer with a soft refusal on a topic they judge
"sensitive" instead of erroring or returning blank, so neither of the above
checks catches it. This is deliberately a model-behavior concern, not a
backend one — every ``LLMProvider`` implementation (Ollama, llama.cpp, and
whatever comes next) wires the same detector into its own retry loop via its
own ``embed()``, rather than each backend inventing its own heuristic.
"""
import math
import re
from collections.abc import Awaitable, Callable
RETRY_BACKOFF_SECS = 1.5
"""Flat delay between retries — gives a still-loading model time to finish
before the next attempt, rather than hammering it with identical requests
back-to-back."""
def next_seed(options: dict | None, attempt: int) -> int:
"""Deterministic seed for retry ``attempt`` (1-indexed).
Attempt 1 is the caller's original request and is never touched by this
function — callers only call it for attempt >= 2. Starts from
``options["seed"]`` if the caller pinned one, else 0, and increments by
``attempt - 1`` so each retry is a new, reproducible value instead of
repeating the exact same request that just failed.
"""
base = 0
if options and isinstance(options.get("seed"), int):
base = options["seed"]
return base + (attempt - 1)
def next_timeout_secs(base_timeout: float, attempt: int) -> float:
"""Escalating per-attempt timeout for retries (1-indexed ``attempt``).
Attempt 1 gets the caller's own ``timeout_secs`` unchanged; each retry
multiplies it by the attempt number. A request that timed out may
genuinely need more time — a slow-to-load or heavily-loaded model, a
large prompt — not just an identical retry under the same budget it
just failed to meet.
"""
return base_timeout * attempt
def record_attempt_info(
attempt_info: dict | None,
*,
seed: int,
attempts: int,
timeout_secs: float,
refusals: int,
) -> None:
"""Populate an optional caller-supplied dict with the retry loop's
final outcome — the seed/timeout actually used, how many attempts it
took, and how many were refusal-triggered.
A plain out-param rather than a return-type change, so it's fully
backward compatible: a caller that doesn't pass ``attempt_info`` sees
no change in behavior at all. ``ChatCompletion`` uses this to expose
the seed actually used as a node output and to build a UI status line
when a retry/refusal happened.
"""
if attempt_info is None:
return
attempt_info.update(
{
"seed": seed,
"attempts": attempts,
"timeout_secs": timeout_secs,
"refusals": refusals,
}
)
OnStatus = Callable[[str], None]
"""A caller-supplied, synchronous, best-effort progress callback — see
``format_retry_status``/``format_recovered_status``. Not async: providers
call it inline mid-retry-loop, and the one real implementation
(``ChatCompletion``'s closure over ``PromptServer.send_progress_text``) is
itself synchronous, so there's nothing to await."""
def format_retry_status(
reason: str, attempt: int, total_attempts: int, seed: int, timeout_secs: float
) -> str:
"""One-line, human-readable status for ``on_status()`` callers — shown
live on the node via ComfyUI's ``PromptServer.send_progress_text``
(see ``ChatCompletion.chat()``). Centralized so every provider's retry
loop describes a retry the same way rather than each inventing its own
wording.
"""
return (
f"⚠ {reason} on attempt {attempt}/{total_attempts} — "
f"retrying with seed={seed}, timeout={timeout_secs:.0f}s"
)
def format_recovered_status(attempt: int, total_attempts: int, seed: int) -> str:
"""Final status shown once a retry loop succeeds after >1 attempt —
lets a live status left over from ``format_retry_status`` resolve to
something other than a stale "retrying..." message."""
return f"✅ Recovered on attempt {attempt}/{total_attempts} (seed={seed})"
# ---------------------------------------------------------------------------
# Refusal/deflection detection
# ---------------------------------------------------------------------------
#
# Hybrid, cheapest-check-first: a fast, free lexical pass catches the blatant
# majority ("I cannot generate...") without ever touching the network; only
# a response that's short and/or hedge-y enough to be genuinely ambiguous
# pays for an embedding call. A long, on-topic response never reaches the
# embedding step at all.
REFUSAL_LEXICAL_PATTERNS: tuple[re.Pattern, ...] = tuple(
re.compile(p, re.IGNORECASE)
for p in (
r"\b(?:I\s*(?:'m|\s+am)?\s*)?(?:cannot|can't|won't|will not)\b[^.]{0,60}?\b"
r"(?:generate|create|produce|write|provide|help|assist|describe|depict|continue)\b",
r"\bI(?:'m|\s+am) (?:(?:not able|unable) to|restricted from)\b",
r"\bI don't feel comfortable\b",
r"\bI'm sorry,?\s*(?:but\s+)?I\s*(?:can't|cannot)\b",
r"\bas an AI\b[^.]{0,60}?\b(?:cannot|can't|unable|not able)\b",
r"\bI must (?:decline|refuse)\b",
r"\bthis (?:request|content|image|prompt) (?:violates|goes against)\b",
r"\bnot something I(?:'m| am)? (?:able to|going to) (?:help with|generate|create|provide)\b",
r"\bI (?:won't|will not) be (?:able to )?(?:generat|creat|provid)",
)
)
_SOFT_HEDGE_KEYWORDS: tuple[str, ...] = (
"sensitive",
"unfortunately",
"i must",
"policy",
"guideline",
"responsible ai",
"not appropriate",
"inappropriate",
"instead, i",
"i'd rather",
"i would rather",
"controversial",
"harmful",
"offensive",
"disturbing",
"explicit content",
"nsfw",
"consider an alternative",
)
_AMBIGUOUS_LENGTH_THRESHOLD = 600
"""Below this many characters, a response is short enough that a soft
refusal is plausible and worth the extra embedding check — chosen well
under this pipeline's normal structured-JSON response sizes (typically
1000+ characters), not a tuned/validated threshold.
Note: ``is_refusal`` no longer applies this gate when the caller supplies
``embed_fn`` (see below) — an explicitly configured embedding model always
runs the check regardless of length/keywords. ``is_ambiguous`` is kept as a
standalone, independently tested heuristic for other callers (e.g. logging
"why was this worth embedding-checking") rather than wired into the hybrid
detector itself."""
REFUSAL_EXEMPLARS: tuple[str, ...] = (
"I cannot generate an image or description involving this topic.",
"I'm not able to help with this request as it involves a politically sensitive subject.",
"I don't feel comfortable creating content related to this figure or event.",
"This request touches on a sensitive political topic that I must decline to engage with.",
"As an AI, I'm unable to produce content depicting this scenario.",
"I'd rather not generate this — it involves content I'm not comfortable with.",
)
def is_lexical_refusal(text: str) -> bool:
"""Cheap, free regex pass — catches the blatant majority of refusals."""
return any(p.search(text) for p in REFUSAL_LEXICAL_PATTERNS)
def is_ambiguous(text: str) -> bool:
"""Whether ``text`` is short/hedge-y enough to be worth the pricier
embedding check, having already failed the free lexical pass.
Deliberately cheap and approximate — false positives here only cost one
extra embedding call, false negatives skip a refusal that a real
similarity check might have caught. Not meant to be a precise signal on
its own, just a gate on when the more expensive check runs at all.
"""
stripped = text.strip()
if not stripped:
return False
if len(stripped) < _AMBIGUOUS_LENGTH_THRESHOLD:
return True
lowered = stripped.lower()
return any(keyword in lowered for keyword in _SOFT_HEDGE_KEYWORDS)
def cosine_similarity(a: list[float], b: list[float]) -> float:
"""Standard cosine similarity, no numpy dependency (comfydv has none)."""
if not a or not b or len(a) != len(b):
return 0.0
dot = sum(x * y for x, y in zip(a, b))
norm_a = math.sqrt(sum(x * x for x in a))
norm_b = math.sqrt(sum(y * y for y in b))
if norm_a == 0.0 or norm_b == 0.0:
return 0.0
return dot / (norm_a * norm_b)
EmbedFn = Callable[[str], Awaitable[list[float] | None]]
_exemplar_embedding_cache: dict[str, list[list[float]]] = {}
async def _exemplar_embeddings(
embed_fn: EmbedFn, cache_key: str, exemplars: tuple[str, ...] = REFUSAL_EXEMPLARS
) -> list[list[float]]:
"""Embed ``exemplars`` once per ``cache_key`` and reuse — the exemplar
set only changes if the caller's custom phrases change (folded into
``cache_key`` by the caller), or the embedding space (i.e. which model
produced the vectors) does.
"""
cached = _exemplar_embedding_cache.get(cache_key)
if cached is not None:
return cached
embeddings = []
for exemplar in exemplars:
vec = await embed_fn(exemplar)
if not vec:
# An embedding call failing for one exemplar almost certainly
# means embeddings aren't usable at all right now (wrong/missing
# embedding model, unreachable server) — bail out rather than
# caching a partial, unusable exemplar set.
return []
embeddings.append(vec)
_exemplar_embedding_cache[cache_key] = embeddings
return embeddings
def _matches_custom_phrase(text: str, custom_phrases: tuple[str, ...]) -> bool:
"""Case-insensitive substring match against user-supplied phrases."""
if not custom_phrases:
return False
lowered = text.lower()
return any(phrase.lower() in lowered for phrase in custom_phrases)
async def is_refusal(
text: str,
*,
embed_fn: EmbedFn | None = None,
embed_cache_key: str = "",
threshold: float = 0.82,
custom_phrases: tuple[str, ...] = (),
) -> bool:
"""Hybrid refusal/deflection detector: free lexical pass first, then an
embedding-similarity fallback whenever the caller has configured one.
``embed_fn`` is supplied by the caller's own ``LLMProvider.embed()`` —
this function has no idea which backend or model produced ``text``, by
design (ADR: refusal detection is a model-behavior concern, not a
backend one). ``embed_fn=None`` (no embedding model configured) degrades
to lexical-only detection rather than erroring; ``embed_fn`` present
means the caller already opted in to the extra cost, so every non-blank,
non-lexically-caught response gets checked — no further length/keyword
gating. Any failure while embedding (unreachable server, no
embedding-capable model loaded) is swallowed the same way — an optional
enhancement failing shouldn't take down the retry loop it's assisting.
``custom_phrases`` lets a caller extend detection at runtime — e.g. a
ComfyUI node field the user edits directly — without touching the
shipped patterns/exemplars. Each phrase is checked two ways: a free
case-insensitive substring match (same cost tier as the lexical pass,
so it runs even with no ``embed_fn`` configured), and, when ``embed_fn``
is present, folded in as additional exemplars for the similarity check
so near-matches (not just exact substrings) of the user's phrases count
too.
"""
if not text or not text.strip():
return False # blank responses are the *other* retry trigger, not this one
if is_lexical_refusal(text):
return True
custom_phrases = tuple(p.strip() for p in custom_phrases if p and p.strip())
if _matches_custom_phrase(text, custom_phrases):
return True
if embed_fn is None:
return False
# embed_fn only exists when the caller explicitly configured an
# embedding_model — that's an opt-in to pay for the check, so run it on
# every non-blank, non-lexically-caught response rather than gating
# further on is_ambiguous. The length/keyword heuristic exists to avoid
# *unwanted* embedding calls when no embedding model is configured (see
# the embed_fn is None branch above); it has no reason to also suppress
# calls once the caller has already asked for them, and doing so was
# exactly what let the subtle/on-topic-looking deflections this feature
# targets slip through undetected.
exemplars = (
REFUSAL_EXEMPLARS + custom_phrases if custom_phrases else REFUSAL_EXEMPLARS
)
exemplar_cache_key = (
f"{embed_cache_key}|custom:{','.join(custom_phrases)}"
if custom_phrases
else embed_cache_key
)
try:
exemplar_vecs = await _exemplar_embeddings(
embed_fn, exemplar_cache_key, exemplars
)
if not exemplar_vecs:
return False
text_vec = await embed_fn(text)
if not text_vec:
return False
except Exception:
return False
return max(cosine_similarity(text_vec, vec) for vec in exemplar_vecs) >= threshold
+34 -20
View File
@@ -21,7 +21,7 @@ import sys
from typing import Any, Dict, List
from aiohttp import web
from jinja2 import exceptions, sandbox
from jinja2 import exceptions, meta, sandbox
# Set up logger for this module
logger = logging.getLogger(__name__)
@@ -75,6 +75,10 @@ class FormatString:
# Create a sandboxed Jinja2 environment for security
jinja_env = sandbox.SandboxedEnvironment()
# Jinja2 ships `tojson` but not its inverse; add one so STRING inputs
# carrying a JSON array/object (ComfyUI has no native list socket type)
# can be parsed back into real Python data, e.g. `{{ hints | fromjson }}`.
jinja_env.filters["fromjson"] = json.loads
# Define additional context
@staticmethod
@@ -252,38 +256,48 @@ class FormatString:
>>> # Test with additional context (should be excluded)
>>> keys = FormatString._extract_keys("Time: {{ datetime.now() }}")
>>> assert keys == []
>>> # Test {% for %} control structures: the loop variable is bound by the
>>> # template itself and must not be treated as a required input, while the
>>> # iterable it draws from must be.
>>> keys = FormatString._extract_keys(
... "{% for hint in extraction_hints %}{{ hint }}{% endfor %}"
... )
>>> assert keys == ['extraction_hints']
-->
"""
variables = []
seen = set()
def add_var(var):
var = var.split("|")[0].split(".")[0].strip()
if var not in seen and var not in FormatString.additional_context:
seen.add(var)
variables.append(var)
# Extract variables from Jinja2 expressions {{ }}
for match in re.finditer(
r"\{\{\s*([\w.]+)(?:\s*\|[\w\s]+)?(?:\.[^\(\)]+\(\))?\s*\}\}", template
):
add_var(match.group(1))
# Extract variables from f-string style { }
# Extract variables from Python str.format() style { }
for match in re.finditer(r"\{(\w+)\}", template):
add_var(match.group(1))
# Extract variables from Jinja2 control structures {% %}
for structure in re.finditer(r"\{%.*?%\}", template):
for var in re.findall(r"\b(\w+)\|\b", structure.group(0)):
if not var.startswith("end") and var not in {
"if",
"else",
"elif",
"for",
"in",
}:
add_var(var)
# Extract variables referenced anywhere in Jinja2 syntax ({{ }} expressions
# and {% %} control structures) by parsing the template with Jinja2 itself
# rather than approximating it with regexes. This is what correctly excludes
# names bound within the template (e.g. the `hint` loop variable in
# `{% for hint in extraction_hints %}`) while still surfacing names the
# template expects the caller to supply (e.g. `extraction_hints`).
try:
template_ast = FormatString.jinja_env.parse(template)
except exceptions.TemplateSyntaxError:
pass
else:
# find_undeclared_variables returns an unordered set; sort by first
# textual occurrence so extraction order is deterministic and matches
# the order the template reads left to right (callers rely on this
# for positional outputs, e.g. two {{ }} variables in sequence).
undeclared = sorted(
meta.find_undeclared_variables(template_ast),
key=lambda name: template.find(name),
)
for var in undeclared:
add_var(var)
return variables
+34
View File
@@ -0,0 +1,34 @@
"""llama.cpp connection node for ComfyUI.
Mirrors comfydv.ollama's OllamaClient exactly (ADR-007's parallel-
implementation pattern) — LlamaCppClient is the only new node this feature
introduces. Every other generic node (ChatCompletion, LLMModelSelector,
LLMLoadModel, LLMUnloadModel) already works with any LLM_CLIENT-typed
provider unchanged.
Deployment prerequisite: llama-server must be launched in router mode
(--models-dir or --models-preset) — see specs/008-llamacpp-integration/quickstart.md.
"""
from ._llm.llamacpp_provider import LlamaCppProvider
class LlamaCppClient:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"host": ("STRING", {"default": "http://localhost:8080"}),
},
"optional": {
"headers": ("OLLAMA_HEADERS",),
},
}
RETURN_TYPES = ("LLM_CLIENT",)
RETURN_NAMES = ("client",)
FUNCTION = "create_client"
CATEGORY = "dv/llamacpp"
def create_client(self, host: str, headers: dict | None = None):
return (LlamaCppProvider(host, headers),)
+625 -214
View File
File diff suppressed because it is too large Load Diff
+36 -5
View File
@@ -1,3 +1,4 @@
import json
import logging
import random
import sys
@@ -7,6 +8,28 @@ from .utils import any_type
logger = logging.getLogger(__name__)
def _preview_text(value) -> str:
"""Best-effort text preview for RandomChoice's arbitrary-typed output.
Mirrors ComfyUI core's own ``PreviewAny`` node's value handling (str/
number passthrough, else JSON, else ``str()``) rather than inventing a
new convention — RandomChoice's output can be anything (an IMAGE
tensor, a LATENT, a plain string), so this only needs to be "good
enough to glance at," not a faithful repr of every type.
"""
if isinstance(value, str):
return value
if isinstance(value, (int, float, bool)):
return str(value)
try:
return json.dumps(value, default=str, indent=2)
except Exception:
try:
return str(value)
except Exception:
return "<value could not be serialized>"
class RandomChoice:
def __init__(self):
pass
@@ -25,15 +48,20 @@ class RandomChoice:
FUNCTION = "random_choice"
OUTPUT_NODE = False
OUTPUT_NODE = True
CATEGORY = "dv/utils"
@classmethod
def IS_CHANGED(s, **kwargs):
return s.random_choice(s, **kwargs)
# Unchanged from before the UI-preview addition: returns the raw
# picked value (not the ui-wrapped dict random_choice() now returns)
# so ComfyUI's change-detection comparison keeps working exactly as
# it did previously.
return s._pick(**kwargs)
def random_choice(self, **kwargs):
@staticmethod
def _pick(**kwargs):
(
random.seed(kwargs.get("seed"))
if kwargs.get("seed")
@@ -41,10 +69,13 @@ class RandomChoice:
)
input = [i for i in kwargs.items() if i[0] != "seed"]
logger.debug("RandomChoice inputs: %s", input)
return random.choice(input)[1]
def random_choice(self, **kwargs):
try:
choice = random.choice(input)[1]
choice = self._pick(**kwargs)
logger.debug("RandomChoice chose: %s", choice)
return (choice,)
return {"ui": {"text": [_preview_text(choice)]}, "result": (choice,)}
except Exception as e:
logger.error("RandomChoice: unexpected error: %s", e)
raise
+99 -29
View File
@@ -1,11 +1,15 @@
/**
* ollama.js — ComfyUI frontend extension for comfydv Ollama nodes.
* ollama.js — ComfyUI frontend extension for comfydv's generic LLM nodes.
*
* Populates the model widget on Ollama nodes from a live call to
* GET /dv/ollama/models?host=<url>.
* Populates the model widget on LLM nodes from a live call to
* GET /dv/ollama/models?host=<url>&backend=<ollama|llamacpp>. Despite the
* file/route name (kept for historical reasons — see MIGRATION_MAP in
* comfydv.ollama), this now serves both backends: which one a given node's
* upstream client is determines the `backend` param (see
* getHostAndBackendFromNode below).
*
* OllamaModelSelector and OllamaLoadModel use a COMBO widget (dropdown).
* OllamaChatCompletion uses a plain STRING widget (accepts wired values).
* LLMModelSelector and LLMLoadModel use a COMBO widget (dropdown).
* ChatCompletion uses a plain STRING widget (accepts wired values).
* The Refresh button works the same way for both: it fetches the live list
* and sets the widget value / updates COMBO options as appropriate.
*/
@@ -13,12 +17,15 @@
import { app } from "../../scripts/app.js";
/** Nodes whose model widget is a COMBO dropdown. */
const OLLAMA_COMBO_NODES = new Set(["OllamaModelSelector", "OllamaLoadModel"]);
const LLM_COMBO_NODES = new Set(["LLMModelSelector", "LLMLoadModel"]);
/** Nodes whose model widget is a plain STRING (accepts wired input). */
const OLLAMA_STRING_MODEL_NODES = new Set(["OllamaChatCompletion"]);
const LLM_STRING_MODEL_NODES = new Set(["ChatCompletion"]);
const OLLAMA_ALL_NODES = new Set([...OLLAMA_COMBO_NODES, ...OLLAMA_STRING_MODEL_NODES]);
const LLM_ALL_NODES = new Set([...LLM_COMBO_NODES, ...LLM_STRING_MODEL_NODES]);
/** Registered client node type -> backend param the /dv/ollama/models route expects. */
const CLIENT_NODE_BACKENDS = { OllamaClient: "ollama", LlamaCppClient: "llamacpp" };
/**
* Fetch model list and update the node's model widget.
@@ -27,9 +34,11 @@ const OLLAMA_ALL_NODES = new Set([...OLLAMA_COMBO_NODES, ...OLLAMA_STRING_MODEL_
* - STRING: sets the value to the first model; keeps existing value if it
* still appears in the live list (user may have typed a valid name).
*/
async function refreshModelWidget(node, host) {
async function refreshModelWidget(node, host, backend) {
try {
const resp = await fetch(`/dv/ollama/models?host=${encodeURIComponent(host)}`);
const resp = await fetch(
`/dv/ollama/models?host=${encodeURIComponent(host)}&backend=${encodeURIComponent(backend)}`
);
if (!resp.ok) return;
const data = await resp.json();
const models = data.models ?? [];
@@ -53,54 +62,115 @@ async function refreshModelWidget(node, host) {
node.setDirtyCanvas(true, false);
} catch (_) {
// Ollama unreachable — leave widget unchanged
// Server unreachable — leave widget unchanged
}
}
/**
* Locate the host string for a node.
* Locate the host and backend for a node's connected LLM client.
*
* First checks the node's own widgets (OllamaClient has a "host" widget).
* Otherwise traverses graph links to find a connected OllamaClient node and
* reads its "host" widget — this is the common case for downstream nodes.
* Traverses graph links to find a connected OllamaClient or LlamaCppClient
* node and reads its "host" widget. Falls back to Ollama's default if
* nothing is wired yet, matching the pre-existing fallback behavior.
*/
function getHostFromNode(node) {
const ownHostWidget = node.widgets?.find(w => w.name === "host");
if (ownHostWidget) return ownHostWidget.value;
function getHostAndBackendFromNode(node) {
for (const input of node.inputs ?? []) {
if (!input.link) continue;
const link = node.graph?.links[input.link];
if (!link) continue;
const sourceNode = node.graph?.getNodeById(link.origin_id);
if (sourceNode?.type === "OllamaClient") {
const backend = sourceNode ? CLIENT_NODE_BACKENDS[sourceNode.type] : undefined;
if (backend) {
const hostWidget = sourceNode.widgets?.find(w => w.name === "host");
if (hostWidget?.value) return hostWidget.value;
if (hostWidget?.value) return { host: hostWidget.value, backend };
}
}
return "http://localhost:11434";
return { host: "http://localhost:11434", backend: "ollama" };
}
app.registerExtension({
name: "comfydv.ollama",
async beforeRegisterNodeDef(nodeType, nodeData) {
if (!OLLAMA_ALL_NODES.has(nodeData.name)) return;
if (!LLM_ALL_NODES.has(nodeData.name)) return;
const onNodeCreated = nodeType.prototype.onNodeCreated;
nodeType.prototype.onNodeCreated = function () {
const result = onNodeCreated?.apply(this, arguments);
const refresh = () => {
const { host, backend } = getHostAndBackendFromNode(this);
refreshModelWidget(this, host, backend);
};
// Add a Refresh button below the model widget
this.addWidget("button", "⟳ Refresh models", null, () => {
const host = getHostFromNode(this);
refreshModelWidget(this, host);
});
this.addWidget("button", "⟳ Refresh models", null, refresh);
// Initial population on node creation
const host = getHostFromNode(this);
refreshModelWidget(this, host);
refresh();
return result;
};
},
});
/**
* Live structured-output dynamic sockets for ChatCompletion.
*
* Mirrors FormatString's live dynamic-output pattern (see format_string.js):
* editing structured_output or output_schema posts to a backend route that
* recomputes ChatCompletion.RETURN_TYPES/RETURN_NAMES (the same
* update_outputs() path chat() itself uses at execution time) and returns
* the resulting output list — applied to this node's sockets immediately,
* so you see the extracted fields appear without having to run the graph
* first. Backend-agnostic: ChatCompletion is the one generic node both
* OllamaProvider and LlamaCppProvider feed.
*/
async function updateStructuredOutputs(node, structuredOutput, outputSchema) {
try {
const resp = await fetch("/dv/ollama/update_structured_outputs", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
unique_id: String(node.id),
structured_output: structuredOutput,
output_schema: outputSchema,
}),
});
if (!resp.ok) return;
const data = await resp.json();
applyOutputs(node, data.outputs ?? []);
} catch (_) {
// Backend unreachable — leave sockets unchanged.
}
}
function applyOutputs(node, outputs) {
node.outputs.length = 0;
outputs.forEach(o => node.addOutput(o.name, o.type));
node.setDirtyCanvas(true, true);
node.graph?.setDirtyCanvas(true, true);
}
app.registerExtension({
name: "comfydv.ollama.structuredOutput",
async beforeRegisterNodeDef(nodeType, nodeData) {
if (nodeData.name !== "ChatCompletion") return;
const onNodeCreated = nodeType.prototype.onNodeCreated;
nodeType.prototype.onNodeCreated = function () {
const result = onNodeCreated?.apply(this, arguments);
const structuredWidget = this.widgets?.find(w => w.name === "structured_output");
const schemaWidget = this.widgets?.find(w => w.name === "output_schema");
if (!structuredWidget || !schemaWidget) return result;
const update = () =>
updateStructuredOutputs(this, structuredWidget.value, schemaWidget.value);
structuredWidget.callback = update;
schemaWidget.callback = update;
return result;
};
+57
View File
@@ -0,0 +1,57 @@
/**
* preview_text.js — read-only output preview for comfydv's OUTPUT_NODE=True
* nodes that return a ComfyUI "ui": {"text": [...]} payload
* (ChatCompletion, FormatString, RandomChoice).
*
* ComfyUI does NOT auto-render an arbitrary node's ui.text — each node type
* that wants one implements its own onExecuted handler. This mirrors core's
* own ``PreviewAny`` node (comfy_extras/nodes_preview_any.py +
* "Comfy.PreviewAny" in the frontend bundle) minus its Markdown/Plaintext
* toggle, which none of these three nodes need.
*/
import { app } from "../../scripts/app.js";
import { ComfyWidgets } from "../../scripts/widgets.js";
const PREVIEW_NODES = new Set(["ChatCompletion", "FormatString", "RandomChoice"]);
app.registerExtension({
name: "comfydv.previewText",
async beforeRegisterNodeDef(nodeType, nodeData) {
if (!PREVIEW_NODES.has(nodeData.name)) return;
const onNodeCreated = nodeType.prototype.onNodeCreated;
nodeType.prototype.onNodeCreated = function () {
const result = onNodeCreated?.apply(this, arguments);
const widget = ComfyWidgets.STRING(
this,
"comfydv_preview_text",
["STRING", { multiline: true }],
app
).widget;
widget.label = "Preview";
widget.options.read_only = true;
// Not a real input — nothing to save/replay in the saved
// workflow JSON, and read-only anyway.
widget.options.serialize = false;
widget.serialize = false;
widget.inputEl.readOnly = true;
return result;
};
const onExecuted = nodeType.prototype.onExecuted;
nodeType.prototype.onExecuted = function (message) {
onExecuted?.apply(this, arguments);
const widget = this.widgets?.find(w => w.name === "comfydv_preview_text");
if (!widget) return;
const text = message?.text ?? "";
widget.value = Array.isArray(text) ? (text.join("\n\n") ?? "") : text;
this.setDirtyCanvas(true, true);
};
},
});
+32 -8
View File
@@ -60,6 +60,15 @@ def pytest_configure(config):
sys.modules["folder_paths"] = MockFolderPaths
# Force "comfydv" to resolve to src/comfydv and get cached in sys.modules now,
# while our sys.path.insert(0, ...) above is still the definitive answer. The
# repo root's own __init__.py (ComfyUI's custom-node entry point) is also a
# valid "comfydv" package from certain sys.path states pytest transiently
# constructs during fixture setup; without this, a later bare `import comfydv`
# (e.g. in the _clear_ollama_caches fixture) can resolve to that root package
# instead, which lacks the _llm submodule and fails with ModuleNotFoundError.
import comfydv # noqa: F401
# ---------------------------------------------------------------------------
# Ollama fixtures (used by @pytest.mark.integration tests)
@@ -68,20 +77,35 @@ def pytest_configure(config):
@pytest.fixture(autouse=True)
def _clear_ollama_caches():
"""Reset comfydv.ollama's module-level LRU caches around every test.
"""Reset the shared LLM provider caches and ChatCompletion's dynamic
RETURN_TYPES/RETURN_NAMES around every test.
Several tests reuse identical client/model/prompt inputs across cases
with different monkeypatched responses — without this, a later test would
silently get an earlier test's cached result instead of exercising its
own fake.
"""
from comfydv.ollama import _CHAT_RESPONSE_CACHE, _MODEL_LIST_CACHE
own fake. RETURN_TYPES/RETURN_NAMES are class-level mutable state (set by
ChatCompletion.update_outputs for structured_output mode) shared across
every test in the module — without resetting them, a structured-output
test would leak its dynamic outputs into unrelated tests that assert the
fixed 3-tuple.
_MODEL_LIST_CACHE.clear()
_CHAT_RESPONSE_CACHE.clear()
Caches live in comfydv._llm.ollama_provider (ADR-007's single source of
truth) — comfydv.ollama's combo-widget helpers (_fetch_models) share the
same cache instance, not a separate copy.
"""
from comfydv._llm.ollama_provider import _CHAT_RESPONSE_CACHE, _MODEL_LIST_CACHE
from comfydv.ollama import ChatCompletion
def _reset():
_MODEL_LIST_CACHE.clear()
_CHAT_RESPONSE_CACHE.clear()
ChatCompletion.RETURN_TYPES = ChatCompletion._BASE_RETURN_TYPES
ChatCompletion.RETURN_NAMES = ChatCompletion._BASE_RETURN_NAMES
ChatCompletion.node_configs.clear()
_reset()
yield
_MODEL_LIST_CACHE.clear()
_CHAT_RESPONSE_CACHE.clear()
_reset()
@pytest.fixture(scope="session")
+148
View File
@@ -0,0 +1,148 @@
"""Guards against the exact bug found while validating spec 008 against a
real ComfyUI dev harness (docker-compose): every `from comfydv._llm.X
import Y`-style absolute self-import inside src/comfydv/ silently broke the
*entire* plugin (every node, not just LLM ones) as soon as ComfyUI actually
loaded it.
ComfyUI's custom_nodes loader imports the plugin via a *relative* chain —
the repo-root __init__.py does `from .src.comfydv import ...`, nesting
comfydv under whatever top-level name the folder has (never `comfydv`
itself). An absolute `from comfydv...` self-import only resolves if `src/`
has separately been placed on sys.path — which conftest.py does for every
other test file in this suite, masking the bug completely. This file
deliberately does NOT rely on that sys.path insertion: it reproduces
ComfyUI's actual nested-relative-import shape in a subprocess.
Confirmed via git history: this predates spec 008 entirely — it was already
broken immediately after PR #17 merged (spec 007), well before llamacpp.py
existed. No test caught it because none exercised this exact loading shape
until the docker harness was run by hand.
"""
import subprocess
import sys
import textwrap
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).parent.parent
@pytest.fixture(autouse=True)
def _clear_ollama_caches():
"""Shadow conftest.py's autouse fixture of the same name for this module
only. That fixture's own setup does `from comfydv._llm.ollama_provider
import ...` — this file's tests are the exact reproduction of an
environment where `comfydv` resolving correctly can't be assumed (that's
the point of the file), so depending on it for an unrelated cache-reset
would make these tests order-dependent on whichever other test file
happens to import `comfydv` "the normal way" first in the session. These
tests touch no OllamaProvider/ChatCompletion state, so there is nothing
to reset."""
yield
_SUBPROCESS_SCRIPT = textwrap.dedent(
"""
import sys
import types
# Minimal ComfyUI stubs — same shape as conftest.py's pytest_configure,
# but this script intentionally runs outside pytest so it isn't reusing
# (or accidentally validated by) that fixture's sys.path setup.
class _InterruptProcessingException(Exception):
pass
comfy_module = types.ModuleType("comfy")
comfy_module.model_management = types.SimpleNamespace(
InterruptProcessingException=_InterruptProcessingException
)
sys.modules["comfy"] = comfy_module
sys.modules["comfy.model_management"] = comfy_module.model_management
class _Routes:
def post(self, path):
return lambda fn: fn
def get(self, path):
return lambda fn: fn
class _PromptServer:
pass
_PromptServer.instance = _PromptServer()
_PromptServer.instance.routes = _Routes()
server_module = types.ModuleType("server")
server_module.PromptServer = _PromptServer
sys.modules["server"] = server_module
folder_paths_module = types.ModuleType("folder_paths")
folder_paths_module.get_output_directory = lambda: "/tmp/comfydv_test"
sys.modules["folder_paths"] = folder_paths_module
# The critical part: put the repo's *parent* directory on sys.path, so
# `import comfydv` resolves to the repo-root __init__.py — exactly how
# ComfyUI resolves a folder under custom_nodes/ — NOT to src/comfydv
# directly (that's what conftest.py's sys.path.insert(0, ".../src")
# does for the rest of this test suite, and why it never caught this).
sys.path.insert(0, sys.argv[1])
import comfydv
required = {"FormatString", "RandomChoice", "CircuitBreaker",
"OllamaClient", "LlamaCppClient", "ChatCompletion"}
missing = required - set(comfydv.NODE_CLASS_MAPPINGS)
if missing:
print(f"MISSING_NODES:{missing}")
sys.exit(1)
print("OK")
"""
)
def test_package_imports_under_comfyui_style_relative_nesting():
"""Reproduces ComfyUI's real loading shape and fails loudly — with the
actual traceback — if any internal module reverts to an absolute
`from comfydv...` self-import."""
result = subprocess.run(
[sys.executable, "-c", _SUBPROCESS_SCRIPT, str(REPO_ROOT.parent)],
cwd=REPO_ROOT,
capture_output=True,
text=True,
timeout=30,
)
assert result.returncode == 0, (
"comfydv failed to import the way ComfyUI actually loads it "
"(relative nesting, not a top-level `comfydv` on sys.path). "
f"This means every node in the plugin would fail to register.\n"
f"--- stdout ---\n{result.stdout}\n--- stderr ---\n{result.stderr}"
)
assert "OK" in result.stdout
def test_no_absolute_self_imports_in_package():
"""Cheap, fast static guard alongside the dynamic test above: no file
under src/comfydv/ should import itself as `comfydv.X` — internal
imports must be relative (`.X` / `..X`) so they resolve regardless of
what the outer package happens to be named at load time."""
import ast
offenders = []
for path in (REPO_ROOT / "src" / "comfydv").rglob("*.py"):
tree = ast.parse(path.read_text(), filename=str(path))
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom):
if node.module and (
node.module == "comfydv" or node.module.startswith("comfydv.")
):
offenders.append(f"{path.relative_to(REPO_ROOT)}:{node.lineno}")
elif isinstance(node, ast.Import):
for alias in node.names:
if alias.name == "comfydv" or alias.name.startswith("comfydv."):
offenders.append(f"{path.relative_to(REPO_ROOT)}:{node.lineno}")
assert not offenders, (
"Absolute self-imports found — use relative imports instead "
f"(they break under ComfyUI's actual loader): {offenders}"
)
+41 -5
View File
@@ -35,9 +35,32 @@ class TestVariableExtraction:
def test_extract_jinja2_with_multiple_filters(self, format_string_class):
"""Test extraction of variables with multiple Jinja2 filters."""
keys = format_string_class._extract_keys("{{ name | upper | trim }}")
# Multiple chained filters may not extract - that's a limitation of the regex
# Just test that it doesn't crash
assert isinstance(keys, list)
assert keys == ["name"]
def test_extract_jinja2_for_loop_excludes_loop_variable(self, format_string_class):
"""The for-loop target (e.g. `hint`) is bound by the template and must
not be treated as a required input; the iterable it draws from must be."""
keys = format_string_class._extract_keys(
"{% for hint in extraction_hints %}- {{ hint }}\n{% endfor %}"
)
assert keys == ["extraction_hints"]
def test_extract_jinja2_if_condition_variable(self, format_string_class):
"""A variable referenced only in an {% if %} condition must still be
detected, even without a matching {{ }} expression elsewhere."""
keys = format_string_class._extract_keys(
"{% if extraction_hints is defined and extraction_hints %}yes{% endif %}"
)
assert keys == ["extraction_hints"]
def test_extract_jinja2_filter_with_arguments(self, format_string_class):
"""A filter called with arguments (e.g. tojson(indent=2)) has parens
in the way of the old regex's anchor to the closing }} — the variable
must still be detected."""
keys = format_string_class._extract_keys(
"{{ scene_manifest | tojson(indent=2) }}"
)
assert keys == ["scene_manifest"]
def test_extract_jinja2_multiple_variables(self, format_string_class):
"""Test extraction of multiple variables from Jinja2 template."""
@@ -178,6 +201,18 @@ class TestJinja2Formatting:
assert result[2] == "John"
assert result[3] == "Doe"
def test_jinja2_fromjson_filter_parses_array(self, format_string_class):
"""ComfyUI has no native list socket, so a STRING input carrying a
JSON array must be parseable back into a real list for iteration."""
result = format_string_class.format_string(
template_type="Jinja2",
template="{% for hint in hints | fromjson %}- {{ hint }}\n{% endfor %}",
save_path="",
unique_id="test-fromjson",
hints='["motion", "camera pan"]',
)["result"]
assert result[0] == "- motion\n- camera pan\n"
def test_jinja2_with_datetime(self, format_string_class):
"""Test Jinja2 formatting with datetime context."""
result = format_string_class.format_string(
@@ -199,10 +234,11 @@ class TestJinja2Formatting:
unique_id="test9",
value=sample_data["value"],
)["result"]
# value is not extracted as a variable because it's used in an expression
assert len(result) == 2 # Just formatted_string, saved_file_path
# value is extracted even though it's used in an expression
assert len(result) == 3 # formatted_string, saved_file_path, value
assert result[0] == "Result: 10"
assert result[1] == ""
assert result[2] == str(sample_data["value"])
class TestInlineDisplay:
+155
View File
@@ -0,0 +1,155 @@
"""
Tests for comfydv.llamacpp.LlamaCppClient — the one new ComfyUI node this
feature introduces. Also proves the adapter pattern end-to-end (US4): the
same generic nodes work unmodified against either provider.
BDD coverage:
../specs/008-llamacpp-integration/features/us1_connect_and_chat.feature
../specs/008-llamacpp-integration/features/us4_swap_backends.feature
"""
from comfydv._llm.llamacpp_provider import LlamaCppProvider
from comfydv._llm.ollama_provider import OllamaProvider
from comfydv.llamacpp import LlamaCppClient
from comfydv.ollama import (
ChatCompletion,
LLMLoadModel,
LLMModelSelector,
LLMUnloadModel,
)
class _FakeProvider:
"""Mirrors tests/test_ollama.py's _FakeProvider — reused here for US4's
swap-backends proof rather than duplicated, since the whole point is
that node behavior doesn't depend on which concrete provider it gets."""
def __init__(self, chat_response="ok"):
self.chat_response = chat_response
self.models = [{"name": "m", "status": "loaded"}]
self.calls: list[tuple] = []
async def list_models(self):
self.calls.append(("list_models",))
return self.models
async def load_model(self, model):
self.calls.append(("load_model", model))
async def unload_model(self, model):
self.calls.append(("unload_model", model))
async def chat(
self,
model,
messages,
options=None,
timeout_secs=300.0,
max_retries=2,
attempt_info=None,
on_status=None,
):
self.calls.append(("chat", model))
if attempt_info is not None:
attempt_info.update(
{
"seed": (options or {}).get("seed", 0),
"attempts": 1,
"timeout_secs": timeout_secs,
"refusals": 0,
}
)
return self.chat_response
def test_client_outputs_llamacpp_provider():
(client,) = LlamaCppClient().create_client("http://localhost:8080")
assert isinstance(client, LlamaCppProvider)
assert client.host == "http://localhost:8080"
def test_client_default_host_matches_llama_server_default_port():
input_types = LlamaCppClient.INPUT_TYPES()
assert input_types["required"]["host"][1]["default"] == "http://localhost:8080"
def test_client_output_type_is_generic_llm_client():
assert LlamaCppClient.RETURN_TYPES == ("LLM_CLIENT",)
def test_client_carries_headers():
(client,) = LlamaCppClient().create_client(
"http://localhost:8080", headers={"Authorization": "Bearer abc"}
)
assert client.headers == {"Authorization": "Bearer abc"}
def test_node_contract():
assert hasattr(LlamaCppClient, "INPUT_TYPES")
assert hasattr(LlamaCppClient, "RETURN_TYPES")
assert hasattr(LlamaCppClient, "FUNCTION")
assert hasattr(LlamaCppClient, "CATEGORY")
assert hasattr(LlamaCppClient, LlamaCppClient.FUNCTION)
def test_registered_in_node_class_mappings():
from comfydv import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS
assert NODE_CLASS_MAPPINGS["LlamaCppClient"] is LlamaCppClient
assert "LlamaCppClient" in NODE_DISPLAY_NAME_MAPPINGS
# ---------------------------------------------------------------------------
# US4 — swap backends without touching downstream nodes
# ---------------------------------------------------------------------------
def _run_workflow(client) -> None:
"""The same node sequence a workflow author would wire up, regardless
of which provider `client` is."""
ChatCompletion().chat(client=client, model="m", prompt="hi")
LLMModelSelector().select_model(client=client, model="m")
LLMLoadModel().load_model(client=client, model="m")
LLMUnloadModel().unload_model(client=client, model="m")
def test_same_workflow_runs_against_either_fake_provider():
"""No node branches on provider type — the same call sequence succeeds
whether client looks like an Ollama-shaped or llama.cpp-shaped provider."""
ollama_like = _FakeProvider(chat_response="ollama says hi")
llamacpp_like = _FakeProvider(chat_response="llamacpp says hi")
# Neither call raises — that's the actual assertion. If ChatCompletion/
# LLMModelSelector/LLMLoadModel/LLMUnloadModel secretly special-cased a
# concrete provider type (isinstance checks, attribute probing beyond
# the protocol), one of these would fail.
_run_workflow(ollama_like)
_run_workflow(llamacpp_like)
# LLMModelSelector is pure passthrough (client is accepted only for
# wiring/typing, never dereferenced), so it makes no provider call.
expected = ["chat", "load_model", "unload_model"]
assert [c[0] for c in ollama_like.calls] == expected
assert [c[0] for c in llamacpp_like.calls] == expected
def test_real_providers_are_interchangeable_client_output():
"""OllamaClient and LlamaCppClient both emit LLM_CLIENT — a workflow
author can wire either one into the same downstream nodes."""
from comfydv.ollama import OllamaClient
(ollama_client,) = OllamaClient().create_client("http://localhost:11434")
(llamacpp_client,) = LlamaCppClient().create_client("http://localhost:8080")
assert isinstance(ollama_client, OllamaProvider)
assert isinstance(llamacpp_client, LlamaCppProvider)
# Both satisfy the same protocol shape — same method names available.
for method in (
"list_models",
"load_model",
"unload_model",
"chat",
"chat_structured",
):
assert callable(getattr(ollama_client, method))
assert callable(getattr(llamacpp_client, method))
+777
View File
@@ -0,0 +1,777 @@
"""
Tests for comfydv._llm.llamacpp_provider.LlamaCppProvider — mirrors
test_ollama_provider.py's structure exactly (ADR-007's parallel-
implementation pattern). Mocks at the provider's own _post_json/_get_json
seam.
BDD coverage:
../specs/008-llamacpp-integration/features/us1_connect_and_chat.feature
../specs/008-llamacpp-integration/features/us2_structured_output.feature
../specs/008-llamacpp-integration/features/us3_model_lifecycle.feature
"""
import pytest
import comfydv._llm.llamacpp_provider as provider_mod
from comfydv._llm.llamacpp_provider import LlamaCppProvider, _fetch_models
from comfydv._llm.ollama_provider import _run_async
from comfydv._llm.provider import Message, ModelStatus
@pytest.fixture(autouse=True)
def _clear_provider_caches():
provider_mod._MODEL_LIST_CACHE.clear()
provider_mod._CHAT_RESPONSE_CACHE.clear()
yield
provider_mod._MODEL_LIST_CACHE.clear()
provider_mod._CHAT_RESPONSE_CACHE.clear()
# ---------------------------------------------------------------------------
# list_models — the "id" field name and nested "status.value" are the two
# details research.md flagged as easy to get wrong by assumption.
# ---------------------------------------------------------------------------
def test_list_models_maps_id_field_to_name(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
return {"data": [{"id": "gemma-3-4b:Q4_K_M", "status": {"value": "loaded"}}]}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
(model,) = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
assert model.name == "gemma-3-4b:Q4_K_M"
def test_list_models_reads_nested_status_value(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
return {
"data": [
{"id": "a", "status": {"value": "sleeping"}},
{"id": "b", "status": {"value": "downloading", "progress": {}}},
]
}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
by_name = {m.name: m for m in models}
assert by_name["a"].status == ModelStatus.SLEEPING
assert by_name["b"].status == ModelStatus.DOWNLOADING
def test_list_models_no_normalization_needed_full_vocabulary(monkeypatch):
"""Unlike OllamaProvider, llama.cpp's status vocabulary is exactly
ModelStatus's full set — every value should pass through untouched."""
async def fake_get(url, *, timeout=5.0, headers=None):
return {
"data": [
{"id": v, "status": {"value": v}}
for v in ["unloaded", "loading", "loaded", "sleeping", "downloading"]
]
}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
assert {m.status for m in models} == set(ModelStatus)
def test_list_models_skips_unrecognized_status(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
return {
"data": [
{"id": "crashed", "status": {"value": "failed", "exit_code": 1}},
{"id": "ok", "status": {"value": "loaded"}},
]
}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
models = _run_async(LlamaCppProvider("http://localhost:8080").list_models())
assert [m.name for m in models] == ["ok"]
def test_list_models_unreachable_returns_empty(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
raise ConnectionError("no route to host")
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
models = _run_async(LlamaCppProvider("http://localhost:19999").list_models())
assert models == []
def test_list_models_non_router_mode_raises_clear_error(monkeypatch):
"""FR-006: a llama-server that IS reachable but wasn't launched with
--models-dir/--models-preset answers GET /models with an HTTP error
(the endpoint doesn't exist outside router mode). That must surface as
a specific, actionable error — not silently degrade to an empty list,
which would be indistinguishable from "server has no models"."""
async def fake_get(url, *, timeout=5.0, headers=None):
raise RuntimeError("Server returned HTTP 404 for http://x/models: not found")
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
with pytest.raises(RuntimeError, match="router mode"):
_run_async(LlamaCppProvider("http://localhost:8080").list_models())
def test_list_models_cached_second_call(monkeypatch):
calls = {"n": 0}
async def fake_get(url, *, timeout=5.0, headers=None):
calls["n"] += 1
return {"data": [{"id": "a", "status": {"value": "loaded"}}]}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
provider = LlamaCppProvider("http://localhost:8080")
_run_async(provider.list_models())
_run_async(provider.list_models())
assert calls["n"] == 1
# ---------------------------------------------------------------------------
# load_model / unload_model
# ---------------------------------------------------------------------------
def test_load_model_posts_to_models_load_with_model_field(monkeypatch):
captured = {}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
captured["url"] = url
captured["payload"] = payload
return {"success": True}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
assert captured["url"] == "http://localhost:8080/models/load"
assert captured["payload"] == {"model": "gemma-3-4b"}
def test_unload_model_posts_to_models_unload_with_model_field(monkeypatch):
captured = {}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
captured["url"] = url
captured["payload"] = payload
return {"success": True}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
assert captured["url"] == "http://localhost:8080/models/unload"
assert captured["payload"] == {"model": "gemma-3-4b"}
def test_load_model_already_running_is_idempotent(monkeypatch):
"""Confirmed live: router mode's /models/load is NOT idempotent at the
wire level — it 400s "model is already running" rather than returning
{"success": true}. The LLMProvider protocol requires load_model() to be
idempotent, so LlamaCppProvider must absorb this itself."""
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append((url, payload))
raise RuntimeError(
"Server returned HTTP 400 for "
f'{url}: {{"error":{{"code":400,"message":"model is already '
'running","type":"invalid_request_error"}}}}'
)
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
# The absence of a raised exception is only meaningful if the request
# actually happened and hit the "already running" branch — assert that
# directly rather than trusting silence alone.
assert calls == [("http://localhost:8080/models/load", {"model": "gemma-3-4b"})]
def test_load_model_other_http_error_still_raises(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
raise RuntimeError(f"Server returned HTTP 500 for {url}: internal error")
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
with pytest.raises(RuntimeError, match="500"):
_run_async(LlamaCppProvider("http://localhost:8080").load_model("gemma-3-4b"))
def test_unload_model_not_running_is_idempotent(monkeypatch):
"""Mirror of the load_model case, confirmed live: /models/unload 400s
"model is not running" on an already-unloaded model."""
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append((url, payload))
raise RuntimeError(
"Server returned HTTP 400 for "
f'{url}: {{"error":{{"code":400,"message":"model is not '
'running","type":"invalid_request_error"}}}}'
)
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
assert calls == [("http://localhost:8080/models/unload", {"model": "gemma-3-4b"})]
def test_unload_model_other_http_error_still_raises(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
raise RuntimeError(f"Server returned HTTP 500 for {url}: internal error")
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
with pytest.raises(RuntimeError, match="500"):
_run_async(LlamaCppProvider("http://localhost:8080").unload_model("gemma-3-4b"))
def test_load_model_empty_raises_before_network(monkeypatch):
def fail_post(*a, **k):
raise AssertionError("must not call _post_json for an empty model name")
monkeypatch.setattr(provider_mod, "_post_json", fail_post)
with pytest.raises(ValueError, match="cannot be empty"):
_run_async(LlamaCppProvider("http://localhost:8080").load_model(""))
def test_unload_model_empty_raises_before_network(monkeypatch):
def fail_post(*a, **k):
raise AssertionError("must not call _post_json for an empty model name")
monkeypatch.setattr(provider_mod, "_post_json", fail_post)
with pytest.raises(ValueError, match="cannot be empty"):
_run_async(LlamaCppProvider("http://localhost:8080").unload_model(" "))
# ---------------------------------------------------------------------------
# chat — OpenAI response shape (choices[0].message.content), not Ollama's
# native shape
# ---------------------------------------------------------------------------
def test_chat_posts_to_v1_chat_completions_and_parses_openai_shape(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
assert url == "http://localhost:8080/v1/chat/completions"
return {"choices": [{"message": {"role": "assistant", "content": "hello"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b", [Message(role="user", content="hi")]
)
)
assert result == "hello"
def test_chat_no_choices_returns_empty_string(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
return {"choices": []}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b", [Message(role="user", content="hi")], max_retries=0
)
)
assert result == ""
def test_chat_second_identical_call_is_cached(monkeypatch):
calls = {"n": 0}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls["n"] += 1
return {"choices": [{"message": {"content": "cached"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
provider = LlamaCppProvider("http://localhost:8080")
messages = [Message(role="user", content="hi")]
r1 = _run_async(provider.chat("m", messages))
r2 = _run_async(provider.chat("m", messages))
assert r1 == r2 == "cached"
assert calls["n"] == 1
# ---------------------------------------------------------------------------
# chat — retry-on-blank-output. Mirrors test_ollama_provider.py's coverage;
# the one llama.cpp-specific detail is that the retry seed must land in the
# top-level OpenAI "seed" field, not nested under "options" (see chat()'s
# comment on why the options passthrough doesn't reach llama-server at all).
# ---------------------------------------------------------------------------
async def _fake_sleep(_secs):
"""No-op stand-in for asyncio.sleep — keeps retry tests instant."""
def test_chat_retries_on_blank_response_and_returns_second_attempt(monkeypatch):
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append(payload)
if len(calls) == 1:
return {"choices": [{"message": {"content": ""}}]}
return {"choices": [{"message": {"content": "real answer"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b", [Message(role="user", content="hi")]
)
)
assert result == "real answer"
assert len(calls) == 2
def test_chat_retry_seed_is_top_level_not_nested_in_options(monkeypatch):
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append(payload)
if len(calls) < 3:
return {"choices": [{"message": {"content": ""}}]}
return {"choices": [{"message": {"content": "ok"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
_run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b", [Message(role="user", content="hi")], max_retries=2
)
)
assert "seed" not in calls[0]
assert calls[1]["seed"] == 1
assert calls[2]["seed"] == 2
def test_chat_disable_thinking_sets_chat_template_kwargs_and_reasoning_effort(
monkeypatch,
):
"""ADR-010: llama-server doesn't recognize a "think" key nested inside
"options" (that's an Ollama-native convention) — it needs its own two
documented request-body toggles instead, and "think" must not leak into
the nested options object llama-server actually does understand."""
captured = {}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
captured.update(payload)
return {"choices": [{"message": {"content": "ok"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b",
[Message(role="user", content="hi")],
options={"temperature": 0.5, "think": False},
)
)
assert captured["chat_template_kwargs"] == {"enable_thinking": False}
assert captured["reasoning_effort"] == "none"
assert captured["options"] == {"temperature": 0.5} # "think" popped out
def test_chat_exhausted_retries_returns_blank_without_raising(monkeypatch):
calls = {"n": 0}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls["n"] += 1
return {"choices": [{"message": {"content": ""}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b", [Message(role="user", content="hi")], max_retries=2
)
)
assert result == ""
assert calls["n"] == 3 # original + 2 retries, per max_retries=2
def test_chat_timeout_escalates_per_retry_attempt(monkeypatch):
timeouts = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
timeouts.append(timeout)
if len(timeouts) < 3:
return {"choices": [{"message": {"content": ""}}]}
return {"choices": [{"message": {"content": "done"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
_run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b",
[Message(role="user", content="hi")],
timeout_secs=50.0,
max_retries=2,
)
)
assert timeouts == [50.0, 100.0, 150.0]
def test_chat_attempt_info_populated_on_success(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
return {"choices": [{"message": {"content": "ok"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
attempt_info: dict = {}
_run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b",
[Message(role="user", content="hi")],
options={"seed": 9},
attempt_info=attempt_info,
)
)
assert attempt_info == {
"seed": 9,
"attempts": 1,
"timeout_secs": 300.0,
"refusals": 0,
}
def test_chat_on_status_called_on_retry_and_recovery(monkeypatch):
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append(payload)
if len(calls) == 1:
return {"choices": [{"message": {"content": "I cannot generate that."}}]}
return {"choices": [{"message": {"content": "a real, on-topic answer"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
statuses = []
_run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b",
[Message(role="user", content="hi")],
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
on_status=statuses.append,
)
)
assert len(statuses) == 2
assert "Refusal/deflection detected" in statuses[0]
assert "Recovered" in statuses[1]
# ---------------------------------------------------------------------------
# chat_structured — zero new logic, delegates to the shared helper unchanged
# ---------------------------------------------------------------------------
def test_chat_structured_builds_v1_base_url_and_delegates(monkeypatch):
from pydantic import BaseModel
class Widget(BaseModel):
name: str
captured = {}
async def fake_chat_structured(**kwargs):
captured.update(kwargs)
return Widget(name="x")
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat_structured(
"gemma-3-4b", [Message(role="user", content="hi")], Widget
)
)
assert result == Widget(name="x")
assert captured["base_url"] == "http://localhost:8080/v1"
assert captured["model"] == "gemma-3-4b"
def test_chat_structured_forwards_options(monkeypatch):
from pydantic import BaseModel
class Widget(BaseModel):
name: str
captured = {}
async def fake_chat_structured(**kwargs):
captured.update(kwargs)
return Widget(name="x")
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
_run_async(
LlamaCppProvider("http://localhost:8080").chat_structured(
"gemma-3-4b",
[Message(role="user", content="hi")],
Widget,
options={"temperature": 0.0},
)
)
assert captured["options"] == {"temperature": 0.0}
def test_chat_structured_forwards_attempt_info(monkeypatch):
from pydantic import BaseModel
class Widget(BaseModel):
name: str
captured = {}
async def fake_chat_structured(**kwargs):
captured.update(kwargs)
return Widget(name="x")
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
attempt_info: dict = {}
_run_async(
LlamaCppProvider("http://localhost:8080").chat_structured(
"gemma-3-4b",
[Message(role="user", content="hi")],
Widget,
attempt_info=attempt_info,
)
)
# Same object passed straight through — the shared helper populates it,
# this provider doesn't need to know its shape.
assert captured["attempt_info"] is attempt_info
def test_chat_structured_forwards_on_status(monkeypatch):
from pydantic import BaseModel
class Widget(BaseModel):
name: str
captured = {}
async def fake_chat_structured(**kwargs):
captured.update(kwargs)
return Widget(name="x")
monkeypatch.setattr("comfydv._llm.chat.chat_structured", fake_chat_structured)
def on_status(msg):
pass
_run_async(
LlamaCppProvider("http://localhost:8080").chat_structured(
"gemma-3-4b",
[Message(role="user", content="hi")],
Widget,
on_status=on_status,
)
)
assert captured["on_status"] is on_status
# ---------------------------------------------------------------------------
# _fetch_models — name-only view used by ComfyUI's /dv/ollama/models?backend=
# llamacpp route (the JS refresh button / node-creation auto-populate).
# Deliberately more forgiving than list_models(): degrades to [] on any
# failure rather than raising on a non-router-mode server, matching the
# combo-widget UX OllamaProvider's own _fetch_models already gives.
# ---------------------------------------------------------------------------
def test_fetch_models_returns_name_only_list(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
return {
"data": [
{"id": "a", "status": {"value": "loaded"}},
{"id": "b", "status": {"value": "unloaded"}},
]
}
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
names = _run_async(_fetch_models("http://localhost:8080"))
assert names == ["a", "b"]
def test_fetch_models_degrades_to_empty_on_non_router_mode(monkeypatch):
"""Unlike list_models() (FR-006), this combo-population view swallows
even the non-router-mode error — a quiet empty dropdown, not a toast."""
async def fake_get(url, *, timeout=5.0, headers=None):
raise RuntimeError("Server returned HTTP 404 for http://x/models: not found")
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
names = _run_async(_fetch_models("http://localhost:8080"))
assert names == []
def test_fetch_models_degrades_to_empty_when_unreachable(monkeypatch):
async def fake_get(url, *, timeout=5.0, headers=None):
raise ConnectionError("no route to host")
monkeypatch.setattr(provider_mod, "_get_json", fake_get)
names = _run_async(_fetch_models("http://localhost:19999"))
assert names == []
# ---------------------------------------------------------------------------
# chat — image input (spec 009, US2; features/us2_both_backends.feature)
# ---------------------------------------------------------------------------
def test_chat_maps_images_to_openai_content_parts(monkeypatch):
"""llama.cpp's OpenAI-compatible endpoint takes images as image_url parts
inside content, not a flat images field (ADR-008)."""
captured = {}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
captured["payload"] = payload
return {"choices": [{"message": {"content": "a red square"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(
LlamaCppProvider("http://localhost:8080").chat(
"m", [Message(role="user", content="describe", images=["QUJD"])]
)
)
assert captured["payload"]["messages"][-1] == {
"role": "user",
"content": [
{"type": "text", "text": "describe"},
{
"type": "image_url",
"image_url": {"url": "data:image/png;base64,QUJD"},
},
],
}
def test_chat_text_only_content_stays_plain_string(monkeypatch):
"""FR-003/SC-004: an image-less message keeps a plain string content,
byte-identical to today (no content-parts, no images key)."""
captured = {}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
captured["payload"] = payload
return {"choices": [{"message": {"content": "ok"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
_run_async(
LlamaCppProvider("http://localhost:8080").chat(
"m", [Message(role="user", content="hi")]
)
)
assert captured["payload"]["messages"] == [{"role": "user", "content": "hi"}]
# ---------------------------------------------------------------------------
# refusal-retry — parity with test_ollama_provider.py's coverage (ADR: this
# is a model-behavior concern, not a backend one — both providers wire the
# same comfydv._llm.retry.is_refusal() into their own retry loop).
# ---------------------------------------------------------------------------
def test_chat_retries_on_lexical_refusal_and_returns_clean_second_attempt(
monkeypatch,
):
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append(payload)
if len(calls) == 1:
return {"choices": [{"message": {"content": "I cannot generate that."}}]}
return {"choices": [{"message": {"content": "a real, on-topic answer"}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
monkeypatch.setattr(provider_mod.asyncio, "sleep", _fake_sleep)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b",
[Message(role="user", content="hi")],
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
)
)
assert result == "a real, on-topic answer"
assert len(calls) == 2
def test_chat_refusal_retry_disabled_returns_refusal_text_unchanged(monkeypatch):
calls = []
async def fake_post(url, payload, *, timeout=120.0, headers=None):
calls.append(payload)
return {"choices": [{"message": {"content": "I cannot generate that."}}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
result = _run_async(
LlamaCppProvider("http://localhost:8080").chat(
"gemma-3-4b", [Message(role="user", content="hi")]
)
)
assert result == "I cannot generate that."
assert len(calls) == 1
def test_embed_returns_vector_from_v1_embeddings(monkeypatch):
captured = {}
async def fake_post(url, payload, *, timeout=120.0, headers=None):
captured["url"] = url
captured["payload"] = payload
return {"data": [{"embedding": [0.4, 0.5, 0.6], "index": 0}]}
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
result = _run_async(
LlamaCppProvider("http://localhost:8080").embed("nomic-embed-text", "hello")
)
assert result == [0.4, 0.5, 0.6]
assert captured["url"] == "http://localhost:8080/v1/embeddings"
assert captured["payload"] == {"model": "nomic-embed-text", "input": "hello"}
def test_embed_returns_none_when_no_embedding_model_configured(monkeypatch):
async def fake_post(url, payload, *, timeout=120.0, headers=None):
raise RuntimeError("llama-server returned HTTP 404 for /v1/embeddings")
monkeypatch.setattr(provider_mod, "_post_json", fake_post)
result = _run_async(
LlamaCppProvider("http://localhost:8080").embed("nomic-embed-text", "hello")
)
assert result is None
+704
View File
@@ -0,0 +1,704 @@
"""
Tests for comfydv._llm.chat.chat_structured — shared pydantic-ai backed
structured output, used by every LLMProvider implementation (ADR-007).
Mocks at the comfydv._llm.chat._build_agent seam (returns a fake agent
exposing an async .run()), mirroring test_ollama.py's existing convention
of monkeypatching the module-level HTTP seam rather than the network itself.
Uses _run_async (same helper comfydv.ollama uses) to drive the coroutine
synchronously, matching this project's existing test style rather than
introducing a pytest-asyncio dependency.
BDD coverage:
../specs/007-llm-provider-abstraction/features/us2_structured_output.feature
"""
from dataclasses import dataclass
import pytest
from pydantic import BaseModel, ValidationError
import comfydv._llm.chat as chat_mod
from comfydv._llm.ollama_provider import _run_async
from comfydv._llm.provider import Message
@pytest.fixture(autouse=True)
def _no_retry_backoff(monkeypatch):
"""Keep the real RETRY_BACKOFF_SECS delay out of this suite's wall-clock
time for every test except the ones that specifically assert on it
(which re-monkeypatch locally, overriding this)."""
async def _instant_sleep(_secs):
pass
monkeypatch.setattr(chat_mod.asyncio, "sleep", _instant_sleep)
class _Widget(BaseModel):
name: str
count: int
@dataclass
class _FakeResult:
output: object
class _FakeAgent:
"""Stand-in for pydantic_ai.Agent — .run() is scripted per test."""
def __init__(self, responses):
self._responses = list(responses)
self.calls = []
async def run(self, prompt, *, message_history=None, model_settings=None):
self.calls.append((prompt, message_history, model_settings))
outcome = self._responses.pop(0)
if isinstance(outcome, Exception):
raise outcome
return _FakeResult(output=outcome)
def _messages(*, system=None, history=None, prompt="hi"):
msgs = []
if system:
msgs.append(Message(role="system", content=system))
for role, content in history or []:
msgs.append(Message(role=role, content=content))
msgs.append(Message(role="user", content=prompt))
return msgs
def test_chat_structured_returns_validated_output(monkeypatch):
fake = _FakeAgent([_Widget(name="a", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
result = _run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(prompt="describe a widget"),
schema=_Widget,
)
)
assert result == _Widget(name="a", count=1)
assert fake.calls[0][0] == "describe a widget"
def test_build_agent_uses_native_output_not_tool_calling(monkeypatch):
"""ADR-009: the Agent must be built with NativeOutput (response_format /
JSON-schema constrained decoding), not pydantic-ai's tool-calling
default. Regression guard against reverting to a bare `output_type=schema`,
which let a thinking-capable model exhaust its token budget on reasoning
and never emit a tool call (confirmed live against a real Ollama server).
Spies on the real pydantic_ai.Agent constructor (only _build_agent's own
seam is mocked in every other test in this file) and asserts on its
output_type argument via the public NativeOutput marker class, rather
than pydantic-ai's private internal schema representation."""
from pydantic_ai import NativeOutput
captured = {}
real_agent_cls = chat_mod.Agent
class _SpyAgent(real_agent_cls):
def __init__(self, *args, **kwargs):
captured.update(kwargs)
super().__init__(*args, **kwargs)
monkeypatch.setattr(chat_mod, "Agent", _SpyAgent)
chat_mod._build_agent(
base_url="http://localhost:11434/v1",
model="llama3",
schema=_Widget,
headers=None,
timeout_secs=300.0,
)
output_type = captured["output_type"]
assert isinstance(output_type, NativeOutput)
assert output_type.outputs == _Widget
def test_chat_structured_retries_on_validation_failure(monkeypatch):
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, _Widget(name="b", count=2)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
result = _run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
max_retries=2,
)
)
assert result == _Widget(name="b", count=2)
assert len(fake.calls) == 2
def test_chat_structured_timeout_escalates_per_retry_attempt(monkeypatch):
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, bad, _Widget(name="b", count=2)])
build_calls = []
def fake_build_agent(**kw):
build_calls.append(kw["timeout_secs"])
return fake
monkeypatch.setattr(chat_mod, "_build_agent", fake_build_agent)
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
timeout_secs=100.0,
max_retries=2,
)
)
assert build_calls == [100.0, 200.0, 300.0]
def test_chat_structured_attempt_info_populated_on_success(monkeypatch):
fake = _FakeAgent([_Widget(name="a", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
attempt_info: dict = {}
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
options={"seed": 5},
attempt_info=attempt_info,
)
)
assert attempt_info == {
"seed": 5,
"attempts": 1,
"timeout_secs": 300.0,
"refusals": 0,
}
def test_chat_structured_attempt_info_populated_on_exhaustion(monkeypatch):
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, bad])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
attempt_info: dict = {}
with pytest.raises(RuntimeError):
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
max_retries=1,
attempt_info=attempt_info,
)
)
assert attempt_info["attempts"] == 2
def test_chat_structured_on_status_called_on_retry_and_recovery(monkeypatch):
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, _Widget(name="b", count=2)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
statuses = []
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
max_retries=2,
on_status=statuses.append,
)
)
assert len(statuses) == 2
assert "Structured output failed" in statuses[0]
assert "attempt 1/3" in statuses[0]
assert "Recovered" in statuses[1]
assert "attempt 2/3" in statuses[1]
def test_chat_structured_exhausted_retries_raises_runtime_error(monkeypatch):
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, bad, bad]) # max_retries=2 -> 3 total attempts
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
with pytest.raises(RuntimeError) as exc_info:
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
max_retries=2,
)
)
message = str(exc_info.value)
assert "llama3" in message
assert "3 attempt(s)" in message
assert len(fake.calls) == 3
def test_chat_structured_max_retries_clamped_to_five(monkeypatch):
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad] * 6)
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
with pytest.raises(RuntimeError, match=r"6 attempt\(s\)"):
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
max_retries=999, # clamped to 5 -> 6 total attempts
)
)
assert len(fake.calls) == 6
def test_chat_structured_forwards_options_as_extra_body(monkeypatch):
"""Regression guard: options (Ollama-native sampling params set via the
OllamaOption* nodes — temperature, seed, num_predict, repeat_penalty,
etc.) must reach the request, not be silently dropped in structured
mode. Forwarded verbatim via pydantic-ai's model_settings.extra_body,
matching the pre-ADR-007 payload shape exactly (no lossy remapping onto
ModelSettings' own standardized field names)."""
fake = _FakeAgent([_Widget(name="a", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
options={"temperature": 0.0, "seed": 42, "num_predict": 128},
)
)
assert fake.calls[0][2] == {
"extra_body": {"options": {"temperature": 0.0, "seed": 42, "num_predict": 128}}
}
def test_chat_structured_no_options_means_no_model_settings(monkeypatch):
fake = _FakeAgent([_Widget(name="a", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
)
)
assert fake.calls[0][2] is None
def test_chat_structured_disable_thinking_sets_chat_template_kwargs(monkeypatch):
"""ADR-010: llama-server's two documented reasoning toggles, applied via
extra_body the same way options is — "think" must not leak into the
nested options.extra_body.options object llama-server's native sampling
params live in."""
fake = _FakeAgent([_Widget(name="a", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:8080/v1",
model="gemma-3-4b",
messages=_messages(),
schema=_Widget,
options={"temperature": 0.0, "think": False},
)
)
assert fake.calls[0][2] == {
"extra_body": {
"options": {"temperature": 0.0},
"chat_template_kwargs": {"enable_thinking": False},
"reasoning_effort": "none",
}
}
def test_chat_structured_enable_thinking_skips_reasoning_effort(monkeypatch):
fake = _FakeAgent([_Widget(name="a", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:8080/v1",
model="gemma-3-4b",
messages=_messages(),
schema=_Widget,
options={"think": True},
)
)
extra_body = fake.calls[0][2]["extra_body"]
assert extra_body["chat_template_kwargs"] == {"enable_thinking": True}
assert "reasoning_effort" not in extra_body
assert "options" not in extra_body # only "think" was in options
def test_chat_structured_requires_last_message_user_role():
with pytest.raises(ValueError, match="role='user'"):
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=[Message(role="system", content="only a system message")],
schema=_Widget,
)
)
# ---------------------------------------------------------------------------
# retry-on-failure seed/backoff — live-verified: a freshly-loaded model's
# first structured-output attempt can fail outright, then behave normally on
# the very next call.
# ---------------------------------------------------------------------------
def test_chat_structured_retry_injects_incrementing_seed(monkeypatch):
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, bad, _Widget(name="c", count=3)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
max_retries=2,
)
)
assert fake.calls[0][2] is None # attempt 1 untouched — no options set
assert fake.calls[1][2]["seed"] == 1
assert fake.calls[2][2]["seed"] == 2
def test_chat_structured_retry_seed_starts_from_pinned_base(monkeypatch):
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, _Widget(name="c", count=3)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
options={"seed": 42},
max_retries=2,
)
)
assert fake.calls[0][2] == {"extra_body": {"options": {"seed": 42}}}
assert fake.calls[1][2]["seed"] == 43 # base(42) + (attempt 2 - 1)
# The nested Ollama-native options.seed must track the same retry seed as
# the top-level one — a backend that honors the nested field over the
# top-level OpenAI "seed" must not keep seeing the stale pinned value.
assert fake.calls[1][2]["extra_body"] == {"options": {"seed": 43}}
def test_chat_structured_retry_does_not_mutate_callers_options_dict(monkeypatch):
"""Regression guard for the fix above: syncing the nested seed must copy,
not mutate, the caller's options dict — otherwise a second call reusing
the same options object would start from the wrong base seed."""
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, _Widget(name="c", count=3)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
caller_options = {"seed": 42}
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
options=caller_options,
max_retries=2,
)
)
assert caller_options == {"seed": 42}
def test_chat_structured_retry_sleeps_between_attempts(monkeypatch):
sleep_calls = []
async def fake_sleep(secs):
sleep_calls.append(secs)
monkeypatch.setattr(chat_mod.asyncio, "sleep", fake_sleep)
bad = ValidationError.from_exception_data("Widget", [])
fake = _FakeAgent([bad, _Widget(name="c", count=3)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
max_retries=2,
)
)
assert sleep_calls == [chat_mod.RETRY_BACKOFF_SECS]
# ---------------------------------------------------------------------------
# refusal-retry — parity with test_ollama_provider.py's coverage (this
# module serves LlamaCppProvider.chat_structured() — see
# comfydv._llm.retry.is_refusal and OllamaOptionRefusalRetry).
# ---------------------------------------------------------------------------
def test_chat_structured_retries_on_refusal_and_returns_clean_second_attempt(
monkeypatch,
):
refused = _Widget(name="I cannot generate that content.", count=1)
clean = _Widget(name="clean", count=2)
fake = _FakeAgent([refused, clean])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
result = _run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
max_retries=2,
)
)
assert result == clean
assert len(fake.calls) == 2
def test_chat_structured_on_status_reports_refusal_reason(monkeypatch):
refused = _Widget(name="I cannot generate that content.", count=1)
clean = _Widget(name="clean", count=2)
fake = _FakeAgent([refused, clean])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
statuses = []
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
max_retries=2,
on_status=statuses.append,
)
)
assert len(statuses) == 2
assert "Refusal/deflection detected" in statuses[0]
assert "Recovered" in statuses[1]
def test_chat_structured_refusal_retry_disabled_returns_refusal_unchanged(
monkeypatch,
):
refused = _Widget(name="I cannot generate that content.", count=1)
fake = _FakeAgent([refused])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
result = _run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
)
)
assert result == refused
assert len(fake.calls) == 1
def test_chat_structured_refusal_retry_exhausted_raises(monkeypatch):
refused = _Widget(name="I cannot help with this.", count=1)
fake = _FakeAgent([refused, refused])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
with pytest.raises(RuntimeError, match="failed validation"):
_run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
options={"refusal_retry": {"enabled": True, "embedding_model": ""}},
max_retries=1,
)
)
def test_chat_structured_refusal_retry_uses_embed_fn(monkeypatch):
from comfydv._llm.retry import REFUSAL_EXEMPLARS
refused = _Widget(name="not today, sorry", count=1)
clean = _Widget(name="a clean value", count=2)
fake = _FakeAgent([refused, clean])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
embed_calls = []
async def fake_embed(text):
embed_calls.append(text)
if "not today" in text or text in REFUSAL_EXEMPLARS:
return [1.0, 0.0]
return [0.0, 1.0]
result = _run_async(
chat_mod.chat_structured(
base_url="http://localhost:11434/v1",
model="llama3",
messages=_messages(),
schema=_Widget,
options={
"refusal_retry": {
"enabled": True,
"embedding_model": "nomic-embed-text",
"threshold": 0.5,
}
},
max_retries=2,
embed_fn=fake_embed,
)
)
assert result == clean
assert embed_calls # embedding path was actually exercised
def test_history_to_messages_preserves_order_and_roles():
from pydantic_ai.messages import ModelRequest, ModelResponse
msgs = _messages(
system="be terse",
history=[("user", "first"), ("assistant", "reply")],
prompt="second",
)
history = chat_mod._history_to_messages(msgs)
# system, user(first), assistant(reply) — "second" is excluded (it's the
# current turn, passed separately as Agent.run()'s user_prompt).
assert len(history) == 3
assert isinstance(history[0], ModelRequest) # system
assert isinstance(history[1], ModelRequest) # user
assert isinstance(history[2], ModelResponse) # assistant
# ---------------------------------------------------------------------------
# Image input (spec 009, US3; features/us3_structured_image.feature)
# ---------------------------------------------------------------------------
def test_chat_structured_attaches_image_to_user_prompt(monkeypatch):
"""The current turn's image rides on Agent.run()'s user_prompt as a
pydantic-ai BinaryContent (ADR-008, research.md Decision 1)."""
import base64
from pydantic_ai.messages import BinaryContent
fake = _FakeAgent([_Widget(name="sq", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
b64 = base64.b64encode(b"PNGDATA").decode()
_run_async(
chat_mod.chat_structured(
base_url="http://x/v1",
model="m",
schema=_Widget,
messages=[Message(role="user", content="describe", images=[b64])],
)
)
prompt = fake.calls[0][0]
assert isinstance(prompt, list)
assert prompt[0] == "describe"
assert isinstance(prompt[1], BinaryContent)
assert prompt[1].data == b"PNGDATA"
assert prompt[1].media_type == "image/png"
def test_chat_structured_text_only_prompt_is_plain_string(monkeypatch):
"""FR-003: an image-less structured call is unchanged — plain-string
user_prompt, exactly as before spec 009."""
fake = _FakeAgent([_Widget(name="a", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
_run_async(
chat_mod.chat_structured(
base_url="http://x/v1",
model="m",
schema=_Widget,
messages=[Message(role="user", content="hi")],
)
)
assert fake.calls[0][0] == "hi"
def test_chat_structured_attaches_image_to_history_user_turn(monkeypatch):
"""A prior user turn that carried an image keeps it in message_history."""
import base64
from pydantic_ai.messages import BinaryContent, UserPromptPart
fake = _FakeAgent([_Widget(name="a", count=1)])
monkeypatch.setattr(chat_mod, "_build_agent", lambda **kw: fake)
b64 = base64.b64encode(b"IMG").decode()
msgs = [
Message(role="user", content="earlier", images=[b64]),
Message(role="assistant", content="ok"),
Message(role="user", content="now"),
]
_run_async(
chat_mod.chat_structured(
base_url="http://x/v1", model="m", schema=_Widget, messages=msgs
)
)
history = fake.calls[0][1]
part = history[0].parts[0]
assert isinstance(part, UserPromptPart)
assert isinstance(part.content, list)
assert part.content[0] == "earlier"
assert isinstance(part.content[1], BinaryContent)
assert part.content[1].data == b"IMG"
+74
View File
@@ -0,0 +1,74 @@
"""
Tests for comfydv._llm — shared LLMProvider protocol and OllamaProvider.
Test layers:
Unit (no marker) — pure Python, no live services
Integration (-m integration) — requires live Ollama at localhost:11434
BDD coverage:
../specs/007-llm-provider-abstraction/features/us1_connect_and_chat.feature
../specs/007-llm-provider-abstraction/features/us2_structured_output.feature
../specs/007-llm-provider-abstraction/features/us3_model_lifecycle.feature
"""
from comfydv._llm.ollama_provider import OllamaProvider
from comfydv._llm.provider import Message, ModelInfo, ModelStatus
def test_ollama_provider_captures_connection_state():
provider = OllamaProvider("http://localhost:11434", headers={"X-Test": "1"})
assert provider.host == "http://localhost:11434"
assert provider.headers == {"X-Test": "1"}
def test_ollama_provider_headers_default_to_none():
provider = OllamaProvider("http://localhost:11434")
assert provider.headers is None
def test_model_status_values():
assert ModelStatus.LOADED == "loaded"
assert ModelStatus.SLEEPING == "sleeping"
assert ModelStatus.DOWNLOADING == "downloading"
def test_model_info_optional_size():
info = ModelInfo(name="llama3", status=ModelStatus.UNLOADED)
assert info.size is None
def test_message_roles():
Message(role="system", content="be terse")
Message(role="user", content="hi")
Message(role="assistant", content="hello")
# --- US1 foundational: Message.images carrier (spec 009, contract T1) ---
def test_message_images_defaults_to_none():
"""A text-only turn carries no images."""
msg = Message(role="user", content="hi")
assert msg.images is None
def test_message_images_round_trips_base64_list():
msg = Message(role="user", content="describe", images=["aGVsbG8=", "d29ybGQ="])
assert msg.images == ["aGVsbG8=", "d29ybGQ="]
def test_message_text_only_dump_omits_images_key():
"""FR-003/SC-004: an image-less message must serialize byte-identically to
today — no stray ``images`` key in the transport payload."""
msg = Message(role="user", content="hi")
assert msg.model_dump(exclude_none=True) == {"role": "user", "content": "hi"}
def test_message_with_images_dump_includes_images_key():
msg = Message(role="user", content="describe", images=["aGVsbG8="])
dumped = msg.model_dump(exclude_none=True)
assert dumped == {
"role": "user",
"content": "describe",
"images": ["aGVsbG8="],
}
+426
View File
@@ -0,0 +1,426 @@
"""Tests for comfydv._llm.retry — shared retry-on-blank-output helpers used
by both providers' chat() and the shared chat_structured() helper, plus the
refusal/deflection detector that rides the same retry-with-a-new-seed
mechanism.
Refusal-detection tests here are pure-logic only — no provider/HTTP
involved. See test_ollama_provider.py, test_llamacpp_provider.py, and
test_llm_chat_structured.py for the retry-loop integration (does a detected
refusal actually trigger a reseeded retry).
"""
import pytest
from comfydv._llm.ollama_provider import _run_async
from comfydv._llm.retry import (
REFUSAL_EXEMPLARS,
cosine_similarity,
format_recovered_status,
format_retry_status,
is_ambiguous,
is_lexical_refusal,
is_refusal,
next_seed,
next_timeout_secs,
record_attempt_info,
)
def test_next_seed_attempt_one_is_zero_by_default():
assert next_seed(None, 1) == 0
def test_next_seed_increments_from_zero_when_unset():
assert next_seed(None, 2) == 1
assert next_seed({}, 3) == 2
def test_next_seed_starts_from_pinned_base():
assert next_seed({"seed": 42}, 1) == 42
assert next_seed({"seed": 42}, 2) == 43
assert next_seed({"seed": 42}, 3) == 44
def test_next_seed_ignores_non_int_seed():
assert next_seed({"seed": "not-an-int"}, 2) == 1
class TestNextTimeoutSecs:
def test_attempt_one_returns_base_timeout_unchanged(self):
assert next_timeout_secs(300.0, 1) == 300.0
def test_escalates_multiplicatively_per_attempt(self):
assert next_timeout_secs(300.0, 2) == 600.0
assert next_timeout_secs(300.0, 3) == 900.0
class TestRecordAttemptInfo:
def test_none_attempt_info_is_a_no_op(self):
# Must not raise — callers that don't care about this metadata pass
# None and should see no behavior change at all.
record_attempt_info(None, seed=1, attempts=2, timeout_secs=600.0, refusals=1)
def test_populates_dict_in_place(self):
info: dict = {}
record_attempt_info(info, seed=7, attempts=3, timeout_secs=900.0, refusals=2)
assert info == {
"seed": 7,
"attempts": 3,
"timeout_secs": 900.0,
"refusals": 2,
}
def test_overwrites_previous_values(self):
# Callers call this once per attempt (or once per return path) —
# a later call must replace, not merge with, an earlier one.
info: dict = {"seed": 1, "attempts": 1, "timeout_secs": 300.0, "refusals": 0}
record_attempt_info(info, seed=2, attempts=2, timeout_secs=600.0, refusals=1)
assert info == {
"seed": 2,
"attempts": 2,
"timeout_secs": 600.0,
"refusals": 1,
}
class TestFormatStatus:
def test_retry_status_includes_reason_attempt_seed_and_timeout(self):
msg = format_retry_status("Blank response", 1, 3, seed=1, timeout_secs=200.0)
assert "Blank response" in msg
assert "attempt 1/3" in msg
assert "seed=1" in msg
assert "timeout=200s" in msg
def test_recovered_status_includes_attempt_and_seed(self):
msg = format_recovered_status(2, 3, seed=1)
assert "attempt 2/3" in msg
assert "seed=1" in msg
# ---------------------------------------------------------------------------
# Refusal/deflection detection
# ---------------------------------------------------------------------------
class TestLexicalDetection:
def test_blatant_refusal_openers_match(self):
blatant = [
"I cannot generate that image for you.",
"I'm not able to help with this request.",
"I am unable to comply with your instruction.",
"Cannot generate this image due to the sensitive content.",
'{"error": "cannot generate this content"}',
"I don't feel comfortable creating this content.",
"I'm sorry, but I can't produce that description.",
"As an AI, I'm unable to depict this scenario.",
"I must decline this request.",
"This request violates my usage policies.",
"That's not something I'm able to help with right now.",
]
for text in blatant:
assert is_lexical_refusal(text), f"expected refusal match: {text!r}"
def test_ordinary_content_does_not_match(self):
ordinary = [
"The subject turns to face the camera and smiles warmly.",
"A person cannot simply walk into Mordor, the guide joked.",
"I can help you plan a birthday party for your dog.",
"",
]
for text in ordinary:
assert not is_lexical_refusal(text), f"unexpected match: {text!r}"
class TestAmbiguityHeuristic:
def test_short_response_is_ambiguous(self):
assert is_ambiguous("Sorry, can't do that one.")
def test_long_response_without_hedge_keywords_is_not_ambiguous(self):
long_text = "The subject rotates smoothly toward the lens. " * 20
assert len(long_text) >= 400
assert not is_ambiguous(long_text)
def test_long_response_with_hedge_keyword_is_ambiguous(self):
long_text = "Unfortunately, " + "this touches on a sensitive area. " * 20
assert len(long_text) >= 400
assert is_ambiguous(long_text)
def test_blank_text_is_not_ambiguous(self):
assert not is_ambiguous(" ")
class TestCosineSimilarity:
def test_identical_vectors_score_one(self):
assert cosine_similarity([1.0, 0.0], [1.0, 0.0]) == pytest.approx(1.0)
def test_orthogonal_vectors_score_zero(self):
assert cosine_similarity([1.0, 0.0], [0.0, 1.0]) == pytest.approx(0.0)
def test_opposite_vectors_score_negative_one(self):
assert cosine_similarity([1.0, 0.0], [-1.0, 0.0]) == pytest.approx(-1.0)
def test_mismatched_lengths_return_zero(self):
assert cosine_similarity([1.0, 0.0], [1.0, 0.0, 0.0]) == 0.0
def test_empty_vectors_return_zero(self):
assert cosine_similarity([], []) == 0.0
class TestIsRefusalHybrid:
def test_blank_text_is_never_a_refusal(self):
assert not _run_async(is_refusal(""))
assert not _run_async(is_refusal(" "))
def test_lexical_match_short_circuits_without_embedding(self):
calls = []
async def embed_fn(text):
calls.append(text)
return [1.0, 0.0]
result = _run_async(
is_refusal("I cannot generate that image for you.", embed_fn=embed_fn)
)
assert result is True
assert calls == [] # never reached the embedding step
def test_long_clean_response_is_still_embedding_checked_but_not_a_refusal(self):
# embed_fn present means the caller opted in to the embedding check
# regardless of length/hedge-keywords (is_ambiguous no longer gates
# this) — a long, on-topic response should still be embedding-
# checked, it just shouldn't score as similar to the refusal
# exemplars.
calls = []
async def embed_fn(text):
calls.append(text)
if text == long_text:
return [0.0, 1.0] # orthogonal to the exemplar vector below
return [1.0, 0.0] # exemplars
long_text = "The subject rotates smoothly toward the lens. " * 20
result = _run_async(is_refusal(long_text, embed_fn=embed_fn))
assert result is False
assert long_text in calls # embedding check DID run, just scored low
def test_no_embed_fn_degrades_to_lexical_only(self):
# Ambiguous (short), no lexical match, no embed_fn -> can't check further
assert not _run_async(is_refusal("Not today, sorry.", embed_fn=None))
def test_ambiguous_response_above_threshold_is_refusal(self):
async def embed_fn(text):
# Exemplars and a near-identical short "refusal-ish" probe get a
# high similarity score; distinguish by a marker substring.
if "PROBE" in text:
return [1.0, 0.0]
return [0.99, 0.14] # cos-sim with [1,0] is ~0.99
result = _run_async(
is_refusal(
"PROBE: not comfortable with this one",
embed_fn=embed_fn,
embed_cache_key="test-model",
threshold=0.8,
)
)
assert result is True
def test_ambiguous_response_below_threshold_is_not_refusal(self):
async def embed_fn(text):
if "PROBE" in text:
return [1.0, 0.0]
return [0.0, 1.0] # orthogonal -> cos-sim 0.0
result = _run_async(
is_refusal(
"PROBE: a short reply",
embed_fn=embed_fn,
embed_cache_key="test-model-2",
threshold=0.8,
)
)
assert result is False
def test_long_json_shaped_soft_refusal_without_hedge_keywords_is_caught(self):
# Regression case: a structured-output-shaped response (>600 chars
# once you count JSON braces/field names) whose deflection doesn't
# use any of the canned hedge keywords used to be invisible to the
# embedding check entirely, because is_ambiguous gated on length
# and keywords. embed_fn now runs unconditionally once configured.
soft_refusal = (
'{"prompt": "'
+ "Let's take this in a different creative direction that everyone can enjoy. "
* 8
+ '"}'
)
assert len(soft_refusal) >= 600
assert not is_lexical_refusal(soft_refusal)
async def embed_fn(text):
if text == soft_refusal:
return [1.0, 0.0]
return [0.97, 0.24] # exemplars: cos-sim with [1,0] is ~0.97
result = _run_async(
is_refusal(
soft_refusal,
embed_fn=embed_fn,
embed_cache_key="regression-model",
threshold=0.8,
)
)
assert result is True
def test_exemplar_embeddings_cached_across_calls(self):
exemplar_calls = {"n": 0}
async def embed_fn(text):
if "PROBE" in text:
return [1.0, 0.0]
exemplar_calls["n"] += 1
return [1.0, 0.0]
_run_async(
is_refusal(
"PROBE: first ambiguous call",
embed_fn=embed_fn,
embed_cache_key="cache-key-shared",
threshold=0.5,
)
)
first_count = exemplar_calls["n"]
assert first_count > 0
_run_async(
is_refusal(
"PROBE: second ambiguous call",
embed_fn=embed_fn,
embed_cache_key="cache-key-shared",
threshold=0.5,
)
)
# Exemplar embeddings reused from cache -> no additional exemplar calls
assert exemplar_calls["n"] == first_count
def test_embed_fn_failure_degrades_to_not_refused(self):
async def failing_embed_fn(text):
raise RuntimeError("server unreachable")
result = _run_async(
is_refusal(
"Not comfortable with this one, sorry.",
embed_fn=failing_embed_fn,
embed_cache_key="unreachable-model",
)
)
assert result is False
def test_embed_fn_returning_none_degrades_to_not_refused(self):
async def none_embed_fn(text):
return None
result = _run_async(
is_refusal(
"Not comfortable with this one, sorry.",
embed_fn=none_embed_fn,
embed_cache_key="no-embeddings-model",
)
)
assert result is False
class TestCustomPhrases:
def test_custom_phrase_substring_match_needs_no_embed_fn(self):
# A phrase the user added at runtime that the shipped lexical
# patterns don't cover — should be caught for free, no embedding
# model required.
result = _run_async(
is_refusal(
"I am restricted from producing that kind of content.",
custom_phrases=("restricted from",),
)
)
assert result is True
def test_custom_phrase_match_is_case_insensitive(self):
result = _run_async(
is_refusal(
"SORRY, THAT'S OFF LIMITS FOR ME.",
custom_phrases=("off limits",),
)
)
assert result is True
def test_unrelated_custom_phrase_does_not_match(self):
result = _run_async(
is_refusal(
"The subject walks calmly toward the horizon.",
custom_phrases=("restricted from", "off limits"),
)
)
assert result is False
def test_blank_and_whitespace_custom_phrases_are_ignored(self):
# A stray empty entry must never become a universal substring match.
result = _run_async(
is_refusal(
"The subject walks calmly toward the horizon.",
custom_phrases=("", " "),
)
)
assert result is False
def test_custom_phrase_folded_into_embedding_exemplars(self):
# No exact substring match, but embed_fn scores the response as
# similar to the custom phrase (not one of the shipped exemplars).
custom = "my creators have limited what I can show you"
async def embed_fn(text):
if text == custom:
return [1.0, 0.0]
if text in REFUSAL_EXEMPLARS:
return [0.0, 1.0] # shipped exemplars score orthogonal
return [0.99, 0.14] # the probe response is near the custom one
result = _run_async(
is_refusal(
"There are limits my creators placed on what I can show.",
embed_fn=embed_fn,
embed_cache_key="custom-exemplar-model",
threshold=0.8,
custom_phrases=(custom,),
)
)
assert result is True
def test_different_custom_phrase_sets_do_not_share_exemplar_cache(self):
# Regression guard: if the exemplar cache key ignored custom_phrases,
# a second call with a different custom phrase set would incorrectly
# reuse the first call's cached (and now stale) exemplar vectors.
calls = []
async def embed_fn(text):
calls.append(text)
return [1.0, 0.0]
_run_async(
is_refusal(
"short reply",
embed_fn=embed_fn,
embed_cache_key="shared-model",
custom_phrases=("phrase one",),
)
)
first_call_count = len(calls)
_run_async(
is_refusal(
"short reply",
embed_fn=embed_fn,
embed_cache_key="shared-model",
custom_phrases=("phrase two",),
)
)
# A fresh custom phrase set re-embeds the exemplars (including the
# new phrase) rather than reusing the first set's cached vectors.
assert len(calls) > first_call_count
+918 -679
View File
File diff suppressed because it is too large Load Diff

Some files were not shown because too many files have changed in this diff Show More