ltx-i2v-pipeline.json (API-format) can't be reliably drag-and-dropped into ComfyUI: its generic API->graph importer never triggers the callbacks ChatCompletion (structured-output sockets) and FormatString (per-template- variable sockets) need to create their dynamic inputs/outputs, so every link touching one silently drops on load. ltx-i2v-pipeline-canvas.json is the fix: loaded against a live ComfyUI+ Ollama instance, repaired node-by-node by calling each node's own dynamic- socket route directly, replayed all 39 links from the API-format file by name (zero errors), added a shared Model Name node (one edit instead of six) and a User Intent node, then saved via ComfyUI's own serializer. Verified programmatically post-save: 52 links, zero missing connections beyond the intentionally-blank optional fields. README updated to point users at the canvas file for UI loading and keep the API-format file for direct /prompt POSTs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA
13 KiB
LTX-2.3 I2V Multi-Agent Pipeline
ltx-i2v-pipeline.json wires the 6-agent prompt-compiler pipeline from
project-management/Work/planning/ltx.md
onto comfydv's existing generic LLM nodes — OllamaClient, ChatCompletion
(structured output, image input) and FormatString (Jinja2 templating). No
new node code was needed; this is a wiring exercise, not a feature.
Loading it
Two files, two purposes:
ltx-i2v-pipeline-canvas.json— the canvas/UI format. Load this one via drag-and-drop or File > Open in ComfyUI. It's fully wired (20 nodes, 52 links) and ready to run — no manual reconnecting required. It also includes two small conveniences not in the API-format file below: a shared "Model Name"PrimitiveStringnode feeding all 6ChatCompletionagents (edit the model in one place), and a "User Intent"PrimitiveStringnode holding the literal starting request text.ltx-i2v-pipeline.json— the ComfyUI API-format workflow ({node_id: {class_type, inputs}}). Use this toPOSTstraight to/promptfor headless/scripted runs. Do not drag-and-drop this one onto the canvas — confirmed live: ComfyUI's generic API→graph importer doesn't trigger the dynamic-socket callbacksChatCompletion(structured outputs) andFormatString(per-template-variable inputs) rely on, so nodes load with only their static fields and every dynamic link is silently dropped.ltx-i2v-pipeline-canvas.jsonwas produced by loading this file, live-repairing exactly that gap against a real ComfyUI+Ollama instance (calling each node's own dynamic-socket routes directly, then replaying every link from this file by name), and saving the result — see "How this was verified" below.
Prerequisites
- Ollama running locally with a model that has both
visionandtoolscapability (checkollama show <model>'scapabilitieslist) — every agent in this pipeline usesstructured_output=True, and some also need vision. Default baked into the workflow:qwen3.5:9b, used for every agent (text and vision alike) — edit themodelfield on eachChatCompletionnode if you have something else installed.lukey03/qwen3.5-9b-abliterated-visionwas tried and rejected: its chat template is degenerate enough that it returns valid-shaped but garbled content regardless of structured-output mechanism (see ADR-009). - Thinking is off by default. Node
18(OllamaOptionDisableThinking,disable_thinking=True) sits at the end of the options chain everyChatCompletionnode reads from. Without it, a thinking-capable model routinely burns its wholemax_tokensbudget on chain-of-thought and never emits the closing JSON —chat_structuredthen fails validation against an empty string. Flip node18'sdisable_thinkingtoFalseif you deliberately want a model to reason before answering; if you do, give it real headroom: nodes16/17setmax_tokens=8192/num_ctx=32768for exactly that case, and Ollama's default context (4096, with--context-shiftsilently evicting old context rather than stopping) is nowhere near enough for these agents' long system prompts on top of reasoning tokens. Node17'snum_ctxreaches Ollama correctly becauseOllamaProvider.chat_structured()calls Ollama's native/api/chat+"format"directly (ADR-009) — an earlier version of this fix tried priming context via a separate call before the real request and that didn't work, because Ollama's OpenAI-compatible endpoint silently reloads the model at its default context on every call, undoing any priming; the native endpoint doesn't have that problem and appliesoptionsand structured output atomically in one request. EveryChatCompletionnode'stimeout_secs=600for the same headroom reason; lower it if your hardware is faster than the machine this was tuned against. - To use
LlamaCppClientinstead ofOllamaClient, swap node1'sclass_typeandhost— everyChatCompletionnode keeps working unchanged, since both emit the sameLLM_CLIENTtype (ADR-007). Note llama-server's context is fixed at process launch (--ctx-size), not a per-request setting — node17'snum_ctxonly affects Ollama. - Replace node
2'simagefilename with your actual starting frame.
Status: confirmed working end-to-end against a live server
Ran to completion (status: success) against a real local Ollama server
(qwen3.5:9b), through the actual ComfyUI node graph, all 6 agents plus
the final-output node. Getting here took two real comfydv bugs, both fixed
and documented in ADR-009:
- Ollama's OpenAI-compatible endpoint silently reloads the model at its
default (tiny) context size on every call, discarding any
options.num_ctx— fixed by switchingOllamaProvider.chat_structured()to Ollama's native/api/chat+"format", which doesn't have that problem. _build_structured_model(src/comfydv/ollama.py) typed non-required schema fields as barepy_typewith aNonedefault, which only covers a field being omitted — an explicitnullin the model's JSON (which models routinely emit) failed pydantic validation. Fixed by typing those fieldspy_type | None.
One remaining quirk, not a wiring bug: qwen3.5:9b sometimes
under-attends to short/simple prompt content — in one full run it reported
Agent 1's user_intent as "no content provided" despite the field being
populated, which cascaded into an empty Director prompt, which the Judge
correctly caught (decision: FAIL) and the Refiner correctly attempted to
repair. That's the multi-agent design working as intended against a bad
upstream extraction — the fix for that is prompt/model tuning on Agent 1,
not a pipeline change. Re-run if you hit it; it isn't consistent.
# isolated single-agent test — much faster to debug than the full graph
python3 -c "
import json
d = json.load(open('ltx-i2v-pipeline.json'))
subset = {k: d[k] for k in ['1','2','16','17','3','4']} # Agent 1 only
json.dump({'prompt': subset}, open('/tmp/agent1_only.json','w'))
"
curl -X POST http://localhost:8188/prompt -H "Content-Type: application/json" \
--data @/tmp/agent1_only.json
Pipeline shape
OllamaClient ─┬─────────────────────────────────────────────────────────┐
LoadImage ────┼──────────┬──────────┬──────────┬──────────┐ │
│ │ │ │ │ │
FormatString→ChatCompletion (Agent 1: Intent Compiler) [no image] │
│ │ │
FormatString→ChatCompletion (Agent 2: Scene Grounder) ←image
│
FormatString→ChatCompletion (Agent 3: Manifest Verifier) ←image
│
intent ─┴─ audited_manifest
FormatString→ChatCompletion (Agent 4: Director) ←image
│
+ intent + manifest ────────┴── candidate_prompt
FormatString→ChatCompletion (Agent 5: Judge) ←image
│
+ everything above ─────────┴── judge_report
FormatString→ChatCompletion (Agent 6: Refiner) [no image]
│
FormatString (Final Output — judge decision + both prompts)
Deliberate adaptations from ltx.md
-
Single round, no retry loop. ltx.md's reference pseudocode runs
for iteration in range(2): judge → refine, short-circuiting on PASS. ComfyUI graphs are DAGs with no native conditional/loop node in this repo (checkedcircuit_breaker.py,random_choice.py— neither fits), so a real retry loop can't be expressed as a static graph. This workflow always runs Judge once and Refiner once. The Final Output node (15) shows the Judge'sdecisionnext to both the Director's candidate prompt and the Refiner's patched prompt — read the decision and use the candidate prompt on PASS, the refined prompt on FAIL. Wire a second Judge/Refiner pair after node14yourself if you want the second round. -
Structured output carries whole objects, not just fields.
ChatCompletion'sresponseoutput is the full JSON object (parsed.model_dump_json()), and — sincestructured_output=Truealso adds one extra named output per top-level schema property — a specific nested object can be pulled out directly by name (e.g. Agent 3'saudited_manifestoutput, used instead of itsresponsewrapper, which also containsverification). Templates use{{ x }}directly rather than ltx.md's{{ x | tojson(indent=2) }}, sincexarrives already JSON-encoded. -
List-valued template variables are JSON strings.
FormatString's dynamic inputs are alwaysSTRING; there's no native list socket. Fields likepreservation_requirementsorextraction_hintsare typed as JSON arrays (e.g.["keep hairstyle"]) and unpacked in-template with thefromjsonfilter your repo'sFormatStringalready ships. Leave them as empty string""to omit the section entirely — falsy-string{% if %}checks guard every optional block, sofromjsonis never called on an empty value. -
output_schemais intentionally shallow.ChatCompletiononly enforces top-level property types (see_build_structured_modelinsrc/comfydv/ollama.py) — it doesn't validate nested structure. The detailed nested shape each agent must produce (e.g. every field insiderequired_camera) is still communicated to the model via the literal JSON example embedded in that agent's system prompt (verbatim from ltx.md), so nothing is lost — theoutput_schemaJSON here just needs to get the top-level field list and types right, which is also allChatCompletionuses it for. -
Required fields exclude anything legitimately blank. A
requiredstring field is forced non-empty byChatCompletion(Field(..., min_length=1)) to catch blank-output failures. Fields that are correctly empty on a non-nominal status — Director'spromptonUNSATISFIABLE, Judge'srefinement_instructiononPASS, Refiner'sprompt/unresolvable_reason— are deliberately left out of each schema'srequiredlist so a legitimate empty string doesn't trigger a retry loop against the model. -
Optional Agent 7 (Targeted Manifest Resolver) is not wired. It only fires on a
MANIFEST_CHALLENGE, which this static graph can't branch on. Add it manually if the Director or Judge start returning that status for your inputs.
How this was verified
ltx-i2v-pipeline-canvas.json was produced against a live ComfyUI+Ollama
instance: loaded ltx-i2v-pipeline.json via the frontend's own
app.loadApiJson, then for every ChatCompletion node called its real
/dv/ollama/update_structured_outputs route (and for every FormatString
node its real updateNodeConfig()) with that node's own schema/template to
get its actual dynamic sockets, then replayed all 39 links from
ltx-i2v-pipeline.json by input name — zero errors — before saving. Every
expected connection was re-checked programmatically (not just visually)
against the saved file: 52 links total, and the only unconnected input
sockets are the intentionally-blank optional ones (headers, history,
the two non-vision agents' unused image input, and the Director's unset
shot-constraint overrides).
Everything in ltx-i2v-pipeline.json itself (the API-format source) was
separately checked against this repo's actual node code (not just read —
executed), with mocked comfy/server/folder_paths modules the way
tests/conftest.py does:
- Every
[node_id, index]link target resolves to a real node. - Every
FormatStringtemplate's variables (via the samejinja_env.parse+meta.find_undeclared_variablesAST extraction_extract_keysuses) exactly match the inputs supplied in the workflow. - Every template renders through the real
FormatString.format_string()with representative values, including the Director'sSHOT CONSTRAINTSblock, which was parsed back withjson.loadsto confirm it's valid JSON even when every optional field is left blank. - Every
output_schemaparses through the real_parse_output_schema/_build_structured_model, and every link that targets a named structured output (e.g.audited_manifest,prompt,decision) was checked against the actual computed(response, updated_history, model_name, *properties)output order for that schema.