Every ChatCompletion node in ltx-i2v-pipeline.json now reads its options through a new OllamaOptionDisableThinking node (id 18, disable_thinking=True) appended to the end of the existing max-tokens/extra-body chain. Without it, a thinking-capable model routinely burns its whole max_tokens budget on chain-of-thought and never emits the closing JSON, so chat_structured fails validation against an empty string — confirmed live against saracen9/amoral-qwen3.5-9b. README's thinking-headroom section updated to match: thinking off by default, headroom advice now framed as what to do if you deliberately re-enable it via node 18. Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
12 KiB
LTX-2.3 I2V Multi-Agent Pipeline
ltx-i2v-pipeline.json wires the 6-agent prompt-compiler pipeline from
project-management/Work/planning/ltx.md
onto comfydv's existing generic LLM nodes — OllamaClient, ChatCompletion
(structured output, image input) and FormatString (Jinja2 templating). No
new node code was needed; this is a wiring exercise, not a feature.
Loading it
This is a ComfyUI API-format workflow ({node_id: {class_type, inputs}}),
not the canvas/UI export format. Recent ComfyUI frontends accept this format
directly via drag-and-drop onto the graph (it auto-lays-out the nodes), or you
can POST it straight to /prompt. This format was chosen deliberately over
hand-authoring the litegraph UI-export format: the latter requires exact
per-node-type widget-value ordering and link/slot bookkeping that's easy to
get subtly wrong by hand and impossible for me to verify without a live
ComfyUI+Ollama instance. Every template, schema, and link index in this file
was verified against the actual node source (see "How this was verified"
below) — only the outer graph-serialization format is unverified against a
real ComfyUI load.
If you'd rather have the literal draggable canvas file, load this one once, arrange the nodes, and use ComfyUI's own "Save (API Format)" vs. regular "Save" to produce one — that guarantees a format your ComfyUI build actually round-trips.
Prerequisites
- Ollama running locally with a model that has both
visionandtoolscapability (checkollama show <model>'scapabilitieslist) — every agent in this pipeline usesstructured_output=True, and some also need vision. Default baked into the workflow:qwen3.5:9b, used for every agent (text and vision alike) — edit themodelfield on eachChatCompletionnode if you have something else installed.lukey03/qwen3.5-9b-abliterated-visionwas tried and rejected: its chat template is degenerate enough that it returns valid-shaped but garbled content regardless of structured-output mechanism (see ADR-009). - Thinking is off by default. Node
18(OllamaOptionDisableThinking,disable_thinking=True) sits at the end of the options chain everyChatCompletionnode reads from. Without it, a thinking-capable model routinely burns its wholemax_tokensbudget on chain-of-thought and never emits the closing JSON —chat_structuredthen fails validation against an empty string. Flip node18'sdisable_thinkingtoFalseif you deliberately want a model to reason before answering; if you do, give it real headroom: nodes16/17setmax_tokens=8192/num_ctx=32768for exactly that case, and Ollama's default context (4096, with--context-shiftsilently evicting old context rather than stopping) is nowhere near enough for these agents' long system prompts on top of reasoning tokens. Node17'snum_ctxreaches Ollama correctly becauseOllamaProvider.chat_structured()calls Ollama's native/api/chat+"format"directly (ADR-009) — an earlier version of this fix tried priming context via a separate call before the real request and that didn't work, because Ollama's OpenAI-compatible endpoint silently reloads the model at its default context on every call, undoing any priming; the native endpoint doesn't have that problem and appliesoptionsand structured output atomically in one request. EveryChatCompletionnode'stimeout_secs=600for the same headroom reason; lower it if your hardware is faster than the machine this was tuned against. - To use
LlamaCppClientinstead ofOllamaClient, swap node1'sclass_typeandhost— everyChatCompletionnode keeps working unchanged, since both emit the sameLLM_CLIENTtype (ADR-007). Note llama-server's context is fixed at process launch (--ctx-size), not a per-request setting — node17'snum_ctxonly affects Ollama. - Replace node
2'simagefilename with your actual starting frame.
Status: confirmed working end-to-end against a live server
Ran to completion (status: success) against a real local Ollama server
(qwen3.5:9b), through the actual ComfyUI node graph, all 6 agents plus
the final-output node. Getting here took two real comfydv bugs, both fixed
and documented in ADR-009:
- Ollama's OpenAI-compatible endpoint silently reloads the model at its
default (tiny) context size on every call, discarding any
options.num_ctx— fixed by switchingOllamaProvider.chat_structured()to Ollama's native/api/chat+"format", which doesn't have that problem. _build_structured_model(src/comfydv/ollama.py) typed non-required schema fields as barepy_typewith aNonedefault, which only covers a field being omitted — an explicitnullin the model's JSON (which models routinely emit) failed pydantic validation. Fixed by typing those fieldspy_type | None.
One remaining quirk, not a wiring bug: qwen3.5:9b sometimes
under-attends to short/simple prompt content — in one full run it reported
Agent 1's user_intent as "no content provided" despite the field being
populated, which cascaded into an empty Director prompt, which the Judge
correctly caught (decision: FAIL) and the Refiner correctly attempted to
repair. That's the multi-agent design working as intended against a bad
upstream extraction — the fix for that is prompt/model tuning on Agent 1,
not a pipeline change. Re-run if you hit it; it isn't consistent.
# isolated single-agent test — much faster to debug than the full graph
python3 -c "
import json
d = json.load(open('ltx-i2v-pipeline.json'))
subset = {k: d[k] for k in ['1','2','16','17','3','4']} # Agent 1 only
json.dump({'prompt': subset}, open('/tmp/agent1_only.json','w'))
"
curl -X POST http://localhost:8188/prompt -H "Content-Type: application/json" \
--data @/tmp/agent1_only.json
Pipeline shape
OllamaClient ─┬─────────────────────────────────────────────────────────┐
LoadImage ────┼──────────┬──────────┬──────────┬──────────┐ │
│ │ │ │ │ │
FormatString→ChatCompletion (Agent 1: Intent Compiler) [no image] │
│ │ │
FormatString→ChatCompletion (Agent 2: Scene Grounder) ←image
│
FormatString→ChatCompletion (Agent 3: Manifest Verifier) ←image
│
intent ─┴─ audited_manifest
FormatString→ChatCompletion (Agent 4: Director) ←image
│
+ intent + manifest ────────┴── candidate_prompt
FormatString→ChatCompletion (Agent 5: Judge) ←image
│
+ everything above ─────────┴── judge_report
FormatString→ChatCompletion (Agent 6: Refiner) [no image]
│
FormatString (Final Output — judge decision + both prompts)
Deliberate adaptations from ltx.md
-
Single round, no retry loop. ltx.md's reference pseudocode runs
for iteration in range(2): judge → refine, short-circuiting on PASS. ComfyUI graphs are DAGs with no native conditional/loop node in this repo (checkedcircuit_breaker.py,random_choice.py— neither fits), so a real retry loop can't be expressed as a static graph. This workflow always runs Judge once and Refiner once. The Final Output node (15) shows the Judge'sdecisionnext to both the Director's candidate prompt and the Refiner's patched prompt — read the decision and use the candidate prompt on PASS, the refined prompt on FAIL. Wire a second Judge/Refiner pair after node14yourself if you want the second round. -
Structured output carries whole objects, not just fields.
ChatCompletion'sresponseoutput is the full JSON object (parsed.model_dump_json()), and — sincestructured_output=Truealso adds one extra named output per top-level schema property — a specific nested object can be pulled out directly by name (e.g. Agent 3'saudited_manifestoutput, used instead of itsresponsewrapper, which also containsverification). Templates use{{ x }}directly rather than ltx.md's{{ x | tojson(indent=2) }}, sincexarrives already JSON-encoded. -
List-valued template variables are JSON strings.
FormatString's dynamic inputs are alwaysSTRING; there's no native list socket. Fields likepreservation_requirementsorextraction_hintsare typed as JSON arrays (e.g.["keep hairstyle"]) and unpacked in-template with thefromjsonfilter your repo'sFormatStringalready ships. Leave them as empty string""to omit the section entirely — falsy-string{% if %}checks guard every optional block, sofromjsonis never called on an empty value. -
output_schemais intentionally shallow.ChatCompletiononly enforces top-level property types (see_build_structured_modelinsrc/comfydv/ollama.py) — it doesn't validate nested structure. The detailed nested shape each agent must produce (e.g. every field insiderequired_camera) is still communicated to the model via the literal JSON example embedded in that agent's system prompt (verbatim from ltx.md), so nothing is lost — theoutput_schemaJSON here just needs to get the top-level field list and types right, which is also allChatCompletionuses it for. -
Required fields exclude anything legitimately blank. A
requiredstring field is forced non-empty byChatCompletion(Field(..., min_length=1)) to catch blank-output failures. Fields that are correctly empty on a non-nominal status — Director'spromptonUNSATISFIABLE, Judge'srefinement_instructiononPASS, Refiner'sprompt/unresolvable_reason— are deliberately left out of each schema'srequiredlist so a legitimate empty string doesn't trigger a retry loop against the model. -
Optional Agent 7 (Targeted Manifest Resolver) is not wired. It only fires on a
MANIFEST_CHALLENGE, which this static graph can't branch on. Add it manually if the Director or Judge start returning that status for your inputs.
How this was verified
Everything except the outer graph-serialization format was checked against
this repo's actual node code (not just read — executed), with mocked
comfy/server/folder_paths modules the way tests/conftest.py does:
- Every
[node_id, index]link target resolves to a real node. - Every
FormatStringtemplate's variables (via the samejinja_env.parse+meta.find_undeclared_variablesAST extraction_extract_keysuses) exactly match the inputs supplied in the workflow. - Every template renders through the real
FormatString.format_string()with representative values, including the Director'sSHOT CONSTRAINTSblock, which was parsed back withjson.loadsto confirm it's valid JSON even when every optional field is left blank. - Every
output_schemaparses through the real_parse_output_schema/_build_structured_model, and every link that targets a named structured output (e.g.audited_manifest,prompt,decision) was checked against the actual computed(response, updated_history, model_name, *properties)output order for that schema.