Files
darth-veitcher-comfydv/workflows/README.md
T
James VeitchandClaude Sonnet 5 2724f535f5 fix(workflows): disable thinking by default in LTX-2.3 pipeline (#34)
Every ChatCompletion node in ltx-i2v-pipeline.json now reads its options
through a new OllamaOptionDisableThinking node (id 18, disable_thinking=True)
appended to the end of the existing max-tokens/extra-body chain. Without it,
a thinking-capable model routinely burns its whole max_tokens budget on
chain-of-thought and never emits the closing JSON, so chat_structured fails
validation against an empty string — confirmed live against
saracen9/amoral-qwen3.5-9b. README's thinking-headroom section updated to
match: thinking off by default, headroom advice now framed as what to do if
you deliberately re-enable it via node 18.


Claude-Session: https://claude.ai/code/session_01YArD9ZjBWKsAvazmS48amA

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 08:15:49 +01:00

12 KiB

LTX-2.3 I2V Multi-Agent Pipeline

ltx-i2v-pipeline.json wires the 6-agent prompt-compiler pipeline from project-management/Work/planning/ltx.md onto comfydv's existing generic LLM nodes — OllamaClient, ChatCompletion (structured output, image input) and FormatString (Jinja2 templating). No new node code was needed; this is a wiring exercise, not a feature.

Loading it

This is a ComfyUI API-format workflow ({node_id: {class_type, inputs}}), not the canvas/UI export format. Recent ComfyUI frontends accept this format directly via drag-and-drop onto the graph (it auto-lays-out the nodes), or you can POST it straight to /prompt. This format was chosen deliberately over hand-authoring the litegraph UI-export format: the latter requires exact per-node-type widget-value ordering and link/slot bookkeping that's easy to get subtly wrong by hand and impossible for me to verify without a live ComfyUI+Ollama instance. Every template, schema, and link index in this file was verified against the actual node source (see "How this was verified" below) — only the outer graph-serialization format is unverified against a real ComfyUI load.

If you'd rather have the literal draggable canvas file, load this one once, arrange the nodes, and use ComfyUI's own "Save (API Format)" vs. regular "Save" to produce one — that guarantees a format your ComfyUI build actually round-trips.

Prerequisites

  • Ollama running locally with a model that has both vision and tools capability (check ollama show <model>'s capabilities list) — every agent in this pipeline uses structured_output=True, and some also need vision. Default baked into the workflow: qwen3.5:9b, used for every agent (text and vision alike) — edit the model field on each ChatCompletion node if you have something else installed. lukey03/qwen3.5-9b-abliterated-vision was tried and rejected: its chat template is degenerate enough that it returns valid-shaped but garbled content regardless of structured-output mechanism (see ADR-009).
  • Thinking is off by default. Node 18 (OllamaOptionDisableThinking, disable_thinking=True) sits at the end of the options chain every ChatCompletion node reads from. Without it, a thinking-capable model routinely burns its whole max_tokens budget on chain-of-thought and never emits the closing JSON — chat_structured then fails validation against an empty string. Flip node 18's disable_thinking to False if you deliberately want a model to reason before answering; if you do, give it real headroom: nodes 16/17 set max_tokens=8192/num_ctx=32768 for exactly that case, and Ollama's default context (4096, with --context-shift silently evicting old context rather than stopping) is nowhere near enough for these agents' long system prompts on top of reasoning tokens. Node 17's num_ctx reaches Ollama correctly because OllamaProvider.chat_structured() calls Ollama's native /api/chat + "format" directly (ADR-009) — an earlier version of this fix tried priming context via a separate call before the real request and that didn't work, because Ollama's OpenAI-compatible endpoint silently reloads the model at its default context on every call, undoing any priming; the native endpoint doesn't have that problem and applies options and structured output atomically in one request. Every ChatCompletion node's timeout_secs=600 for the same headroom reason; lower it if your hardware is faster than the machine this was tuned against.
  • To use LlamaCppClient instead of OllamaClient, swap node 1's class_type and host — every ChatCompletion node keeps working unchanged, since both emit the same LLM_CLIENT type (ADR-007). Note llama-server's context is fixed at process launch (--ctx-size), not a per-request setting — node 17's num_ctx only affects Ollama.
  • Replace node 2's image filename with your actual starting frame.

Status: confirmed working end-to-end against a live server

Ran to completion (status: success) against a real local Ollama server (qwen3.5:9b), through the actual ComfyUI node graph, all 6 agents plus the final-output node. Getting here took two real comfydv bugs, both fixed and documented in ADR-009:

  1. Ollama's OpenAI-compatible endpoint silently reloads the model at its default (tiny) context size on every call, discarding any options.num_ctx — fixed by switching OllamaProvider.chat_structured() to Ollama's native /api/chat + "format", which doesn't have that problem.
  2. _build_structured_model (src/comfydv/ollama.py) typed non-required schema fields as bare py_type with a None default, which only covers a field being omitted — an explicit null in the model's JSON (which models routinely emit) failed pydantic validation. Fixed by typing those fields py_type | None.

One remaining quirk, not a wiring bug: qwen3.5:9b sometimes under-attends to short/simple prompt content — in one full run it reported Agent 1's user_intent as "no content provided" despite the field being populated, which cascaded into an empty Director prompt, which the Judge correctly caught (decision: FAIL) and the Refiner correctly attempted to repair. That's the multi-agent design working as intended against a bad upstream extraction — the fix for that is prompt/model tuning on Agent 1, not a pipeline change. Re-run if you hit it; it isn't consistent.

# isolated single-agent test — much faster to debug than the full graph
python3 -c "
import json
d = json.load(open('ltx-i2v-pipeline.json'))
subset = {k: d[k] for k in ['1','2','16','17','3','4']}  # Agent 1 only
json.dump({'prompt': subset}, open('/tmp/agent1_only.json','w'))
"
curl -X POST http://localhost:8188/prompt -H "Content-Type: application/json" \
  --data @/tmp/agent1_only.json

Pipeline shape

OllamaClient ─┬─────────────────────────────────────────────────────────┐
LoadImage ────┼──────────┬──────────┬──────────┬──────────┐             │
              │          │          │          │          │             │
FormatString→ChatCompletion (Agent 1: Intent Compiler)     [no image]   │
              │                                    │                     │
              FormatString→ChatCompletion (Agent 2: Scene Grounder) ←image
                                    │
              FormatString→ChatCompletion (Agent 3: Manifest Verifier) ←image
                                    │
       intent ─┴─ audited_manifest
              FormatString→ChatCompletion (Agent 4: Director) ←image
                                    │
       + intent + manifest ────────┴── candidate_prompt
              FormatString→ChatCompletion (Agent 5: Judge) ←image
                                    │
       + everything above ─────────┴── judge_report
              FormatString→ChatCompletion (Agent 6: Refiner)  [no image]
                                    │
              FormatString (Final Output — judge decision + both prompts)

Deliberate adaptations from ltx.md

  1. Single round, no retry loop. ltx.md's reference pseudocode runs for iteration in range(2): judge → refine, short-circuiting on PASS. ComfyUI graphs are DAGs with no native conditional/loop node in this repo (checked circuit_breaker.py, random_choice.py — neither fits), so a real retry loop can't be expressed as a static graph. This workflow always runs Judge once and Refiner once. The Final Output node (15) shows the Judge's decision next to both the Director's candidate prompt and the Refiner's patched prompt — read the decision and use the candidate prompt on PASS, the refined prompt on FAIL. Wire a second Judge/Refiner pair after node 14 yourself if you want the second round.

  2. Structured output carries whole objects, not just fields. ChatCompletion's response output is the full JSON object (parsed.model_dump_json()), and — since structured_output=True also adds one extra named output per top-level schema property — a specific nested object can be pulled out directly by name (e.g. Agent 3's audited_manifest output, used instead of its response wrapper, which also contains verification). Templates use {{ x }} directly rather than ltx.md's {{ x | tojson(indent=2) }}, since x arrives already JSON-encoded.

  3. List-valued template variables are JSON strings. FormatString's dynamic inputs are always STRING; there's no native list socket. Fields like preservation_requirements or extraction_hints are typed as JSON arrays (e.g. ["keep hairstyle"]) and unpacked in-template with the fromjson filter your repo's FormatString already ships. Leave them as empty string "" to omit the section entirely — falsy-string {% if %} checks guard every optional block, so fromjson is never called on an empty value.

  4. output_schema is intentionally shallow. ChatCompletion only enforces top-level property types (see _build_structured_model in src/comfydv/ollama.py) — it doesn't validate nested structure. The detailed nested shape each agent must produce (e.g. every field inside required_camera) is still communicated to the model via the literal JSON example embedded in that agent's system prompt (verbatim from ltx.md), so nothing is lost — the output_schema JSON here just needs to get the top-level field list and types right, which is also all ChatCompletion uses it for.

  5. Required fields exclude anything legitimately blank. A required string field is forced non-empty by ChatCompletion (Field(..., min_length=1)) to catch blank-output failures. Fields that are correctly empty on a non-nominal status — Director's prompt on UNSATISFIABLE, Judge's refinement_instruction on PASS, Refiner's prompt/unresolvable_reason — are deliberately left out of each schema's required list so a legitimate empty string doesn't trigger a retry loop against the model.

  6. Optional Agent 7 (Targeted Manifest Resolver) is not wired. It only fires on a MANIFEST_CHALLENGE, which this static graph can't branch on. Add it manually if the Director or Judge start returning that status for your inputs.

How this was verified

Everything except the outer graph-serialization format was checked against this repo's actual node code (not just read — executed), with mocked comfy/server/folder_paths modules the way tests/conftest.py does:

  • Every [node_id, index] link target resolves to a real node.
  • Every FormatString template's variables (via the same jinja_env.parse + meta.find_undeclared_variables AST extraction _extract_keys uses) exactly match the inputs supplied in the workflow.
  • Every template renders through the real FormatString.format_string() with representative values, including the Director's SHOT CONSTRAINTS block, which was parsed back with json.loads to confirm it's valid JSON even when every optional field is left blank.
  • Every output_schema parses through the real _parse_output_schema/ _build_structured_model, and every link that targets a named structured output (e.g. audited_manifest, prompt, decision) was checked against the actual computed (response, updated_history, model_name, *properties) output order for that schema.