diff --git a/CHANGELOG.md b/CHANGELOG.md index d337098..b854dac 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,6 +11,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). - `CHANGELOG.md` (this file) - README: What-is-this, Install, and Quickstart sections - `LLMProvider` protocol (`comfydv._llm`) — a shared adapter boundary so ComfyUI LLM nodes work with any backend that implements it, starting with `OllamaProvider`. Structured output now goes through `pydantic-ai` (ADR-007), superseding the hand-rolled Ollama tool-calling approach. +- **Chat Completion** now accepts an optional `image` input for vision-capable models (VLMs): wire a ComfyUI `IMAGE` and the connected model can describe or reason about it. Works identically on both backends (Ollama multimodal models; llama.cpp launched with `--mmproj`), and composes with structured output and multi-turn history. Images are carried on `Message.images` and translated to each backend's native shape (Ollama's flat `images` array, llama.cpp's OpenAI `image_url` parts, pydantic-ai `BinaryContent` on the structured path) — ADR-008, extending ADR-007's adapter pattern to a second input modality. Text-only workflows are unchanged when no image is wired. ### Changed - **Breaking:** `OllamaChatCompletion` → `ChatCompletion`, `OllamaModelSelector` → `LLMModelSelector`, `OllamaLoadModel` → `LLMLoadModel`, `OllamaUnloadModel` → `LLMUnloadModel`, and the `OLLAMA_CLIENT` socket type → `LLM_CLIENT` — these nodes are now backend-generic. `OllamaClient` is unchanged by name but now outputs an `OllamaProvider` rather than a plain string; existing saved workflows using the old node/socket names need reconnecting (see `comfydv.ollama.MIGRATION_MAP` for the full old→new mapping). diff --git a/README.md b/README.md index ca4a1f8..e1712dc 100644 --- a/README.md +++ b/README.md @@ -114,6 +114,15 @@ A complete graph looks like this: ![Full LLM workflow](docs/assets/ollama_workflow.png) +### Describing images (vision) + +**Chat Completion** has an optional **image** input. Wire any `IMAGE` into it and, with a vision-capable model loaded, the model can describe or reason about the picture — captioning, visual Q&A, reading text in an image, whatever the model supports. + +- **Ollama** — use a multimodal model (e.g. a llava-class model). +- **llama.cpp** — launch `llama-server` with a multimodal projector: `--mmproj ` alongside the model. + +Image input works the same on both backends — same node, same wiring — and composes with everything else: structured output (schema-validated fields pulled straight from the image), multi-turn history, and the option nodes. A batch of images is sent as multiple images on the turn. Leave the input unwired and Chat Completion behaves exactly as before, text only. + ### Manual memory management Single-GPU and memory-constrained setups need explicit control over what's resident in VRAM. **LLM Load Model** pins a model into memory; **LLM Unload Model** evicts it immediately, freeing room for the next model or the rest of your image pipeline. diff --git a/specs/009-vlm-image-input/tasks.md b/specs/009-vlm-image-input/tasks.md index 378bbfc..8afbc63 100644 --- a/specs/009-vlm-image-input/tasks.md +++ b/specs/009-vlm-image-input/tasks.md @@ -85,7 +85,7 @@ handling lives only in the `comfy`-guarded `ollama.py`. ## Phase 6: Polish & Cross-Cutting Concerns -- [ ] T008 [P] Document image input on `ChatCompletion` in `README.md` and add a `CHANGELOG.md` Unreleased entry — note the vision-model / llama.cpp `--mmproj` prerequisite (quickstart.md) +- [x] T008 [P] Document image input on `ChatCompletion` in `README.md` and add a `CHANGELOG.md` Unreleased entry — note the vision-model / llama.cpp `--mmproj` prerequisite (quickstart.md) - [ ] T009 Run the full quality gate green: `uv run ruff check --fix && uv run ruff format && uv run ty check && uv run pytest && beacon doctor --strict` - [-] T010 End-to-end `quickstart.md` validation against a live vision backend (Ollama multimodal model and `llama-server --mmproj`) _Deferred — requires a live vision-capable backend not available in CI/this environment; validate manually before release._