Author SHA1 Message Date
gokayfem 9aeca11c35 Merge pull request #164 from gokayfem/chore/repo-hygiene-audit 2026-07-30 13:20:35 +03:00
gokayfem 79a929c1ca fix: make node reference schemas exact 2026-07-30 13:14:54 +03:00
Gokay Aydogan e2b20cde13 test: add real-weight node smoke test and share the path bootstrap
The offline suite stubs the llama.cpp boundary, so it proves what the nodes
send and how they handle responses, but never that inference works. Adds an
opt-in real-weight script that drives the node classes themselves:

- LLMLoader resolving a real GGUF through ComfyUI folder_paths, and
  returning before any weights load
- LLMSampler generating real text, asserted deterministic for a fixed seed
  at temperature 0 rather than asserting on model knowledge
- StructuredOutput constraining a real model to a generated JSON Schema.
  This is the llama.cpp grammar path, which a stub cannot verify at all.
- LLMOptionalMemoryFreeSimple releasing a real llama.cpp allocation

Verified against ggml-org/Qwen3.5-0.8B-GGUF (563 MB, Q4_0) on Metal with
llama-cpp-python 0.3.34: all checks pass.

Also extracts tests/_bootstrap.py. Four of the six manual scripts imported
the package only when the checkout directory was named ComfyUI_VLM_nodes,
which conftest.py already notes is not safe for worktrees named after a
branch. All six now share one helper.
2026-07-30 11:17:11 +03:00
Gokay Aydogan 04275b57cb docs: point pre-3.3.0 changelog links at commits
Compare links assumed tags that do not exist. Retroactively tagging the
2.x/3.x versions would run current CI against code that predates it, so
tagging starts at 3.3.0 and earlier entries link to the commit that
declared each version.
2026-07-30 03:10:24 +03:00
Gokay Aydogan 2e41de6ac2 docs: add complete node reference and documentation guards
47 of the 78 registered node classes were never named in the README, so
there was no way to look up a node seen on a canvas. Adds a reference of
all 78, grouped by menu category, with the class_type that appears in
workflow JSON.

Guards it with tests so it cannot drift again: every registered node must
appear in the README, declared license must match LICENSE, and the current
version must have a changelog entry. All three were verified to fail when
violated.
2026-07-30 03:08:52 +03:00
Gokay Aydogan 8b4226474a docs: add changelog, contributing guide, and templates
The repo had no releases, no tags, and no changelog despite being on
version 3.3.0, so users could not pin a version, roll back, or tell that a
bug they filed had been fixed.

CHANGELOG.md reconstructs the 2.x/3.x history from the version-to-commit
mapping in git, and cites the issues each change resolved.

CONTRIBUTING.md documents the constraints that are easy to break: no work
at import time, never reorder existing widgets, no forceInput, optional
dependencies must fail only their own node, and never install torch.

Issue templates require the environment detail that the historically
unresolvable reports lacked, and route llama-cpp-python build failures to
the upstream install guide.
2026-07-30 03:08:52 +03:00
Gokay Aydogan 8a432184d3 ci: add ruff lint job and a coverage floor
CI previously ran pytest, compileall, and build, with no linting and no
coverage measurement. Adds a fast lint job and a --cov-fail-under=70 gate
on the Linux/Python 3.13 leg (currently 73%), plus a coverage artifact.
2026-07-30 03:08:52 +03:00
Gokay Aydogan da5f4d5787 style: apply ruff autofixes
Mechanical only: import ordering, typing -> collections.abc imports,
PEP 604 unions, dict.fromkeys, one unused import, and two over-long
tooltip strings rewrapped by hand. No behaviour change.
2026-07-30 03:08:52 +03:00
Gokay Aydogan 45d21d0642 fix: declare Apache-2.0 in package metadata, add tooling config
pyproject.toml declared license = "MIT" while LICENSE has been Apache-2.0
since the initial commit in 2024-01. The MIT string was introduced in
39fc116 one day earlier, so this is a fresh regression, and Apache-2.0 is
the real license: 2.5 years of outside contributions landed under it.

This metadata is published to the Comfy Registry and into any built wheel,
so the wrong license was being advertised downstream.

Also adds ruff and pytest configuration, and requirements-dev.txt for the
lint/coverage tooling. Vendored nodes/joytagger is excluded from lint; the
three ignored rules are documented inline with their reasons.

Bumps version to 3.3.1. Two fixes (b8ae298, c13ee23) landed after 3.3.0
without a version bump, so the publish workflow saw 3.3.0 already on the
registry and skipped them. They reach Registry users with this release.
2026-07-30 03:08:51 +03:00
Gokay Aydogan 8e55c81b34 test: add contract tests for GGUF text and multimodal nodes
nodes/suggest.py (1011 LOC, 11 node classes) and nodes/llavaloader.py
(599 LOC, 6 node classes) had no test references at all, despite carrying
the longest bug history in the pack.

Coverage: suggest.py 0% -> 98%, llavaloader.py 0% -> 99%.

The cases pin the behaviours the historical reports depended on:
- widget order, which Comfy serializes by position (#156)
- sampling kwargs reaching create_chat_completion (#144)
- handle reuse, and teardown on both success and failure (#137)
- structured-output JSON Schema construction and its error paths
- ChatMusician owning the 'respond in ABC notation' instruction (#149)

No llama.cpp wheel, GGUF weights, or GPU are required.
2026-07-30 03:08:51 +03:00
Gökay Aydoğan f06a2a3e6c Merge pull request #163 from gokayfem/codex/fix-smolvlm-llama-cpp-setup
Fix SmolVLM setup dependencies
2026-07-30 02:27:56 +03:00
38 changed files with 2352 additions and 34 deletions
+104
View File
@@ -0,0 +1,104 @@
name: Bug report
description: A node fails, errors, or produces wrong output.
labels: ["bug"]
body:
- type: markdown
attributes:
value: |
Most unresolvable reports are missing the environment details below.
Please run the **VLM Runtime Diagnostics** node and paste its output —
it captures your OS, Python, PyTorch, accelerator backend, and which
optional backends are installed.
- type: input
id: version
attributes:
label: Node pack version
description: From ComfyUI Manager, or the `version` in `pyproject.toml`.
placeholder: "3.3.1"
validations:
required: true
- type: dropdown
id: install
attributes:
label: How did you install it?
options:
- ComfyUI Manager
- Comfy Registry
- git clone into custom_nodes
- Other (describe below)
validations:
required: true
- type: dropdown
id: comfy
attributes:
label: ComfyUI flavour
options:
- ComfyUI Desktop
- ComfyUI Portable (python_embeded)
- Manual install (venv)
- Manual install (conda)
- Cloud / RunPod / other host
validations:
required: true
- type: textarea
id: diagnostics
attributes:
label: VLM Runtime Diagnostics output
description: Add the node to any workflow, run it, and paste the result.
render: text
validations:
required: true
- type: input
id: node
attributes:
label: Which node fails?
placeholder: "LLMSampler, LLavaSamplerSimple, ModernVLM, ..."
validations:
required: true
- type: input
id: model
attributes:
label: Which model / GGUF file?
description: Include the exact filename or Hugging Face repo id.
placeholder: "Qwen 3 VL 4B Instruct, or llava-1.6-mistral-7b.Q4_K_M.gguf"
validations:
required: true
- type: textarea
id: expected
attributes:
label: What did you expect, and what happened instead?
validations:
required: true
- type: textarea
id: traceback
attributes:
label: Full console output
description: |
The complete traceback from the ComfyUI terminal, not just the last
line. Include the startup log if the pack failed to import.
render: shell
validations:
required: true
- type: checkboxes
id: checks
attributes:
label: Before submitting
options:
- label: I updated to the latest version of this node pack and ComfyUI.
required: true
- label: I searched existing open and closed issues.
required: true
- label: >-
If this involves GGUF or `llama-cpp-python`, I installed it with
the arguments for my accelerator from the
[llama-cpp-python install guide](https://github.com/abetlen/llama-cpp-python#installation).
required: false
+11
View File
@@ -0,0 +1,11 @@
blank_issues_enabled: false
contact_links:
- name: llama-cpp-python installation help
url: https://github.com/abetlen/llama-cpp-python#installation
about: >-
Build or GPU-offload failures for GGUF nodes are almost always
llama-cpp-python installation issues. Install the wheel matching your
accelerator first.
- name: ComfyUI Manager and installation problems
url: https://github.com/Comfy-Org/ComfyUI-Manager/issues
about: For problems installing or updating custom nodes in general.
+46
View File
@@ -0,0 +1,46 @@
name: Model or feature request
description: Ask for support for a new VLM/LLM, or a new node.
labels: ["enhancement"]
body:
- type: textarea
id: what
attributes:
label: What would you like added?
validations:
required: true
- type: input
id: model
attributes:
label: Model repository (if requesting a model)
description: A Hugging Face repo id, so the architecture can be checked.
placeholder: "Qwen/Qwen3-VL-8B-Instruct"
- type: dropdown
id: backend
attributes:
label: Which backend would it use?
options:
- transformers (safetensors)
- llama.cpp (GGUF)
- Hosted API
- Not sure
validations:
required: true
- type: textarea
id: why
attributes:
label: What does it let you do that current nodes cannot?
validations:
required: true
- type: checkboxes
id: checks
attributes:
label: Before submitting
options:
- label: >-
I checked the README node reference to confirm this is not already
supported.
required: true
+33
View File
@@ -0,0 +1,33 @@
## What does this change?
<!-- One or two sentences. Link any issue it closes: "Closes #123". -->
## Type of change
- [ ] Bug fix
- [ ] New model support
- [ ] New node
- [ ] Refactor / maintenance
- [ ] Documentation
## Checklist
- [ ] `python -m pytest -q` passes.
- [ ] `python -m ruff check .` passes.
- [ ] Importing the pack still performs no network access, compilation, or
package install.
- [ ] If a node schema changed, existing widget order is preserved (Comfy
serializes widget values by position, so reordering breaks saved
workflows).
- [ ] New optional dependencies fail only the node that needs them, with an
actionable error.
- [ ] `pyproject.toml` `version` is bumped if this is user-visible, and
`CHANGELOG.md` has an entry. Releases only publish on a version change.
## Testing
<!--
Which nodes did you run, on which backend (CUDA / ROCm / Metal / XPU / CPU),
and with which model? Real-weight checks are opt-in:
python tests/manual_model_smoke.py --model "Qwen 3 VL 4B Instruct"
-->
+14
View File
@@ -0,0 +1,14 @@
version: 2
updates:
# Action versions only. Python dependency ranges are deliberately loose
# because ComfyUI owns torch, numpy, and Pillow in the shared environment.
- package-ecosystem: github-actions
directory: "/"
schedule:
interval: monthly
open-pull-requests-limit: 5
commit-message:
prefix: "ci"
groups:
actions:
patterns: ["*"]
+35 -1
View File
@@ -8,6 +8,22 @@ permissions:
contents: read
jobs:
lint:
name: Lint
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v7
- uses: actions/setup-python@v7
with:
python-version: "3.12"
cache: pip
cache-dependency-path: requirements-dev.txt
- name: Install lint tooling
run: python -m pip install -r requirements-dev.txt
- name: Ruff
run: python -m ruff check --output-format github .
test:
name: ${{ matrix.label }}
runs-on: ${{ matrix.os }}
@@ -20,18 +36,22 @@ jobs:
os: ubuntu-latest
python: "3.10"
cpu_index: true
coverage: false
- label: Linux / Python 3.13
os: ubuntu-latest
python: "3.13"
cpu_index: true
coverage: true
- label: Windows / Python 3.12
os: windows-latest
python: "3.12"
cpu_index: true
coverage: false
- label: macOS / Python 3.12
os: macos-14
python: "3.12"
cpu_index: false
coverage: false
steps:
- uses: actions/checkout@v7
- uses: actions/setup-python@v7
@@ -52,10 +72,24 @@ jobs:
- name: Install ComfyUI and node dependencies
run: |
git clone --depth 1 https://github.com/Comfy-Org/ComfyUI.git ../ComfyUI
python -m pip install pytest packaging build
python -m pip install -r requirements-dev.txt
python -m pip install -r ../ComfyUI/requirements.txt -r requirements.txt
- name: Test
if: matrix.coverage == false
run: python -m pytest -q
- name: Test with coverage
if: matrix.coverage == true
run: >-
python -m pytest -q
--cov=nodes --cov-report=term-missing:skip-covered
--cov-report=xml --cov-fail-under=70
- name: Upload coverage report
if: matrix.coverage == true && always()
uses: actions/upload-artifact@v4
with:
name: coverage-xml
path: coverage.xml
if-no-files-found: warn
- name: Compile
run: python -m compileall -q .
- name: Build distribution
+154
View File
@@ -0,0 +1,154 @@
# Changelog
All notable changes to this project are documented here.
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
Versions are published to the [Comfy Registry](https://registry.comfy.org/)
from `pyproject.toml`. A release is only published when `version` changes, so
every user-visible fix needs a version bump.
## [3.3.1] - 2026-07-30
### Fixed
- Package metadata declared `license = "MIT"` while the bundled `LICENSE` has
been Apache-2.0 since the initial commit. Built wheels therefore contained
contradictory MIT metadata and Apache-2.0 license text. The Registry already
referenced the license file and was unaffected. Metadata now says
`Apache-2.0`.
- Moondream 2 and Moondream 3.1 local inference (`b8ae298`).
- SmolVLM setup dependencies (`c13ee23`).
The two fixes above landed on `main` after 3.3.0 without a version bump, so
the Registry publish workflow saw 3.3.0 already published and skipped them.
They reach Registry users for the first time in 3.3.1.
### Added
- Test coverage for the GGUF text and multimodal node families, which
previously had none: `nodes/suggest.py` (0% to 98%) and
`nodes/llavaloader.py` (0% to 99%). The new cases pin the behaviours behind
the pack's longest-running bug reports: widget ordering (#156), sampling
kwarg plumbing (#144), and handle teardown on both success and failure
(#137).
- `ruff` lint gate and a coverage floor in CI, plus `requirements-dev.txt`
for the tooling.
- `CHANGELOG.md`, `CONTRIBUTING.md`, issue and pull request templates, and a
Dependabot configuration.
- A complete node reference in the README covering all 78 registered nodes.
## [3.3.0] - 2026-07-29
### Added
- Moondream Photon support and universal VLM acceleration utilities, including
the image pixel-budget and performance-profile nodes (`102f166`).
## [3.2.0] - 2026-07-29
### Added
- Adaptive video intelligence with temporal reasoning, plus the text workflow
toolkit (join, template, clean, replace, split, JSON extract, inspect)
(`44fefcb`).
## [3.1.0] - 2026-07-29
### Changed
- Hosted LLM and VLM API nodes modernized and hardened, with provider profiles
for OpenAI, Google Gemini, Anthropic, xAI, DeepSeek, and others (`505b324`).
## [3.0.0] - 2026-07-29
### Added
- Unified vision stack: open-vocabulary detection (Grounding DINO, OWLv2,
OmDet), SAM2.1 and SAM3.1 segmentation, tracking, and creator mask tools,
with structured detection/segmentation schemas (`39fc116`).
### Changed
- **Breaking:** detection and segmentation nodes now emit structured data
types rather than loose strings. Workflows wiring these outputs into text
nodes need the new converter utilities.
## [2.3.0] - 2026-07-29
### Added
- Reliable streaming VLM text output (`239c904`).
## [2.2.0] - 2026-07-29
### Changed
- llama.cpp GGUF runtime modernized. `llama-cpp-agent` was removed in favour
of llama-cpp-python's native JSON Schema support, which resolves the
unstable wrapper API behind the `unexpected keyword argument 'temperature'`
crashes (#144).
## [2.1.0] - 2026-07-28
### Added
- Cross-platform runtime support across NVIDIA CUDA, AMD ROCm, Apple Metal,
Intel XPU, and CPU, without replacing ComfyUI's PyTorch (`4c200c4`).
## [2.0.1] - 2026-07-28
### Added
- Small VLM catalog and real-weight model validation evidence
(see `MODEL_VALIDATION.md`) (`460b27a`).
## [2.0.0] - 2026-07-28
### Changed
- **Breaking:** node pack modernized with an explicit GPU lifecycle. Models
now load lazily on first execution and register with ComfyUI's model manager
so they participate in smart VRAM offloading, which addresses models
remaining resident after generation (#137) (`b89f628`).
- **Breaking:** `forceInput` string hacks removed from node schemas. They
corrupted the widget index during serialization and shifted inputs on saved
workflows (#156). Use the native right-click "Convert to Input" instead.
- Import is now failure-isolated: a broken optional model cannot prevent
unrelated nodes from loading (#94, #145).
- `numpy` is no longer pinned. The old `numpy<2.0.0` pin crashed startup on
NumPy 2.x environments (#157).
- Model coverage moved to current releases, including Qwen 3 / 3.5 VL,
SmolVLM2, InternVL, Granite Vision, and Gemma 3 (#148, #151). The
unmaintained InternLM-XComposer2 nodes were dropped (#139).
### Removed
- **Breaking:** `llama-cpp-agent` dependency (see 2.2.0).
- **Breaking:** InternLM-XComposer2 nodes, which depended on an AutoGPTQ stack
that pinned incompatible PyTorch versions (#139).
## 1.0.0 - 1.0.6 (2024-05-20 to 2024-11-03)
Initial packaged releases, predating changelog tracking. This line covered
LLaVA GGUF loaders and samplers, Moondream, Kosmos-2, JoyTag, UForm,
MiniCPM-V, PaLI-Gemma, Florence-2, Molmo, Qwen2-VL, the LLM prompt and
suggestion generators, AudioLDM2, and ChatMusician. See the
[commit history](https://github.com/gokayfem/ComfyUI_VLM_nodes/commits/main)
for detail.
Tagging began at 3.3.0. Earlier versions link to the commit that declared
them, because retroactively tagging them would run current CI against code
that predates it.
[3.3.1]: https://github.com/gokayfem/ComfyUI_VLM_nodes/compare/v3.3.0...v3.3.1
[3.3.0]: https://github.com/gokayfem/ComfyUI_VLM_nodes/releases/tag/v3.3.0
[3.2.0]: https://github.com/gokayfem/ComfyUI_VLM_nodes/commit/44fefcb
[3.1.0]: https://github.com/gokayfem/ComfyUI_VLM_nodes/commit/505b324
[3.0.0]: https://github.com/gokayfem/ComfyUI_VLM_nodes/commit/39fc116
[2.3.0]: https://github.com/gokayfem/ComfyUI_VLM_nodes/commit/239c904
[2.2.0]: https://github.com/gokayfem/ComfyUI_VLM_nodes/commit/0da5070
[2.1.0]: https://github.com/gokayfem/ComfyUI_VLM_nodes/commit/4c200c4
[2.0.1]: https://github.com/gokayfem/ComfyUI_VLM_nodes/commit/460b27a
[2.0.0]: https://github.com/gokayfem/ComfyUI_VLM_nodes/commit/b89f628
+106
View File
@@ -0,0 +1,106 @@
# Contributing
Thanks for helping out. This pack runs inside other people's ComfyUI installs
on five accelerator backends, so a few rules exist to keep it from breaking
them.
## The rules that matter most
**Importing the pack must never download a model, install a package, compile
anything, or allocate VRAM.** Models load on first execution. This is enforced
by `tests/test_nodes.py`, which asserts the source contains no `pip install`,
no `subprocess.run`, and no direct `torch.cuda.empty_cache`.
**Never reorder or insert widgets in an existing node's `INPUT_TYPES`.** Comfy
serializes widget values by position, so a reordered schema silently rebinds
every saved workflow. Add new inputs to `optional` at the end. The widget order
of the long-lived nodes is pinned by tests; if a test fails because you moved a
widget, the test is right.
**Never use `forceInput`.** It corrupts the widget index during serialization.
Users get the same result from the native right-click "Convert to Input".
**An optional dependency must fail only the node that needs it.** Use
`require_module()` from `nodes/runtime.py`, which raises an actionable error at
execution time rather than at import time.
**Do not install or replace `torch`.** ComfyUI's own installer picks the CUDA,
ROCm, XPU, Metal, or CPU build. The same applies to `numpy` and `Pillow`.
## Setting up
```bash
cd ComfyUI/custom_nodes
git clone https://github.com/gokayfem/ComfyUI_VLM_nodes.git
cd ComfyUI_VLM_nodes
python -m pip install -r requirements.txt -r requirements-dev.txt
```
Use ComfyUI's Python. On ComfyUI Portable there is no `activate` script, so
call the interpreter directly:
```
..\..\python_embeded\python.exe -m pip install -r requirements.txt
```
## Running checks
```bash
PYTHONPATH=/path/to/custom_nodes:/path/to/ComfyUI python -m pytest -q
python -m ruff check .
```
`PYTHONPATH` needs the directory *containing* this checkout plus ComfyUI
itself, because the tests import `ComfyUI_VLM_nodes` as a package and the nodes
import ComfyUI's `folder_paths`.
CI additionally enforces a coverage floor on Linux/Python 3.13:
```bash
python -m pytest -q --cov=nodes --cov-fail-under=70
```
Real-weight tests are opt-in because they download multi-gigabyte checkpoints,
and are never run in CI:
```bash
python tests/manual_model_smoke.py --model "Qwen 3 VL 4B Instruct"
python tests/manual_specialized_smoke.py --backend florence-large
python tests/manual_llama_cpp_smoke.py --download
```
## Writing tests
Tests must pass without model weights, without a GPU, and without
`llama-cpp-python`. Stub the model boundary instead: see
`tests/test_suggest.py` and `tests/test_llavaloader.py` for the pattern of
faking `LlamaHandle` and `create_chat_completion` to assert what the node sends
to the backend.
`nodes/joytagger/` is vendored upstream code kept byte-compatible with its
source. It is excluded from lint; please don't reformat it.
## Adding a model
1. Prefer adding an entry to the catalog in `nodes/modern_vlm.py` over a new
node. Most current VLMs work through the shared `transformers` path.
2. If it needs a bespoke loader, follow `nodes/minicpm.py` as the smallest
complete example.
3. Register the module in the `node_list` in `__init__.py`.
4. Record what you actually ran in `MODEL_VALIDATION.md`. Catalog entries that
were never executed against real weights must be marked as such.
5. Add the node to the reference table in `README.md`.
## Releasing
The Comfy Registry publishes from `pyproject.toml`, and only when `version`
changes. A fix merged without a version bump never reaches Registry users. So:
- bump `version` in `pyproject.toml`,
- add a `CHANGELOG.md` entry,
- tag the merge commit `vX.Y.Z`.
## Commit messages
Short imperative subject, one logical change per commit. Reference the issue it
closes in the body.
+187
View File
@@ -345,6 +345,193 @@ Upload the named media to ComfyUI's input directory, adjust the filenames and
labels, then submit the JSON object as the `prompt` value to `/prompt`. These
are API graphs, not frontend workflow-export JSON.
## Node reference
All 78 registered nodes, grouped by their menu category. The **Node ID** is the
`class_type` written into workflow and API JSON — search for that string when
you need to find a node you saw on a canvas.
### Modern VLM
The main entry point for current vision-language models.
| Node | Node ID | Outputs |
| --- | --- | --- |
| Modern VLM (Qwen / SmolVLM2 / LFM / InternVL / Granite / Gemma) | `ModernVLM` | `STRING` |
| Moondream 2 | `Moondream2model` | `STRING` |
### Moondream 3
Moondream 3 / 3.1 in an isolated Photon runtime. Load once, then reuse the
`MOONDREAM31_MODEL` output across the task nodes.
| Node | Node ID | Outputs |
| --- | --- | --- |
| Moondream 3 / 3.1 Loader (Isolated Photon) | `Moondream31Loader` | `MOONDREAM31_MODEL`, `STRING` |
| Moondream 3 / 3.1 Caption | `Moondream31Caption` | `STRING`, `STRING` |
| Moondream 3 / 3.1 Query | `Moondream31Query` | `STRING`, `STRING`, `STRING` |
| Moondream 3 / 3.1 Detect (Image / Video) | `Moondream31Detect` | `VLM_DETECTIONS`, `STRING`, `IMAGE`, `MASK`, `BOUNDING_BOX`, `BOUNDING_BOXES`, `STRING` |
| Moondream 3 / 3.1 Point (Image / Video) | `Moondream31Point` | `VLM_POINTS`, `STRING`, `IMAGE`, `STRING` |
| Moondream 3 Preview SVG Segment (Image / Video) | `Moondream31Segment` | `VLM_DETECTIONS`, `STRING`, `STRING`, `MASK`, `IMAGE`, `IMAGE`, `IMAGE`, `BOUNDING_BOX`, `BOUNDING_BOXES`, `STRING` |
### Florence-2
| Node | Node ID | Outputs |
| --- | --- | --- |
| Florence-2 Multitask Vision | `Florence2` | `STRING`, `STRING`, `MASK`, `IMAGE` |
### Vision: detection, segmentation, tracking
Open-vocabulary detection and video segmentation. These emit the structured
`VLM_DETECTIONS` / `VLM_POINTS` / `VLM_TRACKS` types rather than loose strings.
| Node | Node ID | Outputs |
| --- | --- | --- |
| VLM Open-Vocabulary Detection | `VLMOpenVocabularyDetection` | `VLM_DETECTIONS`, `STRING`, `IMAGE`, `MASK`, `BOUNDING_BOX`, `BOUNDING_BOXES` |
| VLM SAM2.1 Video Segmentation | `VLMSAM2VideoSegmentation` | `VLM_TRACKS`, `STRING`, `MASK`, `MASK`, `IMAGE` |
| VLM SAM3 Track Adapter | `VLMSAM3TrackAdapter` | `VLM_TRACKS`, `SAM3_TRACK_DATA` |
| VLM Track Detections | `VLMTrackDetections` | `VLM_TRACKS` |
| VLM Track Report | `VLMTrackReport` | `STRING`, `STRING` |
| JoyTag | `Joytag` | `STRING` |
### Vision: spatial reasoning
| Node | Node ID | Outputs |
| --- | --- | --- |
| VLM Spatial Prompt Builder | `VLMSpatialPromptBuilder` | `STRING` |
| VLM Structured Spatial Parser | `VLMStructuredSpatialParser` | `VLM_DETECTIONS`, `VLM_POINTS`, `STRING` |
### Vision: detection utilities
Converters and filters between structured detections and ordinary Comfy types.
| Node | Node ID | Outputs |
| --- | --- | --- |
| Filter VLM Detections | `VLMFilterDetections` | `VLM_DETECTIONS` |
| Select VLM Detection | `VLMSelectDetection` | `VLM_DETECTIONS` |
| Crop VLM Detections | `VLMCropDetections` | `IMAGE`, `STRING` |
| Render VLM Detections | `VLMRenderDetections` | `IMAGE` |
| VLM Detection Centers | `VLMDetectionsToPoints` | `VLM_POINTS`, `STRING` |
| VLM Detections from JSON | `VLMDetectionsFromJSON` | `VLM_DETECTIONS` |
| VLM Detections to JSON | `VLMDetectionsToJSON` | `STRING` |
| VLM Detections to Bounding Boxes | `VLMDetectionsToBoundingBoxes` | `BOUNDING_BOXES`, `STRING` |
| VLM Detections to Masks | `VLMDetectionsToMasks` | `MASK`, `MASK`, `STRING`, `MASK`, `IMAGE`, `IMAGE`, `IMAGE` |
### Vision: mask tools
| Node | Node ID | Outputs |
| --- | --- | --- |
| VLM Mask Processor | `VLMMaskProcessor` | `MASK`, `MASK`, `MASK`, `IMAGE` |
| VLM Mask Composite | `VLMMaskComposite` | `IMAGE`, `IMAGE`, `IMAGE`, `IMAGE` |
### Video intelligence
Adaptive frame selection and temporal reasoning for long videos.
| Node | Node ID | Outputs |
| --- | --- | --- |
| VLM Adaptive Frame Sampler | `VLMAdaptiveFrameSampler` | `IMAGE`, `VLM_VIDEO_SELECTION`, `STRING`, `STRING` |
| VLM Video Reasoning Prompt | `VLMVideoReasoningPrompt` | `STRING`, `STRING` |
| VLM Video Temporal Reasoner | `VLMVideoTemporalReasoner` | `STRING`, `VLM_EVENTS`, `VLM_VIDEO_SELECTION`, `IMAGE`, `STRING`, `STRING`, `STRING`, `STRING` |
| VLM Temporal Events From JSON | `VLMEventsFromVideoJSON` | `VLM_EVENTS`, `STRING`, `STRING` |
| VLM Persistent Scene State | `VLMBuildSceneState` | `VLM_SCENE_STATE`, `STRING`, `STRING` |
| VLM Track-Aware Semantic Crops | `VLMTrackAwareCrops` | `IMAGE`, `STRING` |
### LLM (local GGUF)
llama.cpp text models. `LLM Loader (GGUF)` produces the `CUSTOM` model handle
the samplers consume; the *Managed Cache* variants own their own handle and can
release it after each run.
| Node | Node ID | Outputs |
| --- | --- | --- |
| LLM Loader (GGUF) | `LLMLoader` | `CUSTOM` |
| LLM Sampler | `LLMSampler` | `STRING` |
| LLM Prompt Generator | `LLMPromptGenerator` | `STRING` |
| LLM (Managed Cache) | `LLMOptionalMemoryFreeSimple` | `STRING` |
| LLM (Managed Cache, Advanced) | `LLMOptionalMemoryFreeAdvanced` | `STRING` |
| Structured Output | `StructuredOutput` | `STRING` |
| Structured Keyword Extraction | `KeywordExtraction` | `STRING` |
| Structured Prompt Generator | `LLavaPromptGenerator` | `STRING` |
| Creative Art Prompt Generator | `CreativeArtPromptGenerator` | `STRING` |
| Prompt Suggester | `Suggester` | `STRING` |
### LLaVA (local GGUF multimodal)
Vision models through llama.cpp. These need both a GGUF and its vision
projector (mmproj).
| Node | Node ID | Outputs |
| --- | --- | --- |
| LLaVA Loader | `LLava Loader Simple` | `CUSTOM` |
| LLaVA Vision Projector Loader | `LlavaClipLoader` | `CUSTOM` |
| LLaVA Sampler | `LLavaSamplerSimple` | `STRING` |
| LLaVA Sampler (Advanced) | `LLavaSamplerAdvanced` | `STRING` |
| LLaVA (Managed Cache) | `LLavaOptionalMemoryFreeSimple` | `STRING` |
| LLaVA (Managed Cache, Advanced) | `LLavaOptionalMemoryFreeAdvanced` | `STRING` |
### Hosted APIs
| Node | Node ID | Outputs |
| --- | --- | --- |
| Hosted VLM API (Secure) | `HostedVLMAPI` | `STRING`, `STRING`, `INT` |
| Hosted LLM API (Secure) | `PromptGenerateAPI` | `STRING` |
### Text toolkit
Dependency-free string handling, so a VLM response can be shaped without an
extra node pack.
| Node | Node ID | Outputs |
| --- | --- | --- |
| Text | `SimpleText` | `STRING`, `INT`, `INT`, `INT` |
| Text Join | `VLMTextJoin` | `STRING`, `STRING`, `INT` |
| Text Template | `VLMTextTemplate` | `STRING`, `STRING`, `STRING` |
| Text Clean | `VLMTextClean` | `STRING`, `STRING` |
| Text Replace | `VLMTextReplace` | `STRING`, `INT`, `STRING` |
| Text Split / Batch | `VLMTextSplit` | `STRING`, `STRING`, `INT` |
| Text Inspector | `VLMTextInspect` | `STRING`, `INT`, `INT`, `INT`, `INT`, `INT`, `STRING`, `STRING` |
| View Text (Streaming) | `ViewText` | `STRING`, `INT`, `INT`, `INT`, `STRING` |
| JSON Extract | `VLMJSONExtract` | `STRING`, `BOOLEAN`, `STRING`, `STRING` |
| JSON to Text | `JsonToText` | `STRING`, `STRING`, `INT` |
### Performance and diagnostics
Run **VLM Runtime Diagnostics** before reporting a bug — it reports your
device, backend, and which optional packages are installed.
| Node | Node ID | Outputs |
| --- | --- | --- |
| VLM Runtime Diagnostics | `VLMRuntimeDiagnostics` | `STRING` |
| VLM Performance Profile | `VLMPerformanceProfile` | `INT`, `FLOAT`, `INT`, `INT`, `BOOLEAN`, `STRING` |
| VLM Image Pixel Budget | `VLMImagePixelBudget` | `IMAGE`, `INT`, `INT`, `STRING` |
### Audio
| Node | Node ID | Outputs |
| --- | --- | --- |
| AudioLDM2 | `AudioLDM2Node` | `*`, `INT`, `AUDIO` |
| Chat Musician | `ChatMusician` | `STRING`, `*`, `INT`, `AUDIO` |
| PlayMusic Node | `PlayMusic` | `*` |
| Save Audio | `SaveAudioNode` | — |
### Legacy model loaders
Kept for existing workflows. New graphs should prefer **Modern VLM**, which
covers most of these architectures through one interface.
| Node | Node ID | Outputs |
| --- | --- | --- |
| Qwen2-VL | `Qwen2VLNode` | `STRING` |
| MiniCPM-V 2.6 (GGUF) | `MiniCPMNode` | `STRING` |
| Molmo Vision-Language Model | `MolmoNode` | `STRING` |
| PaLI-Gemma (Official Segmentation) | `Paligemma` | `STRING`, `MASK`, `IMAGE` |
| Kosmos-2 | `Kosmos2model` | `STRING` |
| MC-LLaVA | `MCLLaVAModel` | `STRING` |
| UForm Gen2 Qwen | `UformGen2QwenNode` | `STRING` |
| MoonDream (Moondream 2) | `MoonDream` | `STRING` |
| [Legacy] Modern VLM Compatibility | `LegacyModernVLM` | `STRING` |
## Install
Install through ComfyUI Manager, or clone into `ComfyUI/custom_nodes` and run:
-1
View File
@@ -15,7 +15,6 @@ from typing import Any
import torch
import torch.nn.functional as functional
RESIZE_QUALITY = (
"Fast (area)",
"Quality (bicubic)",
+1 -2
View File
@@ -4,11 +4,10 @@ from __future__ import annotations
from pathlib import Path
import folder_paths
import numpy as np
import torch
import folder_paths
from .runtime import (
CachedModelNode,
execution_device,
+1 -1
View File
@@ -5,8 +5,8 @@ from __future__ import annotations
import colorsys
import hashlib
import math
from collections.abc import Iterable, Mapping
from dataclasses import dataclass
from typing import Iterable, Mapping
import numpy as np
import torch
+3 -2
View File
@@ -9,13 +9,14 @@ Custom endpoints can read only ``CUSTOM_API_KEY``.
from __future__ import annotations
import ipaddress
import io
import ipaddress
import json
import os
import re
from collections.abc import Callable
from dataclasses import dataclass, replace
from typing import Any, Callable
from typing import Any
from urllib.parse import quote, quote_plus, urlsplit, urlunsplit
import torch
+3 -2
View File
@@ -8,8 +8,9 @@ small and large VLM families while keeping downloads and VRAM allocation lazy.
from __future__ import annotations
import threading
from collections.abc import Callable
from dataclasses import dataclass
from typing import Any, Callable
from typing import Any
import torch
@@ -25,8 +26,8 @@ from .runtime import (
model_device,
move_inputs,
normalize_hf_model_id,
require_quantization_backend,
require_module,
require_quantization_backend,
reserve_external_vram,
snapshot_download,
tensor_batch_to_pil,
+1 -1
View File
@@ -14,8 +14,8 @@ from .runtime import (
external_device_map,
inference_context,
model_device,
require_quantization_backend,
require_module,
require_quantization_backend,
reserve_external_vram,
snapshot_download,
tensor_batch_to_pil,
+1 -2
View File
@@ -24,15 +24,14 @@ from .runtime import (
normalize_hf_model_id,
pil_mask_to_tensor,
pil_to_tensor,
require_quantization_backend,
require_module,
require_quantization_backend,
reserve_external_vram,
snapshot_download,
tensor_batch_to_pil,
torch_dtype,
)
PALIGEMMA_MODELS = [
"gokaygokay/sd3-long-captioner-v2",
"google/paligemma-3b-ft-refcoco-seg-896",
+1 -1
View File
@@ -51,4 +51,4 @@ Optional: If asked to create a random prompt create one.
# Define the system message
system_msg_simple = """
You are an helpful asistant. Answer optional questions or help the user for their optional queries.
"""
"""
+1 -2
View File
@@ -17,15 +17,14 @@ from .runtime import (
inference_context,
model_device,
move_inputs,
require_quantization_backend,
require_module,
require_quantization_backend,
reserve_external_vram,
snapshot_download,
tensor_batch_to_pil,
torch_dtype,
)
QWEN2_VL_MODELS = {
"Qwen2-VL-2B": "Qwen/Qwen2-VL-2B-Instruct",
"Qwen2-VL-7B": "Qwen/Qwen2-VL-7B-Instruct",
+11 -3
View File
@@ -18,11 +18,12 @@ import os
import platform
import re
import threading
from collections.abc import Callable, Iterable, Mapping
from contextlib import nullcontext
from dataclasses import dataclass
from importlib import metadata
from pathlib import Path
from typing import Any, Callable, Iterable, Mapping
from typing import Any
import folder_paths
import numpy as np
@@ -680,7 +681,10 @@ def llama_runtime_input_types() -> dict[str, tuple[Any, ...]]:
"min": 1,
"max": 8192,
"step": 1,
"tooltip": "Logical prompt batch. Lower this if context loading runs out of memory.",
"tooltip": (
"Logical prompt batch. Lower this if context loading runs "
"out of memory."
),
},
),
"n_ubatch": (
@@ -697,7 +701,11 @@ def llama_runtime_input_types() -> dict[str, tuple[Any, ...]]:
list(LLAMA_FLASH_ATTENTION_CHOICES),
{
"default": "Auto",
"tooltip": "Auto enables llama.cpp flash attention only with accelerator offload and safely retries without it when unsupported.",
"tooltip": (
"Auto enables llama.cpp flash attention only with "
"accelerator offload and safely retries without it when "
"unsupported."
),
},
),
"use_mmap": (
+1 -1
View File
@@ -258,7 +258,7 @@ class Sam2VideoPredictor:
):
raise ValueError("seed_mask must have shape [objects, height, width].")
object_ids = list(range(1, masks_for_seed.shape[0] + 1))
labels = {object_id: None for object_id in object_ids}
labels = dict.fromkeys(object_ids)
if not object_ids:
raise ValueError(
"Connect detections, a BOUNDING_BOX, or at least one seed mask."
-1
View File
@@ -14,7 +14,6 @@ import re
import unicodedata
from typing import Any
TEXT_CATEGORY = "VLM Nodes/Text"
CREATE_CATEGORY = f"{TEXT_CATEGORY}/Create"
TRANSFORM_CATEGORY = f"{TEXT_CATEGORY}/Transform"
+2 -2
View File
@@ -9,7 +9,7 @@ from __future__ import annotations
import json
import re
from typing import Any, Literal, Optional
from typing import Any
import folder_paths
import torch
@@ -68,7 +68,7 @@ class ArtisticTechniques(BaseModel):
class ImageryTheme(BaseModel):
core_subject: str
additional_elements: Optional[list[str]] = None
additional_elements: list[str] | None = None
class VisualStyle(BaseModel):
+1 -1
View File
@@ -8,8 +8,8 @@ keeps the baseline portable across CUDA, ROCm, MPS, XPU, and CPU systems.
from __future__ import annotations
import math
from collections.abc import Iterable
from dataclasses import dataclass, field
from typing import Iterable
import numpy as np
from scipy.optimize import linear_sum_assignment
+2 -1
View File
@@ -4,8 +4,9 @@ from __future__ import annotations
import json
import math
from collections.abc import Iterable
from dataclasses import replace
from typing import Any, Iterable
from typing import Any
import numpy as np
import torch
+36 -2
View File
@@ -1,10 +1,10 @@
[project]
name = "comfyui_vlm_nodes"
version = "3.3.0"
version = "3.3.1"
description = "Production-ready local and API vision-language nodes for ComfyUI"
readme = "README.md"
requires-python = ">=3.10"
license = "MIT"
license = "Apache-2.0"
license-files = ["LICENSE"]
dependencies = [
"accelerate>=1.1,<2",
@@ -52,6 +52,40 @@ gguf = [
Repository = "https://github.com/gokayfem/ComfyUI_VLM_nodes"
Issues = "https://github.com/gokayfem/ComfyUI_VLM_nodes/issues"
[tool.pytest.ini_options]
testpaths = ["tests"]
# manual_*.py download multi-gigabyte checkpoints and are run by hand.
python_files = ["test_*.py"]
addopts = "-ra --strict-markers --strict-config"
filterwarnings = ["default"]
[tool.ruff]
line-length = 100
# nodes/joytagger is vendored upstream code kept byte-compatible with its
# source, including tab indentation. Reformatting it would break that.
extend-exclude = ["nodes/joytagger"]
[tool.ruff.lint]
select = ["E", "F", "W", "I", "UP", "C4", "B", "SIM"]
ignore = [
# zip(strict=) changes behaviour when lengths differ; enabling it needs a
# per-call-site audit rather than a blanket flag.
"B905",
# The streaming closures in modern_vlm are started and joined inside the
# same loop iteration, so the loop variable cannot change under them.
"B023",
# try/except/pass around optional backends stays readable as-is;
# contextlib.suppress would hide which dependency is being probed.
"SIM105",
]
[tool.ruff.lint.per-file-ignores]
# Prompt templates are data. Rewrapping them changes the model input.
"nodes/prompts.py" = ["E501"]
# The manual smoke scripts must bootstrap the package onto sys.path before
# they can import from it.
"tests/manual_*.py" = ["E402"]
[tool.comfy]
PublisherId = "gokayfem"
DisplayName = "ComfyUI VLM Nodes"
+8
View File
@@ -0,0 +1,8 @@
# Development and CI tooling. Not needed to run the nodes in ComfyUI.
# Install with ComfyUI's Python alongside requirements.txt:
# python -m pip install -r requirements.txt -r requirements-dev.txt
build>=1.2,<2
packaging>=24
pytest>=8,<9
pytest-cov>=5,<8
ruff>=0.14,<1
+50
View File
@@ -0,0 +1,50 @@
"""Make this checkout importable from the manual smoke scripts.
The manual scripts run as `python tests/manual_*.py`, outside pytest, so they
do not get `conftest.py`. Without this they only import when the checkout
directory happens to be named `ComfyUI_VLM_nodes`, which is true in a normal
ComfyUI install but not in a git worktree named after a feature branch.
Usage, before importing anything from the package:
from _bootstrap import bootstrap
bootstrap()
"""
from __future__ import annotations
import importlib.util
import sys
from pathlib import Path
PACKAGE = "ComfyUI_VLM_nodes"
REPOSITORY = Path(__file__).resolve().parents[1]
def bootstrap() -> None:
"""Put the repository and ComfyUI on sys.path, then load this checkout."""
for candidate in (
REPOSITORY.parent,
REPOSITORY.parent / "ComfyUI",
REPOSITORY.parents[1],
):
if candidate.exists():
sys.path.insert(0, str(candidate))
if PACKAGE in sys.modules or REPOSITORY.name == PACKAGE:
return
# Load this checkout explicitly so the script can never pass by silently
# importing a sibling clone with the canonical directory name.
specification = importlib.util.spec_from_file_location(
PACKAGE,
REPOSITORY / "__init__.py",
submodule_search_locations=[str(REPOSITORY)],
)
if specification is None or specification.loader is None:
raise RuntimeError(f"Could not load package from {REPOSITORY}.")
package = importlib.util.module_from_spec(specification)
sys.modules[PACKAGE] = package
specification.loader.exec_module(package)
+5 -2
View File
@@ -11,9 +11,12 @@ from __future__ import annotations
import json
from transformers import AutoConfig, AutoProcessor
from _bootstrap import bootstrap
from ComfyUI_VLM_nodes.nodes.modern_vlm import MODEL_CATALOG
bootstrap()
from ComfyUI_VLM_nodes.nodes.modern_vlm import MODEL_CATALOG # noqa: E402
from transformers import AutoConfig, AutoProcessor # noqa: E402
def main() -> int:
+5 -1
View File
@@ -11,7 +11,11 @@ import json
import time
from pathlib import Path
from ComfyUI_VLM_nodes.nodes.runtime import (
from _bootstrap import bootstrap
bootstrap()
from ComfyUI_VLM_nodes.nodes.runtime import ( # noqa: E402
LlamaHandle,
default_llama_threads,
hf_download,
+178
View File
@@ -0,0 +1,178 @@
"""Opt-in real-weight smoke test for the GGUF *node classes*.
`manual_llama_cpp_smoke.py` proves the shared `LlamaHandle` runtime loads and
generates. This script goes one level up and drives the actual ComfyUI node
classes end to end against real weights, which covers the parts the offline
suite deliberately stubs:
* `LLMLoader` resolving a real file through ComfyUI's `folder_paths`
* `LLMSampler` producing real text from real sampling arguments
* `StructuredOutput` constraining a real model to a generated JSON Schema —
the llama.cpp grammar path, which cannot be verified with a stub
* `LLMOptionalMemoryFreeSimple` releasing a real llama.cpp allocation
Never run in CI: it downloads weights and needs `llama-cpp-python`.
Example:
python tests/manual_llm_node_smoke.py --download
python tests/manual_llm_node_smoke.py --model /models/qwen.gguf
"""
from __future__ import annotations
import argparse
import json
import shutil
import time
from pathlib import Path
from _bootstrap import bootstrap
bootstrap()
import folder_paths # noqa: E402
from ComfyUI_VLM_nodes.nodes.runtime import hf_download, model_root # noqa: E402
from ComfyUI_VLM_nodes.nodes.suggest import ( # noqa: E402
LLMLoader,
LLMOptionalMemoryFreeSimple,
LLMSampler,
StructuredOutput,
)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("--model", type=Path)
parser.add_argument("--download", action="store_true")
parser.add_argument("--repo", default="ggml-org/Qwen3.5-0.8B-GGUF")
parser.add_argument("--filename", default="Qwen3.5-0.8B-Q4_0.gguf")
parser.add_argument("--n-gpu-layers", type=int, default=-1)
return parser.parse_args()
def stage_model(args: argparse.Namespace) -> str:
"""Put the GGUF where ComfyUI's folder_paths can enumerate it."""
if args.model is None:
if not args.download:
raise SystemExit(
"Pass --model /path/to/model.gguf, or allow the small default "
"download with --download."
)
source = hf_download(args.repo, args.filename, "llm-node-smoke")
else:
source = args.model.resolve()
if not source.is_file():
raise SystemExit(f"{source} is not a file.")
destination = model_root() / source.name
if not destination.exists():
destination.parent.mkdir(parents=True, exist_ok=True)
shutil.copy2(source, destination)
# The loader nodes offer whatever folder_paths enumerates, so the staged
# file has to actually show up there. Staging happens before the first
# get_filename_list call in this process, so there is no cache to clear.
listed = folder_paths.get_filename_list("LLavacheckpoints")
if source.name not in listed:
raise SystemExit(
f"{source.name} is not enumerated in LLavacheckpoints: {listed}"
)
return source.name
def main() -> None:
args = parse_args()
checkpoint = stage_model(args)
results: dict[str, object] = {"checkpoint": checkpoint}
# 1. The loader must hand back a lazy handle that has not loaded yet.
started = time.perf_counter()
(model,) = LLMLoader().load_llm_checkpoint(
ckpt_name=checkpoint,
max_ctx=2048,
gpu_layers=args.n_gpu_layers,
n_threads=4,
)
results["loader_returned_without_loading"] = model._llm is None
results["loader_seconds"] = round(time.perf_counter() - started, 3)
# 2. Real generation through the real sampler node.
#
# Deliberately no assertion on what the model *says*: at 0.8B/Q4 the answer
# is often factually wrong, and that is model quality, not node
# correctness. What the node owns is that generation happens and that its
# sampling arguments actually reach llama.cpp — so assert determinism for a
# fixed seed at temperature 0 instead.
def sample(seed: int) -> tuple[str, float]:
started = time.perf_counter()
(text,) = LLMSampler().generate_text_advanced(
system_msg="You answer with a single short sentence.",
prompt="Name the largest planet in the solar system.",
model=model,
max_tokens=48,
temperature=0.0,
top_p=0.95,
top_k=40,
frequency_penalty=0.0,
presence_penalty=0.0,
repeat_penalty=1.1,
seed=seed,
)
return text, round(time.perf_counter() - started, 3)
text, elapsed = sample(42)
repeat, _ = sample(42)
results["sampler_seconds"] = elapsed
results["sampler_text"] = text
results["sampler_produced_text"] = bool(text.strip())
results["sampler_deterministic_for_fixed_seed"] = text == repeat
# 3. The grammar-constrained path. A stub cannot prove this works.
started = time.perf_counter()
(value,) = StructuredOutput().keyword_extract(
prompt="The photograph shows a calm, empty beach at sunrise.",
model=model,
temperature=0.0,
attribute_name="mood",
attribute_type="Category",
attribute_description="The overall mood of the described scene.",
categories="calm, tense, joyful, melancholy",
)
results["structured_seconds"] = round(time.perf_counter() - started, 3)
results["structured_value"] = value
# The whole point of the schema is that the model cannot answer off-menu.
results["structured_respected_enum"] = value in {
"calm",
"tense",
"joyful",
"melancholy",
}
model.close()
# 4. A managed-cache node must really release its allocation.
node = LLMOptionalMemoryFreeSimple()
(cached_text,) = node.generate_text(
ckpt_name=checkpoint,
max_ctx=2048,
gpu_layers=args.n_gpu_layers,
n_threads=4,
prompt="Say the word: ready",
temperature=0.0,
unload=True,
)
results["managed_cache_text"] = cached_text
results["managed_cache_released"] = node._handle is None and node._key is None
checks = {
key: value for key, value in results.items() if isinstance(value, bool)
}
results["ALL_CHECKS_PASSED"] = all(checks.values())
print(json.dumps(results, ensure_ascii=False, indent=2))
if not results["ALL_CHECKS_PASSED"]:
failed = [key for key, value in checks.items() if not value]
raise SystemExit(f"Failed checks: {failed}")
if __name__ == "__main__":
main()
+7 -1
View File
@@ -14,8 +14,14 @@ import json
import time
import torch
from _bootstrap import bootstrap
from ComfyUI_VLM_nodes.nodes.modern_vlm import MODEL_CATALOG, ModernVLMPredictor
bootstrap()
from ComfyUI_VLM_nodes.nodes.modern_vlm import ( # noqa: E402
MODEL_CATALOG,
ModernVLMPredictor,
)
def test_image() -> torch.Tensor:
+2
View File
@@ -12,7 +12,9 @@ import json
import time
import torch
from _bootstrap import bootstrap
bootstrap()
BACKENDS = (
"florence-base",
+110
View File
@@ -0,0 +1,110 @@
"""Keep the shipped documentation honest about what the pack actually contains.
The README node reference drifted to 47 undocumented nodes before these checks
existed, and the packaged license metadata disagreed with the LICENSE file.
Both are cheap to assert and expensive to notice by hand.
"""
from __future__ import annotations
import re
from pathlib import Path
import ComfyUI_VLM_nodes as package
REPOSITORY = Path(package.__file__).parent
def read(name: str) -> str:
return (REPOSITORY / name).read_text(encoding="utf-8")
def project_field(name: str) -> str:
"""Read a top-level [project] string field.
Deliberately regex-based rather than tomllib: this suite also runs on
Python 3.10, which has no tomllib in the standard library.
"""
match = re.search(rf'^{name}\s*=\s*"([^"]+)"', read("pyproject.toml"), re.M)
assert match is not None, f"pyproject.toml has no {name} field."
return match.group(1)
def test_every_registered_node_appears_in_the_readme():
readme = read("README.md")
documented = set(re.findall(r"`([^`]+)`", readme))
missing = sorted(set(package.NODE_CLASS_MAPPINGS) - documented)
assert not missing, (
"These nodes are registered but never named in README.md. "
f"Add them to the node reference: {missing}"
)
def test_node_reference_matches_registered_output_types():
row_pattern = re.compile(
r"^\|[^|]+\|\s*`(?P<node_id>[^`]+)`\s*\|(?P<outputs>[^|]*)\|$",
re.M,
)
documented = {
match.group("node_id"): tuple(
re.findall(r"`([^`]+)`", match.group("outputs"))
)
for match in row_pattern.finditer(read("README.md"))
}
mismatches = {}
for node_id, node_class in package.NODE_CLASS_MAPPINGS.items():
expected = tuple(
"*" if output is any else str(output)
for output in node_class.RETURN_TYPES
)
if documented.get(node_id) != expected:
mismatches[node_id] = {
"documented": documented.get(node_id),
"registered": expected,
}
assert not mismatches, (
"README.md output schemas do not match the registered RETURN_TYPES: "
f"{mismatches}"
)
def test_declared_license_matches_the_license_file():
declared = project_field("license")
license_text = read("LICENSE")
if "Apache License" in license_text:
expected = "Apache-2.0"
elif "MIT License" in license_text:
expected = "MIT"
else:
raise AssertionError("Could not identify the license in LICENSE.")
assert declared == expected, (
f"pyproject.toml declares {declared!r} but LICENSE is {expected}. "
"This metadata is embedded in built distribution artifacts."
)
def test_changelog_documents_the_current_version():
version = project_field("version")
changelog = read("CHANGELOG.md")
assert f"[{version}]" in changelog, (
f"pyproject version {version} has no CHANGELOG.md entry. The Comfy "
"Registry only publishes on a version change, so every release needs "
"one."
)
def test_contributor_and_security_docs_are_present():
for name in ("CONTRIBUTING.md", "SECURITY.md", "CHANGELOG.md", "LICENSE"):
assert (REPOSITORY / name).is_file(), f"{name} is missing."
def test_issue_templates_are_valid_and_request_diagnostics():
template_dir = REPOSITORY / ".github" / "ISSUE_TEMPLATE"
bug_report = (template_dir / "bug_report.yml").read_text(encoding="utf-8")
# Environment detail is what the historically unresolvable reports lacked.
assert "VLMRuntimeDiagnostics" in bug_report or "Diagnostics" in bug_report
assert "Node pack version" in bug_report
-1
View File
@@ -6,7 +6,6 @@ from types import SimpleNamespace
import pytest
import torch
from ComfyUI_VLM_nodes.nodes import hosted_api
+496
View File
@@ -0,0 +1,496 @@
"""Contract tests for the llama.cpp multimodal nodes in ``nodes/llavaloader.py``.
Covers batch handling, the vision message envelope, projector wiring, and the
cached-handle lifecycle. No llama.cpp wheel, mmproj, or GGUF weights required.
"""
from __future__ import annotations
import base64
from pathlib import Path
import pytest
import torch
from ComfyUI_VLM_nodes.nodes import llavaloader
from ComfyUI_VLM_nodes.nodes.runtime import LlamaHandle, LlavaClipConfig
MODEL_FILE = "llava.gguf"
CLIP_FILE = "mmproj.gguf"
class FakeLlama:
def __init__(self, contents: list[str] | None = None):
self.contents = contents or ["a description"]
self.calls: list[dict] = []
def create_chat_completion(self, **kwargs):
index = min(len(self.calls), len(self.contents) - 1)
self.calls.append(kwargs)
return {"choices": [{"message": {"content": self.contents[index]}}]}
class FakeHandle:
instances: list[FakeHandle] = []
def __init__(self, model_path, **kwargs):
self.model_path = model_path
self.kwargs = kwargs
self.closed = False
self.llama = FakeLlama()
FakeHandle.instances.append(self)
def ensure_loaded(self):
return self.llama
def close(self):
self.closed = True
@pytest.fixture
def resolved_paths(monkeypatch):
root = Path("/models/LLavacheckpoints")
monkeypatch.setattr(llavaloader, "resolve_model_path", lambda name: root / name)
return root
@pytest.fixture
def fake_handles(monkeypatch):
FakeHandle.instances = []
monkeypatch.setattr(llavaloader, "LlamaHandle", FakeHandle)
return FakeHandle
def image_batch(count: int = 1, size: int = 4) -> torch.Tensor:
"""A ComfyUI BHWC float image batch."""
return torch.rand(count, size, size, 3)
# --------------------------------------------------------------------------
# Widget ordering (see issue #156).
# --------------------------------------------------------------------------
def test_llava_sampler_simple_widget_order_is_frozen():
assert list(llavaloader.LLavaSamplerSimple.INPUT_TYPES()["required"]) == [
"image",
"prompt",
"model",
"temperature",
]
def test_llava_sampler_advanced_widget_order_is_frozen():
assert list(llavaloader.LLavaSamplerAdvanced.INPUT_TYPES()["required"]) == [
"image",
"system_msg",
"prompt",
"model",
"max_tokens",
"temperature",
"top_p",
"top_k",
"frequency_penalty",
"presence_penalty",
"repeat_penalty",
"seed",
]
def test_llava_loader_widget_order_is_frozen():
schema = llavaloader.LLavaLoader.INPUT_TYPES()
assert list(schema["required"]) == [
"ckpt_name",
"max_ctx",
"gpu_layers",
"n_threads",
"clip",
]
def test_optional_memory_free_simple_widget_order_is_frozen():
schema = llavaloader.LLavaOptionalMemoryFreeSimple.INPUT_TYPES()
assert list(schema["required"]) == [
"ckpt_name",
"clip_name",
"max_ctx",
"gpu_layers",
"n_threads",
"image",
"prompt",
"temperature",
"unload",
]
assert list(schema["optional"])[0] == "handler"
def test_every_llava_node_declares_a_callable_function_and_return_types():
for name, node_class in llavaloader.NODE_CLASS_MAPPINGS.items():
assert isinstance(node_class.RETURN_TYPES, tuple), name
assert node_class.RETURN_TYPES, name
assert callable(getattr(node_class, node_class.FUNCTION, None)), name
assert node_class.CATEGORY.startswith("VLM Nodes"), name
def test_display_names_cover_every_registered_node():
assert set(llavaloader.NODE_CLASS_MAPPINGS) == set(
llavaloader.NODE_DISPLAY_NAME_MAPPINGS
)
# --------------------------------------------------------------------------
# Vision message envelope.
# --------------------------------------------------------------------------
def test_vision_messages_place_the_image_before_the_text():
messages = llavaloader._vision_messages("sys", "what is this?", "data:image/png;b")
assert messages[0] == {"role": "system", "content": "sys"}
content = messages[1]["content"]
assert messages[1]["role"] == "user"
# llama.cpp vision handlers require the image part first.
assert content[0]["type"] == "image_url"
assert content[0]["image_url"]["url"] == "data:image/png;b"
assert content[1] == {"type": "text", "text": "what is this?"}
def test_run_batch_sends_a_png_data_uri_per_image():
llama = FakeLlama()
llavaloader._run_batch(
image_batch(1), llama, system_msg="sys", prompt="p", temperature=0.1
)
(call,) = llama.calls
url = call["messages"][1]["content"][0]["image_url"]["url"]
assert url.startswith("data:image/png;base64,")
# The payload must be real decodable PNG bytes.
decoded = base64.b64decode(url.split(",", 1)[1])
assert decoded.startswith(b"\x89PNG\r\n\x1a\n")
def test_run_batch_calls_the_model_once_per_batch_item():
llama = FakeLlama(["first", "second", "third"])
text = llavaloader._run_batch(
image_batch(3), llama, system_msg="sys", prompt="p", temperature=0.1
)
assert len(llama.calls) == 3
# Every batch item must survive into the response.
assert "first" in text
assert "second" in text
assert "third" in text
assert "--- Image 1 ---" in text
assert "--- Image 3 ---" in text
def test_run_batch_returns_bare_text_for_a_single_image():
llama = FakeLlama(["only one"])
text = llavaloader._run_batch(
image_batch(1), llama, system_msg="sys", prompt="p", temperature=0.1
)
assert text == "only one"
def test_run_batch_forwards_generation_kwargs_unchanged():
llama = FakeLlama()
llavaloader._run_batch(
image_batch(1),
llama,
system_msg="sys",
prompt="p",
max_tokens=32,
temperature=0.3,
top_p=0.7,
top_k=10,
seed=99,
)
(call,) = llama.calls
assert call["max_tokens"] == 32
assert call["temperature"] == 0.3
assert call["top_p"] == 0.7
assert call["top_k"] == 10
assert call["seed"] == 99
def test_sampler_simple_returns_a_single_string_output():
llama = FakeLlama(["a cat on a mat"])
result = llavaloader.LLavaSamplerSimple().generate_text(
image=image_batch(1), prompt="describe", model=llama, temperature=0.1
)
assert result == ("a cat on a mat",)
def test_sampler_advanced_uses_the_supplied_system_message():
llama = FakeLlama()
llavaloader.LLavaSamplerAdvanced().generate_text_advanced(
image=image_batch(1),
system_msg="answer in French",
prompt="describe",
model=llama,
max_tokens=16,
temperature=0.1,
top_p=0.9,
top_k=5,
frequency_penalty=0.0,
presence_penalty=0.0,
repeat_penalty=1.0,
seed=7,
)
(call,) = llama.calls
assert call["messages"][0] == {"role": "system", "content": "answer in French"}
# --------------------------------------------------------------------------
# Projector / clip wiring.
# --------------------------------------------------------------------------
def test_clip_factory_uses_the_config_create_hook():
config = LlavaClipConfig(Path("/models/mmproj.gguf"), "LLaVA 1.6")
assert llavaloader._clip_factory(config) == config.create
def test_clip_factory_accepts_a_precreated_handler():
sentinel = object()
factory = llavaloader._clip_factory(sentinel)
# Workflows saved before handler selection passed the handler itself.
assert factory() is sentinel
def test_make_handle_derives_the_projector_from_the_clip_config(resolved_paths):
config = LlavaClipConfig(resolved_paths / CLIP_FILE, "LLaVA 1.5")
handle = llavaloader._make_handle(MODEL_FILE, 4096, -1, 4, config)
assert isinstance(handle, LlamaHandle)
assert handle.projector_path == resolved_paths / CLIP_FILE
assert handle.chat_handler_factory == config.create
assert handle.n_ctx == 4096
# Still lazy: no llama.cpp object was constructed.
assert handle._llm is None
def test_make_handle_keeps_an_explicit_projector_override(resolved_paths):
config = LlavaClipConfig(resolved_paths / CLIP_FILE, "LLaVA 1.5")
override = Path("/models/other-mmproj.gguf")
handle = llavaloader._make_handle(
MODEL_FILE,
4096,
-1,
4,
config,
runtime_options={"projector_path": override},
)
assert handle.projector_path == override
def test_llava_loader_does_not_load_weights(resolved_paths):
config = LlavaClipConfig(resolved_paths / CLIP_FILE, "LLaVA 1.5")
(handle,) = llavaloader.LLavaLoader().load_llava_checkpoint(
ckpt_name=MODEL_FILE,
max_ctx=2048,
gpu_layers=10,
n_threads=8,
clip=config,
)
assert isinstance(handle, LlamaHandle)
assert handle._llm is None
assert handle.n_gpu_layers == 10
assert handle.n_threads == 8
def test_clip_loader_returns_a_frozen_config_with_the_chosen_handler(resolved_paths):
(config,) = llavaloader.LlavaClipLoader().load_clip_checkpoint(
CLIP_FILE, handler="MiniCPM-V 2.6"
)
assert isinstance(config, LlavaClipConfig)
assert config.model_path == resolved_paths / CLIP_FILE
assert config.handler == "MiniCPM-V 2.6"
def test_clip_config_rejects_an_unknown_handler(resolved_paths):
config = LlavaClipConfig(resolved_paths / CLIP_FILE, "Not A Handler")
with pytest.raises((ValueError, RuntimeError)) as error:
config.create()
# Either an unknown-handler rejection or a missing-wheel report is correct;
# a silent fallback to the wrong prompt format is not.
assert "handler" in str(error.value).lower() or "llama" in str(error.value).lower()
def test_clip_loader_defaults_to_the_embedded_gguf_chat_template():
handler = llavaloader.LlavaClipLoader.INPUT_TYPES()["optional"]["handler"]
choices, options = handler[0], handler[1]
assert options["default"] == "Auto (GGUF chat template)"
assert options["default"] in choices
assert "LLaVA 1.5" in choices
# --------------------------------------------------------------------------
# Cached-handle lifecycle (issue #137: "model never unloads").
# --------------------------------------------------------------------------
def test_cached_llava_reuses_one_handle_for_identical_settings(
fake_handles, resolved_paths
):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
first = node._model(MODEL_FILE, CLIP_FILE, 4096, -1, 4)
second = node._model(MODEL_FILE, CLIP_FILE, 4096, -1, 4)
assert first is second
assert len(fake_handles.instances) == 1
def test_cached_llava_rebuilds_when_the_projector_changes(fake_handles, resolved_paths):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
first = node._model(MODEL_FILE, CLIP_FILE, 4096, -1, 4)
second = node._model(MODEL_FILE, "other-mmproj.gguf", 4096, -1, 4)
assert first is not second
assert first.closed is True
assert len(fake_handles.instances) == 2
def test_cached_llava_rebuilds_when_the_handler_changes(fake_handles, resolved_paths):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
node._model(MODEL_FILE, CLIP_FILE, 4096, -1, 4, handler="LLaVA 1.5")
node._model(MODEL_FILE, CLIP_FILE, 4096, -1, 4, handler="LLaVA 1.6")
assert len(fake_handles.instances) == 2
def test_cached_llava_unload_releases_the_handle(fake_handles, resolved_paths):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
handle = node._model(MODEL_FILE, CLIP_FILE, 4096, -1, 4)
node._maybe_unload(False)
assert handle.closed is False
node._maybe_unload(True)
assert handle.closed is True
assert node._handle is None
assert node._key is None
def test_cached_llava_unload_is_safe_before_any_load(fake_handles, resolved_paths):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
# Must not raise when nothing was ever loaded.
node._maybe_unload(True)
assert node._handle is None
def _memory_free_kwargs(**overrides):
kwargs = {
"ckpt_name": MODEL_FILE,
"clip_name": CLIP_FILE,
"max_ctx": 4096,
"gpu_layers": -1,
"n_threads": 4,
"image": image_batch(1),
"prompt": "describe this",
"temperature": 0.1,
"unload": False,
}
kwargs.update(overrides)
return kwargs
def test_memory_free_simple_generates_through_the_cached_handle(
fake_handles, resolved_paths
):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
(text,) = node.generate_text(**_memory_free_kwargs())
assert text == "a description"
(handle,) = fake_handles.instances
assert handle.closed is False
(call,) = handle.llama.calls
assert call["temperature"] == 0.1
assert call["messages"][1]["content"][1]["text"] == "describe this"
def test_memory_free_simple_processes_every_image_in_the_batch(
fake_handles, resolved_paths
):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
node.generate_text(**_memory_free_kwargs(image=image_batch(2)))
(handle,) = fake_handles.instances
assert len(handle.llama.calls) == 2
def test_memory_free_simple_unloads_after_generating_when_asked(
fake_handles, resolved_paths
):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
node.generate_text(**_memory_free_kwargs(unload=True))
(handle,) = fake_handles.instances
assert handle.closed is True
assert node._handle is None
def test_memory_free_simple_unloads_even_when_generation_fails(
fake_handles, resolved_paths, monkeypatch
):
"""Issue #137: a failed generation must not strand the model in VRAM."""
def explode(*args, **kwargs):
raise RuntimeError("llama.cpp exploded")
monkeypatch.setattr(llavaloader, "_run_batch", explode)
node = llavaloader.LLavaOptionalMemoryFreeSimple()
with pytest.raises(RuntimeError, match="exploded"):
node.generate_text(**_memory_free_kwargs(unload=True))
(handle,) = fake_handles.instances
assert handle.closed is True
assert node._handle is None
def test_memory_free_advanced_forwards_the_system_message_and_sampling(
fake_handles, resolved_paths
):
node = llavaloader.LLavaOptionalMemoryFreeAdvanced()
(text,) = node.generate_text_advanced(
ckpt_name=MODEL_FILE,
clip_name=CLIP_FILE,
max_ctx=4096,
gpu_layers=-1,
n_threads=4,
image=image_batch(1),
system_msg="answer in German",
prompt="describe",
max_tokens=64,
temperature=0.4,
top_p=0.85,
top_k=25,
frequency_penalty=0.0,
presence_penalty=0.0,
repeat_penalty=1.05,
seed=5,
unload=False,
)
assert text == "a description"
(handle,) = fake_handles.instances
(call,) = handle.llama.calls
assert call["messages"][0] == {"role": "system", "content": "answer in German"}
assert call["max_tokens"] == 64
assert call["temperature"] == 0.4
assert call["top_p"] == 0.85
assert call["top_k"] == 25
assert call["seed"] == 5
def test_cached_llava_key_is_insensitive_to_runtime_option_ordering(
fake_handles, resolved_paths
):
node = llavaloader.LLavaOptionalMemoryFreeSimple()
node._model(MODEL_FILE, CLIP_FILE, 4096, -1, 4, n_batch=256, main_gpu=1)
node._model(MODEL_FILE, CLIP_FILE, 4096, -1, 4, main_gpu=1, n_batch=256)
assert len(fake_handles.instances) == 1
+1 -2
View File
@@ -1,9 +1,8 @@
import json
from pathlib import Path
import pytest
import ComfyUI_VLM_nodes as package
import pytest
from ComfyUI_VLM_nodes.nodes import simpletext
+735
View File
@@ -0,0 +1,735 @@
"""Contract tests for the GGUF text/LLM nodes in ``nodes/suggest.py``.
These nodes carry the pack's longest bug history (widget-index drift, unexpected
sampling kwargs, JSON that never parsed), so the assertions below pin the
behaviours those reports depended on rather than the models themselves. No
llama.cpp wheel and no GGUF weights are required.
"""
from __future__ import annotations
import json
from pathlib import Path
from types import SimpleNamespace
import pytest
from ComfyUI_VLM_nodes.nodes import suggest
from ComfyUI_VLM_nodes.nodes.runtime import LlamaHandle
MODEL_FILE = "some-model.gguf"
class FakeLlama:
"""Records the kwargs llama.cpp would have received."""
def __init__(self, content: str = "generated text"):
self.content = content
self.calls: list[dict] = []
def create_chat_completion(self, **kwargs):
self.calls.append(kwargs)
return {"choices": [{"message": {"content": self.content}}]}
class FakeHandle:
"""Stands in for LlamaHandle so caching can be observed without weights."""
instances: list[FakeHandle] = []
def __init__(self, model_path, **kwargs):
self.model_path = model_path
self.kwargs = kwargs
self.closed = False
self.llama = FakeLlama()
FakeHandle.instances.append(self)
def ensure_loaded(self):
return self.llama
def close(self):
self.closed = True
@pytest.fixture
def resolved_model(monkeypatch):
"""Bypass folder_paths so no real GGUF has to exist on disk."""
path = Path("/models/LLavacheckpoints") / MODEL_FILE
monkeypatch.setattr(suggest, "resolve_model_path", lambda name: path)
return path
@pytest.fixture
def fake_handles(monkeypatch):
FakeHandle.instances = []
monkeypatch.setattr(suggest, "LlamaHandle", FakeHandle)
return FakeHandle
# --------------------------------------------------------------------------
# Widget ordering. Comfy serializes widget values by position, so a reordered
# INPUT_TYPES silently rebinds every saved workflow (issue #156).
# --------------------------------------------------------------------------
def test_llm_sampler_widget_order_is_frozen():
assert list(suggest.LLMSampler.INPUT_TYPES()["required"]) == [
"system_msg",
"prompt",
"model",
"max_tokens",
"temperature",
"top_p",
"top_k",
"frequency_penalty",
"presence_penalty",
"repeat_penalty",
"seed",
]
def test_llm_prompt_generator_widget_order_is_frozen():
assert list(suggest.LLMPromptGenerator.INPUT_TYPES()["required"]) == [
"prompt",
"model",
"max_tokens",
"temperature",
"top_p",
"top_k",
"frequency_penalty",
"presence_penalty",
"repeat_penalty",
]
def test_llm_loader_widget_order_is_frozen():
schema = suggest.LLMLoader.INPUT_TYPES()
assert list(schema["required"]) == [
"ckpt_name",
"max_ctx",
"gpu_layers",
"n_threads",
]
# chat_format must stay ahead of the shared runtime widgets.
assert list(schema["optional"])[0] == "chat_format"
def test_structured_output_widget_order_is_frozen():
assert list(suggest.StructuredOutput.INPUT_TYPES()["required"]) == [
"prompt",
"model",
"temperature",
"attribute_name",
"attribute_type",
"attribute_description",
"categories",
]
def test_every_suggest_node_declares_a_callable_function_and_return_types():
for name, node_class in suggest.NODE_CLASS_MAPPINGS.items():
assert isinstance(node_class.RETURN_TYPES, tuple), name
assert node_class.RETURN_TYPES, name
assert callable(getattr(node_class, node_class.FUNCTION, None)), name
assert node_class.CATEGORY.startswith("VLM Nodes"), name
def test_display_names_cover_every_registered_node():
assert set(suggest.NODE_CLASS_MAPPINGS) == set(suggest.NODE_DISPLAY_NAME_MAPPINGS)
# --------------------------------------------------------------------------
# Sampling kwargs. Issue #144 was an "unexpected keyword argument" crash, so
# the plumbing from node widget to create_chat_completion is asserted directly.
# --------------------------------------------------------------------------
def test_llm_sampler_forwards_every_sampling_argument():
llama = FakeLlama("a description")
result = suggest.LLMSampler().generate_text_advanced(
system_msg="be terse",
prompt="describe a cat",
model=llama,
max_tokens=64,
temperature=0.7,
top_p=0.8,
top_k=20,
frequency_penalty=0.1,
presence_penalty=0.2,
repeat_penalty=1.3,
seed=1234,
)
assert result == ("a description",)
(call,) = llama.calls
assert call["messages"] == [
{"role": "system", "content": "be terse"},
{"role": "user", "content": "describe a cat"},
]
assert call["max_tokens"] == 64
assert call["temperature"] == 0.7
assert call["top_p"] == 0.8
assert call["top_k"] == 20
assert call["frequency_penalty"] == 0.1
assert call["presence_penalty"] == 0.2
assert call["repeat_penalty"] == 1.3
assert call["seed"] == 1234
# No response_format unless a structured node asked for one.
assert "response_format" not in call
def test_chat_unwraps_a_lazy_handle_before_generating(fake_handles, resolved_model):
handle = FakeHandle(resolved_model)
text = suggest._chat(handle, prompt="hi", system="sys")
assert text == "generated text"
def test_llama_chat_content_rejects_an_empty_completion():
llama = FakeLlama(content=" ")
with pytest.raises(RuntimeError, match="empty response"):
suggest.LLMSampler().generate_text_advanced(
system_msg="s",
prompt="p",
model=llama,
max_tokens=8,
temperature=0.1,
top_p=0.9,
top_k=1,
frequency_penalty=0.0,
presence_penalty=0.0,
repeat_penalty=1.0,
seed=1,
)
# --------------------------------------------------------------------------
# Lazy loading. Building a loader node must not touch llama.cpp or the GGUF.
# --------------------------------------------------------------------------
def test_llm_loader_builds_a_lazy_handle_without_loading_weights(resolved_model):
(handle,) = suggest.LLMLoader().load_llm_checkpoint(
ckpt_name=MODEL_FILE,
max_ctx=8192,
gpu_layers=20,
n_threads=6,
)
assert isinstance(handle, LlamaHandle)
assert handle.model_path == resolved_model
assert handle.n_ctx == 8192
assert handle.n_gpu_layers == 20
assert handle.n_threads == 6
# Nothing was loaded: the llama.cpp object is still absent.
assert handle._llm is None
def test_llm_loader_treats_blank_chat_format_as_the_embedded_template(resolved_model):
(blank,) = suggest.LLMLoader().load_llm_checkpoint(
MODEL_FILE, 2048, -1, 4, chat_format=" "
)
assert blank.chat_format is None
(explicit,) = suggest.LLMLoader().load_llm_checkpoint(
MODEL_FILE, 2048, -1, 4, chat_format=" chatml "
)
assert explicit.chat_format == "chatml"
def test_llm_loader_forwards_advanced_runtime_options(resolved_model):
(handle,) = suggest.LLMLoader().load_llm_checkpoint(
MODEL_FILE,
2048,
-1,
4,
n_batch=256,
n_ubatch=128,
flash_attention="Disabled",
use_mmap=False,
split_mode="Single GPU",
main_gpu=2,
tensor_split="0.6,0.4",
)
assert handle.n_batch == 256
assert handle.n_ubatch == 128
assert handle.flash_attention == "Disabled"
assert handle.use_mmap is False
assert handle.split_mode == "Single GPU"
assert handle.main_gpu == 2
assert handle.tensor_split == [0.6, 0.4]
# --------------------------------------------------------------------------
# Structured output.
# --------------------------------------------------------------------------
def _capture_chat(monkeypatch, payload):
"""Replace _chat so the generated JSON Schema can be inspected."""
recorded: dict = {}
def fake_chat(model, **kwargs):
recorded.update(kwargs)
return payload
monkeypatch.setattr(suggest, "_chat", fake_chat)
return recorded
@pytest.mark.parametrize(
("declared", "expected"),
[
("str", "string"),
("int", "integer"),
("float", "number"),
("bool", "boolean"),
],
)
def test_structured_output_maps_scalar_types_to_json_schema(
monkeypatch, declared, expected
):
recorded = _capture_chat(monkeypatch, json.dumps({"result": "value"}))
suggest.StructuredOutput().keyword_extract(
prompt="p",
model=object(),
temperature=0.1,
attribute_name="result",
attribute_type=declared,
attribute_description="a description",
categories="",
)
schema = recorded["response_format"]["schema"]
assert schema["properties"]["result"]["type"] == expected
assert schema["properties"]["result"]["description"] == "a description"
assert schema["required"] == ["result"]
assert schema["additionalProperties"] is False
assert recorded["response_format"]["type"] == "json_object"
def test_structured_output_builds_an_enum_for_categories(monkeypatch):
recorded = _capture_chat(monkeypatch, json.dumps({"mood": "calm"}))
(value,) = suggest.StructuredOutput().keyword_extract(
prompt="p",
model=object(),
temperature=0.1,
attribute_name=" mood ",
attribute_type="Category",
attribute_description="",
categories=" calm , tense ,, bright ",
)
schema = recorded["response_format"]["schema"]
assert schema["properties"]["mood"]["enum"] == ["calm", "tense", "bright"]
assert schema["properties"]["mood"]["type"] == "string"
assert value == "calm"
def test_structured_output_serializes_non_string_values(monkeypatch):
_capture_chat(monkeypatch, json.dumps({"count": 7}))
(value,) = suggest.StructuredOutput().keyword_extract(
prompt="p",
model=object(),
temperature=0.1,
attribute_name="count",
attribute_type="int",
attribute_description="",
categories="",
)
assert value == "7"
def test_structured_output_rejects_an_empty_attribute_name():
with pytest.raises(ValueError, match="attribute_name cannot be empty"):
suggest.StructuredOutput().keyword_extract(
prompt="p",
model=object(),
temperature=0.1,
attribute_name=" ",
attribute_type="str",
attribute_description="",
categories="",
)
def test_structured_output_rejects_a_category_without_values():
with pytest.raises(ValueError, match="at least one comma-separated value"):
suggest.StructuredOutput().keyword_extract(
prompt="p",
model=object(),
temperature=0.1,
attribute_name="mood",
attribute_type="Category",
attribute_description="",
categories=" , ",
)
def test_structured_chat_reports_unparseable_json_with_a_bounded_excerpt(monkeypatch):
_capture_chat(monkeypatch, "x" * 900)
with pytest.raises(RuntimeError, match="did not return valid JSON") as error:
suggest.KeywordExtraction().keyword_extract(
prompt="p", model=object(), temperature=0.1
)
# The raw completion is truncated so a runaway response cannot flood the log.
assert len(str(error.value)) < 600
def test_keyword_extraction_returns_the_raw_json_document(monkeypatch):
payload = json.dumps(
{
"main_character": ["cat"],
"artform": ["photo"],
"photo_type": ["portrait"],
"color_with_objects": ["black cat"],
"digital_artform": [],
"background": ["studio"],
"lighting": ["soft"],
}
)
_capture_chat(monkeypatch, payload)
(raw,) = suggest.KeywordExtraction().keyword_extract(
prompt="a cat", model=object(), temperature=0.1
)
assert json.loads(raw)["main_character"] == ["cat"]
def test_llava_prompt_generator_returns_only_the_prompt_field(monkeypatch):
_capture_chat(monkeypatch, json.dumps({"prompt": "a moody portrait"}))
(text,) = suggest.LLavaPromptGenerator().generate_prompts(
prompt="p", model=object(), temperature=0.1
)
assert text == "a moody portrait"
def test_creative_art_prompt_generator_prefers_the_narrative(monkeypatch):
_capture_chat(
monkeypatch,
json.dumps(
{
"techniques": {"preferred": ["ink"], "avoided": []},
"theme": {"core_subject": "a harbour"},
"style": {"desired": ["muted"], "undesired": []},
"creative_descriptions": [{"description": "a quiet harbour at dawn"}],
}
),
)
(text,) = suggest.CreativeArtPromptGenerator().create_creative_art_prompts(
prompt="p", model=object(), temperature=0.1
)
assert text == "a quiet harbour at dawn"
def test_creative_art_prompt_generator_composes_a_fallback_without_narratives(
monkeypatch,
):
_capture_chat(
monkeypatch,
json.dumps(
{
"techniques": {"preferred": ["ink", "wash"], "avoided": []},
"theme": {"core_subject": "a harbour"},
"style": {"desired": ["muted", "grainy"], "undesired": []},
"creative_descriptions": [],
}
),
)
(text,) = suggest.CreativeArtPromptGenerator().create_creative_art_prompts(
prompt="p", model=object(), temperature=0.1
)
assert text == (
"a harbour. Techniques: ink, wash. Visual style: muted, grainy."
)
def test_suggester_switches_instruction_on_the_randomize_toggle(monkeypatch):
payload = json.dumps(
{f"suggestion{index}": f"idea {index}" for index in range(1, 6)}
)
similar = _capture_chat(monkeypatch, payload)
suggest.Suggester().generate_suggestions(
prompt="p", model=object(), temperature=0.1, randomize=True
)
assert "close, useful variations" in similar["system"]
different = _capture_chat(monkeypatch, payload)
suggest.Suggester().generate_suggestions(
prompt="p", model=object(), temperature=0.1, randomize=False
)
assert "deliberately different" in different["system"]
# --------------------------------------------------------------------------
# Handle caching. Issue #137 was "model never unloads"; these pin the reuse
# and teardown rules of the optional-memory-free nodes.
# --------------------------------------------------------------------------
def test_cached_llm_reuses_one_handle_for_identical_settings(
fake_handles, resolved_model
):
node = suggest.LLMOptionalMemoryFreeSimple()
first = node._model(MODEL_FILE, 2048, -1, 4)
second = node._model(MODEL_FILE, 2048, -1, 4)
assert first is second
assert len(fake_handles.instances) == 1
def test_cached_llm_closes_the_previous_handle_when_settings_change(
fake_handles, resolved_model
):
node = suggest.LLMOptionalMemoryFreeSimple()
first = node._model(MODEL_FILE, 2048, -1, 4)
second = node._model(MODEL_FILE, 4096, -1, 4)
assert first is not second
assert first.closed is True
assert second.closed is False
assert len(fake_handles.instances) == 2
def test_cached_llm_rebuilds_when_an_advanced_runtime_option_changes(
fake_handles, resolved_model
):
node = suggest.LLMOptionalMemoryFreeSimple()
node._model(MODEL_FILE, 2048, -1, 4, n_batch=512)
node._model(MODEL_FILE, 2048, -1, 4, n_batch=256)
assert len(fake_handles.instances) == 2
def test_cached_llm_unload_releases_the_handle(fake_handles, resolved_model):
node = suggest.LLMOptionalMemoryFreeSimple()
handle = node._model(MODEL_FILE, 2048, -1, 4)
node._maybe_unload(False)
assert handle.closed is False
assert node._handle is handle
node._maybe_unload(True)
assert handle.closed is True
assert node._handle is None
assert node._key is None
def test_any_type_never_reports_a_type_mismatch():
assert (suggest.ANY != "IMAGE") is False
assert (suggest.ANY != "STRING") is False
# --------------------------------------------------------------------------
# ChatMusician. Issue #149's workaround was for users to append "respond in
# ABC notation starting with X:1" themselves; the node now owns that.
# --------------------------------------------------------------------------
def _chat_musician_kwargs():
return {
"max_tokens": 256,
"temperature": 0.2,
"top_p": 0.9,
"top_k": 40,
"frequency_penalty": 0.0,
"presence_penalty": 0.0,
"repeat_penalty": 1.1,
"seed": 42,
"sample_rate": 44100,
}
ABC_TUNE = "X:1\nT:Test\nM:4/4\nK:C\nCDEF|GABc|"
def test_chat_musician_asks_for_abc_notation_without_user_help(monkeypatch):
recorded = _capture_chat(monkeypatch, ABC_TUNE)
monkeypatch.setattr(suggest, "require_module", lambda *a, **k: _fake_symusic())
suggest.ChatMusician().chat_musician(
prompt="a waltz", model=object(), **_chat_musician_kwargs()
)
assert "ABC notation" in recorded["prompt"]
assert "X:" in recorded["prompt"]
assert "a waltz" in recorded["prompt"]
assert "ABC notation" in recorded["system"]
def test_chat_musician_rejects_a_response_without_an_abc_header(monkeypatch):
_capture_chat(monkeypatch, "Sure! Here is a lovely tune for you.")
with pytest.raises(RuntimeError, match="did not contain ABC notation"):
suggest.ChatMusician().chat_musician(
prompt="p", model=object(), **_chat_musician_kwargs()
)
def test_chat_musician_strips_preamble_before_the_abc_header(monkeypatch):
_capture_chat(monkeypatch, "Here you go:\n\n" + ABC_TUNE)
monkeypatch.setattr(suggest, "require_module", lambda *a, **k: _fake_symusic())
abc, _legacy, _rate, _audio = suggest.ChatMusician().chat_musician(
prompt="p", model=object(), **_chat_musician_kwargs()
)
assert abc.startswith("X:1")
assert "Here you go" not in abc
def test_chat_musician_returns_comfy_audio_and_legacy_layouts(monkeypatch):
_capture_chat(monkeypatch, ABC_TUNE)
monkeypatch.setattr(suggest, "require_module", lambda *a, **k: _fake_symusic())
_abc, legacy, rate, audio = suggest.ChatMusician().chat_musician(
prompt="p", model=object(), **_chat_musician_kwargs()
)
# Comfy AUDIO is [batch, channels, samples].
assert audio["waveform"].shape == (1, 2, 100)
assert audio["sample_rate"] == 44100
assert rate == 44100
# soundfile-compatible legacy output is [samples, channels].
assert legacy.shape == (100, 2)
def _fake_symusic():
"""A symusic stand-in so the AUDIO contract is testable without the wheel."""
import numpy as np
class Synthesizer:
def __init__(self, sample_rate):
self.sample_rate = sample_rate
def render(self, score, stereo=True):
return np.zeros((2, 100), dtype=np.float32)
class Score:
@staticmethod
def from_abc(abc):
return SimpleNamespace(abc=abc)
return SimpleNamespace(Score=Score, Synthesizer=Synthesizer)
def _memory_free_kwargs(**overrides):
kwargs = {
"ckpt_name": MODEL_FILE,
"max_ctx": 2048,
"gpu_layers": -1,
"n_threads": 4,
"prompt": "write a haiku",
"temperature": 0.2,
"unload": False,
}
kwargs.update(overrides)
return kwargs
def test_memory_free_simple_generates_through_the_cached_handle(
fake_handles, resolved_model
):
node = suggest.LLMOptionalMemoryFreeSimple()
(text,) = node.generate_text(**_memory_free_kwargs())
assert text == "generated text"
(handle,) = fake_handles.instances
assert handle.closed is False
assert node._handle is handle
(call,) = handle.llama.calls
assert call["temperature"] == 0.2
assert call["messages"][1]["content"] == "write a haiku"
def test_memory_free_simple_unloads_after_generating_when_asked(
fake_handles, resolved_model
):
node = suggest.LLMOptionalMemoryFreeSimple()
node.generate_text(**_memory_free_kwargs(unload=True))
(handle,) = fake_handles.instances
assert handle.closed is True
assert node._handle is None
def test_memory_free_simple_unloads_even_when_generation_fails(
fake_handles, resolved_model, monkeypatch
):
"""Issue #137: a failed generation must not strand the model in VRAM."""
def explode(*args, **kwargs):
raise RuntimeError("llama.cpp exploded")
monkeypatch.setattr(suggest, "_chat", explode)
node = suggest.LLMOptionalMemoryFreeSimple()
with pytest.raises(RuntimeError, match="exploded"):
node.generate_text(**_memory_free_kwargs(unload=True))
(handle,) = fake_handles.instances
assert handle.closed is True
assert node._handle is None
def test_memory_free_advanced_forwards_every_sampling_argument(
fake_handles, resolved_model
):
node = suggest.LLMOptionalMemoryFreeAdvanced()
signature = suggest.LLMOptionalMemoryFreeAdvanced.INPUT_TYPES()["required"]
assert "system_msg" in signature
(text,) = node.generate_text_advanced(
ckpt_name=MODEL_FILE,
max_ctx=2048,
gpu_layers=-1,
n_threads=4,
system_msg="be brief",
prompt="a haiku",
max_tokens=48,
temperature=0.5,
top_p=0.8,
top_k=15,
frequency_penalty=0.1,
presence_penalty=0.2,
repeat_penalty=1.2,
seed=11,
unload=False,
)
assert text == "generated text"
(handle,) = fake_handles.instances
(call,) = handle.llama.calls
assert call["messages"][0] == {"role": "system", "content": "be brief"}
assert call["max_tokens"] == 48
assert call["temperature"] == 0.5
assert call["top_p"] == 0.8
assert call["top_k"] == 15
assert call["seed"] == 11
def test_cached_llm_key_is_insensitive_to_runtime_option_ordering(
fake_handles, resolved_model
):
node = suggest.LLMOptionalMemoryFreeSimple()
node._model(MODEL_FILE, 2048, -1, 4, n_batch=256, main_gpu=1)
node._model(MODEL_FILE, 2048, -1, 4, main_gpu=1, n_batch=256)
# Keyword order must not invalidate the cache and reload the GGUF.
assert len(fake_handles.instances) == 1
def test_schema_helper_emits_a_json_schema_for_a_pydantic_model():
schema = suggest._schema(suggest.PromptGen)
assert schema["properties"]["prompt"]["type"] == "string"
assert schema["required"] == ["prompt"]
def test_response_content_extraction_matches_the_runtime_helper():
response = {"choices": [{"message": {"content": " text "}}]}
assert suggest._response_content(response) == "text"
def test_stub_handle_matches_the_real_handle_api():
"""Guard the stub: LlamaHandle must keep the API these tests rely on."""
assert callable(getattr(LlamaHandle, "ensure_loaded", None))
assert callable(getattr(LlamaHandle, "close", None))
-1
View File
@@ -5,7 +5,6 @@ from pathlib import Path
import pytest
import torch
from ComfyUI_VLM_nodes.nodes.video_intelligence import (
NODE_CLASS_MAPPINGS,
VLMAdaptiveFrameSampler,