Second LLMProvider (ADR-007), backed by llama-server's router mode. Only
one new ComfyUI node — LlamaCppClient, emitting the same LLM_CLIENT socket
Ollama Client does. ChatCompletion, LLMModelSelector, LLMLoadModel,
LLMUnloadModel work unmodified once wired to it — the actual proof the
provider abstraction generalizes, not just Ollama-shaped in practice.
API shape verified live against ggml-org/llama.cpp's tools/server/README.md
(research.md) rather than assumed: model identifier field is "id" (not
"name"), status is a nested {"value": "..."} object. list_models() needs
no normalization — llama.cpp's status vocabulary is exactly ModelStatus's
full set, unlike Ollama's narrower one.
38 new tests (test_llamacpp_provider.py, test_llamacpp.py), including a
dedicated US4 test proving no generic node branches on provider type —
the same call sequence succeeds against an Ollama-shaped or llama.cpp-
shaped fake provider.
README/docs updated for the new node this time, not left stale.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj