11 Commits
Author SHA1 Message Date
MNeMoNiCuZ c8abe83cac Updated LLM Request node with improved UI, support for A Thousand Words VLM server, and other tweaks. 2026-09-27 20:52:56 +02:00
MNeMoNiCuZ ea6e3f5ce1 Added LLM Request node 2026-09-27 02:37:07 +02:00
Claude 2f8c933c28 Fix review round 2 (new features): CLI isolation, stderr race, custom panel
- Codex runs with --ignore-user-config (no MCP servers or hooks from
  config.toml; sign-in unaffected) and also disables hooks, multi_agent,
  view_image and goals; an images-only request gets a default prompt.
- Claude Code runs with --setting-sources "" so user/project settings
  (hooks) don't apply.
- CLI reader threads are joined before reading stderr, so login errors are
  no longer occasionally lost.
- The custom-endpoint panel keeps unsaved edits across status refreshes
  (Test, finished runs); the model tooltip mentions CLI defaults.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 22:24:08 +00:00
Claude 1c890958f4 Codex: refuse to run unless its shell tools are confirmed off (review round 2)
If `codex features list` failed or no longer listed shell_tool/unified_exec,
Codex ran with every tool enabled; a prompt could then have it read .env or
the custom-endpoint store into its reply. The probe now fails closed and a
failed probe is not cached.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 22:20:15 +00:00
Claude 73a066631b Harden custom endpoint and CLI model names (review round 1)
- Changing a custom endpoint's address or protocol without re-entering
  the key clears the saved key: the id is in every shared workflow, so it
  must not be enough to point a saved key at another server. The store
  file is created 0600 from the start.
- CLI model names must match a strict pattern: on Windows an npm-installed
  CLI is a .cmd run through cmd.exe, where a crafted model name from a
  shared workflow could inject commands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 22:12:01 +00:00
Claude 58f2e0cbbc Add local Claude Code / Codex subscriptions, Sanctum, custom endpoint; allow empty inputs
- Empty user_input is allowed: instructions (system message or preset)
  alone are sent as the request. Only all-empty (no text, no images) fails.
- Local Claude Code Subscription: runs `claude -p` with stream-json in/out,
  all tools and MCP off, no session persistence, system prompt via file,
  images supported. Local Codex Subscription: runs `codex exec --json` in a
  read-only sandbox in a temp folder, with the tool features this Codex
  version knows disabled; images via -i. API-key variables are removed from
  the CLI environment so the subscription login is used. Paths overridable
  via CLAUDE_CLI / CODEX_CLI in .env.
- Sanctum endpoint (SANCTUM_URL / SANCTUM_API_KEY), listed last.
- Custom Endpoint - WARNING: address, protocol and key entered on the node
  are stored on this machine (git-ignored CustomEndpoints.local.json); the
  workflow keeps only a random id. The key is never sent back to the
  browser. Panel and docs spell out the risks.
- Generic endpoint options: ensure_path, model_optional, keep_order,
  pin_last; OpenAI model lists show collection labels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 22:09:20 +00:00
Claude 97394e8c8d Fix review round 5 (backend): auto-pick, parameter matching, output caps
- Auto-picking a model skips embedding/reranker models and only applies
  to local/network endpoints (an Ollama endpoint in the cloud included).
- Rejected parameters are matched as whole words in snake_case or
  camelCase, so "top_p" no longer matches "stopped".
- A "max_tokens: N > LIMIT" rejection (older Claude models' output caps)
  lowers max_tokens to the limit, keeping any thinking budget inside it,
  and retries once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 21:18:43 +00:00
Claude 61d4bf9dbb Fix review round 4 (frontend/docs): empty model on local servers, status refresh
- An empty model on a local or network endpoint without a default (LM
  Studio, llama.cpp, vLLM…) uses the first model the server lists, making
  "no setup needed" true; cloud endpoints still ask for one. Tooltip,
  docs and the status hint say so.
- The status chips refresh after each run, so a key added to .env shows
  as set; the Test dot stays "busy" until the test finishes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 21:13:15 +00:00
Claude 592427d8a1 Fix review round 3 (frontend): reply survives undo, stale Test, doc accuracy
- Undo/redo and workflow tab switches rebuild the node without replaying
  onExecuted; the panel now restores its last reply from app.nodeOutputs.
- A slow ⚡ Test result is ignored if the endpoint changed or a run started
  meanwhile.
- Docs and the status hint describe Ollama's empty-model choice correctly
  (a model in memory first) and scope the preset greying to the classic
  node view.
- An empty error body reports "(no error message)".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 21:02:32 +00:00
Claude b5013c5409 Fix review round 1: current Claude thinking API, stream fallbacks, UI
- Claude: model-aware requests. Current models (Sonnet 5, Opus 4.6+,
  Fable) get adaptive thinking + output_config.effort and no sampling
  params where they were removed; Haiku 4.5 and older keep budget_tokens.
  "none" sends thinking disabled where accepted, lowest effort otherwise.
  Thinking display is "summarized" so the thinking output isn't empty.
  Default max_tokens 16000. Dropping a rejected budget also returns the
  budget from max_tokens.
- Streaming: a server that ignores stream:true and answers with JSON is
  parsed as JSON; a stream that ends with nothing raises instead of
  returning an empty success.
- A literal </think> mid-sentence no longer cuts off the reply.
- UI: presets refetch when a name is unknown (after R), Copy falls back
  to execCommand outside secure contexts and reports failure, and the
  default model shows in the status line (single-line widgets have no
  placeholder).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 20:45:38 +00:00
Claude 9d2a2b4cb9 Add Universal LLM API node; move all LLM secrets to .env
New node ✨🧠 Universal LLM API (MNeMiC_LLMAPI) talks to any LLM through
three protocols: OpenAI-compatible (OpenAI, Gemini, Grok, Groq, OpenRouter,
Mistral, DeepSeek, LM Studio, llama.cpp, vLLM…), Anthropic, and native
Ollama. Twelve endpoints are built in; users add their own in the
git-ignored nodes/llm/UserEndpoints.json.

Secrets never enter a workflow: the node stores only the endpoint and model
name. URLs, keys and headers are resolved on the backend from the endpoint
config and .env (${VAR} / ${VAR:-fallback}), keys never reach the browser,
and any key a server echoes back is masked in outputs, errors and logs.

Features: live streaming preview on the node, searchable model browser,
connection test, preset viewer, vision (image batch), reasoning control
mapped per provider, <think> separation into a thinking output, JSON mode,
retries with backoff, cancel mid-stream, automatic retry without parameters
an endpoint rejects, and Ollama keep_alive/num_ctx plus a ComfyUI VRAM free.

Shared infrastructure: utils/env_manager.py is now the single secret store
(hot-reloads .env, redaction, placeholder detection) and
utils/prompt_presets.py loads the presets the Groq nodes and the new node
share. The Groq nodes use both, so .env and preset edits apply without a
restart.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MmQWqMnuUXU4XwF6xyaDyM
2026-09-25 19:51:33 +00:00