Compare commits

...
Author SHA1 Message Date
Jedrzej KosinskiandAmp 04ca70a853 Add VIDEO input encoder to close the encode-side gap
The proxy_node has accepted VIDEO outputs since PR #13 (server-side
`make_video_envelope` + client-side `decode_video_envelope`), but
the inbound encoding direction has been a no-op: `_encode_one` in
`proxy_node.py` only dispatched AUDIO + IMAGE + MASK tensors, so any
remote node declaring an `IO.Video.Input` would receive the raw
`VideoInput` object as a non-JSON-serialisable Python value and the
upstream serialisation would either drop it or fail.

`serialization.py:468` made the gap explicit with a placeholder
comment ("encode lands when a remote node accepts VIDEO inputs"),
and several server-side providers shipped against the
`VIDEO_MP4_BASE64` capability + `input_serialization={"video":
"mp4_base64"}` declarations on the assumption that the client
encoder would land soon — most notably `GrokVideoEditProvider`
and `GrokVideoExtendProvider`, whose end-to-end VIDEO upload path
was broken until now.

This change wires the encode direction:

* `serialization.py` gains `encode_video_input(video)` — mirrors
  the AUDIO encoder above. Re-encodes the `VideoInput` to an
  in-memory mp4/H.264 byte buffer via the upstream
  `comfy_api_nodes/util/conversions.video_to_base64_string` (same
  call shape: `video.save_to(buf, format=MP4, codec=H264)`) and
  base64-encodes the result. Populates `duration_s` from
  `video.get_duration()` so server-side providers (e.g.
  `_validate_video_duration_envelope` in `grok.py`) can range-check
  without demuxing the MP4.

* `serialization.py` gains `is_video_input(value)` — duck-types on
  the `save_to` + `get_duration` method pair (both defined on
  `comfy_api.latest._input.VideoInput` and present on every
  concrete subclass). Avoids importing `VideoInput` at module
  load (which would pull torch via the `comfy_api` tree). Both
  helpers added to `__all__`.

* `proxy_node.py:_encode_one` gains a VIDEO branch right after the
  AUDIO branch (and before the tensor-rank dispatch), mirroring the
  AUDIO pattern at the same call site.

* `proxy_node.py` `_inputs_to_envelopes` docstring updates the
  duck-typing-rules list to mention the new VIDEO branch.

Symmetric to the (still-open) MODEL_3D INPUT gap noted at
`serialization.py:484`. AUDIO INPUT was already wired (ElevenLabs
nodes exercise it end-to-end); IMAGE / MASK have always been
wired; VIDEO INPUT closes today's last input-direction envelope
gap for the existing capability vocabulary.

Verified by importing `serialization` with a stubbed
`comfy_remote_nodes` package: `is_video_input` returns False for
None / dict / audio-shaped dict and True for an object exposing
both `save_to` + `get_duration`. `encode_video_input` + the
proxy_node integration are exercised end-to-end by the server-side
`Wan2VideoContinuationApi_RNP` provider landing in the matching
comfy-rnp-server PR (it consumes `first_clip: VIDEO` and uploads
the resolved bytes to Comfy storage), and also unblocks the
shipped-but-broken Grok VideoEdit/VideoExtend providers.

Co-authored-by: Amp <amp@ampcode.com>
Amp-Thread-ID: https://ampcode.com/threads/T-019e35d1-3c68-74ce-9918-fe9cacd74276
2026-05-17 05:07:57 -07:00
Jedrzej Kosinski ed5823e8e7 Merge pull request #18 from Comfy-Org/revert/task-handle-decoder
Revert "Add client-side cross-node task_handle decoder" (PR #17)
2026-05-16 22:08:11 -07:00
Jedrzej KosinskiandAmp 22e0742d47 Revert "Add client-side cross-node task_handle decoder" (PR #17)
The task_handle envelope was overcooked for what the RNP/1 prototype
actually needs. The goal of the prototype is to move the existing
ComfyUI partner nodes to execute server-side without changing what
they do — and the upstream Tripo / Kling task chaining already works
fine over a plain ``str``:

* Tripo uses ``IO.Custom("MODEL_TASK_ID")`` / ``IO.Custom("RIG_TASK_ID")``
  / ``IO.Custom("RETARGET_TASK_ID")`` sockets, which is purely a
  client-side graph-validation mechanism. The value on the noodle is
  just a string — already covered by ``Capability.IO_OPAQUE``, which
  preserves the type string for connection validity and round-trips
  the value untouched.
* Kling uses plain ``IO.String`` for ``video_id`` — covered trivially
  by the existing string handling.

The ``TaskHandle`` dataclass + ``vendor`` / ``parent_chain`` /
``origin_node_id`` machinery were forward-looking infrastructure for
future cascading-replay semantics that the upstream partner nodes
don't have. Out of scope for "move execution server-side, change
nothing else".

Reverts merge commit 90ecd32 (PR #17: ``feat/task-handle-decoder``).
``Capability.IO_TASK_HANDLE``, the ``task_handle`` HEAVY_TYPES entry,
``TaskHandle`` dataclass, ``decode_task_handle_envelope``,
``encode_task_handle``, the dispatcher entry, the
``CLIENT_CAPABILITIES`` advertisement, and the smoke test are all
removed. Matching server-side revert in comfy-rnp-server.

Co-authored-by: Amp <amp@ampcode.com>
Amp-Thread-ID: https://ampcode.com/threads/T-019e3431-c416-742b-973c-e262ab13f2ff
2026-05-16 22:07:22 -07:00
Jedrzej Kosinski 90ecd3218c Merge pull request #17 from Comfy-Org/feat/task-handle-decoder
Add client-side cross-node task_handle decoder + TaskHandle dataclass
2026-05-16 21:34:42 -07:00
Jedrzej KosinskiandAmp 1f136526cf Add client-side cross-node task_handle decoder + TaskHandle dataclass
Mirrors comfy-rnp-server PR #39 on the client. Adds the surfaces
needed for chained partner provider nodes (Tripo Texture/Refine/
Rig/Retarget/Conversion, Kling VideoExtend, future partner chained
nodes) to pass vendor-native task / video / job IDs through a
workflow with their lineage metadata intact.

Surfaces:

* ``protocol.Capability.IO_TASK_HANDLE = "io:task_handle"`` and
  ``"task_handle"`` added to ``protocol.HEAVY_TYPES`` so
  ``is_envelope`` recognises the new envelope.
* ``serialization.TaskHandle`` dataclass with the wire-shape fields
  (``vendor``, ``kind``, ``native_id``, ``origin_node_id``,
  ``parent_chain``). ``__str__`` surfaces ``native_id`` for log /
  UI preview without leaking the lineage chain.
* ``serialization.decode_task_handle_envelope`` validates every
  field (mirroring server-side ``_validate_task_handle_ref`` so a
  bad envelope surfaces the same error message on both sides) and
  returns the dataclass. Rejects:
  - wrong encoding (only ``vendor_inline`` recognised);
  - missing / empty top-level fields;
  - comma-union ``kind`` strings (the input-acceptance string used
    by some provider nodes, NOT an emitted handle kind);
  - non-list ``parent_chain``;
  - malformed parent_chain entries (missing fields, non-dict,
    comma-union kind in lineage).
* ``serialization.encode_task_handle`` re-encodes the dataclass
  back to wire shape — symmetric counterpart of the decoder so a
  handle decoded from one provider's output and fed into another
  provider's input round-trips losslessly without ``proxy_node``
  having to know anything provider-specific.
* ``decode_envelope`` dispatcher routes ``type="task_handle"`` to
  the new decoder so the generic ``_deserialize_output`` path picks
  it up.
* ``client.CLIENT_CAPABILITIES`` advertises
  ``Capability.IO_TASK_HANDLE`` so server-side capability-gated
  descriptors (whole-descriptor gating — partial-socket synthesis
  is rejected because a chain that breaks meaningfully if either
  end is missing is worse than no chain at all) negotiate cleanly.

The dataclass return (NOT a plain ``str``) is deliberate: a plain
string would silently drop ``parent_chain`` when the value is
passed through the workflow, breaking the lineage the server uses
for future cascading-replay semantics (deferred to a follow-up PR
alongside server-side replay-cache storage — see PR #39 for the
replay design notes).

Smoke test ``notes/run_task_handle_decoder.py`` mirrors the pattern
used by ``run_model3d_bundle_decoder.py``: stubs out the ``client``
import for the relative ``from . import client as rnp_client`` at
the top of serialization.py, then exercises the 9 contract points
end-to-end. End-to-end check imports the server's
``make_task_handle_envelope`` from the comfy-rnp-server worktree
and verifies the wire shape round-trips losslessly through
decode -> encode.

Co-authored-by: Amp <amp@ampcode.com>
Amp-Thread-ID: https://ampcode.com/threads/T-019e3431-c416-742b-973c-e262ab13f2ff
2026-05-16 21:34:05 -07:00
Jedrzej Kosinski e5db513585 Merge pull request #16 from Comfy-Org/feat/model-3d-bundle-decoder
Add client-side multi-file 3D bundle decoder + BundledFile3D
2026-05-16 21:16:09 -07:00
2 changed files with 60 additions and 1 deletions
+4
View File
@@ -1546,6 +1546,8 @@ async def _encode_inputs(
* ``torch.Tensor`` rank 3 with last-dim != 3/4 → mask envelope (B,H,W).
Rank-2 (H,W) tensors are also treated as masks.
* ``dict`` with ``waveform`` + ``sample_rate`` → audio envelope.
* ``VideoInput`` subclass (duck-typed via ``save_to`` +
``get_duration``) → mp4-base64 video envelope.
* Everything else passes through as-is (already JSON-serializable).
When ``max_inline_bytes`` is set, any envelope whose inline payload
@@ -1694,6 +1696,8 @@ def _encode_one(
) -> Any:
if serialization.is_audio_input(value):
return serialization.encode_audio_input(value)
if serialization.is_video_input(value):
return serialization.encode_video_input(value)
if not serialization._is_torch_tensor(value):
return value
rank = value.dim()
+56 -1
View File
@@ -465,9 +465,44 @@ def decode_audio_envelope(envelope: dict[str, Any]) -> dict[str, Any]:
# ---------------------------------------------------------------------------
# Video decode (encode lands when a remote node accepts VIDEO inputs)
# Video encode / decode
# ---------------------------------------------------------------------------
def encode_video_input(video: Any) -> dict[str, Any]:
"""Encode a ComfyUI VIDEO (``VideoInput`` subclass) as an mp4-base64
video envelope.
Mirrors the AUDIO encoder above: re-encodes the Video object to an
in-memory mp4/H.264 byte buffer (matches upstream
``comfy_api_nodes/util/conversions.py`` ``video_to_base64_string``)
and base64-encodes it. Populates ``duration_s`` metadata when the
Video object exposes ``get_duration()`` so server-side providers
(e.g. ``GrokVideoExtendProvider._validate_video_duration_envelope``)
can range-check without demuxing the MP4.
"""
from comfy_api_nodes.util.conversions import video_to_base64_string
b64 = video_to_base64_string(video)
extra: dict[str, Any] = {}
duration_s: float | None = None
try:
getter = getattr(video, "get_duration", None)
if callable(getter):
d = getter()
if d is not None:
duration_s = float(d)
except Exception:
duration_s = None
if duration_s is not None:
extra["duration_s"] = duration_s
return {
"type": "video",
"encoding": "mp4_base64",
"data": b64,
**extra,
}
def decode_video_envelope(envelope: dict[str, Any]) -> Any:
"""Decode a video envelope into a ComfyUI Video object (mp4 inline)."""
encoding = envelope.get("encoding")
@@ -837,6 +872,24 @@ def is_audio_input(value: Any) -> bool:
)
def is_video_input(value: Any) -> bool:
"""Best-effort check for a ComfyUI ``VideoInput`` subclass.
Duck-types on the ``save_to`` + ``get_duration`` method pair: both
are defined on the ``VideoInput`` abstract base in
``comfy_api.latest._input`` and present on every concrete subclass
(``VideoFromFile`` / ``VideoFromComponents``). Avoids importing
``VideoInput`` at module load (which would pull torch via the
``comfy_api`` tree).
"""
if isinstance(value, dict):
return False
return (
callable(getattr(value, "save_to", None))
and callable(getattr(value, "get_duration", None))
)
def _is_torch_tensor(value: Any) -> bool:
try:
import torch
@@ -853,6 +906,7 @@ __all__ = [
"decode_mask_envelope",
"encode_audio_input",
"decode_audio_envelope",
"encode_video_input",
"decode_video_envelope",
"decode_model3d_envelope",
"_bundled_file3d_class",
@@ -861,4 +915,5 @@ __all__ = [
"is_image_tensor",
"is_mask_tensor",
"is_audio_input",
"is_video_input",
]