Compare commits

..
Author SHA1 Message Date
Jedrzej KosinskiandAmp 04ca70a853 Add VIDEO input encoder to close the encode-side gap
The proxy_node has accepted VIDEO outputs since PR #13 (server-side
`make_video_envelope` + client-side `decode_video_envelope`), but
the inbound encoding direction has been a no-op: `_encode_one` in
`proxy_node.py` only dispatched AUDIO + IMAGE + MASK tensors, so any
remote node declaring an `IO.Video.Input` would receive the raw
`VideoInput` object as a non-JSON-serialisable Python value and the
upstream serialisation would either drop it or fail.

`serialization.py:468` made the gap explicit with a placeholder
comment ("encode lands when a remote node accepts VIDEO inputs"),
and several server-side providers shipped against the
`VIDEO_MP4_BASE64` capability + `input_serialization={"video":
"mp4_base64"}` declarations on the assumption that the client
encoder would land soon — most notably `GrokVideoEditProvider`
and `GrokVideoExtendProvider`, whose end-to-end VIDEO upload path
was broken until now.

This change wires the encode direction:

* `serialization.py` gains `encode_video_input(video)` — mirrors
  the AUDIO encoder above. Re-encodes the `VideoInput` to an
  in-memory mp4/H.264 byte buffer via the upstream
  `comfy_api_nodes/util/conversions.video_to_base64_string` (same
  call shape: `video.save_to(buf, format=MP4, codec=H264)`) and
  base64-encodes the result. Populates `duration_s` from
  `video.get_duration()` so server-side providers (e.g.
  `_validate_video_duration_envelope` in `grok.py`) can range-check
  without demuxing the MP4.

* `serialization.py` gains `is_video_input(value)` — duck-types on
  the `save_to` + `get_duration` method pair (both defined on
  `comfy_api.latest._input.VideoInput` and present on every
  concrete subclass). Avoids importing `VideoInput` at module
  load (which would pull torch via the `comfy_api` tree). Both
  helpers added to `__all__`.

* `proxy_node.py:_encode_one` gains a VIDEO branch right after the
  AUDIO branch (and before the tensor-rank dispatch), mirroring the
  AUDIO pattern at the same call site.

* `proxy_node.py` `_inputs_to_envelopes` docstring updates the
  duck-typing-rules list to mention the new VIDEO branch.

Symmetric to the (still-open) MODEL_3D INPUT gap noted at
`serialization.py:484`. AUDIO INPUT was already wired (ElevenLabs
nodes exercise it end-to-end); IMAGE / MASK have always been
wired; VIDEO INPUT closes today's last input-direction envelope
gap for the existing capability vocabulary.

Verified by importing `serialization` with a stubbed
`comfy_remote_nodes` package: `is_video_input` returns False for
None / dict / audio-shaped dict and True for an object exposing
both `save_to` + `get_duration`. `encode_video_input` + the
proxy_node integration are exercised end-to-end by the server-side
`Wan2VideoContinuationApi_RNP` provider landing in the matching
comfy-rnp-server PR (it consumes `first_clip: VIDEO` and uploads
the resolved bytes to Comfy storage), and also unblocks the
shipped-but-broken Grok VideoEdit/VideoExtend providers.

Co-authored-by: Amp <amp@ampcode.com>
Amp-Thread-ID: https://ampcode.com/threads/T-019e35d1-3c68-74ce-9918-fe9cacd74276
2026-05-17 05:07:57 -07:00
Jedrzej Kosinski ed5823e8e7 Merge pull request #18 from Comfy-Org/revert/task-handle-decoder
Revert "Add client-side cross-node task_handle decoder" (PR #17)
2026-05-16 22:08:11 -07:00
2 changed files with 60 additions and 1 deletions
+4
View File
@@ -1546,6 +1546,8 @@ async def _encode_inputs(
* ``torch.Tensor`` rank 3 with last-dim != 3/4 → mask envelope (B,H,W).
Rank-2 (H,W) tensors are also treated as masks.
* ``dict`` with ``waveform`` + ``sample_rate`` → audio envelope.
* ``VideoInput`` subclass (duck-typed via ``save_to`` +
``get_duration``) → mp4-base64 video envelope.
* Everything else passes through as-is (already JSON-serializable).
When ``max_inline_bytes`` is set, any envelope whose inline payload
@@ -1694,6 +1696,8 @@ def _encode_one(
) -> Any:
if serialization.is_audio_input(value):
return serialization.encode_audio_input(value)
if serialization.is_video_input(value):
return serialization.encode_video_input(value)
if not serialization._is_torch_tensor(value):
return value
rank = value.dim()
+56 -1
View File
@@ -465,9 +465,44 @@ def decode_audio_envelope(envelope: dict[str, Any]) -> dict[str, Any]:
# ---------------------------------------------------------------------------
# Video decode (encode lands when a remote node accepts VIDEO inputs)
# Video encode / decode
# ---------------------------------------------------------------------------
def encode_video_input(video: Any) -> dict[str, Any]:
"""Encode a ComfyUI VIDEO (``VideoInput`` subclass) as an mp4-base64
video envelope.
Mirrors the AUDIO encoder above: re-encodes the Video object to an
in-memory mp4/H.264 byte buffer (matches upstream
``comfy_api_nodes/util/conversions.py`` ``video_to_base64_string``)
and base64-encodes it. Populates ``duration_s`` metadata when the
Video object exposes ``get_duration()`` so server-side providers
(e.g. ``GrokVideoExtendProvider._validate_video_duration_envelope``)
can range-check without demuxing the MP4.
"""
from comfy_api_nodes.util.conversions import video_to_base64_string
b64 = video_to_base64_string(video)
extra: dict[str, Any] = {}
duration_s: float | None = None
try:
getter = getattr(video, "get_duration", None)
if callable(getter):
d = getter()
if d is not None:
duration_s = float(d)
except Exception:
duration_s = None
if duration_s is not None:
extra["duration_s"] = duration_s
return {
"type": "video",
"encoding": "mp4_base64",
"data": b64,
**extra,
}
def decode_video_envelope(envelope: dict[str, Any]) -> Any:
"""Decode a video envelope into a ComfyUI Video object (mp4 inline)."""
encoding = envelope.get("encoding")
@@ -837,6 +872,24 @@ def is_audio_input(value: Any) -> bool:
)
def is_video_input(value: Any) -> bool:
"""Best-effort check for a ComfyUI ``VideoInput`` subclass.
Duck-types on the ``save_to`` + ``get_duration`` method pair: both
are defined on the ``VideoInput`` abstract base in
``comfy_api.latest._input`` and present on every concrete subclass
(``VideoFromFile`` / ``VideoFromComponents``). Avoids importing
``VideoInput`` at module load (which would pull torch via the
``comfy_api`` tree).
"""
if isinstance(value, dict):
return False
return (
callable(getattr(value, "save_to", None))
and callable(getattr(value, "get_duration", None))
)
def _is_torch_tensor(value: Any) -> bool:
try:
import torch
@@ -853,6 +906,7 @@ __all__ = [
"decode_mask_envelope",
"encode_audio_input",
"decode_audio_envelope",
"encode_video_input",
"decode_video_envelope",
"decode_model3d_envelope",
"_bundled_file3d_class",
@@ -861,4 +915,5 @@ __all__ = [
"is_image_tensor",
"is_mask_tensor",
"is_audio_input",
"is_video_input",
]