added minimax h3 utilities

This commit is contained in:
IAMCCS
2026-08-09 02:09:44 +02:00
parent 934f481de2
commit 9e8fc5832d
12 changed files with 2711 additions and 254 deletions
+1
View File
@@ -47,6 +47,7 @@ audio preprocessing, VAE decode, and video combine nodes. Before sharing or
testing a SuperNode workflow, check the dedicated requirements document:
- [IAMCCS SuperNodes Requirements](SUPERNODES_REQUIREMENTS.md)
- [MiniMax H3 Shotboard Workflow Requirements](docs/MINIMAX_H3_WORKFLOW_REQUIREMENTS.md)
- [AudioBoard + BusOut Guide](AUDIOBOARD_BUSOUT_GUIDE.md)
## 🆕 Added new LTX-2.3 nodes for v2v, au+img2vid (instructions: patreon.com/IAMCCS)
+3
View File
@@ -1,5 +1,8 @@
# IAMCCS SuperNodes Requirements
> MiniMax H3 Shotboard, Turbo, LTX/Wan finishing, RIFE and RTX VSR have their
> own dependency matrix: [MiniMax H3 Shotboard Workflow Requirements](docs/MINIMAX_H3_WORKFLOW_REQUIREMENTS.md).
The IAMCCS SuperNodes are workflow wrappers. They do not replace the underlying
ComfyUI, LTXV, audio, VAE, stitching, and helper nodes; they orchestrate them.
If one dependency is missing or outdated, the SuperNode may load in the graph but
+143
View File
@@ -0,0 +1,143 @@
# IAMCCS MiniMax H3 Shotboard Workflow Requirements
This document covers the IAMCCS MiniMax H3 Shotboard production workflows,
including the native H3 route, Turbo sampling, live preview, LTX/Wan finishing,
RIFE interpolation and optional RTX Video Super Resolution delivery.
## Base runtime
- A current ComfyUI build with native MiniMax H3 AV conditioning/sampling and
current LTX audio-video nodes.
- A recent NVIDIA driver and a PyTorch/CUDA build compatible with the selected
attention extensions. Reinstall compiled attention wheels after changing the
PyTorch/CUDA build.
- `IAMCCS-nodes` installed once in `custom_nodes`. Remove or move duplicate and
backup copies outside `custom_nodes`, otherwise ComfyUI can register stale
classes or report import failures.
- NVIDIA CUDA GPU for the supplied accelerated graphs. Twelve GB VRAM is
supported through dynamic weight offload; more VRAM reduces offload and wait
time. At least 32 GB system RAM is practical, while 64 GB is recommended for
H3 plus an LTX/Wan finishing pass.
## Required node packs for the supplied H3 graphs
### IAMCCS-nodes
Provides the Shotboard and its workflow-facing wrappers:
- `IAMCCS_MiniMaxH3ShotPlanner`
- `IAMCCS_MiniMaxH3AtomicModelRouter`
- `IAMCCS_MiniMaxH3AtomicConditioningBackend`
- `IAMCCS_MiniMaxH3GenerationBackendV2`
- `IAMCCS_MiniMaxH3PostUpscaleControlV2`
- `IAMCCS_MiniMaxH3DeliveryRouterV2`
- `IAMCCS_MiniMaxH3SegmentQueueLoop`
- `IAMCCS_MiniMaxH3SequentialLTXLoaderV2`
- `IAMCCS_MiniMaxH3OptionalLTXDetailerLoRA`
- `IAMCCS_MiniMaxH3RTX4KPost`
- `IAMCCS_Prompter` in the Prompter editions
Repository: https://github.com/IAMCCS/IAMCCS-nodes
### ComfyUI-GGUF
Required when either the H3 diffusion model or Qwen3-VL text encoder is loaded
from GGUF. The supplied graphs use `UnetLoaderGGUFAdvanced` and
`CLIPLoaderGGUF`.
Repository: https://github.com/city96/ComfyUI-GGUF
### ComfyUI-KJNodes
Used by the supplied graphs for MiniMax H3 Sage/low-VRAM patches, image resize,
TAEH3 preview override, KJ VAE loading and the LTX spatiotemporal tiled decode.
Repository: https://github.com/kijai/ComfyUI-KJNodes
### MiniMax H3 Turbo
Required only when the Shotboard Turbo route is enabled. It supplies the Turbo
LoRA loader and `MiniMaxH3TurboSampler` used by the reference Turbo graphs.
Repository: https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
## Optional finishing and acceleration packs
- LTX Video nodes: recent ComfyUI contains the native LTX path used by the
current graph. `ComfyUI-LTXVideo` remains useful for compatible extended LTX
workflows: https://github.com/Lightricks/ComfyUI-LTXVideo
- Wan finishing: install the node pack required by the selected Wan branch;
the IAMCCS reference environment uses
https://github.com/kijai/ComfyUI-WanVideoWrapper
- RIFE frame interpolation:
https://github.com/Fannovel16/ComfyUI-Frame-Interpolation
- Sol Attention: https://github.com/kijai/ComfyUI-SolAttn_triton
- Spectrum for MiniMax H3:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
- MiniMax H3 Adaptive Cache:
https://github.com/FFFFFFpy/ComfyUI-MiniMaxH3-AdaptiveCache
These accelerators are alternatives or composable options only where the
Shotboard/backend explicitly reports them as active. A node merely present in
the graph does not accelerate an execution path that is bypassed.
## Optional RTX 4K delivery
The RTX final pass uses `RTXVideoSuperResolution` from:
https://github.com/BetaDoggo/comfyui-rtx-simple
It additionally requires NVIDIA VFX. Install it in the exact Python environment
that launches ComfyUI:
```text
python -m pip install nvidia-vfx --extra-index-url https://pypi.nvidia.com/
```
RTX VSR is an optional final delivery stage. It does not replace the LTX
generative finishing pass, and it must stay bypassed when `RTX final 4K` is off
in the Shotboard settings.
## Model families expected by the workflow
- One MiniMax H3 T2VA/I2VA/FL2VA model and, for REF2VA, the compatible REF2VA
model. Full, pruned INT8 and GGUF variants can be routed when their loader is
compatible with ComfyUI's native H3 model type.
- The matching Qwen3-VL MiniMax H3 text/vision encoder.
- MiniMax H3 video VAE and audio VAE.
- TAEH3 decoder for the live denoise preview.
- Turbo LoRA matching the chosen base model when Turbo is enabled.
- For LTX finishing: the selected LTX diffusion model, Gemma text encoder,
text projection, video VAE and audio VAE. A finishing LoRA is optional; the
Shotboard selector intentionally exposes all compatible installed LTX LoRAs.
- The selected latent/upscale model when the LTX graph uses a latent upres stage.
## Low-VRAM execution contract
The H3 backend uses a phased memory contract:
1. Qwen3-VL conditioning runs GPU-first.
2. On GPUs up to 17 GB, IAMCCS supplies a temporary activation reserve so
ComfyUI dynamically offloads some Qwen weights instead of filling VRAM with
the entire encoder.
3. FL2VA keyframes or REF2VA media are encoded by their VAE.
4. `unload_all_models()`, model cleanup and CUDA cache cleanup run before the H3
sampler requests the diffusion model.
5. A full CPU conditioning retry is used only after a genuine CUDA OOM.
On a cold run, look for a log line containing `dynamic_reserve`, followed by
`conditioning complete`, `pre-sampler barrier`, and only then
`Requested to load MiniMaxH3`. If a warm Queue reuses cached conditioning, the
text-encoder load lines may legitimately be absent.
## Installation validation
After restarting ComfyUI:
1. Open the workflow and confirm there are no red or `UNKNOWN` nodes.
2. Queue a short native H3 take with upscale, RIFE and RTX disabled.
3. Confirm the preview tap updates and the sampler advances.
4. Test LTX/Wan, RIFE and RTX as separate finishing checks before combining
them in a long multi-segment render.
5. Confirm the final saver uses numbered filenames so no previous take is
overwritten.
File diff suppressed because it is too large Load Diff
+490
View File
@@ -0,0 +1,490 @@
# SPDX-FileCopyrightText: 2026 Carmine Cristallo Scalzi (IAMCCS)
# SPDX-License-Identifier: GPL-3.0-or-later
"""MiniMax H3 reference transport for IAMCCS CineLinX.
The node deliberately keeps REF2VA reference media outside the Shotboard
timeline. The Shotboard remains the source of prompt, duration and shot
timing, while this node publishes optional image/video/audio references and
their roles to the isolated H3 backend through the existing CineLinX cable.
"""
from __future__ import annotations
import copy
import json
from typing import Any
import torch
from .iamccs_supernodes_linx import SUPERNODE_LINX_TYPE, build_stage_linx_payload
from .iamccs_minimax_h3_shotboard_core import build_shotplan, plan_json
CATEGORY = "IAMCCS/MiniMax H3"
STAGE_NAME = "iamccs_cine_info_h3"
RESOURCE_PREFIX = "iamccs_minimax_h3_"
IMAGE_ROLES = ["subject_identity", "keyframe", "composition", "style", "disabled"]
VIDEO_ROLES = ["off", "motion_camera", "temporal_structure", "video_edit", "continuation"]
AUDIO_ROLES = ["off", "voice_timbre", "rhythm_timing", "audio_reuse", "sound_reference"]
TASK_OVERRIDES = ["from_shotboard", "t2va", "i2va", "fl2va", "ref2va"]
def _resources(cine_linx: Any) -> dict[str, Any]:
if not isinstance(cine_linx, dict):
return {}
resources = cine_linx.get("resources")
return resources if isinstance(resources, dict) else {}
def _has_h3_plan(cine_linx: Any) -> bool:
resources = _resources(cine_linx)
for key in ("iamccs_minimax_h3_shotplan", "minimax_h3_shotplan", "shotplan"):
value = resources.get(key)
if isinstance(value, dict) and value.get("schema") == "iamccs.minimax_h3.shotplan":
return True
return False
def _h3_plan(cine_linx: Any) -> dict[str, Any]:
resources = _resources(cine_linx)
outputs = cine_linx.get("outputs", {}) if isinstance(cine_linx, dict) else {}
payload = resources.get("cine_payload") if isinstance(resources.get("cine_payload"), dict) else {}
for value in (
resources.get("iamccs_minimax_h3_shotplan"),
resources.get("minimax_h3_shotplan"),
resources.get("shotplan"),
outputs.get("shotplan") if isinstance(outputs, dict) else None,
payload.get("minimax_h3_shotplan"),
payload.get("shotplan"),
):
if isinstance(value, dict) and value.get("schema") == "iamccs.minimax_h3.shotplan":
return value
raise ValueError("IAMCCS Cine Info H3 did not find a valid MiniMax H3 shotplan")
def _json_dict(value: Any) -> dict[str, Any]:
if isinstance(value, dict):
return copy.deepcopy(value)
try:
parsed = json.loads(str(value or "{}"))
except (TypeError, ValueError, json.JSONDecodeError):
return {}
return parsed if isinstance(parsed, dict) else {}
def _routed_timeline(cine_linx: Any) -> dict[str, Any]:
"""Return only a TakeRouter-owned timeline; never infer a different take."""
resources = _resources(cine_linx)
for value in (
resources.get("cine_take_router_timeline_data"),
resources.get("cine_take_router_timeline_json"),
):
timeline = _json_dict(value)
if timeline:
return timeline
return {}
def _copy_runtime_contract(source: dict[str, Any], target: dict[str, Any]) -> None:
"""Preserve non-timeline H3 controls while rebuilding the selected take."""
for key in (
"prompter_injection",
"performance_profile",
"sampling",
"turbo",
"reference_resize",
"upscale_settings",
"control_contract",
"edition",
):
if key in source:
target[key] = copy.deepcopy(source[key])
previous_performance = source.get("performance") if isinstance(source.get("performance"), dict) else {}
max_frames = max(
[int(chunk.get("frame_count", 0) or 0) for chunk in target.get("chunks", []) if isinstance(chunk, dict)]
or [0]
)
width = int(target.get("width", 960) or 960)
height = int(target.get("height", 544) or 544)
performance = copy.deepcopy(previous_performance)
performance["max_chunk_frames"] = max_frames
performance["relative_native_load_vs_960x544x124"] = round(
(float(width) * float(height) * max(1, max_frames)) / (960.0 * 544.0 * 124.0),
3,
)
if performance:
target["performance"] = performance
def _pack_recompiled_plan(
cine_linx: dict[str, Any],
plan: dict[str, Any],
timeline: dict[str, Any],
report: str,
) -> dict[str, Any]:
slots = plan.get("slots") if isinstance(plan.get("slots"), list) else []
chunks = plan.get("chunks") if isinstance(plan.get("chunks"), list) else []
audio_segments = timeline.get("audioSegments")
if not isinstance(audio_segments, list):
audio_segments = timeline.get("audio_segments")
if not isinstance(audio_segments, list):
audio_segments = []
global_prompt = str(plan.get("global_prompt", "") or "")
local_prompts = " | ".join(
str(slot.get("prompt", "")).strip()
for slot in slots
if isinstance(slot, dict) and str(slot.get("prompt", "")).strip()
)
segment_lengths = ",".join(
str(int(chunk.get("frame_count", 0) or 0))
for chunk in chunks
if isinstance(chunk, dict)
)
plan_text = plan_json(plan)
prompt_map_text = json.dumps(plan.get("prompt_map", []), ensure_ascii=False, indent=2)
timeline_text = json.dumps(timeline, ensure_ascii=False)
previous_payload = _resources(cine_linx).get("cine_payload")
payload = copy.deepcopy(previous_payload) if isinstance(previous_payload, dict) else {}
payload.update({
"backend_mode": "minimax_h3_multitimeline",
"pipeline_kind": "minimax_h3",
"global_prompt": global_prompt,
"local_prompts": local_prompts,
"segment_lengths": segment_lengths,
"duration_seconds": float(plan.get("effective_duration_seconds", 0.0) or 0.0),
"effective_duration_seconds": float(plan.get("effective_duration_seconds", 0.0) or 0.0),
"frame_rate": int(plan.get("fps", 24) or 24),
"width": int(plan.get("width", 0) or 0),
"height": int(plan.get("height", 0) or 0),
"timeline_data": copy.deepcopy(timeline),
"visual_segments": copy.deepcopy(slots),
"audioSegments": copy.deepcopy(audio_segments),
"minimax_h3_shotplan": plan,
})
outputs = {
"shotplan": plan,
"shotplan_json": plan_text,
"prompt_map_json": prompt_map_text,
"total_segments": int(plan.get("total_segments", len(chunks)) or 0),
"effective_duration": float(plan.get("effective_duration_seconds", 0.0) or 0.0),
"global_prompt": global_prompt,
"local_prompts": local_prompts,
"segment_lengths": segment_lengths,
"timeline_data": timeline_text,
"audio_timeline_json": json.dumps({"audioSegments": audio_segments}, ensure_ascii=False),
"report": report,
}
resources = {
"cine_payload": payload,
"cine_global_prompt": global_prompt,
"cine_local_prompts": local_prompts,
"cine_segment_lengths": segment_lengths,
"cine_duration_seconds": float(plan.get("effective_duration_seconds", 0.0) or 0.0),
"cine_frame_rate": int(plan.get("fps", 24) or 24),
"cine_width": int(plan.get("width", 0) or 0),
"cine_height": int(plan.get("height", 0) or 0),
"cine_timeline_data_json": timeline_text,
"cine_visual_segments_json": json.dumps(slots, ensure_ascii=False),
"cine_audio_timeline_json": json.dumps({"audioSegments": audio_segments}, ensure_ascii=False),
"iamccs_minimax_h3_shotplan": plan,
"iamccs_minimax_h3_shotplan_json": plan_text,
"iamccs_minimax_h3_prompt_map": copy.deepcopy(plan.get("prompt_map", [])),
"iamccs_minimax_h3_prompt_map_json": prompt_map_text,
"iamccs_minimax_h3_total_segments": int(plan.get("total_segments", len(chunks)) or 0),
"iamccs_minimax_h3_effective_duration": float(plan.get("effective_duration_seconds", 0.0) or 0.0),
"iamccs_minimax_h3_multitimeline_report": report,
}
out = build_stage_linx_payload(
cine_linx,
stage_name="MiniMax H3 MultiTimeline Recompile",
stage_kind="minimax_h3_take_router_recompile",
payload=payload,
report=report,
outputs=outputs,
resources=resources,
policies={
"minimax_h3_timeline_truth": "cine_take_router_timeline_data",
"minimax_h3_take_fallback": "forbidden",
},
downstream_stages=("IAMCCS Cine Info H3", "MiniMax H3 backend"),
requires={"resources": ["cine_take_router_timeline_data", "iamccs_minimax_h3_shotplan"]},
)
out["mode"] = "minimax_h3_multitimeline"
return out
def _recompile_routed_h3_plan(cine_linx: dict[str, Any]) -> tuple[dict[str, Any], str]:
"""Rebuild the H3 plan only when a strict TakeRouter timeline is present."""
timeline = _routed_timeline(cine_linx)
if not timeline:
return cine_linx, "single_timeline"
source = _h3_plan(cine_linx)
global_prompt = str(timeline.get("global_prompt", timeline.get("prompt", source.get("global_prompt", ""))) or "")
duration = float(
timeline.get("duration_seconds", timeline.get("duration", source.get("requested_duration_seconds", 10.0)))
or 10.0
)
rebuilt = build_shotplan(
timeline_data=timeline,
global_prompt=global_prompt,
duration_seconds=duration,
task_mode=str(source.get("task_mode", "auto_from_timeline") or "auto_from_timeline"),
audio_mode=str(source.get("audio_mode", "h3_native_generated") or "h3_native_generated"),
prompt_mapping=str(source.get("prompt_mapping", "global_plus_local") or "global_plus_local"),
upscale_mode=str(source.get("upscale_mode", "off") or "off"),
width=int(source.get("width", 960) or 960),
height=int(source.get("height", 544) or 544),
acceleration=str(source.get("acceleration", "native") or "native"),
ref_image_size=str(source.get("ref_image_size", "match") or "match"),
text_encoder_device=str(source.get("text_encoder_device", "auto") or "auto"),
reference_roles=copy.deepcopy(source.get("reference_roles", [])),
reference_video_role=str(source.get("reference_video_role", "off") or "off"),
reference_audio_role=str(source.get("reference_audio_role", "off") or "off"),
sol_conditioning=str(source.get("sol_conditioning", "exact_kv") or "exact_kv"),
spectrum_profile=str(source.get("spectrum_profile", "conservative_3060") or "conservative_3060"),
vram_clean_before_decode=bool(source.get("vram_clean_before_decode", True)),
rife_mode=str(source.get("rife_mode", "off") or "off"),
upscale_enabled=bool(source.get("upscale_enabled", False)),
)
_copy_runtime_contract(source, rebuilt)
multi = timeline.get("multiGeneration") if isinstance(timeline.get("multiGeneration"), dict) else {}
timeline_id = str(multi.get("activeTimelineId") or timeline.get("activeTimelineId") or "unknown")
take_index = int(multi.get("activeTake") or timeline.get("activeTake") or 0)
rebuilt["multitimeline"] = {
"routed": True,
"timeline_id": timeline_id,
"take_index": take_index,
"source": "IAMCCS_TakeRouter",
"fallback": "forbidden",
}
report = (
"MiniMax H3 MultiTimeline recompiled | "
f"take={take_index} | timeline={timeline_id} | chunks={rebuilt.get('total_segments', 0)} | "
f"duration={float(rebuilt.get('effective_duration_seconds', 0.0)):.3f}s | "
"sampler/Turbo/resolution preserved"
)
return _pack_recompiled_plan(cine_linx, rebuilt, timeline, report), report
def _shape(value: Any) -> list[int]:
if not torch.is_tensor(value):
return []
return [int(item) for item in value.shape]
def _audio_meta(value: Any) -> dict[str, Any]:
if not isinstance(value, dict) or not torch.is_tensor(value.get("waveform")):
return {"connected": False}
waveform = value["waveform"]
sample_rate = int(value.get("sample_rate", 32000) or 32000)
samples = int(waveform.shape[-1]) if waveform.ndim else 0
return {
"connected": True,
"shape": [int(item) for item in waveform.shape],
"sample_rate": sample_rate,
"duration_seconds": round(samples / max(1, sample_rate), 4),
}
def _clean_previous_h3_info(cine_linx: dict[str, Any]) -> dict[str, Any]:
"""Remove a previous H3-info stage without copying tensor payloads."""
cleaned = dict(cine_linx)
resources = dict(_resources(cine_linx))
for key in list(resources):
if key.startswith(RESOURCE_PREFIX) and (
key.startswith(f"{RESOURCE_PREFIX}ref_")
or key in {
f"{RESOURCE_PREFIX}cine_info",
f"{RESOURCE_PREFIX}reference_manifest",
f"{RESOURCE_PREFIX}reference_manifest_json",
}
):
resources.pop(key, None)
cleaned["resources"] = resources
return cleaned
class IAMCCS_CineInfoH3:
"""Attach MiniMax H3 REF2VA media to CineLinX, outside the timeline."""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"cine_linx": (SUPERNODE_LINX_TYPE,),
"task_override": (TASK_OVERRIDES, {"default": "from_shotboard"}),
"reference_role_1": (IMAGE_ROLES, {"default": "subject_identity"}),
"reference_role_2": (IMAGE_ROLES, {"default": "subject_identity"}),
"reference_role_3": (IMAGE_ROLES, {"default": "composition"}),
"reference_role_4": (IMAGE_ROLES, {"default": "style"}),
"reference_video_role": (VIDEO_ROLES, {"default": "off"}),
"reference_audio_role": (AUDIO_ROLES, {"default": "off"}),
"ref_image_size": (["match", "max"], {"default": "match"}),
"reference_resize_policy": (
["canvas_crop", "canvas_pad", "total_pixels", "off"],
{"default": "canvas_crop"},
),
"reference_resize_megapixels": (
"FLOAT",
{"default": 0.5, "min": 0.1, "max": 2.0, "step": 0.05},
),
"reference_resize_filter": (
["area", "bilinear", "bicubic", "nearest-exact"],
{"default": "area"},
),
},
"optional": {
"reference_image_1": ("IMAGE",),
"reference_image_2": ("IMAGE",),
"reference_image_3": ("IMAGE",),
"reference_image_4": ("IMAGE",),
"reference_video": ("IMAGE",),
"reference_video_audio": ("AUDIO",),
"reference_audio": ("AUDIO",),
},
}
RETURN_TYPES = (SUPERNODE_LINX_TYPE,)
RETURN_NAMES = ("cine_linx",)
FUNCTION = "attach"
CATEGORY = CATEGORY
def attach(
self,
cine_linx,
task_override,
reference_role_1,
reference_role_2,
reference_role_3,
reference_role_4,
reference_video_role,
reference_audio_role,
ref_image_size,
reference_resize_policy,
reference_resize_megapixels,
reference_resize_filter,
reference_image_1=None,
reference_image_2=None,
reference_image_3=None,
reference_image_4=None,
reference_video=None,
reference_video_audio=None,
reference_audio=None,
):
if not isinstance(cine_linx, dict):
raise ValueError("IAMCCS Cine Info H3 requires a valid cine_linx input")
if not _has_h3_plan(cine_linx):
raise ValueError(
"IAMCCS Cine Info H3 did not find a MiniMax H3 shotplan. "
"Connect MiniMax H3 Shotboard to IAMCCS CineInfo, then connect its cine_linx output here."
)
cine_linx, timeline_mode = _recompile_routed_h3_plan(cine_linx)
images = [reference_image_1, reference_image_2, reference_image_3, reference_image_4]
roles = [reference_role_1, reference_role_2, reference_role_3, reference_role_4]
image_items = []
for index, (image, role) in enumerate(zip(images, roles), start=1):
image_items.append({
"slot": index,
"label": f"<Picture {index}>",
"role": str(role),
"connected": bool(torch.is_tensor(image)),
"shape": _shape(image),
})
manifest = {
"schema": "iamccs.minimax_h3.cine_info",
"schema_version": 1,
"reference_source": "cine_info_h3_only",
"task_override": str(task_override),
"image_references": image_items,
"video_reference": {
"connected": bool(torch.is_tensor(reference_video)),
"shape": _shape(reference_video),
"role": str(reference_video_role),
"audio": _audio_meta(reference_video_audio),
},
"audio_reference": {
**_audio_meta(reference_audio),
"role": str(reference_audio_role),
},
"ref_image_size": str(ref_image_size),
"reference_resize": {
"policy": str(reference_resize_policy),
"megapixels": float(reference_resize_megapixels),
"filter": str(reference_resize_filter),
"multiple_of": 32,
"downscale_only": True,
},
}
active_images = sum(1 for item in image_items if item["connected"] and item["role"] != "disabled")
active_video = bool(torch.is_tensor(reference_video) and str(reference_video_role) != "off")
active_audio = bool(
(isinstance(reference_audio, dict) and str(reference_audio_role) != "off")
or isinstance(reference_video_audio, dict)
)
manifest["active_reference_count"] = active_images + int(active_video) + int(active_audio)
manifest_json = json.dumps(manifest, ensure_ascii=False, indent=2)
report = (
"IAMCCS Cine Info H3 | references outside Shotboard timeline | "
f"task={task_override} | images={active_images}/4 | video={'on' if active_video else 'off'} | "
f"audio={'on' if active_audio else 'off'} | ref_size={ref_image_size} | "
f"resize={reference_resize_policy}:{float(reference_resize_megapixels):.2f}MP/{reference_resize_filter} | "
f"timeline={timeline_mode}"
)
config = {
"schema": manifest["schema"],
"schema_version": manifest["schema_version"],
"reference_source": manifest["reference_source"],
"task_override": str(task_override),
"reference_roles": [str(item) for item in roles],
"reference_video_role": str(reference_video_role),
"reference_audio_role": str(reference_audio_role),
"ref_image_size": str(ref_image_size),
"reference_resize": dict(manifest["reference_resize"]),
}
base = _clean_previous_h3_info(cine_linx)
resources = {
f"{RESOURCE_PREFIX}cine_info": config,
f"{RESOURCE_PREFIX}reference_manifest": manifest,
f"{RESOURCE_PREFIX}reference_manifest_json": manifest_json,
f"{RESOURCE_PREFIX}ref_image_1": reference_image_1,
f"{RESOURCE_PREFIX}ref_image_2": reference_image_2,
f"{RESOURCE_PREFIX}ref_image_3": reference_image_3,
f"{RESOURCE_PREFIX}ref_image_4": reference_image_4,
f"{RESOURCE_PREFIX}ref_video": reference_video,
f"{RESOURCE_PREFIX}ref_video_audio": reference_video_audio,
f"{RESOURCE_PREFIX}ref_audio": reference_audio,
}
out_linx = build_stage_linx_payload(
base,
stage_name=STAGE_NAME,
stage_kind="minimax_h3_reference_transport",
payload=config,
report=report,
slot_map={"cine_linx": "MiniMax H3 backend cine_linx"},
downstream_stages=("IAMCCS MiniMax H3 Atomic Model Router", "IAMCCS MiniMax H3 Atomic Conditioning"),
policies={
"reference_media_location": "cine_info_h3_not_shotboard_timeline",
"shotboard_owns": "prompt_duration_timeline",
"reference_precedence": "explicit_backend_socket_then_cine_info_h3",
},
outputs={"minimax_h3_reference_manifest_json": manifest_json, "report": report},
resources=resources,
requires={"resources": ["iamccs_minimax_h3_shotplan"]},
)
return (out_linx,)
NODE_CLASS_MAPPINGS = {
"IAMCCS_CineInfoH3": IAMCCS_CineInfoH3,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"IAMCCS_CineInfoH3": "IAMCCS Cine Info H3 - REF2VA Inputs",
}
+296 -45
View File
@@ -408,16 +408,19 @@ def _encode_images(images: torch.Tensor, audio: dict[str, Any] | None, fps: floa
raise RuntimeError("ffmpeg non trovato: impossibile salvare i segmenti MiniMax H3")
if not torch.is_tensor(images) or images.ndim != 4 or images.shape[0] < 1:
raise ValueError("images deve essere un batch IMAGE [T,H,W,C]")
if int(images.shape[-1]) < 3:
raise ValueError(f"images deve avere almeno tre canali RGB, shape ricevuta: {tuple(images.shape)}")
output.parent.mkdir(parents=True, exist_ok=True)
with tempfile.TemporaryDirectory(prefix="minimax_h3_segment_") as temp:
temp_path = Path(temp)
for index, frame in enumerate(images):
_save_frame(temp_path / f"frame_{index:05d}.png", frame)
wav_path = temp_path / "audio.wav"
has_audio = _write_wav(audio, wav_path)
height = int(images.shape[1])
width = int(images.shape[2])
command = [
ffmpeg, "-nostdin", "-n", "-framerate", f"{float(fps):.6f}",
"-i", str(temp_path / "frame_%05d.png"),
ffmpeg, "-hide_banner", "-loglevel", "error", "-nostdin", "-n",
"-f", "rawvideo", "-pix_fmt", "rgb24", "-video_size", f"{width}x{height}",
"-framerate", f"{float(fps):.6f}", "-i", "pipe:0",
]
if has_audio:
command += ["-i", str(wav_path)]
@@ -425,9 +428,42 @@ def _encode_images(images: torch.Tensor, audio: dict[str, Any] | None, fps: floa
if has_audio:
command += ["-c:a", "aac", "-b:a", "192k", "-ar", "48000", "-shortest"]
command += ["-movflags", "+faststart", str(output)]
result = subprocess.run(command, capture_output=True, text=True, stdin=subprocess.DEVNULL)
if result.returncode != 0:
raise RuntimeError(f"ffmpeg segment encode failed: {result.stderr.strip() or result.stdout.strip()}")
process = subprocess.Popen(
command,
stdin=subprocess.PIPE,
stdout=subprocess.DEVNULL,
stderr=subprocess.PIPE,
)
try:
if process.stdin is None:
raise RuntimeError("ffmpeg raw-video stdin non disponibile")
total_frames = int(images.shape[0])
for index, frame in enumerate(images):
rgb = (
frame[..., :3]
.detach()
.to(device="cpu", dtype=torch.float32)
.nan_to_num(nan=0.0, posinf=1.0, neginf=0.0)
.clamp_(0.0, 1.0)
.mul_(255.0)
.round_()
.to(dtype=torch.uint8)
.contiguous()
.numpy()
)
process.stdin.write(rgb.tobytes())
if index == 0 or (index + 1) % 24 == 0 or index + 1 == total_frames:
LOG.info("MiniMax H3 streaming video encode | %d/%d frames", index + 1, total_frames)
process.stdin.close()
error_bytes = process.stderr.read() if process.stderr is not None else b""
return_code = process.wait()
except Exception:
process.kill()
process.wait()
raise
if return_code != 0:
error = error_bytes.decode("utf-8", errors="replace").strip()
raise RuntimeError(f"ffmpeg segment encode failed: {error or f'exit {return_code}'}")
def _concat_videos(paths: list[Path], output: Path) -> None:
@@ -475,7 +511,7 @@ def _next_numbered_render_id(output_folder: Path, base_name: str, requested_rend
escaped_base = re.escape(_safe_name(base_name, "segment"))
escaped_root = re.escape(root)
pattern = re.compile(
rf"^{escaped_base}_{escaped_root}(?:_(\d{{4,}}))?_(?:full|seg_\d{{4,}})\.mp4$",
rf"^{escaped_base}_{escaped_root}(?:_(\d{{4,}}))?_(?:native_)?(?:full|seg_\d{{4,}})\.mp4$",
re.IGNORECASE,
)
used_numbers: set[int] = set()
@@ -495,9 +531,12 @@ def _next_numbered_render_id(output_folder: Path, base_name: str, requested_rend
next_number = max(used_numbers, default=0) + 1
while True:
candidate = f"{root}_{next_number:04d}"
segment_collision = any(output_folder.glob(f"{_safe_name(base_name, 'segment')}_{candidate}_seg_*.mp4"))
final_collision = (output_folder / f"{_safe_name(base_name, 'segment')}_{candidate}_full.mp4").exists()
if not segment_collision and not final_collision:
safe_base = _safe_name(base_name, "segment")
segment_collision = any(output_folder.glob(f"{safe_base}_{candidate}_seg_*.mp4"))
native_segment_collision = any(output_folder.glob(f"{safe_base}_{candidate}_native_seg_*.mp4"))
final_collision = (output_folder / f"{safe_base}_{candidate}_full.mp4").exists()
native_final_collision = (output_folder / f"{safe_base}_{candidate}_native_full.mp4").exists()
if not segment_collision and not native_segment_collision and not final_collision and not native_final_collision:
return candidate
next_number += 1
@@ -567,7 +606,7 @@ class IAMCCS_MiniMaxH3GGUFLoader:
"spectrum_debug": ("BOOLEAN", {"default": False}),
},
"optional": {
"text_encoder_device": (["cpu_safe_12gb", "auto"], {"default": "cpu_safe_12gb"}),
"text_encoder_device": (["auto", "cpu_safe_12gb"], {"default": "auto"}),
},
}
@@ -586,7 +625,7 @@ class IAMCCS_MiniMaxH3GGUFLoader:
acceleration,
spectrum_history,
spectrum_debug,
text_encoder_device="cpu_safe_12gb",
text_encoder_device="auto",
):
if str(unet_name).startswith("NO_") or str(clip_name).startswith("NO_"):
raise FileNotFoundError("GGUF H3 UNET/CLIP non disponibili. Attendi la fine dei download e riavvia ComfyUI.")
@@ -602,15 +641,10 @@ class IAMCCS_MiniMaxH3GGUFLoader:
)[0]
clip_cls = _node_class("CLIPLoaderGGUF")
clip = clip_cls().load_clip(clip_name, type="minimax")[0]
if str(text_encoder_device) == "cpu_safe_12gb":
# Qwen3-VL 32B Q2 is ~8 GiB before temporary dequant buffers.
# Running it on a 12 GiB GPU fails before H3 sampling begins.
# Use ComfyUI's native device retargeter instead of mutating the
# GGUF patcher internals so future model-management changes remain
# compatible.
from comfy_extras.nodes_multigpu import SelectCLIPDeviceNode
clip = SelectCLIPDeviceNode.execute(clip=clip, device="cpu")[0]
requested_text_encoder_device = str(text_encoder_device or "auto").lower()
text_encoder_device = "auto"
if requested_text_encoder_device == "cpu_safe_12gb":
text_encoder_device = "auto(gpu-first; migrated legacy cpu_safe_12gb)"
vae_cls = _node_class("VAELoader")
video_vae = vae_cls().load_vae(video_vae_name)[0]
audio_vae = vae_cls().load_vae(audio_vae_name)[0]
@@ -660,9 +694,27 @@ class IAMCCS_MiniMaxH3ShotPlanner:
name for name in folder_paths.get_filename_list("loras")
if "minimax" in name.lower() and "h3" in name.lower() and "turbo" in name.lower()
]
# The LTX finishing slot intentionally exposes the complete ComfyUI
# LoRA registry. Some useful LTX LoRAs (including community Crisp
# variants) do not carry reliable "ltx/detail/enhance" tokens in the
# filename or parent folder. Runtime compatibility remains the user's
# choice; keeping the historical field name preserves old workflows.
installed_ltx_detailer_loras = list(folder_paths.get_filename_list("loras"))
crisp_ltx_loras = sorted(
(
name for name in installed_ltx_detailer_loras
if "ltx" in name.lower() and "crisp" in name.lower()
),
key=lambda name: (
0 if Path(name).name.lower() == "ltx2.3_crisp_enhance.safetensors" else 1,
name.lower(),
),
)
preferred_ltx_detailer_lora = crisp_ltx_loras[0] if crisp_ltx_loras else ""
# Only expose files that really exist. The web migration converts old
# saved missing filenames to the empty/Base-H3 choice before validation.
turbo_loras = list(dict.fromkeys(("", *installed_turbo_loras)))
ltx_detailer_loras = list(dict.fromkeys(("", *installed_ltx_detailer_loras)))
if "res_multistep" in samplers:
samplers.remove("res_multistep")
samplers.insert(0, "res_multistep")
@@ -700,7 +752,11 @@ class IAMCCS_MiniMaxH3ShotPlanner:
"img_compression": ("INT", {"default": 0, "min": 0, "max": 100, "step": 1}),
# H3-native backend controls. The dedicated Shotboard UI
# renders these above the timeline and hides the raw widgets.
"acceleration": (["auto_3060", "native", "h3_sage", "sage", "sage_sol", "spectrum", "sage_spectrum"], {"default": "auto_3060"}),
"acceleration": ([
"low_vram_auto", "native", "h3_sage", "sol_low_vram", "sol_adaptive_safe",
"sol_adaptive_balanced", "adaptive_safe", "spectrum", "sage_spectrum",
"auto_3060", "sage", "sage_sol",
], {"default": "low_vram_auto"}),
"ref_image_size": (["match", "max"], {"default": "match"}),
"reference_role_1": (["subject_identity", "keyframe", "composition", "style", "disabled"], {"default": "subject_identity"}),
"reference_role_2": (["subject_identity", "keyframe", "composition", "style", "disabled"], {"default": "subject_identity"}),
@@ -708,8 +764,8 @@ class IAMCCS_MiniMaxH3ShotPlanner:
"reference_role_4": (["subject_identity", "keyframe", "composition", "style", "disabled"], {"default": "style"}),
"reference_video_role": (["off", "motion_camera", "temporal_structure", "video_edit", "continuation"], {"default": "off"}),
"reference_audio_role": (["off", "voice_timbre", "rhythm_timing", "audio_reuse", "sound_reference"], {"default": "off"}),
"sol_conditioning": (["exact_kv", "exact_kv_and_rows"], {"default": "exact_kv"}),
"spectrum_profile": (["conservative_3060", "conservative_quality", "aggressive"], {"default": "conservative_3060"}),
"sol_conditioning": (["exact_kv_and_rows", "exact_kv"], {"default": "exact_kv_and_rows"}),
"spectrum_profile": (["low_vram", "quality", "aggressive", "conservative_3060", "conservative_quality"], {"default": "low_vram"}),
"vram_clean_before_decode": ("BOOLEAN", {"default": True}),
"rife_mode": (["off", "rife_48fps", "rife_60fps"], {"default": "off"}),
"upscale_enabled": ("BOOLEAN", {"default": False}),
@@ -723,14 +779,16 @@ class IAMCCS_MiniMaxH3ShotPlanner:
"upscale_sage": ("BOOLEAN", {"default": True}),
"upscale_seed_offset": ("INT", {"default": 10000, "min": 0, "max": 0xFFFFFFFFFFFFFFFF, "step": 1}),
"wan_upscale_denoise": ("FLOAT", {"default": 0.2, "min": 0.0, "max": 1.0, "step": 0.01}),
# Qwen3-VL 32B GGUF temporarily expands quantized tensors while
# encoding. A 12 GiB GPU cannot hold the full encoder plus its
# dequantization buffers, so the atomic backend defaults to CPU.
"text_encoder_device": (["cpu_safe_12gb", "auto"], {"default": "cpu_safe_12gb"}),
# ComfyUI automatic placement is GPU-first. The atomic backend
# retries on CPU only after a genuine CUDA out-of-memory error.
"text_encoder_device": (["auto", "cpu_safe_12gb"], {"default": "auto"}),
# These values are intentionally owned by the Shotboard. The
# generation node keeps legacy widgets only as a compatibility
# fallback for shotplans created before schema v3.
"performance_profile": (["rtx3060_draft", "rtx3060_balanced", "rtx3060_turbo", "h3_native_quality", "custom"], {"default": "rtx3060_balanced"}),
"performance_profile": ([
"low_vram_draft", "low_vram_balanced", "low_vram_turbo", "h3_native_quality", "custom",
"rtx3060_draft", "rtx3060_balanced", "rtx3060_turbo",
], {"default": "low_vram_balanced"}),
"seed": ("INT", {"default": 42, "min": 0, "max": 0xFFFFFFFFFFFFFFFF}),
"seed_stride": ("INT", {"default": 1, "min": 0, "max": 0xFFFFFFFFFFFFFFFF, "step": 1}),
"steps": ("INT", {"default": 16, "min": 1, "max": 100, "step": 1}),
@@ -751,6 +809,18 @@ class IAMCCS_MiniMaxH3ShotPlanner:
"reference_resize_policy": (["canvas_crop", "canvas_pad", "total_pixels", "off"], {"default": "canvas_crop"}),
"reference_resize_megapixels": ("FLOAT", {"default": 0.5, "min": 0.1, "max": 2.0, "step": 0.05}),
"reference_resize_filter": (["area", "bilinear", "bicubic", "nearest-exact"], {"default": "area"}),
# Optional LTX finishing controls. All installed LoRAs remain
# selectable; an installed LTX Crisp variant is preselected.
# The enable switch still owns whether the LoRA is applied.
"ltx_detailer_enabled": ("BOOLEAN", {"default": False}),
"ltx_detailer_lora_name": (
ltx_detailer_loras,
{"default": preferred_ltx_detailer_lora},
),
"ltx_detailer_strength": ("FLOAT", {"default": 0.6, "min": 0.0, "max": 2.0, "step": 0.05}),
"ltx_4k_enabled": ("BOOLEAN", {"default": False}),
"ltx_4k_quality": (["ULTRA", "HIGH", "MEDIUM", "LOW"], {"default": "ULTRA"}),
"ltx_seam_safe": ("BOOLEAN", {"default": True}),
},
"optional": {
"cine_linx": (
@@ -791,7 +861,7 @@ class IAMCCS_MiniMaxH3ShotPlanner:
image_resize_method="crop",
image_multiple_of=32,
img_compression=0,
acceleration="auto_3060",
acceleration="low_vram_auto",
ref_image_size="match",
reference_role_1="subject_identity",
reference_role_2="subject_identity",
@@ -799,8 +869,8 @@ class IAMCCS_MiniMaxH3ShotPlanner:
reference_role_4="style",
reference_video_role="off",
reference_audio_role="off",
sol_conditioning="exact_kv",
spectrum_profile="conservative_3060",
sol_conditioning="exact_kv_and_rows",
spectrum_profile="low_vram",
vram_clean_before_decode=True,
rife_mode="off",
upscale_enabled=False,
@@ -810,8 +880,8 @@ class IAMCCS_MiniMaxH3ShotPlanner:
upscale_sage=True,
upscale_seed_offset=10000,
wan_upscale_denoise=0.2,
text_encoder_device="cpu_safe_12gb",
performance_profile="rtx3060_balanced",
text_encoder_device="auto",
performance_profile="low_vram_balanced",
seed=42,
seed_stride=1,
steps=16,
@@ -827,6 +897,12 @@ class IAMCCS_MiniMaxH3ShotPlanner:
reference_resize_policy="canvas_crop",
reference_resize_megapixels=0.5,
reference_resize_filter="area",
ltx_detailer_enabled=False,
ltx_detailer_lora_name="",
ltx_detailer_strength=0.6,
ltx_4k_enabled=False,
ltx_4k_quality="ULTRA",
ltx_seam_safe=True,
cine_linx=None,
):
global_prompt, timeline_data, prompter_injection = apply_prompter_to_minimax(
@@ -839,6 +915,14 @@ class IAMCCS_MiniMaxH3ShotPlanner:
image_width = _h3_legal_dimension(image_width, width)
image_height = _h3_legal_dimension(image_height, height)
reference_resize_megapixels = _finite_float(reference_resize_megapixels, 0.5, 0.1, 2.0)
selected_ltx_detailer = str(ltx_detailer_lora_name or "").strip()
ltx_detailer_requested = bool(ltx_detailer_enabled)
ltx_detailer_available = _model_file_available("loras", selected_ltx_detailer)
effective_ltx_detailer = ltx_detailer_requested and ltx_detailer_available
effective_ltx_4k = bool(ltx_4k_enabled) and bool(upscale_enabled) and str(upscale_mode) == "ltx23"
ltx_4k_quality = str(ltx_4k_quality or "ULTRA").upper()
if ltx_4k_quality not in {"ULTRA", "HIGH", "MEDIUM", "LOW"}:
ltx_4k_quality = "ULTRA"
requested_turbo_mode = str(turbo_mode or "off")
selected_turbo_lora = str(turbo_lora_name or "").strip()
@@ -912,17 +996,39 @@ class IAMCCS_MiniMaxH3ShotPlanner:
"sage": bool(upscale_sage),
"seed_offset": int(upscale_seed_offset),
"wan_denoise": float(wan_upscale_denoise),
"ltx_detailer_requested": ltx_detailer_requested,
"ltx_detailer_enabled": effective_ltx_detailer,
"ltx_detailer_available": ltx_detailer_available,
"ltx_detailer_lora_name": selected_ltx_detailer,
"ltx_detailer_strength": float(ltx_detailer_strength),
"ltx_4k_enabled": effective_ltx_4k,
"ltx_4k_quality": ltx_4k_quality,
"ltx_seam_safe": bool(ltx_seam_safe),
"ltx_vae_encode_temporal_size": 500 if bool(ltx_seam_safe) else 64,
"ltx_vae_encode_temporal_overlap": 4 if bool(ltx_seam_safe) else 8,
"ltx_vae_decode_temporal_size": 64 if bool(ltx_seam_safe) else 16,
"ltx_vae_decode_temporal_overlap": 4 if bool(ltx_seam_safe) else 1,
"ltx_vae_decode_spatial_overlap": 4 if bool(ltx_seam_safe) else 1,
"source": "shotboard",
}
chunk_frames = [int(chunk.get("frame_count", 0) or 0) for chunk in plan.get("chunks", [])]
max_chunk_frames = max(chunk_frames, default=0)
native_load = (float(width) * float(height) * max(1, max_chunk_frames)) / (960.0 * 544.0 * 124.0)
warnings: list[str] = []
if str(performance_profile).startswith("rtx3060") and max_chunk_frames > 124:
if ltx_detailer_requested and not ltx_detailer_available:
warnings.append(f"Optional LTX detailer unavailable ({selected_ltx_detailer or 'no LoRA selected'}); continuing without it")
if bool(ltx_4k_enabled) and not effective_ltx_4k:
warnings.append("RTX VSR 4K is available only when LTX 2.3 upscale is enabled")
if effective_ltx_4k:
warnings.append("4K delivery uses LTX at half delivery resolution, then NVIDIA RTX VSR 2x; expect high system-RAM usage")
if str(text_encoder_device).lower() == "cpu_safe_12gb":
warnings.append("Legacy CPU-safe text encoder setting migrated to GPU-first auto with CPU fallback only after CUDA OOM")
low_vram_profile = str(performance_profile).startswith(("low_vram", "rtx3060"))
if low_vram_profile and max_chunk_frames > 124:
warnings.append("Low VRAM: trim this timeline box to 124 frames or less; use a following box for continuation")
if str(performance_profile).startswith("rtx3060") and int(width) * int(height) > 960 * 544:
if low_vram_profile and int(width) * int(height) > 960 * 544:
warnings.append("Low VRAM: generate at 960x544 or below, then upscale for a 1280-class delivery")
if str(acceleration) == "sage_sol":
if str(acceleration) in {"sage_sol", "sol_low_vram", "sol_adaptive_safe", "sol_adaptive_balanced"}:
warnings.append("Sol-Attn is experimental, has a slower first compile, and is not validated for every Low VRAM configuration")
if str(acceleration) in {"spectrum", "sage_spectrum"} and effective_steps < 14:
warnings.append("Spectrum saves few transformer calls below 14 steps because warmup and final native steps remain mandatory")
@@ -936,6 +1042,8 @@ class IAMCCS_MiniMaxH3ShotPlanner:
warnings.append("Early/non-ckpt500 Turbo is normally used at 8-10 steps")
if effective_turbo_mode == "ckpt500_6_8" and not 6 <= effective_steps <= 8:
warnings.append("Turbo ckpt500 is normally used at 6-8 steps")
if str(acceleration) in {"adaptive_safe", "sol_adaptive_safe", "sol_adaptive_balanced"}:
warnings.append("Adaptive Cache is approximate; use Safe for faces, hands, dialogue and lip sync")
if turbo_enabled and str(acceleration) in {"spectrum", "sage_spectrum"}:
warnings.append("Spectrum has little room to forecast at Turbo step counts; Sage-only is the Low VRAM default")
if turbo_enabled and str(turbo_sampler_mode) == "res_multistep_stock" and int(steps) < 10:
@@ -951,7 +1059,7 @@ class IAMCCS_MiniMaxH3ShotPlanner:
"conditioning": ["width", "height", "timeline trim", "prompt mapping", "references", "audio mode"],
"sampling": ["seed", "steps", "sampler", "scheduler", "denoise", "H3 shifts", "acceleration", "Turbo LoRA", "Turbo audio sampler"],
"reference_preprocess": ["resize policy", "target megapixels", "filter", "multiple of 32"],
"delivery": ["VRAM clean", "RIFE", "upscale enabled", "upscale mode", "upscale target", "upscale prompt", "upscale seed"],
"delivery": ["VRAM clean", "RIFE", "upscale enabled", "upscale mode", "upscale target", "upscale prompt", "upscale seed", "LTX seam-safe VAE", "LTX detailer LoRA", "optional RTX VSR 4K"],
"transport": "one IAMCCS_SUPERNODE_LINX cable; the private H3 plan stays inside CineLinX",
}
injection_summary = str(prompter_injection.get("actual_target", "none")) if prompter_injection.get("applied") else "none"
@@ -963,9 +1071,11 @@ class IAMCCS_MiniMaxH3ShotPlanner:
f"load={native_load:.2f}x | sampler={effective_steps}x{sampler_name}+{scheduler} | acceleration={acceleration} | "
f"turbo={effective_turbo_mode}:{selected_turbo_lora or 'none'}@{float(turbo_strength):.2f}/{turbo_sampler_mode} | "
f"ref_resize={reference_resize_policy}:{reference_resize_megapixels:.2f}MP/{reference_resize_filter} | "
f"ref_size={ref_image_size} | text_encoder={text_encoder_device} | "
f"ref_size={ref_image_size} | text_encoder={plan.get('text_encoder_device', 'auto')} | "
f"RIFE={rife_mode} | upscale={'on' if upscale_enabled else 'off'}:{plan['upscale_mode']} "
f"->{int(upscale_width)}x{int(upscale_height)} sage={'on' if upscale_sage else 'off'} "
f"ltx_detailer={'on' if effective_ltx_detailer else 'off'}:{selected_ltx_detailer or 'none'}@{float(ltx_detailer_strength):.2f} "
f"ltx_seam_safe={'on' if ltx_seam_safe else 'off'} ltx_4k={'on' if effective_ltx_4k else 'off'}:{ltx_4k_quality} "
f"wan_denoise={float(wan_upscale_denoise):.2f} | "
f"prompter={injection_summary} | warnings={'; '.join(warnings) if warnings else 'none'}"
)
@@ -1398,6 +1508,133 @@ class IAMCCS_MiniMaxH3BridgeLoad:
raise FileNotFoundError(f"MiniMax H3 bridge non trovato: {bridge_path}")
class IAMCCS_MiniMaxH3NativeCheckpointSave:
"""Persist the native H3 result before any optional upscale branch.
The node is deliberately a pass-through dependency. Downstream LTX, Wan,
or RTX processing cannot begin until the native segment has been encoded,
so an upscale failure never discards the expensive H3 render.
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"images": ("IMAGE",),
"audio": ("AUDIO",),
"current_segment": ("INT", {"forceInput": True}),
"total_segments": ("INT", {"forceInput": True}),
"fps": ("INT", {"forceInput": True}),
"trim_head_frames": ("INT", {"forceInput": True}),
},
"optional": {
"filename_prefix": ("STRING", {"default": "IAMCCS/MiniMaxH3/segment"}),
"merge_segments": ("BOOLEAN", {"default": True}),
"keep_segments": ("BOOLEAN", {"default": True}),
"render_id": ("STRING", {"default": "minimax_h3_render"}),
},
}
RETURN_TYPES = ("IMAGE", "AUDIO", "STRING", "STRING")
RETURN_NAMES = ("native_frames", "native_audio", "resolved_render_id", "report")
FUNCTION = "checkpoint"
OUTPUT_NODE = True
CATEGORY = CATEGORY
@classmethod
def IS_CHANGED(cls, *args, **kwargs):
# Saving is intentional on every queued render, even when ComfyUI can
# reuse the surrounding graph cache.
return float("nan")
def checkpoint(
self,
images,
audio,
current_segment,
total_segments,
fps,
trim_head_frames,
filename_prefix="IAMCCS/MiniMaxH3/segment",
merge_segments=True,
keep_segments=True,
render_id="minimax_h3_render",
):
if not torch.is_tensor(images) or images.ndim != 4 or int(images.shape[0]) < 1:
raise ValueError("MiniMax H3 native checkpoint expects a non-empty IMAGE frame batch")
if not isinstance(audio, dict):
raise ValueError("MiniMax H3 native checkpoint expects the H3 AUDIO output")
current_segment = int(current_segment)
total_segments = int(total_segments)
fps = max(1, int(fps))
if current_segment < 0 or total_segments < 1 or current_segment >= total_segments:
raise ValueError(f"Native checkpoint segment index is invalid: {current_segment + 1}/{total_segments}")
output_folder, base_name = _output_location(filename_prefix)
requested_render_id = _safe_name(str(render_id or "").strip(), "minimax_h3_render")
active_render_id = (
_next_numbered_render_id(output_folder, base_name, requested_render_id)
if current_segment == 0
else requested_render_id
)
trim_count = max(0, int(trim_head_frames or 0))
images_to_save = images
audio_to_save = audio
if trim_count and int(images.shape[0]) > trim_count:
images_to_save = images[trim_count:, ...]
audio_to_save = _trim_audio_frames(audio, trim_count, fps)
segment_name = f"{base_name}_{active_render_id}_native_seg_{current_segment + 1:04d}.mp4"
segment_path = output_folder / segment_name
_require_new_output_path(segment_path)
_encode_images(images_to_save, audio_to_save, fps, segment_path)
messages = [f"Native checkpoint saved: {segment_name}"]
preview_path = segment_path
if current_segment + 1 >= total_segments and bool(merge_segments):
segment_paths = [
output_folder / f"{base_name}_{active_render_id}_native_seg_{index + 1:04d}.mp4"
for index in range(total_segments)
]
final_name = f"{base_name}_{active_render_id}_native_full.mp4"
final_path = output_folder / final_name
_require_new_output_path(final_path)
_concat_videos(segment_paths, final_path)
preview_path = final_path
messages.append(f"Native full video saved: {final_name}")
if not bool(keep_segments):
for path in segment_paths:
path.unlink(missing_ok=True)
LOG.info(
"MiniMax H3 native checkpoint complete | render=%s | segment=%d/%d | fps=%d",
active_render_id,
current_segment + 1,
total_segments,
fps,
)
subfolder = os.path.relpath(
preview_path.parent,
folder_paths.get_output_directory(),
).replace("\\", "/")
preview = {
"filename": preview_path.name,
"subfolder": "" if subfolder == "." else subfolder,
"type": "output",
}
report = " | ".join(messages)
return {
"ui": {
"text": messages,
"images": [preview],
"animated": (True,),
},
"result": (images, audio, active_render_id, report),
}
class IAMCCS_MiniMaxH3SegmentQueueLoop:
@classmethod
def INPUT_TYPES(cls):
@@ -1418,6 +1655,7 @@ class IAMCCS_MiniMaxH3SegmentQueueLoop:
"keep_segments": ("BOOLEAN", {"default": True}),
"render_id": ("STRING", {"default": "minimax_h3_render"}),
"segment_base_name": ("STRING", {"default": ""}),
"resolved_render_id": ("STRING", {"forceInput": True}),
},
"hidden": {"prompt": "PROMPT", "unique_id": "UNIQUE_ID", "extra_pnginfo": "EXTRA_PNGINFO"},
}
@@ -1442,6 +1680,7 @@ class IAMCCS_MiniMaxH3SegmentQueueLoop:
keep_segments=True,
render_id="minimax_h3_render",
segment_base_name="",
resolved_render_id="",
prompt=None,
unique_id=None,
extra_pnginfo=None,
@@ -1464,11 +1703,15 @@ class IAMCCS_MiniMaxH3SegmentQueueLoop:
)
output_folder, resolved_base_name = _output_location(filename_prefix)
active_base_name = _safe_name(str(segment_base_name or "").strip(), resolved_base_name)
active_render_id = (
_next_numbered_render_id(output_folder, active_base_name, requested_render_id)
if current_segment == 0
else requested_render_id
)
locked_render_id = str(resolved_render_id or "").strip()
if locked_render_id:
active_render_id = _safe_name(locked_render_id, requested_render_id)
else:
active_render_id = (
_next_numbered_render_id(output_folder, active_base_name, requested_render_id)
if current_segment == 0
else requested_render_id
)
if current_segment == 0:
LOG.info("MiniMax H3 nuovo render numerato: %s", active_render_id)
segment_name = f"{active_base_name}_{active_render_id}_seg_{current_segment + 1:04d}.mp4"
@@ -1708,6 +1951,7 @@ NODE_CLASS_MAPPINGS = {
"IAMCCS_MiniMaxH3Backend": IAMCCS_MiniMaxH3Backend,
"IAMCCS_MiniMaxH3RenderBackend": IAMCCS_MiniMaxH3RenderBackend,
"IAMCCS_MiniMaxH3BridgeLoad": IAMCCS_MiniMaxH3BridgeLoad,
"IAMCCS_MiniMaxH3NativeCheckpointSave": IAMCCS_MiniMaxH3NativeCheckpointSave,
"IAMCCS_MiniMaxH3SegmentQueueLoop": IAMCCS_MiniMaxH3SegmentQueueLoop,
"IAMCCS_MiniMaxH3AudioConcat": IAMCCS_MiniMaxH3AudioConcat,
"IAMCCS_MiniMaxH3AudioPolicy": IAMCCS_MiniMaxH3AudioPolicy,
@@ -1724,6 +1968,7 @@ NODE_DISPLAY_NAME_MAPPINGS = {
"IAMCCS_MiniMaxH3Backend": "MiniMax H3 Shotboard Backend",
"IAMCCS_MiniMaxH3RenderBackend": "MiniMax H3 Render Backend (Sampler + AV Decode)",
"IAMCCS_MiniMaxH3BridgeLoad": "MiniMax H3 Last-Frame Bridge",
"IAMCCS_MiniMaxH3NativeCheckpointSave": "MiniMax H3 Native Checkpoint Save",
"IAMCCS_MiniMaxH3SegmentQueueLoop": "MiniMax H3 Segment Queue + Concat",
"IAMCCS_MiniMaxH3AudioConcat": "MiniMax H3 Audio Chunk Concat",
"IAMCCS_MiniMaxH3AudioPolicy": "MiniMax H3 Audio Policy",
@@ -1738,6 +1983,12 @@ from .iamccs_minimax_h3_atomic_backend import (
NODE_CLASS_MAPPINGS as _ATOMIC_NODE_CLASS_MAPPINGS,
NODE_DISPLAY_NAME_MAPPINGS as _ATOMIC_NODE_DISPLAY_NAME_MAPPINGS,
)
from .iamccs_minimax_h3_cine_info import (
NODE_CLASS_MAPPINGS as _CINE_INFO_H3_NODE_CLASS_MAPPINGS,
NODE_DISPLAY_NAME_MAPPINGS as _CINE_INFO_H3_NODE_DISPLAY_NAME_MAPPINGS,
)
NODE_CLASS_MAPPINGS.update(_ATOMIC_NODE_CLASS_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(_ATOMIC_NODE_DISPLAY_NAME_MAPPINGS)
NODE_CLASS_MAPPINGS.update(_CINE_INFO_H3_NODE_CLASS_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(_CINE_INFO_H3_NODE_DISPLAY_NAME_MAPPINGS)
+152 -14
View File
@@ -267,6 +267,118 @@ def _normalise_slots(timeline: dict[str, Any], duration_seconds: float, fallback
]
def _timeline_h3_bridges(timeline: dict[str, Any]) -> list[dict[str, Any]]:
"""Return the dedicated MiniMax bridge contract, when the UI supplied it."""
for key in ("h3_bridges", "h3Bridges"):
value = timeline.get(key)
if isinstance(value, list):
return [dict(item) for item in value if isinstance(item, dict)]
nested = timeline.get("timeline")
if isinstance(nested, dict):
return _timeline_h3_bridges(nested)
return []
def _normalise_flf_bridge_slots(
timeline: dict[str, Any],
slots: list[dict[str, Any]],
duration_seconds: float,
) -> list[dict[str, Any]]:
"""Convert N image anchors into N-1 MiniMax first/last-frame chunks.
The Shotboard renders the local prompt from the centre of one image box to
the centre of the next. Those centre distances determine the *relative*
duration of the FLF chunks, while the first and last centres are normalised
to the full requested timeline duration. Consequently two image anchors
on a ten-second board still produce one ten-second FLF chunk; with three or
more anchors, resizing or moving a box changes the proportional timing of
the adjacent chunks without losing the requested total duration.
"""
anchors = [slot for slot in slots if _text(slot.get("image"))]
if len(anchors) < 2:
return slots
ui_bridges = _timeline_h3_bridges(timeline)
bridge_by_pair: dict[tuple[str, str], dict[str, Any]] = {}
for bridge in ui_bridges:
pair = (
_text(_first_value(bridge, ("from_segment_id", "fromSegmentId", "from_id"))),
_text(_first_value(bridge, ("to_segment_id", "toSegmentId", "to_id"))),
)
if pair[0] and pair[1]:
bridge_by_pair[pair] = bridge
centres = [
float(slot["start_seconds"]) + float(slot["requested_frame_count"]) / H3_FPS / 2.0
for slot in anchors
]
gaps = [max(1.0 / H3_FPS, centres[index + 1] - centres[index]) for index in range(len(centres) - 1)]
gap_total = sum(gaps) or float(len(gaps))
requested_total = max(H3_MIN_FRAMES, int(round(max(0.01, _float(duration_seconds, 10.0)) * H3_FPS)))
if requested_total > H3_MAX_TRAINED_FRAMES * len(gaps):
raise ValueError(
f"La timeline FLF richiede {requested_total} frame ma {len(gaps)} ponti H3 possono contenerne "
f"al massimo {H3_MAX_TRAINED_FRAMES * len(gaps)}. Aggiungi keyframe o riduci la durata."
)
raw_lengths = [requested_total * gap / gap_total for gap in gaps]
requested_lengths = [max(H3_MIN_FRAMES, int(math.floor(value))) for value in raw_lengths]
remainder = requested_total - sum(requested_lengths)
order = sorted(
range(len(raw_lengths)),
key=lambda index: raw_lengths[index] - math.floor(raw_lengths[index]),
reverse=remainder > 0,
)
step = 1 if remainder > 0 else -1
for offset in range(abs(remainder)):
index = order[offset % len(order)]
if step < 0 and requested_lengths[index] <= H3_MIN_FRAMES:
continue
requested_lengths[index] += step
bridge_slots: list[dict[str, Any]] = []
cursor = 0.0
for index, (first, last) in enumerate(zip(anchors, anchors[1:])):
requested_frames = requested_lengths[index]
if requested_frames > H3_MAX_TRAINED_FRAMES:
raise ValueError(
f"Il ponte FLF '{first['label']} -> {last['label']}' richiede {requested_frames} frame: "
f"avvicina i centri dei box o aggiungi un keyframe (massimo {H3_MAX_TRAINED_FRAMES})."
)
frame_count = align_h3_frames(requested_frames)
if frame_count > H3_MAX_TRAINED_FRAMES:
raise ValueError(
f"Il ponte FLF '{first['label']} -> {last['label']}' diventa {frame_count} frame dopo "
f"l'allineamento H3 17k+5: riduci leggermente la durata relativa del ponte."
)
ui_bridge = bridge_by_pair.get((_text(first.get("id")), _text(last.get("id"))), {})
local_prompt = _text(_first_value(ui_bridge, ("prompt", "local_prompt", "relay_prompt"))) or _text(first.get("prompt"))
audio_prompt = _text(_first_value(ui_bridge, ("audio_prompt", "sound_prompt"))) or _text(first.get("audio_prompt"))
bridge_slots.append(
{
"id": _text(ui_bridge.get("id")) or f"flf_bridge_{index + 1}",
"label": _text(ui_bridge.get("label")) or f"{first['label']} -> {last['label']}",
"type": "image",
"start_seconds": cursor,
"requested_frame_count": requested_frames,
"frame_count": frame_count,
"duration_seconds": frame_count / H3_FPS,
"image": _text(first.get("image")),
"explicit_last_image": _text(last.get("image")),
"prompt": local_prompt,
"audio_prompt": audio_prompt,
"transition": "start" if index == 0 else "h3_keyframe_chain",
"use_keyframe": True,
"from_anchor_id": _text(first.get("id")),
"to_anchor_id": _text(last.get("id")),
"visual_start_frame": int(round(centres[index] * H3_FPS)),
"visual_end_frame": int(round(centres[index + 1] * H3_FPS)),
}
)
cursor += frame_count / H3_FPS
return bridge_slots
def _compose_prompt(
*,
global_prompt: str,
@@ -341,12 +453,12 @@ def build_shotplan(
height: int = 768,
acceleration: str = "native",
ref_image_size: str = "match",
text_encoder_device: str = "cpu_safe_12gb",
text_encoder_device: str = "auto",
reference_roles: list[str] | tuple[str, ...] | None = None,
reference_video_role: str = "off",
reference_audio_role: str = "off",
sol_conditioning: str = "exact_kv",
spectrum_profile: str = "conservative_3060",
sol_conditioning: str = "exact_kv_and_rows",
spectrum_profile: str = "low_vram",
vram_clean_before_decode: bool = True,
rife_mode: str = "off",
upscale_enabled: bool = False,
@@ -382,19 +494,26 @@ def build_shotplan(
raise ValueError("aspect ratio H3 deve essere compreso tra 2:5 e 5:2")
acceleration = _text(acceleration).lower() or "native"
if acceleration not in {"auto_3060", "native", "h3_sage", "sage", "sage_sol", "spectrum", "sage_spectrum"}:
if acceleration not in {
"auto_3060", "low_vram_auto", "native", "h3_sage", "sage", "sage_sol", "sol_low_vram",
"adaptive_safe", "sol_adaptive_safe", "sol_adaptive_balanced", "spectrum", "sage_spectrum",
}:
raise ValueError(f"accelerazione H3 non valida: {acceleration}")
ref_image_size = _text(ref_image_size).lower() or "match"
if ref_image_size not in {"match", "max"}:
raise ValueError(f"ref_image_size H3 non valido: {ref_image_size}")
text_encoder_device = _text(text_encoder_device).lower() or "cpu_safe_12gb"
text_encoder_device = _text(text_encoder_device).lower() or "auto"
if text_encoder_device not in {"cpu_safe_12gb", "auto"}:
raise ValueError(f"device text encoder H3 non valido: {text_encoder_device}")
sol_conditioning = _text(sol_conditioning).lower() or "exact_kv"
# Old boards remain loadable, but CPU is now an OOM-only fallback handled
# by the atomic conditioning backend rather than a forced placement mode.
if text_encoder_device == "cpu_safe_12gb":
text_encoder_device = "auto"
sol_conditioning = _text(sol_conditioning).lower() or "exact_kv_and_rows"
if sol_conditioning not in {"exact_kv", "exact_kv_and_rows"}:
raise ValueError(f"Sol-Attn conditioning non valido: {sol_conditioning}")
spectrum_profile = _text(spectrum_profile).lower() or "conservative_3060"
if spectrum_profile not in {"conservative_3060", "conservative_quality", "aggressive"}:
spectrum_profile = _text(spectrum_profile).lower() or "low_vram"
if spectrum_profile not in {"conservative_3060", "low_vram", "conservative_quality", "quality", "aggressive"}:
raise ValueError(f"profilo Spectrum non valido: {spectrum_profile}")
rife_mode = _text(rife_mode).lower() or "off"
if rife_mode not in {"off", "rife_48fps", "rife_60fps"}:
@@ -413,6 +532,20 @@ def build_shotplan(
fallback_duration = min(H3_MAX_TRAINED_FRAMES / H3_FPS, max(H3_MIN_FRAMES / H3_FPS, 10.0))
slots = _normalise_slots(timeline, duration_seconds, fallback_duration)
requested_task_mode = _text(task_mode).lower() or "auto_from_timeline"
auto_task_mode = requested_task_mode in {"auto", "auto_from_timeline"}
explicit_flf_mode = requested_task_mode in {"flf", "fflf", "fl2va"}
explicit_i2v_mode = requested_task_mode in {"i2v", "i2va"}
image_slots = [slot for slot in slots if _text(slot.get("image"))]
legacy_explicit_last = len(image_slots) == 1 and bool(_text(image_slots[0].get("explicit_last_image")))
flf_anchor_mode = bool(
(explicit_flf_mode and len(image_slots) >= 2)
or (auto_task_mode and len(image_slots) >= 2)
)
if flf_anchor_mode:
timeline_duration = _float(timeline.get("duration_seconds"), duration_seconds)
slots = _normalise_flf_bridge_slots(timeline, slots, timeline_duration)
i2v_hard_cut_mode = bool(explicit_i2v_mode and len(image_slots) > 1)
chunks: list[dict[str, Any]] = []
prompt_map: list[dict[str, Any]] = []
@@ -420,9 +553,9 @@ def build_shotplan(
for slot_index, slot in enumerate(slots):
frame_count = int(slot["frame_count"])
hard_cut_start = slot_index > 0 and slot["transition"] == "hard_cut"
hard_cut_start = slot_index > 0 and (slot["transition"] == "hard_cut" or i2v_hard_cut_mode)
next_slot = slots[slot_index + 1] if slot_index + 1 < len(slots) else None
next_is_cut = bool(next_slot and next_slot["transition"] == "hard_cut")
next_is_cut = bool(next_slot and (next_slot["transition"] == "hard_cut" or i2v_hard_cut_mode))
next_anchor = ""
if next_slot and not next_is_cut:
next_anchor = _text(next_slot.get("image"))
@@ -491,23 +624,25 @@ def build_shotplan(
)
unique_frames_total += frame_count - overlap
image_count = sum(1 for slot in slots if slot.get("image"))
reference_image_paths = _timeline_image_paths(timeline)[:4]
image_count = len(reference_image_paths) or sum(1 for slot in slots if slot.get("image"))
return {
"schema": "iamccs.minimax_h3.shotplan",
"schema_version": 5,
"schema_version": 6,
"source_timeline_schema": _text(timeline.get("schema")),
"fps": H3_FPS,
"width": resolved_width,
"height": resolved_height,
"task_mode": task_mode,
"generation_mode": task_mode,
"continuation_mode": "timeline_keyframe_adjacency",
"continuation_mode": "flf_image_center_bridges" if flf_anchor_mode else ("i2v_hard_cuts" if i2v_hard_cut_mode else "timeline_keyframe_adjacency"),
"audio_mode": audio_mode,
"prompt_mapping": prompt_mapping,
"acceleration": acceleration,
"ref_image_size": ref_image_size,
"text_encoder_device": text_encoder_device,
"reference_roles": roles,
"reference_image_paths": reference_image_paths,
"reference_video_role": _text(reference_video_role).lower() or "off",
"reference_audio_role": _text(reference_audio_role).lower() or "off",
"sol_conditioning": sol_conditioning,
@@ -516,7 +651,10 @@ def build_shotplan(
"rife_mode": rife_mode,
"upscale_enabled": bool(_bool(upscale_enabled, False)),
"upscale_mode": active_upscale_mode,
"chunk_policy": "one_timeline_box_one_h3_chunk",
"chunk_policy": "n_keyframes_n_minus_one_flf_bridges" if flf_anchor_mode else ("one_i2v_box_one_hard_cut_chunk" if i2v_hard_cut_mode else "one_timeline_box_one_h3_chunk"),
"flf_anchor_mode": flf_anchor_mode,
"i2v_hard_cut_mode": i2v_hard_cut_mode,
"legacy_explicit_last": legacy_explicit_last,
"chunk_max_frames": H3_MAX_TRAINED_FRAMES,
"global_prompt": _text(global_prompt),
"slots": slots,
+188 -24
View File
@@ -2,10 +2,11 @@
"""Structured MiniMax H3 prompt editor and CineLinX injection contract.
The browser editor stores only user-authored project data. This backend is
deliberately deterministic: it formats the selected MiniMax prompt structure,
validates the character budget, and carries an injection request through the
standard IAMCCS CineLinX socket. No API key or network service is required.
The browser editor stores only user-authored project data. The deterministic
path formats MiniMax prompt sections and carries an injection request through
CineLinX. The optional assistant is implemented locally in this module and can
call Ollama or a user-selected compatible provider without wrapping another
custom-node package.
"""
from __future__ import annotations
@@ -24,8 +25,10 @@ from typing import Any
SUPERNODE_LINX_TYPE = "IAMCCS_SUPERNODE_LINX"
CATEGORY = "IAMCCS/MiniMax H3/Prompting"
PROJECT_SCHEMA = "iamccs.minimax_h3.prompter_project"
PROJECT_VERSION = 1
PROJECT_VERSION = 2
H3_ABSOLUTE_CHAR_LIMIT = 7000
AI_IMAGE_LIMIT = 4
AI_IMAGE_MAX_BYTES = 16 * 1024 * 1024
MODE_SECTIONS: dict[str, tuple[tuple[str, str], ...]] = {
@@ -121,6 +124,9 @@ def default_project() -> dict[str, Any]:
"injection_target": "global",
"writing_mode": "guided",
"merge_policy": "replace",
"ai_direction": "",
"ai_scope": "active_field",
"ai_visual_roles": {},
"sections": copy.deepcopy(DEFAULT_SECTIONS),
}
@@ -145,6 +151,10 @@ def _safe_project(value: Any) -> dict[str, Any]:
sections = source.get("sections")
if isinstance(sections, dict):
project["sections"].update({str(key): str(value or "") for key, value in sections.items()})
project["ai_direction"] = str(project.get("ai_direction") or "")
project["ai_scope"] = str(project.get("ai_scope") or "active_field")
visual_roles = project.get("ai_visual_roles")
project["ai_visual_roles"] = visual_roles if isinstance(visual_roles, dict) else {}
project["schema"] = PROJECT_SCHEMA
project["schema_version"] = PROJECT_VERSION
return project
@@ -227,24 +237,88 @@ def _merge_text(existing: str, incoming: str, policy: str) -> str:
return f"{old}\n\n{new}"
def _assistant_instruction(task_mode: str, sections: dict[str, str]) -> tuple[str, str]:
def _normalise_ai_images(value: Any) -> list[dict[str, str]]:
images: list[dict[str, str]] = []
for item in value if isinstance(value, list) else []:
if not isinstance(item, dict) or len(images) >= AI_IMAGE_LIMIT:
continue
data = str(item.get("data") or "").strip()
if data.startswith("data:") and "," in data:
header, data = data.split(",", 1)
guessed = header[5:].split(";", 1)[0]
else:
guessed = ""
data = re.sub(r"\s+", "", data)
if not data:
continue
estimated_bytes = (len(data) * 3) // 4
if estimated_bytes > AI_IMAGE_MAX_BYTES:
raise ValueError(f"AI reference image exceeds {AI_IMAGE_MAX_BYTES // (1024 * 1024)} MB")
mime_type = str(item.get("mime_type") or guessed or "image/png").strip().lower()
if not mime_type.startswith("image/"):
mime_type = "image/png"
role = str(item.get("role") or "reference").strip().lower()
if role not in {"opening", "closing", "identity", "composition", "style", "reference"}:
role = "reference"
images.append({
"data": data,
"mime_type": mime_type,
"name": str(item.get("name") or f"Picture {len(images) + 1}").strip(),
"role": role,
"slot": str(item.get("slot") or len(images) + 1),
})
return images
def _assistant_instruction(
task_mode: str,
sections: dict[str, str],
user_direction: str = "",
target_keys: Any = None,
images: Any = None,
) -> tuple[str, str]:
mode = str(task_mode or "t2va").lower()
if mode not in MODE_SECTIONS:
mode = "t2va"
allowed = [key for key, _label in MODE_SECTIONS[mode]]
filled = {key: str(sections.get(key, "") or "").strip() for key in allowed}
filled = {key: value for key, value in filled.items() if value}
rough = {key: str(sections.get(key, "") or "").strip() for key in allowed}
filled = {key: value for key, value in rough.items() if value}
selected = [str(key) for key in (target_keys if isinstance(target_keys, list) else []) if str(key) in allowed]
if not selected:
selected = list(filled)
if not selected:
raise ValueError("Select a MiniMax prompt section or write a rough idea before calling the AI")
visuals = _normalise_ai_images(images)
mode_rules = {
"t2va": "Build the requested event from text. Keep the action chronological, filmable and compatible with one continuous audiovisual clip.",
"i2va": "Treat <Picture 1> as the exact opening-frame authority. Animate from it without redesigning identity, wardrobe, composition or screen geography.",
"fl2va": "Treat the opening and closing pictures as exact boundary frames. Describe one physically continuous path from the first frame to the last; do not solve the transition with a cut, dissolve or unrelated redesign.",
"ref2va": "Use explicit <Picture N>, <Video N>, <Audio N> and <Subject N> references. State what each reference contributes and what must be ignored; preserve the lowercase REF2VA section semantics.",
}[mode]
system = (
"You are a professional MiniMax H3 audiovisual prompt editor. Rewrite the user's rough ideas "
"into precise, filmable English for MiniMax H3. Return one JSON object only. Its keys must be "
f"drawn from {allowed}. Rewrite only keys supplied by the user and do not fill blank sections. "
"Preserve intent, identity facts, reference labels, exact dialogue and requested timing. Do not "
"invent extra characters, dialogue, brands, camera cuts or story events. Use chronological physical "
"action, one coherent camera language, explicit continuity, and separate production sound from "
"non-diegetic music. For REF2VA retain the lowercase section semantics and labels such as "
"<Picture 1>, <Video 1>, <Audio 1> and <Subject 1>. JSON values must be plain strings."
"You are the autonomous IAMCCS MiniMax H3 prompt editor. Improve the user's own direction; do not replace it with a different story. "
"Return one JSON object only, with plain-string values and no markdown. Valid keys are "
f"{allowed}. Return only the selected keys {selected}; never create a blank or unselected section. "
"Write concise production-ready English optimized for MiniMax H3 audiovisual generation. Preserve exact identity facts, reference tags, requested timing, language and quoted dialogue unless the user explicitly asks to change them. "
"Use chronological visible action, realistic body mechanics, stable screen geography and one coherent camera language. Prefer one motivated camera move over a list of conflicting moves. "
"Separate diegetic ambience, dialogue and contact effects from non-diegetic score. Use <Subject N> consistently and keep dialogue inside <d>[Language] ...</d> with stable speaker labels such as (S1) when those tags are present. "
"Do not invent extra characters, products, dialogue, scene changes, cuts, subtitles or logos. Turn negative wishes into concrete continuity safeguards, not vague quality adjectives. "
f"Mode rule: {mode_rules} "
"When images are attached, analyze only the contribution named by each image role. An opening image governs the first frame; a closing image governs the last frame; identity, composition and style images govern only those named attributes. "
"Never mention unavailable media or claim to have seen a detail that is not visible."
)
user = json.dumps({"task_mode": mode, "rough_sections": filled}, ensure_ascii=False, indent=2)
if len(system) > 24000:
raise RuntimeError("MiniMax assistant system prompt exceeds the 7000-token safety envelope")
user = json.dumps({
"task_mode": mode,
"selected_sections": selected,
"user_direction": str(user_direction or "").strip(),
"rough_sections": {key: rough[key] for key in selected},
"visual_context": [
{"slot": item["slot"], "name": item["name"], "role": item["role"]}
for item in visuals
],
}, ensure_ascii=False, indent=2)
return system, user
@@ -273,6 +347,25 @@ def _http_json(url: str, payload: dict[str, Any], headers: dict[str, str], timeo
return parsed
def _http_get_json(url: str, timeout: float = 10.0) -> dict[str, Any]:
request = urllib.request.Request(str(url), headers={"Accept": "application/json"}, method="GET")
try:
with urllib.request.urlopen(request, timeout=max(2.0, min(30.0, float(timeout)))) as response:
raw = response.read().decode("utf-8", errors="replace")
except urllib.error.HTTPError as exc:
detail = exc.read().decode("utf-8", errors="replace")[:1200]
raise RuntimeError(f"Ollama HTTP {exc.code}: {detail}") from exc
except urllib.error.URLError as exc:
raise RuntimeError(f"Ollama connection failed: {exc.reason}") from exc
try:
parsed = json.loads(raw)
except json.JSONDecodeError as exc:
raise RuntimeError("Ollama returned invalid JSON") from exc
if not isinstance(parsed, dict):
raise RuntimeError("Ollama returned an unsupported response")
return parsed
def _extract_json_object(text: str) -> dict[str, str]:
clean = re.sub(r"^\s*```(?:json)?\s*|\s*```\s*$", "", str(text or "").strip(), flags=re.I | re.S)
start = clean.find("{")
@@ -297,12 +390,16 @@ def rewrite_sections_with_ai(
sections: dict[str, str],
temperature: float = 0.35,
timeout: float = 120.0,
user_direction: str = "",
target_keys: Any = None,
images: Any = None,
) -> tuple[dict[str, str], dict[str, Any]]:
provider = str(provider or "ollama").strip().lower()
model = str(model or "").strip()
if not model:
raise ValueError("Select an AI model before rewriting")
system, user = _assistant_instruction(task_mode, sections)
visual_inputs = _normalise_ai_images(images)
system, user = _assistant_instruction(task_mode, sections, user_direction, target_keys, visual_inputs)
api_key = str(api_key or "").strip()
if not api_key:
api_key = {
@@ -320,7 +417,14 @@ def rewrite_sections_with_ai(
"model": model,
"stream": False,
"format": "json",
"messages": [{"role": "system", "content": system}, {"role": "user", "content": user}],
"messages": [
{"role": "system", "content": system},
{
"role": "user",
"content": user,
**({"images": [item["data"] for item in visual_inputs]} if visual_inputs else {}),
},
],
"options": {"temperature": float(temperature)},
},
{},
@@ -331,13 +435,22 @@ def rewrite_sections_with_ai(
root = str(base_url or "https://api.openai.com/v1").rstrip("/")
url = root if root.endswith("/chat/completions") else f"{root}/chat/completions"
headers = {"Authorization": f"Bearer {api_key}"} if api_key else {}
openai_user: Any = user
if visual_inputs:
openai_user = [{"type": "text", "text": user}] + [
{
"type": "image_url",
"image_url": {"url": f"data:{item['mime_type']};base64,{item['data']}"},
}
for item in visual_inputs
]
result = _http_json(
url,
{
"model": model,
"temperature": float(temperature),
"response_format": {"type": "json_object"},
"messages": [{"role": "system", "content": system}, {"role": "user", "content": user}],
"messages": [{"role": "system", "content": system}, {"role": "user", "content": openai_user}],
},
headers,
timeout,
@@ -349,11 +462,16 @@ def rewrite_sections_with_ai(
encoded_model = urllib.parse.quote(model, safe="-._")
suffix = f"/models/{encoded_model}:generateContent"
url = f"{root}{suffix}?key={urllib.parse.quote(api_key)}"
gemini_parts: list[dict[str, Any]] = [{"text": user}]
gemini_parts.extend(
{"inlineData": {"mimeType": item["mime_type"], "data": item["data"]}}
for item in visual_inputs
)
result = _http_json(
url,
{
"systemInstruction": {"parts": [{"text": system}]},
"contents": [{"role": "user", "parts": [{"text": user}]}],
"contents": [{"role": "user", "parts": gemini_parts}],
"generationConfig": {"temperature": float(temperature), "responseMimeType": "application/json"},
},
{},
@@ -365,6 +483,19 @@ def rewrite_sections_with_ai(
elif provider == "anthropic":
root = str(base_url or "https://api.anthropic.com/v1").rstrip("/")
url = root if root.endswith("/messages") else f"{root}/messages"
anthropic_user: Any = user
if visual_inputs:
anthropic_user = [
{
"type": "image",
"source": {
"type": "base64",
"media_type": item["mime_type"],
"data": item["data"],
},
}
for item in visual_inputs
] + [{"type": "text", "text": user}]
result = _http_json(
url,
{
@@ -372,7 +503,7 @@ def rewrite_sections_with_ai(
"max_tokens": 4096,
"temperature": float(temperature),
"system": system,
"messages": [{"role": "user", "content": user}],
"messages": [{"role": "user", "content": anthropic_user}],
},
{"x-api-key": api_key, "anthropic-version": "2023-06-01"},
timeout,
@@ -384,7 +515,11 @@ def rewrite_sections_with_ai(
rewritten = _extract_json_object(content)
allowed = {key for key, _label in MODE_SECTIONS.get(str(task_mode).lower(), MODE_SECTIONS["t2va"])}
supplied = {key for key, value in sections.items() if key in allowed and str(value or "").strip()}
filtered = {key: value for key, value in rewritten.items() if key in supplied and value}
requested = {str(key) for key in target_keys} if isinstance(target_keys, list) else supplied
requested = requested & allowed
if not requested:
requested = supplied
filtered = {key: value for key, value in rewritten.items() if key in requested and value}
if not filtered:
raise RuntimeError("The AI did not return any valid filled MiniMax section")
return filtered, {
@@ -392,6 +527,12 @@ def rewrite_sections_with_ai(
"model": model,
"rewritten_sections": sorted(filtered),
"preserved_blank_sections": sorted(allowed - supplied),
"selected_sections": sorted(requested),
"visual_references": [
{"slot": item["slot"], "name": item["name"], "role": item["role"]}
for item in visual_inputs
],
"system_prompt_characters": len(system),
}
@@ -546,7 +687,7 @@ class IAMCCS_Prompter:
"default": "",
"multiline": True,
"forceInput": True,
"tooltip": "Optional complete draft from H3_Promptor. In Assistant Fill mode it fills only empty structured boxes.",
"tooltip": "Optional structured draft from any text source. In Assistant Fill mode it fills only empty structured boxes.",
},
),
},
@@ -656,6 +797,26 @@ def _register_prompter_routes() -> None:
routes = PromptServer.instance.routes
@routes.get("/iamccs/prompter/ollama/models")
async def iamccs_prompter_ollama_models(request):
try:
base_url = str(request.query.get("base_url") or "http://127.0.0.1:11434").rstrip("/")
payload = await asyncio.to_thread(_http_get_json, f"{base_url}/api/tags", 10.0)
models = []
for item in payload.get("models") if isinstance(payload.get("models"), list) else []:
if not isinstance(item, dict):
continue
name = str(item.get("name") or item.get("model") or "").strip()
if name:
models.append({
"name": name,
"size": int(item.get("size") or 0),
"modified_at": str(item.get("modified_at") or ""),
})
return web.json_response({"ok": True, "models": models})
except Exception as exc:
return web.json_response({"ok": False, "error": str(exc)}, status=400)
@routes.post("/iamccs/prompter/rewrite")
async def iamccs_prompter_rewrite(request):
try:
@@ -673,6 +834,9 @@ def _register_prompter_routes() -> None:
{str(key): str(value or "") for key, value in sections.items()},
float(payload.get("temperature", 0.35)),
float(payload.get("timeout", 120.0)),
str(payload.get("user_direction", "")),
payload.get("target_keys"),
payload.get("images"),
)
return web.json_response({"ok": True, "sections": rewritten, "report": report})
except Exception as exc:
+82 -11
View File
@@ -7,6 +7,8 @@ without adding a graph dependency.
from __future__ import annotations
import gc
import logging
import math
import os
import sys
@@ -16,6 +18,17 @@ from typing import Tuple
import torch
import torch.nn.functional as F
try:
import comfy.model_management as _model_management # type: ignore
from comfy.utils import ProgressBar as _ProgressBar # type: ignore
except ImportError: # Keep the resize helpers importable outside ComfyUI.
_model_management = None
_ProgressBar = None
_LOG = logging.getLogger("IAMCCS.RTXVFX")
RTX_AUTOMATIC_CHUNK_SIZE = 8
RTX_QUALITY_LEVELS = [
"VSR Medium",
@@ -344,6 +357,20 @@ def _safe_cuda_device_index(device: int) -> int:
return 0 if value < 0 or (count and value >= count) else value
def _release_comfy_models_for_rtx() -> None:
"""Give the native RTX runtime an empty CUDA workspace before it starts."""
if _model_management is not None:
_model_management.unload_all_models()
try:
_model_management.cleanup_models()
except Exception:
pass
_model_management.soft_empty_cache()
if torch.cuda.is_available():
torch.cuda.empty_cache()
gc.collect()
def apply_rtx_vfx(
images: torch.Tensor,
mode: str = "VSR Medium",
@@ -356,6 +383,7 @@ def apply_rtx_vfx(
device: int = 0,
ratio_preset: str = "16:9",
resize_method: str = "Center Crop (Fill)",
chunk_size: int = RTX_AUTOMATIC_CHUNK_SIZE,
) -> torch.Tensor:
"""Apply Deno RTX Video Super Resolution semantics directly to IMAGE frames."""
if not torch.cuda.is_available():
@@ -378,13 +406,37 @@ def apply_rtx_vfx(
target_width, target_height = _target_size(
int(source_width), int(source_height), mode, resize_type, float(scale), float(megapixels), int(width), int(height), alignment, ratio_preset
)
# RTX accepts float32 RGB frames, but keeping a complete float32 4K batch
# on CUDA can consume tens of GiB. Detach the source before unloading any
# previous diffusion stack, then retain completed frames on CPU/float16.
source_device = str(images.device)
source_dtype = str(images.dtype)
source = images[..., :3].detach().to(device="cpu")
_release_comfy_models_for_rtx()
VideoSuperRes = _import_video_super_res()
quality = getattr(VideoSuperRes.QualityLevel, _quality_attr(mode))
device_index = _safe_cuda_device_index(device)
cuda_device = torch.device(f"cuda:{device_index}")
out_device = images.device
out_dtype = images.dtype
output = torch.empty((int(batch), int(target_height), int(target_width), 3), device=out_device, dtype=out_dtype)
effective_chunk_size = max(1, min(int(chunk_size or RTX_AUTOMATIC_CHUNK_SIZE), int(batch)))
output = torch.empty(
(int(batch), int(target_height), int(target_width), 3),
device="cpu",
dtype=torch.float16,
)
progress = _ProgressBar(int(batch)) if _ProgressBar is not None else None
_LOG.info(
"IAMCCS Exporter RTX VFX chunked start | %s/%s -> cpu/float16 | "
"frames=%d | chunk=%d | %dx%d -> %dx%d",
source_device,
source_dtype,
int(batch),
effective_chunk_size,
int(source_width),
int(source_height),
int(target_width),
int(target_height),
)
with torch.inference_mode():
try:
@@ -400,12 +452,31 @@ def apply_rtx_vfx(
effect.output_width = int(target_width)
effect.output_height = int(target_height)
effect.load()
for index in range(int(batch)):
frame = images[index, :, :, :3].to(device=cuda_device, dtype=torch.float32).permute(2, 0, 1).contiguous()
if not _same_size_only(mode):
frame = _fit_frame_to_target_aspect(frame, int(target_width), int(target_height), resize_method)
result = effect.run(frame)
enhanced = torch.from_dlpack(result.image).clone().permute(1, 2, 0).contiguous()
output[index].copy_(enhanced.clamp(0.0, 1.0).to(device=out_device, dtype=out_dtype))
del frame, enhanced
for chunk_start in range(0, int(batch), effective_chunk_size):
chunk_end = min(int(batch), chunk_start + effective_chunk_size)
rtx_input = source[chunk_start:chunk_end].to(dtype=torch.float32).contiguous()
rtx_input = torch.nan_to_num(
rtx_input,
nan=0.0,
posinf=1.0,
neginf=0.0,
).clamp_(0.0, 1.0)
for local_index in range(int(rtx_input.shape[0])):
output_index = chunk_start + local_index
frame = rtx_input[local_index].to(device=cuda_device).permute(2, 0, 1).contiguous()
if not _same_size_only(mode):
frame = _fit_frame_to_target_aspect(frame, int(target_width), int(target_height), resize_method)
result = effect.run(frame)
enhanced = torch.from_dlpack(result.image).clone().permute(1, 2, 0).contiguous()
output[output_index].copy_(enhanced.clamp(0.0, 1.0).to(device="cpu", dtype=torch.float16))
del result, frame, enhanced
del rtx_input
if torch.cuda.is_available():
torch.cuda.empty_cache()
gc.collect()
if progress is not None:
progress.update_absolute(chunk_end, int(batch))
_LOG.info("IAMCCS Exporter RTX VFX progress | %d/%d frames", chunk_end, int(batch))
del source
gc.collect()
return output
+2
View File
@@ -893,6 +893,7 @@ class IAMCCS_ShotboarderAudVidExporterPRO:
"rtx_device": int(rtx_device or 0),
"rtx_ratio_preset": str(rtx_ratio_preset or "16:9"),
"rtx_resize_method": str(rtx_resize_method or "Center Crop (Fill)"),
"rtx_memory_mode": "automatic_chunk_8_cpu_float16" if rtx_active else "off",
}
if isinstance(cine_linx, dict):
metadata["cine_linx_type"] = str(cine_linx.get("type", ""))
@@ -1041,6 +1042,7 @@ class IAMCCS_ShotboarderAudVidExporterPRO:
"audio_edl_status": direct_audio_edl_status,
"visual_roll_dedup_active": roll_visual_dedup_active,
"visual_roll_dedup_status": roll_visual_dedup_status,
"rtx_memory_mode": "automatic_chunk_8_cpu_float16" if rtx_active else "off",
"codec_contract": f"{profile['video_args']} + {audio_config['args']}",
"video_lossless": bool(profile.get("lossless")),
"audio_lossless": effective_audio_lossless,
+453 -47
View File
@@ -5,8 +5,47 @@ import { app } from "../../scripts/app.js";
import { api } from "../../scripts/api.js";
console.info("[IAMCCS MiniMax H3] Dedicated Shotboard V3-parity UI loaded.");
const CINE_VERSION = "2026-08-07-minimax-h3-fullscreen-toolbar-fit-v8";
const CINE_VERSION = "2026-08-08-minimax-h3-native-upscale-2x-sync-v12";
const MINIMAX_CINE_LINX_TYPE = "IAMCCS_SUPERNODE_LINX";
const H3_NATIVE_RESOLUTION_PRESETS = Object.freeze([
{ width: 768, height: 448, label: "H · ≈16:9 · 768×448 · Draft" },
{ width: 960, height: 544, label: "H · ≈16:9 · 960×544 · Balanced" },
{ width: 1024, height: 576, label: "H · 16:9 · 1024×576" },
{ width: 1280, height: 736, label: "H · 720-source legal · 1280×736" },
{ width: 1344, height: 768, label: "H · ≈16:9 · 1344×768 · H3 quality" },
{ width: 1024, height: 768, label: "H · 4:3 · 1024×768" },
{ width: 1152, height: 768, label: "H · 3:2 · 1152×768" },
{ width: 1216, height: 640, label: "H · DCI ≈1.90 · 1216×640" },
{ width: 1024, height: 512, label: "SCOPE · 2.00 · 1024×512" },
{ width: 1120, height: 512, label: "SCOPE · ≈2.20 · 1120×512" },
{ width: 1152, height: 480, label: "SCOPE · ≈2.39 · 1152×480" },
{ width: 1536, height: 640, label: "SCOPE · ≈2.39 · 1536×640 · Quality" },
{ width: 448, height: 768, label: "V · ≈9:16 · 448×768 · Draft" },
{ width: 544, height: 960, label: "V · ≈9:16 · 544×960 · Balanced" },
{ width: 576, height: 1024, label: "V · 9:16 · 576×1024" },
{ width: 768, height: 1344, label: "V · ≈9:16 · 768×1344 · H3 quality" },
{ width: 768, height: 960, label: "V · 4:5 · 768×960" },
{ width: 640, height: 960, label: "V · 2:3 · 640×960" },
{ width: 768, height: 768, label: "H/V · 1:1 · 768×768" },
]);
const H3_UPSCALE_RESOLUTION_PRESETS = Object.freeze([
{ width: 1280, height: 720, label: "H · HD · 1280×720" },
{ width: 1920, height: 1080, label: "H · FHD · 1920×1080" },
{ width: 1998, height: 1080, label: "H · DCI Flat 2K · 1998×1080" },
{ width: 2048, height: 1080, label: "H · DCI 2K · 2048×1080" },
{ width: 2560, height: 1440, label: "H · QHD · 2560×1440" },
{ width: 3840, height: 2160, label: "H · UHD · 3840×2160" },
{ width: 3996, height: 2160, label: "H · DCI Flat 4K · 3996×2160" },
{ width: 4096, height: 2160, label: "H · DCI 4K · 4096×2160" },
{ width: 2048, height: 858, label: "SCOPE · DCI 2K 2.39 · 2048×858" },
{ width: 2560, height: 1080, label: "SCOPE · UW ≈2.37 · 2560×1080" },
{ width: 3840, height: 1608, label: "SCOPE · UHD ≈2.39 · 3840×1608" },
{ width: 4096, height: 1716, label: "SCOPE · DCI 4K 2.39 · 4096×1716" },
{ width: 720, height: 1280, label: "V · HD · 720×1280" },
{ width: 1080, height: 1920, label: "V · FHD · 1080×1920" },
{ width: 1440, height: 2560, label: "V · QHD · 1440×2560" },
{ width: 2160, height: 3840, label: "V · UHD · 2160×3840" },
]);
const SHOTBOARD_V3_RIGID_WIDTH = 2360;
const SHOTBOARD_V3_OPEN_HEIGHT = 760;
const SHOTBOARD_V3_COLLAPSED_HEIGHT = 560;
@@ -317,12 +356,18 @@ function repairMiniMaxH3WidgetState(node, serialized = null) {
return [];
};
const ltxDetailerWidget = getWidget(node, "ltx_detailer_lora_name");
const preferredCrispLtxLora = comboValues(ltxDetailerWidget).find((value) => {
const name = String(value || "").toLowerCase();
return name.includes("ltx") && name.includes("crisp");
}) || "";
const comboDefaults = {
task_mode: "auto_from_timeline",
audio_mode: "h3_native_generated",
prompt_mapping: "global_plus_local",
upscale_mode: "off",
acceleration: "auto_3060",
acceleration: "low_vram_auto",
ref_image_size: "match",
reference_role_1: "subject_identity",
reference_role_2: "subject_identity",
@@ -330,11 +375,11 @@ function repairMiniMaxH3WidgetState(node, serialized = null) {
reference_role_4: "style",
reference_video_role: "off",
reference_audio_role: "off",
sol_conditioning: "exact_kv",
spectrum_profile: "conservative_3060",
sol_conditioning: "exact_kv_and_rows",
spectrum_profile: "low_vram",
rife_mode: "off",
text_encoder_device: "cpu_safe_12gb",
performance_profile: "rtx3060_balanced",
text_encoder_device: "auto",
performance_profile: "low_vram_balanced",
sampler_name: "res_multistep",
scheduler: "simple",
turbo_mode: "off",
@@ -342,7 +387,13 @@ function repairMiniMaxH3WidgetState(node, serialized = null) {
turbo_sampler_mode: "audio_fixed",
reference_resize_policy: "canvas_crop",
reference_resize_filter: "area",
ltx_detailer_lora_name: preferredCrispLtxLora,
ltx_4k_quality: "ULTRA",
};
const textEncoderWidget = getWidget(node, "text_encoder_device");
if (String(textEncoderWidget?.value || "").toLowerCase() === "cpu_safe_12gb") {
assign("text_encoder_device", "auto", "legacy-cpu-mode-migrated-to-gpu-first");
}
Object.entries(comboDefaults).forEach(([name, defaultValue]) => {
const item = getWidget(node, name);
if (!item) return;
@@ -351,6 +402,12 @@ function repairMiniMaxH3WidgetState(node, serialized = null) {
const preferred = fallbackValue(name, defaultValue);
assign(name, values.includes(preferred) ? preferred : (values.includes(defaultValue) ? defaultValue : values[0]), "invalid-choice");
});
// Existing workflows commonly serialized the old empty value. Populate
// Crisp once it is available, while preserving every non-empty LoRA the
// user explicitly selected from the complete dropdown.
if (!String(ltxDetailerWidget?.value || "").trim() && preferredCrispLtxLora) {
assign("ltx_detailer_lora_name", preferredCrispLtxLora, "installed-crisp-default");
}
const numberRules = {
duration_seconds: [10, 0.01, 36000, false],
@@ -369,6 +426,7 @@ function repairMiniMaxH3WidgetState(node, serialized = null) {
shift_audio: [3, 0.01, 100, false],
turbo_strength: [1, 0, 4, false],
reference_resize_megapixels: [0.5, 0.1, 2, false],
ltx_detailer_strength: [0.6, 0, 2, false],
};
Object.entries(numberRules).forEach(([name, [defaultValue, min, max, integer]]) => {
const item = getWidget(node, name);
@@ -4174,6 +4232,12 @@ function renderShotboardLite(node) {
"reference_resize_policy",
"reference_resize_megapixels",
"reference_resize_filter",
"ltx_detailer_enabled",
"ltx_detailer_lora_name",
"ltx_detailer_strength",
"ltx_4k_enabled",
"ltx_4k_quality",
"ltx_seam_safe",
].forEach((name) => hideWidget(getWidget(node, name)));
let rows = parseJsonWidget(node, defaultLiteRows).map(normalizeLiteRow);
@@ -7078,6 +7142,12 @@ function renderShotboardV3(node) {
"reference_resize_policy",
"reference_resize_megapixels",
"reference_resize_filter",
"ltx_detailer_enabled",
"ltx_detailer_lora_name",
"ltx_detailer_strength",
"ltx_4k_enabled",
"ltx_4k_quality",
"ltx_seam_safe",
];
// The MiniMax board renders these controls in its own settings boxes
// above the timeline. Keep the underlying Comfy widgets serializable,
@@ -8032,6 +8102,7 @@ function renderShotboardV3(node) {
const videoPath = isVideo ? videoPathForSegment(seg) : "";
const singleStrength = isText ? 0 : Math.max(0, Math.min(1, Number(seg.guideStrength ?? seg.guide_strength ?? seg.force ?? seg.strength ?? defaultForceWidget?.value ?? 0.25)));
return {
id: String(seg.id || `${isText ? "text" : "shot"}_${index + 1}`),
type: isVideo ? "video" : isText ? "text" : "image",
second: Number((startFrame / fps).toFixed(3)),
frame: startFrame,
@@ -8287,6 +8358,11 @@ function renderShotboardV3(node) {
.filter((seg) => String(seg.type || "image") !== "audio" && !seg.placeholder)
.slice()
.sort((a, b) => Number(a.start || 0) - Number(b.start || 0));
const h3Bridges = buildH3BridgeContract(timeline.segments || []);
const h3AnchorMode = h3Bridges.length ? "image_box_centres" : "off";
const h3EditMode = h3Bridges.length
? "flf_n_keyframes_n_minus_one_chunks"
: (h3I2vHardCutMode(timeline.segments || []) ? "i2v_independent_hard_cuts" : "single_or_text");
const directorPrompts = [];
const directorLengths = [];
let cursor = 0;
@@ -8353,6 +8429,9 @@ function renderShotboardV3(node) {
segments: JSON.parse(JSON.stringify(timeline.segments || [])),
motionSegments: isShotboardV4 ? JSON.parse(JSON.stringify(timeline.motionSegments || [])) : [],
rows: JSON.parse(JSON.stringify(timelineRows)),
h3_anchor_mode: h3AnchorMode,
h3_edit_mode: h3EditMode,
h3_bridges: JSON.parse(JSON.stringify(h3Bridges)),
global_prompt: String(promptArea?.value || promptWidget?.value || ""),
prompt: String(promptArea?.value || promptWidget?.value || ""),
director_local_prompts: effectiveDirectorPrompts.join(" | "),
@@ -8506,17 +8585,20 @@ function renderShotboardV3(node) {
duration_seconds: effectiveDurationSeconds,
frame_rate: 24,
task_mode: String(getWidget(node, "task_mode")?.value || "auto_from_timeline"),
h3_anchor_mode: h3AnchorMode,
h3_edit_mode: h3EditMode,
h3_bridges: h3Bridges,
continuation_mode: "timeline_keyframe_adjacency",
audio_mode: String(getWidget(node, "audio_mode")?.value || "h3_native_generated"),
prompt_mapping: String(getWidget(node, "prompt_mapping")?.value || "global_plus_local"),
acceleration: String(getWidget(node, "acceleration")?.value || "sage"),
acceleration: String(getWidget(node, "acceleration")?.value || "low_vram_auto"),
ref_image_size: String(getWidget(node, "ref_image_size")?.value || "match"),
text_encoder_device: String(getWidget(node, "text_encoder_device")?.value || "cpu_safe_12gb"),
text_encoder_device: "auto",
reference_roles: [1, 2, 3, 4].map((index) => String(getWidget(node, `reference_role_${index}`)?.value || (index <= 2 ? "subject_identity" : index === 3 ? "composition" : "style"))),
reference_video_role: String(getWidget(node, "reference_video_role")?.value || "off"),
reference_audio_role: String(getWidget(node, "reference_audio_role")?.value || "off"),
sol_conditioning: String(getWidget(node, "sol_conditioning")?.value || "exact_kv"),
spectrum_profile: String(getWidget(node, "spectrum_profile")?.value || "conservative_3060"),
sol_conditioning: String(getWidget(node, "sol_conditioning")?.value || "exact_kv_and_rows"),
spectrum_profile: String(getWidget(node, "spectrum_profile")?.value || "low_vram"),
vram_clean_before_decode: Boolean(getWidget(node, "vram_clean_before_decode")?.value ?? true),
rife_mode: String(getWidget(node, "rife_mode")?.value || "off"),
upscale_enabled: Boolean(getWidget(node, "upscale_enabled")?.value ?? false),
@@ -8527,6 +8609,12 @@ function renderShotboardV3(node) {
upscale_sage: Boolean(getWidget(node, "upscale_sage")?.value ?? true),
upscale_seed_offset: Number(getWidget(node, "upscale_seed_offset")?.value || 10000),
wan_upscale_denoise: Number(getWidget(node, "wan_upscale_denoise")?.value ?? 0.2),
ltx_detailer_enabled: Boolean(getWidget(node, "ltx_detailer_enabled")?.value ?? false),
ltx_detailer_lora_name: String(getWidget(node, "ltx_detailer_lora_name")?.value || ""),
ltx_detailer_strength: Number(getWidget(node, "ltx_detailer_strength")?.value ?? 0.6),
ltx_4k_enabled: Boolean(getWidget(node, "ltx_4k_enabled")?.value ?? false),
ltx_4k_quality: String(getWidget(node, "ltx_4k_quality")?.value || "ULTRA"),
ltx_seam_safe: Boolean(getWidget(node, "ltx_seam_safe")?.value ?? true),
width: Number(getWidget(node, "width")?.value || 960),
height: Number(getWidget(node, "height")?.value || 544),
h3_backend_contract: {
@@ -8536,11 +8624,15 @@ function renderShotboardV3(node) {
maximum_frames_per_box: 362,
first_last_keyframes: true,
generated_last_becomes_next_first: true,
flf_chunk_policy: "n_keyframes_n_minus_one_chunks",
flf_prompt_span: "image_box_centre_to_next_image_box_centre",
flf_timing_policy: "centre_distances_normalised_to_total_duration",
i2v_multi_image_policy: "one_box_one_i2v_chunk_hard_cut",
text_relay_conditioning: false,
atomic_task_model_conditioning: true,
acceleration: String(getWidget(node, "acceleration")?.value || "sage"),
acceleration: String(getWidget(node, "acceleration")?.value || "low_vram_auto"),
ref_image_size: String(getWidget(node, "ref_image_size")?.value || "match"),
text_encoder_device: String(getWidget(node, "text_encoder_device")?.value || "cpu_safe_12gb"),
text_encoder_device: "auto",
vram_clean_before_decode: Boolean(getWidget(node, "vram_clean_before_decode")?.value ?? true),
native_bridge_before_post: true,
lazy_upscale_enabled: Boolean(getWidget(node, "upscale_enabled")?.value ?? false),
@@ -8551,6 +8643,12 @@ function renderShotboardV3(node) {
Number(getWidget(node, "upscale_width")?.value || 1920),
Number(getWidget(node, "upscale_height")?.value || 1080),
],
ltx_detailer_enabled: Boolean(getWidget(node, "ltx_detailer_enabled")?.value ?? false),
ltx_detailer_lora_name: String(getWidget(node, "ltx_detailer_lora_name")?.value || ""),
ltx_detailer_strength: Number(getWidget(node, "ltx_detailer_strength")?.value ?? 0.6),
ltx_seam_safe: Boolean(getWidget(node, "ltx_seam_safe")?.value ?? true),
ltx_4k_enabled: Boolean(getWidget(node, "ltx_4k_enabled")?.value ?? false),
ltx_4k_quality: String(getWidget(node, "ltx_4k_quality")?.value || "ULTRA"),
rife_mode: String(getWidget(node, "rife_mode")?.value || "off"),
},
image_paths: refPaths(),
@@ -9804,12 +9902,14 @@ function renderShotboardV3(node) {
const requestedFrames = Math.max(1, Math.round(seconds * 24));
const alignedFrames = Math.max(5, Math.min(362, Math.ceil(Math.max(0, requestedFrames - 5) / 17) * 17 + 5));
const load = (width * height * alignedFrames) / (960 * 544 * 124);
const acceleration = String(getWidget(node, "acceleration")?.value || "auto_3060");
const profile = String(getWidget(node, "performance_profile")?.value || "rtx3060_balanced");
const risky = profile.startsWith("rtx3060") && load > 1.15;
const acceleration = String(getWidget(node, "acceleration")?.value || "low_vram_auto");
const profile = String(getWidget(node, "performance_profile")?.value || "low_vram_balanced");
const risky = (profile.startsWith("low_vram") || profile.startsWith("rtx3060")) && load > 1.15;
performanceBand.style.borderColor = risky ? "rgba(210,112,87,.65)" : "rgba(111,146,151,.38)";
performanceBand.style.color = risky ? "#ffc0ac" : "#a9c8c4";
const accelerationLabel = acceleration === "auto_3060" ? "Low VRAM Auto / H3 Sage" : acceleration;
const accelerationLabel = ["low_vram_auto", "auto_3060"].includes(acceleration)
? "Low VRAM Auto / H3 Sage + exact chunks"
: acceleration;
performanceBand.textContent = `Low VRAM load preview: ${load.toFixed(2)}x versus 960x544 / 124f | ${width}x${height} / ~${alignedFrames} aligned frames | ${accelerationLabel}${risky ? " | HEAVY: shorten this box or use a smaller canvas, then upscale" : " | within the selected Low VRAM canvas envelope"}`;
}
const makeSettingsGroup = (kicker, title, description, advanced = false) => {
@@ -9836,18 +9936,43 @@ function renderShotboardV3(node) {
const shotSettings = makeSettingsGroup("01", "SHOT & CANVAS", "Timeline trim is the chunk authority. H3 keeps 24 fps and the 17k+5 frame grid.");
let settingsTarget = shotSettings;
const settingsControls = new Map();
let nativeResolutionSelect = null;
let upscaleResolutionSelect = null;
let syncResolutionPresetControls = () => {};
const setDeckValue = (name, value) => {
setWidgetValue(node, name, value);
const control = settingsControls.get(name);
try { control?._iamccsSetValue?.(value); } catch {}
if (control && "value" in control) control.value = String(value);
if (["width", "height", "upscale_width", "upscale_height"].includes(name)) {
syncResolutionPresetControls();
}
};
const syncUpscaleToNative2x = (nativeWidth, nativeHeight, { notice = true } = {}) => {
const width = Math.max(256, Math.round(Number(nativeWidth) || 0));
const height = Math.max(256, Math.round(Number(nativeHeight) || 0));
const targetWidth = Math.min(7680, width * 2);
const targetHeight = Math.min(4320, height * 2);
setDeckValue("upscale_width", targetWidth);
setDeckValue("upscale_height", targetHeight);
syncResolutionPresetControls();
if (notice) {
const exact = targetWidth === width * 2 && targetHeight === height * 2;
showTimelineNotice(
exact
? `Upscale synced 2x: ${width}x${height} -> ${targetWidth}x${targetHeight}; aspect ratio and framing preserved.`
: `Upscale synced within delivery limits: ${targetWidth}x${targetHeight}.`,
exact ? "ok" : "info",
);
}
return { width: targetWidth, height: targetHeight };
};
const h3LegalDimension = (value, fallback) => {
const numeric = Number(value);
const safe = Number.isFinite(numeric) ? numeric : Number(fallback);
return Math.max(256, Math.min(5760, Math.ceil(Math.max(256, safe) / 32) * 32));
};
const setH3Canvas = (rawWidth, rawHeight, notice = "") => {
const setH3Canvas = (rawWidth, rawHeight, notice = "", { syncUpscale = true } = {}) => {
const oldWidth = Number(getWidget(node, "width")?.value || 960);
const oldHeight = Number(getWidget(node, "height")?.value || 544);
const legalWidth = h3LegalDimension(rawWidth, oldWidth);
@@ -9856,8 +9981,9 @@ function renderShotboardV3(node) {
setDeckValue("height", legalHeight);
setWidgetValue(node, "image_width", legalWidth);
setWidgetValue(node, "image_height", legalHeight);
if (syncUpscale) syncUpscaleToNative2x(legalWidth, legalHeight, { notice: !notice });
if (notice && (legalWidth !== oldWidth || legalHeight !== oldHeight)) {
showTimelineNotice(`${notice}: H3 canvas ${legalWidth}x${legalHeight} (32-pixel legal grid).`, "ok");
showTimelineNotice(`${notice}: H3 canvas ${legalWidth}x${legalHeight}; upscale synced to ${legalWidth * 2}x${legalHeight * 2}.`, "ok");
}
return { width: legalWidth, height: legalHeight };
};
@@ -9874,6 +10000,74 @@ function renderShotboardV3(node) {
const scale = Math.min(1, maxWidth / width, maxHeight / height);
setH3Canvas(width * scale, height * scale, "Image resolution adapted");
};
const resolutionPresetKey = (width, height) => `${Math.round(Number(width) || 0)}x${Math.round(Number(height) || 0)}`;
const matchedResolutionPresetKey = (width, height, presets) => {
const key = resolutionPresetKey(width, height);
return presets.some((preset) => resolutionPresetKey(preset.width, preset.height) === key) ? key : "custom";
};
const addResolutionPresetSetting = (label, presets, kind) => {
const isNative = kind === "native";
const widthName = isNative ? "width" : "upscale_width";
const heightName = isNative ? "height" : "upscale_height";
const wrap = document.createElement("label");
wrap.style.cssText = `grid-column:span 2;display:flex;flex-direction:column;gap:4px;color:${purple.muted};font-size:10px;font-weight:800;text-align:center;`;
wrap.title = isNative
? "H3 native canvases are aligned to a 32-pixel grid. H = horizontal, V = vertical, SCOPE = cinema-wide. Custom remains available."
: "Exact delivery targets for the selected post-upscale route. H = horizontal, V = vertical, SCOPE = cinema-wide. Custom remains available.";
const span = document.createElement("span");
span.textContent = label;
const choices = [
{ value: "custom", label: "CUSTOM · manual width × height" },
...presets.map((preset) => ({
value: resolutionPresetKey(preset.width, preset.height),
label: preset.label,
})),
];
const ctrl = makeChoiceSelect(
matchedResolutionPresetKey(getWidget(node, widthName)?.value, getWidget(node, heightName)?.value, presets),
choices,
(value) => {
if (value === "custom") {
showTimelineNotice(`${label}: manual Width / Height controls are active.`, "info");
return;
}
const preset = presets.find((candidate) => resolutionPresetKey(candidate.width, candidate.height) === value);
if (!preset) return;
if (isNative) {
setH3Canvas(preset.width, preset.height);
} else {
setDeckValue("upscale_width", preset.width);
setDeckValue("upscale_height", preset.height);
}
syncResolutionPresetControls();
showTimelineNotice(`${label}: ${preset.label}.`, "ok");
writeTimeline();
refreshPerformanceBand();
draw();
},
);
styleValueControls(ctrl);
settingsControls.set(`__${kind}_resolution_preset`, ctrl);
wrap.append(span, ctrl);
settingsTarget.appendChild(wrap);
return ctrl;
};
syncResolutionPresetControls = () => {
if (nativeResolutionSelect) {
nativeResolutionSelect.value = matchedResolutionPresetKey(
getWidget(node, "width")?.value,
getWidget(node, "height")?.value,
H3_NATIVE_RESOLUTION_PRESETS,
);
}
if (upscaleResolutionSelect) {
upscaleResolutionSelect.value = matchedResolutionPresetKey(
getWidget(node, "upscale_width")?.value,
getWidget(node, "upscale_height")?.value,
H3_UPSCALE_RESOLUTION_PRESETS,
);
}
};
const addSetting = (label, name, step, min, targetOverride = null, compact = false) => {
const target = targetOverride || settingsTarget;
const wrap = document.createElement("label");
@@ -9897,6 +10091,12 @@ function renderShotboardV3(node) {
}
if (name === "width") setWidgetValue(node, "image_width", nextValue);
if (name === "height") setWidgetValue(node, "image_height", nextValue);
if (name === "width" || name === "height") {
syncUpscaleToNative2x(
name === "width" ? nextValue : getWidget(node, "width")?.value,
name === "height" ? nextValue : getWidget(node, "height")?.value,
);
}
if (name === "duration_seconds") {
setDurationSeconds(nextValue, "duration_control");
enforceDurationMinimum();
@@ -9904,6 +10104,9 @@ function renderShotboardV3(node) {
if (name === "frame_rate") {
setFrameRateValue(nextValue, "fps_control");
}
if (["width", "height", "upscale_width", "upscale_height"].includes(name)) {
syncResolutionPresetControls();
}
writeTimeline();
refreshPerformanceBand();
draw();
@@ -9932,9 +10135,15 @@ function renderShotboardV3(node) {
h3FpsWrap.append(h3FpsLabel, h3FpsValue);
settingsTarget.appendChild(h3FpsWrap);
setWidgetValue(node, "frame_rate", 24);
nativeResolutionSelect = addResolutionPresetSetting("Native format", H3_NATIVE_RESOLUTION_PRESETS, "native");
addSetting("Width", "width", "32", "256");
addSetting("Height", "height", "32", "256");
setH3Canvas(getWidget(node, "width")?.value || 960, getWidget(node, "height")?.value || 544);
setH3Canvas(
getWidget(node, "width")?.value || 960,
getWidget(node, "height")?.value || 544,
"",
{ syncUpscale: false },
);
const addSelectSetting = (label, name, options) => {
const wrap = document.createElement("label");
wrap.style.cssText = `display:flex;flex-direction:column;gap:4px;color:${purple.muted};font-size:10px;font-weight:800;text-align:center;`;
@@ -10049,7 +10258,7 @@ function renderShotboardV3(node) {
{ value: "fl2va", label: "FL2VA / FFLF" },
{ value: "ref2va", label: "REF2VA" },
]);
addStaticH3Setting("Chunk source", "Timeline trim", "Each visual box is exactly one H3 chunk. Drag or trim its length on the meter; the planner never splits it automatically.");
addStaticH3Setting("Chunk source", "Mode-aware timeline", "FL2VA/Auto with two or more images uses N keyframes -> N-1 centre-to-centre prompt bridges. I2VA keeps every image box as an independent hard-cut shot.");
addStaticH3Setting("H3 frames", "17k+5 / max 362", "The requested box length is aligned upward to H3's required 17k+5 temporal grid and must remain at or below 362 frames.");
addWidgetChoiceSetting("H3 audio", "audio_mode", [
{ value: "h3_native_generated", label: "Native generated" },
@@ -10065,17 +10274,24 @@ function renderShotboardV3(node) {
settingsTarget = makeSettingsGroup("02", "GENERATION", "Shotboard-owned sampler values. Backend widgets are compatibility fallbacks only.");
const applyPerformanceProfile = (profile) => {
const presets = {
rtx3060_draft: { width: 768, height: 448, steps: 12, acceleration: "auto_3060" },
rtx3060_balanced: { width: 960, height: 544, steps: 16, acceleration: "auto_3060" },
rtx3060_turbo: { width: 960, height: 544, steps: 8, acceleration: "auto_3060", turbo_mode: "early_8_10", turbo_strength: 1.0, turbo_sampler_mode: "audio_fixed", reference_resize_policy: "canvas_crop", reference_resize_megapixels: 0.5, reference_resize_filter: "area", sampler_name: "res_multistep", scheduler: "simple", shift_video: 12, shift_audio: 3 },
h3_native_quality: { width: 1344, height: 768, steps: 20, acceleration: "sage" },
low_vram_draft: { width: 768, height: 448, steps: 12, acceleration: "low_vram_auto" },
low_vram_balanced: { width: 960, height: 544, steps: 16, acceleration: "low_vram_auto" },
low_vram_turbo: { width: 960, height: 544, steps: 8, acceleration: "low_vram_auto", turbo_mode: "early_8_10", turbo_strength: 1.0, turbo_sampler_mode: "audio_fixed", reference_resize_policy: "canvas_crop", reference_resize_megapixels: 0.5, reference_resize_filter: "area", sampler_name: "res_multistep", scheduler: "simple", shift_video: 12, shift_audio: 3 },
rtx3060_draft: { width: 768, height: 448, steps: 12, acceleration: "low_vram_auto" },
rtx3060_balanced: { width: 960, height: 544, steps: 16, acceleration: "low_vram_auto" },
rtx3060_turbo: { width: 960, height: 544, steps: 8, acceleration: "low_vram_auto", turbo_mode: "early_8_10", turbo_strength: 1.0, turbo_sampler_mode: "audio_fixed", reference_resize_policy: "canvas_crop", reference_resize_megapixels: 0.5, reference_resize_filter: "area", sampler_name: "res_multistep", scheduler: "simple", shift_video: 12, shift_audio: 3 },
h3_native_quality: { width: 1344, height: 768, steps: 20, acceleration: "h3_sage" },
};
const preset = presets[String(profile)] || null;
if (!preset) return;
Object.entries(preset).forEach(([name, value]) => setDeckValue(name, value));
setDeckValue("image_width", preset.width);
setDeckValue("image_height", preset.height);
syncUpscaleToNative2x(preset.width, preset.height, { notice: false });
const profileLabels = {
low_vram_draft: "Low VRAM draft",
low_vram_balanced: "Low VRAM balanced",
low_vram_turbo: "Low VRAM Turbo",
rtx3060_draft: "Low VRAM draft",
rtx3060_balanced: "Low VRAM balanced",
rtx3060_turbo: "Low VRAM Turbo",
@@ -10085,20 +10301,22 @@ function renderShotboardV3(node) {
showTimelineNotice(`Applied ${profileLabels[String(profile)] || "Low VRAM"} canvas/sampler profile. Timeline trims were not changed.`, "info");
};
addWidgetChoiceSetting("Hardware profile", "performance_profile", [
{ value: "rtx3060_draft", label: "Low VRAM draft" },
{ value: "rtx3060_balanced", label: "Low VRAM balanced" },
{ value: "rtx3060_turbo", label: "Low VRAM Turbo" },
{ value: "low_vram_draft", label: "Low VRAM draft" },
{ value: "low_vram_balanced", label: "Low VRAM balanced" },
{ value: "low_vram_turbo", label: "Low VRAM Turbo" },
{ value: "h3_native_quality", label: "H3 native quality" },
{ value: "custom", label: "Custom" },
], applyPerformanceProfile);
addWidgetChoiceSetting("Acceleration", "acceleration", [
{ value: "auto_3060", label: "Low VRAM Auto / H3 Sage" },
{ value: "low_vram_auto", label: "Low VRAM Auto / H3 Sage + exact chunks" },
{ value: "native", label: "Native" },
{ value: "h3_sage", label: "H3 Memory-Efficient Sage" },
{ value: "sage", label: "Sage" },
{ value: "sage_sol", label: "Sage + Sol / exp." },
{ value: "h3_sage", label: "H3 Sage + exact chunks" },
{ value: "sol_low_vram", label: "Sol + exact Low VRAM / exp." },
{ value: "sol_adaptive_safe", label: "Sol + Adaptive Safe / exp." },
{ value: "sol_adaptive_balanced", label: "Sol + Adaptive Balanced / exp." },
{ value: "adaptive_safe", label: "Adaptive Safe / exp." },
{ value: "spectrum", label: "Spectrum / exp." },
{ value: "sage_spectrum", label: "Sage + Spectrum" },
{ value: "sage_spectrum", label: "H3 Sage + Spectrum / exp." },
]);
addSetting("Seed", "seed", "1", "0");
addSetting("Seed stride", "seed_stride", "1", "0");
@@ -10121,7 +10339,7 @@ function renderShotboardV3(node) {
setDeckValue("shift_video", 12);
setDeckValue("turbo_sampler_mode", "audio_fixed");
setDeckValue("shift_audio", 3);
setDeckValue("acceleration", "auto_3060");
setDeckValue("acceleration", "low_vram_auto");
}
};
addWidgetChoiceSetting("Turbo mode", "turbo_mode", [
@@ -10155,14 +10373,13 @@ function renderShotboardV3(node) {
]);
addStaticH3Setting("Turbo audio", "Separate AV clocks", "Audio-fixed uses Larryvrh's sampler: video shift 12 and audio shift 3. Stock RES keeps the user-selected 4-6 audio shift but can distort audio at very low steps.");
settingsTarget = makeSettingsGroup("03", "REFERENCES & EXPERIMENTAL", "Reference semantics plus Sol and Spectrum quality/speed trade-offs.", true);
settingsTarget = makeSettingsGroup("03", "REFERENCES & EXPERIMENTAL", "Reference semantics plus Sol, Adaptive Cache and Spectrum quality/speed trade-offs.", true);
addWidgetChoiceSetting("Ref image size", "ref_image_size", [
{ value: "match", label: "Match canvas" },
{ value: "max", label: "Max / costly" },
]);
addWidgetChoiceSetting("Text encoder", "text_encoder_device", [
{ value: "cpu_safe_12gb", label: "CPU safe / Low VRAM" },
{ value: "auto", label: "Auto / high VRAM" },
{ value: "auto", label: "Auto / GPU-first, CPU only after OOM" },
]);
[1, 2, 3, 4].forEach((index) => addWidgetChoiceSetting(`Ref ${index} role`, `reference_role_${index}`, [
{ value: "subject_identity", label: "Subject" },
@@ -10186,12 +10403,12 @@ function renderShotboardV3(node) {
{ value: "sound_reference", label: "Sound reference" },
]);
addWidgetChoiceSetting("Sol sink", "sol_conditioning", [
{ value: "exact_kv", label: "Exact KV / faster" },
{ value: "exact_kv_and_rows", label: "Exact audio rows" },
{ value: "exact_kv", label: "Exact KV / faster" },
]);
addWidgetChoiceSetting("Spectrum", "spectrum_profile", [
{ value: "conservative_3060", label: "Low VRAM" },
{ value: "conservative_quality", label: "Quality / RAM high" },
{ value: "low_vram", label: "Low VRAM / degree 1" },
{ value: "quality", label: "Quality / RAM high" },
{ value: "aggressive", label: "Aggressive" },
]);
addSelectSetting("Legacy board resize", "image_resize_method", ["crop", "pad", "keep proportion", "stretch"]);
@@ -10209,13 +10426,40 @@ function renderShotboardV3(node) {
{ value: "ltx23", label: "LTX 2.3" },
{ value: "wan22_5b", label: "Wan 2.2 5B" },
]);
upscaleResolutionSelect = addResolutionPresetSetting("Upscale delivery", H3_UPSCALE_RESOLUTION_PRESETS, "upscale");
addSetting("Upscale width", "upscale_width", "8", "256");
addSetting("Upscale height", "upscale_height", "8", "256");
syncResolutionPresetControls();
addWidgetBoolSetting("Upscale Sage", "upscale_sage");
addSetting("Upscale seed offset", "upscale_seed_offset", "1", "0");
addSetting("Wan denoise", "wan_upscale_denoise", "0.01", "0");
addWidgetTextSetting("Upscale prompt (optional)", "upscale_prompt", "Leave empty to reuse the selected H3 global/chunk prompt.");
addStaticH3Setting("Upscale models", "Connected lazy branch", "Model and LoRA files are selected in the connected LTX/Wan workflow branch.");
addWidgetBoolSetting("LTX seam-safe VAE", "ltx_seam_safe");
addWidgetBoolSetting("LTX finishing LoRA", "ltx_detailer_enabled");
const ltxDetailerValues = getWidget(node, "ltx_detailer_lora_name")?.options?.values;
addSelectSetting("LTX LoRA — Crisp preset / click to change", "ltx_detailer_lora_name", Array.isArray(ltxDetailerValues) && ltxDetailerValues.length ? ltxDetailerValues : [""]);
addSetting("LTX LoRA strength", "ltx_detailer_strength", "0.05", "0");
addWidgetBoolSetting("RTX VSR final 4K", "ltx_4k_enabled");
settingsControls.get("ltx_4k_enabled")?.addEventListener("change", (event) => {
const enabled4K = String(event?.target?.value || "off") === "on";
if (!enabled4K) {
if (Number(getWidget(node, "upscale_width")?.value || 0) >= 3000) setDeckValue("upscale_width", 1920);
if (Number(getWidget(node, "upscale_height")?.value || 0) >= 1600) setDeckValue("upscale_height", 1080);
syncSettingsFromWidgets();
return;
}
setDeckValue("upscale_enabled", true);
setDeckValue("upscale_mode", "ltx23");
setDeckValue("upscale_width", 3840);
setDeckValue("upscale_height", 2160);
});
addWidgetChoiceSetting("RTX VSR quality", "ltx_4k_quality", [
{ value: "ULTRA", label: "Ultra" },
{ value: "HIGH", label: "High" },
{ value: "MEDIUM", label: "Medium" },
{ value: "LOW", label: "Low" },
]);
addStaticH3Setting("LTX delivery path", "Seam-safe / exact size", "A legal multiple-of-32 canvas is used inside LTX, then the result is resized to the exact requested dimensions. Optional 4K runs NVIDIA RTX VSR 2x only after the LTX stage. Crisp/enhance LoRAs work as finishing LoRAs; the official IC Detailer needs an IC-guided latent pipeline for its full effect.");
refreshPerformanceBand();
if (isShotboardV4) {
@@ -11132,6 +11376,68 @@ function renderShotboardV3(node) {
function isTimelineImageSegment(seg) {
return String(seg?.type || "image") === "image" && !seg?.placeholder;
}
function isTimelineImageSlot(seg) {
return String(seg?.type || "image") === "image";
}
const h3TaskModeValue = () => String(getWidget(node, "task_mode")?.value || "auto_from_timeline").trim().toLowerCase();
const h3ImageSlots = (items = timeline.segments || []) => (items || [])
.filter(isTimelineImageSlot)
.slice()
.sort((a, b) => Number(a.start || 0) - Number(b.start || 0));
const h3ImageAnchors = (items = timeline.segments || []) => (items || [])
.filter(isTimelineImageSegment)
.slice()
.sort((a, b) => Number(a.start || 0) - Number(b.start || 0));
const h3FlfAnchorMode = (items = timeline.segments || []) => {
const mode = h3TaskModeValue();
const count = h3ImageAnchors(items).length;
return count >= 2 && (["flf", "fflf", "fl2va"].includes(mode) || ["auto", "auto_from_timeline"].includes(mode));
};
const h3FlfLayoutMode = (items = timeline.segments || []) => {
const mode = h3TaskModeValue();
const slots = h3ImageSlots(items);
if (slots.length < 2) return false;
if (["flf", "fflf", "fl2va"].includes(mode)) return true;
if (!["auto", "auto_from_timeline"].includes(mode)) return false;
// In Auto, show the future FLF bridge immediately after the user adds
// an empty image slot. The backend contract still waits for two real
// images, so a placeholder can never become conditioning by mistake.
return h3ImageAnchors(items).length >= 1;
};
const h3NewImageSlotIsFlf = () => {
const mode = h3TaskModeValue();
if (["flf", "fflf", "fl2va"].includes(mode)) return true;
return ["auto", "auto_from_timeline"].includes(mode) && h3ImageSlots(timeline.segments || []).length >= 1;
};
const h3I2vHardCutMode = (items = timeline.segments || []) => {
const mode = h3TaskModeValue();
return h3ImageAnchors(items).length > 1 && ["i2v", "i2va"].includes(mode);
};
const buildH3BridgeContract = (items = timeline.segments || []) => {
const anchors = h3ImageAnchors(items);
if (!h3FlfAnchorMode(anchors)) return [];
return anchors.slice(0, -1).map((from, index) => {
const to = anchors[index + 1];
const startFrame = Math.max(0, Number(from.start || 0) + Number(from.length || 1) / 2);
const endFrame = Math.max(startFrame + 1, Number(to.start || 0) + Number(to.length || 1) / 2);
const prompt = String(from.prompt ?? from.local_prompt ?? from.relay_prompt ?? "");
return {
id: `h3_bridge_${String(from.id || index + 1)}_${String(to.id || index + 2)}`,
label: `FLF ${index + 1} -> ${index + 2}`,
from_segment_id: String(from.id || ""),
to_segment_id: String(to.id || ""),
from_ref: Math.max(1, Math.round(Number(from.ref || index + 1))),
to_ref: Math.max(1, Math.round(Number(to.ref || index + 2))),
visual_start_frame: Number(startFrame.toFixed(3)),
visual_end_frame: Number(endFrame.toFixed(3)),
visual_length_frames: Number((endFrame - startFrame).toFixed(3)),
prompt,
local_prompt: prompt,
enabled: Boolean(prompt.trim() && from.relay_manual_off !== true && from.promptrelay_manual_off !== true),
timing_policy: "image_box_centres_normalised_to_timeline",
};
});
};
function isActionBridgeRelaySegment(seg) {
return String(seg?.type || "") === "text" && Boolean(seg?.actionBridgeSourceId);
}
@@ -11571,6 +11877,7 @@ function renderShotboardV3(node) {
} else {
ensureDurationForFrames(Math.round(Number(source.start || 0) + Number(source.length || 1)) + slotLength);
}
const flfSlot = h3NewImageSlotIsFlf();
const slot = {
id: newId("slot"),
type: "image",
@@ -11581,8 +11888,8 @@ function renderShotboardV3(node) {
label: "empty_slot",
prompt: "",
note: "",
camera: "cut to",
transition: "hard_cut",
camera: flfSlot ? "continuous dolly-in" : "cut to",
transition: flfSlot ? "continuous_motion" : "hard_cut",
guideStrength: Number(defaultForceWidget?.value || 0.25),
imageLockStrength: Number(defaultForceWidget?.value || 0.25),
defaultForceSource: Number(defaultForceWidget?.value || 0.25),
@@ -11908,6 +12215,7 @@ function renderShotboardV3(node) {
ensureDurationForFrames(endOfSegments(sorted) + 1);
}
const splitStart = Math.round(Number(source.start || 0) + Number(source.length || 1));
const flfSlot = kind === "image" && h3NewImageSlotIsFlf();
const placeholder = kind === "text" ? textRelaySegment(splitStart, splitLen) : {
id: newId("slot"),
type: "image",
@@ -11918,8 +12226,8 @@ function renderShotboardV3(node) {
label: "empty_slot",
prompt: "",
note: "",
camera: "continuous dolly-in",
transition: "continuous_motion",
camera: flfSlot ? "continuous dolly-in" : "cut to",
transition: flfSlot ? "continuous_motion" : "hard_cut",
guideStrength: Number(defaultForceWidget?.value || 0.25),
imageLockStrength: Number(defaultForceWidget?.value || 0.25),
defaultForceSource: Number(defaultForceWidget?.value || 0.25),
@@ -11936,6 +12244,7 @@ function renderShotboardV3(node) {
}
const length = defaultLen();
ensureDurationForFrames(Math.round(start) + length);
const flfSlot = kind === "image" && h3NewImageSlotIsFlf();
const placeholder = kind === "text" ? textRelaySegment(start, length) : {
id: newId("slot"),
type: "image",
@@ -11946,8 +12255,8 @@ function renderShotboardV3(node) {
label: "empty_slot",
prompt: "",
note: "",
camera: "continuous dolly-in",
transition: "continuous_motion",
camera: flfSlot ? "continuous dolly-in" : "cut to",
transition: flfSlot ? "continuous_motion" : "hard_cut",
guideStrength: Number(defaultForceWidget?.value || 0.25),
imageLockStrength: Number(defaultForceWidget?.value || 0.25),
defaultForceSource: Number(defaultForceWidget?.value || 0.25),
@@ -12795,7 +13104,8 @@ function renderShotboardV3(node) {
);
block.appendChild(rail);
}
if (!isAudio) {
const promptOwnedByFlfBridge = !isAudio && isTimelineImageSlot(seg) && h3FlfLayoutMode(activeVisualSegments());
if (!isAudio && !promptOwnedByFlfBridge) {
const caption = document.createElement("textarea");
caption.value = String(seg.prompt || "");
caption.placeholder = "Action in this segment...";
@@ -12856,6 +13166,20 @@ function renderShotboardV3(node) {
protectControlDrag(caption);
block.appendChild(caption);
}
else if (promptOwnedByFlfBridge) {
const anchors = h3ImageSlots(activeVisualSegments());
const anchorIndex = anchors.findIndex((item) => String(item.id || "") === String(seg.id || ""));
const role = document.createElement("div");
const pending = Boolean(seg.placeholder || !segmentReferencePath(seg));
role.textContent = pending
? (anchorIndex === anchors.length - 1 ? "FINAL FLF KEYFRAME — ADD IMAGE" : `FLF KEYFRAME ${anchorIndex + 1} — ADD IMAGE`)
: (anchorIndex === anchors.length - 1 ? "FINAL FLF KEYFRAME" : `FLF KEYFRAME ${anchorIndex + 1}`);
role.title = anchorIndex === anchors.length - 1
? "This image closes the previous FLF bridge. It does not create an extra chunk."
: "The editable local prompt is displayed between this image centre and the next image centre.";
role.style.cssText = `position:absolute;left:${innerLeft}px;right:${innerRight}px;top:${promptTop}px;height:${promptHeight}px;display:flex;align-items:center;justify-content:center;box-sizing:border-box;border:1px dashed rgba(223,164,81,.38);border-radius:5px;background:rgba(22,18,14,.28);color:rgba(244,229,196,.56);font:9px/1.2 monospace;font-weight:900;letter-spacing:.06em;text-align:center;pointer-events:none;`;
block.appendChild(role);
}
const label = document.createElement("div");
label.style.cssText = `position:absolute;left:${innerLeft}px;top:${truthRailHeight + 4}px;right:${topRightSafe}px;color:#fff;font-size:10px;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;text-shadow:0 1px 2px #000;`;
label.textContent = `${frameLabel(seg.start)} - ${frameLabel(Number(seg.start || 0) + Number(seg.length || 0))}`;
@@ -13395,6 +13719,87 @@ function renderShotboardV3(node) {
});
}
function drawH3PromptModeOverlays(segments) {
const total = Math.max(1, getTotalFrames());
const fps = getFps();
const anchors = h3ImageAnchors(segments);
if (h3FlfLayoutMode(segments)) {
const layoutAnchors = h3ImageSlots(segments);
layoutAnchors.slice(0, -1).forEach((from, index) => {
const to = layoutAnchors[index + 1];
const fromCenter = Math.max(0, Number(from.start || 0) + Number(from.length || 1) / 2);
const toCenter = Math.max(fromCenter + 1, Number(to.start || 0) + Number(to.length || 1) / 2);
const bridge = document.createElement("div");
bridge.className = "iamccs-h3-flf-centre-prompt-bridge";
bridge.style.cssText = [
"position:absolute",
`left:${(fromCenter / total) * 100}%`,
`width:${Math.max(1, ((toCenter - fromCenter) / total) * 100)}%`,
"min-width:0",
"top:154px",
`height:${74 + timelineExtraH}px`,
"box-sizing:border-box",
"padding:19px 6px 5px",
"border:1px solid rgba(223,164,81,.86)",
"border-radius:7px",
"background:linear-gradient(90deg,rgba(58,43,25,.96),rgba(18,48,50,.96))",
"box-shadow:0 6px 18px rgba(0,0,0,.52),inset 0 1px 0 rgba(255,255,255,.13)",
"z-index:72",
"overflow:visible",
].join(";");
bridge.title = `FLF chunk ${index + 1}: ${String(from.label || `keyframe ${index + 1}`)} -> ${String(to.label || `keyframe ${index + 2}`)}. The visual span follows the two image-box centres.`;
const header = document.createElement("div");
header.textContent = `FLF ${index + 1} -> ${index + 2} | CENTRE TO CENTRE | ${((toCenter - fromCenter) / fps).toFixed(2)}s WEIGHT`;
header.style.cssText = "position:absolute;left:6px;right:6px;top:4px;height:12px;color:#F4D49E;font:8px/1 monospace;font-weight:950;letter-spacing:.025em;text-align:center;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;pointer-events:none;";
const prompt = document.createElement("textarea");
prompt.value = String(from.prompt ?? from.local_prompt ?? from.relay_prompt ?? "");
prompt.placeholder = `Local MiniMax prompt for FLF ${index + 1} -> ${index + 2}`;
prompt.spellcheck = false;
prompt.dataset.iamccsV3SegmentId = String(from.id || "");
prompt.dataset.iamccsV3Key = "prompt";
prompt.style.cssText = `width:100%;height:100%;box-sizing:border-box;padding:6px 8px;border:1px solid rgba(118,181,177,.62);border-radius:5px;background:#F4EFE6;color:#111;font:${promptFontSize(10)}/1.25 monospace;font-weight:750;resize:none;overflow:auto;outline:none;pointer-events:auto;`;
prompt.onpointerdown = (event) => event.stopPropagation();
prompt.onclick = (event) => event.stopPropagation();
prompt.ondblclick = (event) => event.stopPropagation();
prompt.oninput = () => {
markPromptFieldEdited(prompt);
from.prompt = prompt.value;
from.local_prompt = prompt.value;
from.relay_prompt = prompt.value;
from.use_prompt = Boolean(String(prompt.value || "").trim());
if (from.use_prompt) {
from.relay_manual_off = false;
from.promptrelay_manual_off = false;
}
syncSegmentTextPeers(from.id, "prompt", prompt.value, prompt);
syncSegmentRelayPeers(from.id, Boolean(from.use_prompt), null);
writeTimeline({ force: true });
};
prompt.onchange = flushTimelineWrite;
prompt.onblur = flushTimelineWrite;
protectControlDrag(prompt);
bridge.append(header, prompt);
imageTrack.appendChild(bridge);
});
return;
}
if (h3I2vHardCutMode(anchors)) {
anchors.slice(1).forEach((current, index) => {
const previous = anchors[index];
const previousEnd = Number(previous.start || 0) + Number(previous.length || 1);
const currentStart = Number(current.start || 0);
const frame = Math.max(0, Math.min(total, Math.abs(currentStart - previousEnd) <= 1 ? currentStart : (previousEnd + currentStart) / 2));
const marker = document.createElement("div");
marker.className = "iamccs-h3-i2v-hard-cut-marker";
marker.innerHTML = "<span>HARD CUT</span>";
marker.style.cssText = `position:absolute;left:${(frame / total) * 100}%;top:145px;bottom:8px;width:2px;transform:translateX(-1px);background:#E56B5D;box-shadow:0 0 0 1px rgba(0,0,0,.65),0 0 10px rgba(229,107,93,.55);z-index:74;pointer-events:none;`;
const label = marker.querySelector("span");
if (label) label.style.cssText = "position:absolute;left:50%;top:2px;transform:translate(-50%,-100%);padding:2px 5px;border:1px solid rgba(255,178,166,.85);border-radius:4px;background:#52251F;color:#FFE9E3;font:8px/1 monospace;font-weight:950;white-space:nowrap;";
imageTrack.appendChild(marker);
});
}
}
function drawStepTransitionBridges(segments) {
const total = Math.max(1, getTotalFrames());
const visualSegments = (segments || [])
@@ -14844,6 +15249,7 @@ function renderShotboardV3(node) {
}
visualSegments.forEach((seg) => imageTrack.appendChild(makeBlock(seg, false)));
if (!visualSegments.length) imageTrack.appendChild(makeImagePlaceholderBlock());
drawH3PromptModeOverlays(visualSegments);
drawVisualEdgeHandles(visualSegments);
drawIcLoraTrack(motionSegmentsForDraw);
audioSegmentsForDraw.forEach((seg) => audioTracks.appendChild(makeBlock(seg, true)));
+169 -19
View File
@@ -223,12 +223,15 @@ function safeProject(raw) {
try { parsed = JSON.parse(String(raw || "{}")); } catch {}
return {
schema: "iamccs.minimax_h3.prompter_project",
schema_version: 1,
schema_version: 2,
project_name: String(parsed.project_name || "Untitled H3 Prompt"),
task_mode: MODE_META[parsed.task_mode] ? parsed.task_mode : "t2va",
injection_target: ["global", "local_auto", "local_1", "local_2", "local_3"].includes(parsed.injection_target) ? parsed.injection_target : "global",
writing_mode: ["manual", "guided", "assistant_fill"].includes(parsed.writing_mode) ? parsed.writing_mode : "guided",
merge_policy: ["replace", "append"].includes(parsed.merge_policy) ? parsed.merge_policy : "replace",
ai_direction: String(parsed.ai_direction || ""),
ai_scope: String(parsed.ai_scope || "active_field"),
ai_visual_roles: parsed.ai_visual_roles && typeof parsed.ai_visual_roles === "object" ? { ...parsed.ai_visual_roles } : {},
sections: parsed.sections && typeof parsed.sections === "object" ? { ...parsed.sections } : {},
};
}
@@ -320,8 +323,11 @@ function mountPrompter(node) {
.iamccs-pr-ai.show{display:grid}.iamccs-pr-ai-title{color:#9fc9ef;font-size:10px;font-weight:800;letter-spacing:.08em;text-transform:uppercase}
.iamccs-pr-ai-row{display:grid;grid-template-columns:1fr 1fr;gap:6px}.iamccs-pr-ai label{display:grid;gap:3px;color:#8999aa;font-size:9px;font-weight:700}
.iamccs-pr-ai input,.iamccs-pr-ai select{width:100%;height:29px;border:1px solid #35485b;border-radius:5px;background:#0e151d;color:#e7eef5;padding:0 6px;font-size:10px}
.iamccs-pr-ai textarea{width:100%;min-height:66px;resize:vertical;border:1px solid #35485b;border-radius:5px;background:#0e151d;color:#e7eef5;padding:7px;font:10px/1.4 Inter,Segoe UI,sans-serif}
.iamccs-pr-ai-status{min-height:28px;color:#91a4b5;font-size:9px;line-height:1.35}.iamccs-pr-ai-status.ok{color:#8fd1aa}.iamccs-pr-ai-status.error{color:#ed9c92}
.iamccs-pr-ai .iamccs-pr-btn{width:100%;border-color:#6094c0;background:#274866;color:#eef7ff}
.iamccs-pr-ai-modelrow{display:grid;grid-template-columns:minmax(0,1fr) 30px;gap:5px}.iamccs-pr-ai-modelrow .iamccs-pr-btn{height:29px;padding:0!important}
.iamccs-pr-ai-images{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:5px}.iamccs-pr-ai-image{display:grid;grid-template-columns:44px minmax(0,1fr);gap:5px;padding:4px;border:1px solid #304255;border-radius:6px;background:#0c141c;min-width:0}.iamccs-pr-ai-thumb{width:44px;height:44px;object-fit:cover;border-radius:4px;background:#202832}.iamccs-pr-ai-image-meta{display:grid;gap:3px;min-width:0}.iamccs-pr-ai-image-name{overflow:hidden;text-overflow:ellipsis;white-space:nowrap;color:#aebdca;font-size:8px}.iamccs-pr-ai-image select{height:24px!important;font-size:8px!important}.iamccs-pr-ai-file{display:none}
.iamccs-pr-example-select{height:30px;max-width:146px;border:1px solid #3b4350;border-radius:6px;background:#171b23;color:#e9edf2;padding:0 6px;font:600 10px Inter,Segoe UI,sans-serif}
.iamccs-pr-inject{width:100%;height:38px!important;margin:0 0 7px;background:linear-gradient(135deg,#d3a447,#8d5c20)!important;border:1px solid #f0ca7d!important;color:#171109!important;font-size:12px!important;font-weight:900!important;letter-spacing:.06em;box-shadow:0 5px 14px #0007}
.iamccs-pr-inject-status{min-height:30px;margin-bottom:12px;padding:7px;border:1px solid #303944;border-radius:6px;background:#151b22;color:#91a0ae;font-size:9px;line-height:1.35}.iamccs-pr-inject-status.ok{border-color:#3f7957;color:#9fe0b7}.iamccs-pr-inject-status.error{border-color:#824b45;color:#efaaa1}
@@ -407,13 +413,24 @@ function mountPrompter(node) {
const assistantHint = el("div", "iamccs-pr-hint", "AI Rewrite treats every filled box as your rough idea, then rewrites those same boxes into MiniMax H3-ready English in one request. Blank boxes stay blank and your project remains editable before queueing.");
left.appendChild(assistantHint);
const aiPanel = el("div", "iamccs-pr-ai");
aiPanel.appendChild(el("div", "iamccs-pr-ai-title", "Assistant engine"));
aiPanel.appendChild(el("div", "iamccs-pr-ai-title", "Autonomous MiniMax assistant"));
const aiScope = el("select");
const aiDirection = el("textarea");
aiDirection.placeholder = "Your direction for the AI: what to preserve, emphasize, simplify or change. The rough idea remains in the selected prompt field.";
const aiScopeLabel = el("label", "", "Improve target"); aiScopeLabel.appendChild(aiScope);
const aiDirectionLabel = el("label", "", "User direction (applies only when AI Rewrite is active)"); aiDirectionLabel.appendChild(aiDirection);
aiPanel.append(aiScopeLabel, aiDirectionLabel);
const aiProvider = el("select");
aiProvider.innerHTML = `<option value="ollama">Ollama / local</option><option value="openai_compatible">OpenAI-compatible</option><option value="gemini">Google Gemini</option><option value="anthropic">Anthropic</option>`;
const aiBaseUrl = el("input");
aiBaseUrl.placeholder = "Provider base URL";
const aiModel = el("input");
aiModel.placeholder = "Model name";
const aiModelList = el("datalist");
aiModelList.id = `iamccs-prompter-models-${node.id || Math.random().toString(16).slice(2)}`;
aiModel.setAttribute("list", aiModelList.id);
const refreshModelsBtn = button("↻");
refreshModelsBtn.title = "Read the models installed in Ollama";
const aiApiKey = el("input");
aiApiKey.type = "password";
aiApiKey.autocomplete = "off";
@@ -426,20 +443,28 @@ function mountPrompter(node) {
aiTemperature.value = "0.35";
const aiRow1 = el("div", "iamccs-pr-ai-row");
const providerLabel = el("label", "", "Provider"); providerLabel.appendChild(aiProvider);
const modelLabel = el("label", "", "Model"); modelLabel.appendChild(aiModel);
const modelLabel = el("label", "", "Model");
const modelRow = el("div", "iamccs-pr-ai-modelrow"); modelRow.append(aiModel, refreshModelsBtn, aiModelList); modelLabel.appendChild(modelRow);
aiRow1.append(providerLabel, modelLabel);
const aiRow2 = el("div", "iamccs-pr-ai-row");
const urlLabel = el("label", "", "Base URL"); urlLabel.appendChild(aiBaseUrl);
const tempLabel = el("label", "", "Creativity"); tempLabel.appendChild(aiTemperature);
aiRow2.append(urlLabel, tempLabel);
const keyLabel = el("label", "", "API key (never saved)"); keyLabel.appendChild(aiApiKey);
const rewriteBtn = button("Rewrite filled fields with AI");
const aiStatus = el("div", "iamccs-pr-ai-status", "Ollama works locally. Cloud keys are sent only to the selected provider and are not stored in the workflow.");
aiPanel.append(aiRow1, aiRow2, keyLabel, rewriteBtn, aiStatus);
const aiImageInput = el("input", "iamccs-pr-ai-file");
aiImageInput.type = "file";
aiImageInput.accept = "image/png,image/jpeg,image/webp";
aiImageInput.multiple = true;
const addAIImagesBtn = button("Add up to 4 AI image references");
const aiImages = el("div", "iamccs-pr-ai-images");
const rewriteBtn = button("Improve selected prompt with AI");
const aiStatus = el("div", "iamccs-pr-ai-status", "Ollama is local. Choose the field to improve; cloud keys are never stored in the workflow.");
aiPanel.append(aiRow1, aiRow2, keyLabel, addAIImagesBtn, aiImageInput, aiImages, rewriteBtn, aiStatus);
left.appendChild(aiPanel);
const center = el("main", "iamccs-pr-center");
let activePromptArea = null;
let activePromptKey = null;
const tagDeck = el("section", "iamccs-pr-tagdeck");
const tagHead = el("div", "iamccs-pr-taghead");
tagHead.append(el("div", "iamccs-pr-tagtitle", "MINIMAX H3 PROMPT TAGS"));
@@ -523,6 +548,8 @@ function mountPrompter(node) {
const commit = () => {
project.project_name = nameInput.value.trim() || "Untitled H3 Prompt";
project.merge_policy = policy.value;
project.ai_direction = aiDirection.value;
project.ai_scope = aiScope.value || "active_field";
setWidget(node, "project_data", JSON.stringify(project));
setWidget(node, "task_mode", project.task_mode);
setWidget(node, "injection_target", project.injection_target);
@@ -532,7 +559,7 @@ function mountPrompter(node) {
};
const aiDefaults = {
ollama: { baseUrl: "http://127.0.0.1:11434", model: "qwen3:8b" },
ollama: { baseUrl: "http://127.0.0.1:11434", model: "" },
openai_compatible: { baseUrl: "https://api.openai.com/v1", model: "gpt-4.1-mini" },
gemini: { baseUrl: "https://generativelanguage.googleapis.com/v1beta", model: "gemini-2.5-flash" },
anthropic: { baseUrl: "https://api.anthropic.com/v1", model: "claude-sonnet-4-5" },
@@ -551,14 +578,85 @@ function mountPrompter(node) {
temperature: Number(aiTemperature.value || 0.35),
};
};
aiProvider.onchange = () => {
const aiVisualFiles = [];
const visualRolesForTarget = () => {
project.ai_visual_roles = project.ai_visual_roles && typeof project.ai_visual_roles === "object" ? project.ai_visual_roles : {};
const key = project.injection_target || "global";
project.ai_visual_roles[key] = project.ai_visual_roles[key] && typeof project.ai_visual_roles[key] === "object" ? project.ai_visual_roles[key] : {};
return project.ai_visual_roles[key];
};
const readFileDataUrl = (file) => new Promise((resolve, reject) => {
const reader = new FileReader();
reader.onload = () => resolve(String(reader.result || ""));
reader.onerror = () => reject(reader.error || new Error("Unable to read image"));
reader.readAsDataURL(file);
});
const renderAIImages = () => {
aiImages.replaceChildren();
const roles = visualRolesForTarget();
aiVisualFiles.forEach((item, index) => {
const slot = String(index + 1);
const card = el("div", "iamccs-pr-ai-image");
const thumb = el("img", "iamccs-pr-ai-thumb");
thumb.src = item.dataUrl;
const meta = el("div", "iamccs-pr-ai-image-meta");
meta.appendChild(el("div", "iamccs-pr-ai-image-name", `Picture ${slot} · ${item.file.name}`));
const role = el("select");
["ignore", "opening", "closing", "identity", "composition", "style", "reference"].forEach((value) => {
const option = document.createElement("option"); option.value = value; option.textContent = value; role.appendChild(option);
});
role.value = String(roles[slot] || (index === 0 ? "opening" : index === 1 ? "closing" : "reference"));
role.onchange = () => { visualRolesForTarget()[slot] = role.value; commit(); };
meta.appendChild(role);
card.append(thumb, meta);
aiImages.appendChild(card);
});
};
const loadOllamaModels = async ({ quiet = false } = {}) => {
if (aiProvider.value !== "ollama") return [];
refreshModelsBtn.disabled = true;
if (!quiet) aiStatus.textContent = "Reading installed Ollama models…";
try {
const response = await api.fetchApi(`/iamccs/prompter/ollama/models?base_url=${encodeURIComponent(aiBaseUrl.value.trim() || "http://127.0.0.1:11434")}`);
const data = await response.json();
if (!response.ok || !data?.ok) throw new Error(data?.error || `HTTP ${response.status}`);
const names = (data.models || []).map((item) => String(item.name || "")).filter(Boolean);
aiModelList.replaceChildren(...names.map((name) => {
const option = document.createElement("option"); option.value = name; return option;
}));
if ((!aiModel.value.trim() || !names.includes(aiModel.value.trim())) && names.length) aiModel.value = names[0];
persistAI();
aiStatus.className = "iamccs-pr-ai-status ok";
aiStatus.textContent = names.length ? `${names.length} Ollama model(s) available. Selected: ${aiModel.value}.` : "Ollama is reachable but has no installed models.";
return names;
} catch (error) {
aiStatus.className = "iamccs-pr-ai-status error";
aiStatus.textContent = `Ollama unavailable: ${error?.message || error}`;
return [];
} finally {
refreshModelsBtn.disabled = false;
}
};
aiProvider.onchange = async () => {
const selected = aiDefaults[aiProvider.value] || {};
aiBaseUrl.value = selected.baseUrl || "";
aiModel.value = selected.model || "";
aiApiKey.value = "";
persistAI();
if (aiProvider.value === "ollama") await loadOllamaModels();
};
[aiBaseUrl, aiModel, aiTemperature].forEach((control) => control.addEventListener("change", persistAI));
refreshModelsBtn.onclick = () => loadOllamaModels();
addAIImagesBtn.onclick = () => aiImageInput.click();
aiImageInput.onchange = async () => {
const files = Array.from(aiImageInput.files || []).filter((file) => /^image\//.test(file.type)).slice(0, 4);
aiVisualFiles.splice(0, aiVisualFiles.length);
for (const file of files) aiVisualFiles.push({ file, dataUrl: await readFileDataUrl(file) });
renderAIImages();
aiImageInput.value = "";
aiStatus.className = "iamccs-pr-ai-status";
aiStatus.textContent = `${aiVisualFiles.length} temporary AI image reference(s). Assign roles for ${project.injection_target}; images are not saved inside the workflow.`;
};
const renderPreview = () => {
const prompt = composePrompt(project);
@@ -576,6 +674,7 @@ function mountPrompter(node) {
const renderSections = () => {
center.replaceChildren();
activePromptArea = null;
activePromptKey = null;
center.appendChild(tagDeck);
const meta = MODE_META[project.task_mode];
meta.sections.forEach(([key, label, tip], index) => {
@@ -585,8 +684,10 @@ function mountPrompter(node) {
const state = el("div", "iamccs-pr-state");
head.appendChild(state);
const area = el("textarea", "iamccs-pr-text");
area.dataset.sectionKey = key;
area.addEventListener("focus", () => {
activePromptArea = area;
activePromptKey = key;
tagHint.textContent = `Active field: ${label}`;
});
area.value = String(project.sections?.[key] || "");
@@ -611,6 +712,21 @@ function mountPrompter(node) {
renderPreview();
};
const populateAIScope = () => {
const previous = String(project.ai_scope || aiScope.value || "active_field");
aiScope.replaceChildren();
const choices = [
["active_field", "Active prompt field"],
["all_filled", "All filled fields"],
...MODE_META[project.task_mode].sections.map(([key, label]) => [key, `Section · ${label}`]),
];
choices.forEach(([value, label]) => {
const option = document.createElement("option"); option.value = value; option.textContent = label; aiScope.appendChild(option);
});
aiScope.value = choices.some(([value]) => value === previous) ? previous : "active_field";
project.ai_scope = aiScope.value;
};
const renderControls = () => {
nameInput.value = project.project_name;
policy.value = project.merge_policy;
@@ -618,8 +734,11 @@ function mountPrompter(node) {
exampleSelect.title = project.task_mode === "t2va" ? "Choose a cinematic T2V prompt project" : "T2V cinematic projects are available in T2VA mode";
targetButtons.forEach((item, key) => item.classList.toggle("active", key === project.injection_target));
writingButtons.forEach((item, key) => item.classList.toggle("active", key === project.writing_mode));
populateAIScope();
aiDirection.value = String(project.ai_direction || "");
root.classList.toggle("mode-manual", project.writing_mode === "manual");
aiPanel.classList.toggle("show", project.writing_mode === "assistant_fill");
renderAIImages();
targetHint.textContent = project.injection_target === "local_auto"
? "The MiniMax Shotboard reads its timeline, selects the first empty local slot among 1–3, and appends to Local 3 only when all three already contain text."
: project.injection_target === "global"
@@ -639,22 +758,46 @@ function mountPrompter(node) {
};
rewriteBtn.onclick = async () => {
const filled = Object.fromEntries(
MODE_META[project.task_mode].sections
.map(([key]) => [key, String(project.sections?.[key] || "").trim()])
.filter(([, value]) => value)
const allSections = Object.fromEntries(
MODE_META[project.task_mode].sections.map(([key]) => [key, String(project.sections?.[key] || "").trim()])
);
if (!Object.keys(filled).length) {
let targetKeys = [];
if (aiScope.value === "all_filled") {
targetKeys = Object.entries(allSections).filter(([, value]) => value).map(([key]) => key);
} else if (aiScope.value === "active_field") {
if (activePromptKey) targetKeys = [activePromptKey];
} else if (Object.prototype.hasOwnProperty.call(allSections, aiScope.value)) {
targetKeys = [aiScope.value];
}
const direction = aiDirection.value.trim();
const hasRoughText = targetKeys.some((key) => String(allSections[key] || "").trim());
if (!targetKeys.length) {
aiStatus.className = "iamccs-pr-ai-status error";
aiStatus.textContent = "Write a rough idea in at least one field first.";
aiStatus.textContent = aiScope.value === "active_field" ? "Click the prompt field you want the AI to improve first." : "No filled field is available for this target.";
return;
}
if (!hasRoughText && !direction && !aiVisualFiles.length) {
aiStatus.className = "iamccs-pr-ai-status error";
aiStatus.textContent = "Write a rough idea in the selected field or in User direction first.";
return;
}
project.ai_direction = direction;
project.ai_scope = aiScope.value;
persistAI();
commit();
rewriteBtn.disabled = true;
rewriteBtn.textContent = "Rewriting MiniMax fields…";
aiStatus.className = "iamccs-pr-ai-status";
aiStatus.textContent = `Sending ${Object.keys(filled).length} filled section(s) to ${aiProvider.options[aiProvider.selectedIndex]?.text || aiProvider.value}.`;
aiStatus.textContent = `Sending ${targetKeys.join(", ")} to ${aiProvider.options[aiProvider.selectedIndex]?.text || aiProvider.value}.`;
try {
const roles = visualRolesForTarget();
const imagePayload = aiVisualFiles.map((item, index) => ({
slot: index + 1,
name: item.file.name,
role: String(roles[String(index + 1)] || (index === 0 ? "opening" : index === 1 ? "closing" : "reference")),
mime_type: item.file.type || "image/png",
data: item.dataUrl,
})).filter((item) => item.role !== "ignore");
const response = await api.fetchApi("/iamccs/prompter/rewrite", {
method: "POST",
headers: { "Content-Type": "application/json" },
@@ -664,7 +807,10 @@ function mountPrompter(node) {
model: aiModel.value.trim(),
api_key: aiApiKey.value,
task_mode: project.task_mode,
sections: filled,
sections: allSections,
target_keys: targetKeys,
user_direction: direction,
images: imagePayload,
temperature: Number(aiTemperature.value || 0.35),
timeout: 180,
}),
@@ -672,20 +818,21 @@ function mountPrompter(node) {
const data = await response.json();
if (!response.ok || !data?.ok) throw new Error(data?.error || `HTTP ${response.status}`);
Object.entries(data.sections || {}).forEach(([key, value]) => {
if (Object.prototype.hasOwnProperty.call(filled, key)) project.sections[key] = String(value || "");
if (targetKeys.includes(key)) project.sections[key] = String(value || "");
});
renderControls();
renderSections();
commit();
aiStatus.className = "iamccs-pr-ai-status ok";
aiStatus.textContent = `Rewritten: ${(data.report?.rewritten_sections || Object.keys(data.sections || {})).join(", ")}. Review the fields, then save or queue.`;
const visualCount = Number(data.report?.visual_references?.length || 0);
aiStatus.textContent = `Improved: ${(data.report?.rewritten_sections || Object.keys(data.sections || {})).join(", ")}${visualCount ? ` with ${visualCount} visual reference(s)` : ""}. Review, then inject.`;
} catch (error) {
aiStatus.className = "iamccs-pr-ai-status error";
aiStatus.textContent = `Rewrite failed: ${error?.message || error}`;
} finally {
aiApiKey.value = "";
rewriteBtn.disabled = false;
rewriteBtn.textContent = "Rewrite filled fields with AI";
rewriteBtn.textContent = "Improve selected prompt with AI";
}
};
@@ -745,6 +892,8 @@ function mountPrompter(node) {
commit();
};
});
aiScope.onchange = () => { project.ai_scope = aiScope.value; commit(); };
aiDirection.addEventListener("input", () => { project.ai_direction = aiDirection.value; commit(); });
nameInput.addEventListener("input", commit);
policy.addEventListener("change", commit);
exampleBtn.onclick = () => loadExample(project.task_mode);
@@ -810,6 +959,7 @@ function mountPrompter(node) {
renderControls();
renderSections();
commit();
if (aiProvider.value === "ollama") setTimeout(() => loadOllamaModels({ quiet: true }), 0);
}
app.registerExtension({