update h3 prompter node
@@ -11,3 +11,9 @@ QWEN_IMAGE_POC_PLAN.md
|
||||
TESTING_CHECKLIST.md
|
||||
ZIMAGE_SUPPORT_NOTES.md
|
||||
tests/
|
||||
|
||||
# Local test outputs, logs, extracted frames and temporary videos
|
||||
/test/
|
||||
|
||||
# Local agent instructions
|
||||
/AGENTS.md
|
||||
|
||||
|
Before Width: | Height: | Size: 846 KiB |
|
Before Width: | Height: | Size: 1.9 MiB |
|
Before Width: | Height: | Size: 550 KiB |
|
Before Width: | Height: | Size: 682 KiB |
|
Before Width: | Height: | Size: 957 KiB |
|
Before Width: | Height: | Size: 985 KiB |
|
Before Width: | Height: | Size: 2.1 MiB |
|
Before Width: | Height: | Size: 46 KiB |
|
Before Width: | Height: | Size: 66 KiB |
@@ -2,6 +2,20 @@
|
||||
|
||||
This is a custom node that collects the tools I use frequently.
|
||||
|
||||
### MiniMax H3 camera render
|
||||
|
||||
Enable **Camera render** in the camera panel to expose a `camera_render` IMAGE output.
|
||||
On node execution it renders the full panel camera timeline at **1024×1024, 24 fps**
|
||||
as a CPU float32 image batch. Connect it to Save Image for a PNG sequence or a video
|
||||
combine node (24 fps). The existing `length` output is its frame count.
|
||||
Rendering is off by default and writes no files itself. Allow roughly 1.5 GiB RAM
|
||||
per five seconds of output; long sequences require proportionally more RAM.
|
||||
It renders the preview mannequin, ground grid and white edges, without editor overlays.
|
||||
Shot cuts, Move interpolation, composition and roll match the preview. Text-only
|
||||
Motion presets and Qwen/user-text target changes are not 3D-solved and therefore
|
||||
are not animated in this output. This checkbox applies to the entire timeline and
|
||||
is separate from each item's prompt-camera enable checkbox.
|
||||
|
||||
https://github.com/toyxyz/ComfyUI_toyxyz_test_nodes/assets/8006000/8536e96a-514a-48b2-b1aa-8eccbd3fa853
|
||||
|
||||
(This video is at 4x speed)
|
||||
@@ -62,6 +76,17 @@ Direct Webcam capture workflow (without webcam app)
|
||||
|
||||
## Minimax-H3-prompter
|
||||
|
||||
### Camera Advanced path controls
|
||||
|
||||
- **Direction** selects the destination viewpoint; **Orbit route** selects how a Move reaches it in camera-relative coordinates. Choose camera-left or camera-right explicitly when the route matters. Shortest path preserves the previous behavior, with a leftward tie at 180 degrees. Identical directions do not imply a full revolution. The route control is inactive on a Shot opening, which establishes a new take.
|
||||
- Preview and compiled camera text share the signed orbit route. A profile-to-opposite-profile half-circle names the intervening front or rear view, rather than relying on the destination alone.
|
||||
- Continuous paths interpolate distance, elevation, azimuth and framing target with shape-preserving shared waypoint tangents. Reversing components slow at the boundary; the view does not reset. Camera height is derived from those quantities, so height need not be monotonic when both framing target and angle change.
|
||||
- **Composition** selects the screen position of the framed subject region: center, left/right thirds, upper/lower thirds, or four corner-third positions. It is independent of orbital Direction. The preview interpolates an off-axis framing offset rather than inventing subject movement or a physical orbit. Distance can increase to retain the requested body range near an edge; this can make the subject smaller. Camera prose includes the continuous reframing and resulting destination placement.
|
||||
- The Full camera timeline UI box has been removed. Verified camera sentences and diagnostic metadata remain in the raw plan. Diagnostics are not rendered into the final camera sentence. Ground-plane conflicts are reported, not silently corrected. The fast-orbit warning is a planning heuristic, not a measured model limit.
|
||||
- Geometry remains a standing mannequin proxy with a square preview, not a constraint on output aspect ratio or a simulation of arbitrary subject poses. Text-only video generation can still miss paths or insert cuts despite valid geometry.
|
||||
|
||||
Camera checks: `python -m unittest test_advanced_camera_prompt.py` and `node test_camera_geometry.mjs`.
|
||||
|
||||
<img width="2190" height="1624" alt="image" src="https://github.com/user-attachments/assets/fc97abbe-d8ee-498b-8b5c-f248663cb749" />
|
||||
|
||||
|
||||
@@ -80,8 +105,8 @@ audio references. Supported modes are `Auto`, `T2VA`, `I2VA`, `FL2VA`, `L2VA`, a
|
||||
1. Select a mode, duration, and model. `Auto` chooses a mode from the reference layout.
|
||||
2. Describe each shot naturally in **Prompt**, including actions, camera direction, dialogue,
|
||||
visible text, sound, and music.
|
||||
3. Optionally use the preset menu below Prompt. **Camera** provides angle, motion, shot framing, motion-amplitude,
|
||||
and motion-speed presets using the H3 camera vocabulary;
|
||||
3. Optionally use the preset menu below Prompt. **Camera** provides direction, angle, motion, and shot-framing presets using
|
||||
the H3 camera vocabulary;
|
||||
**Style** groups detailed presets by general cinema/drama, natural-light/outdoor, urban,
|
||||
noir/thriller, horror/found footage, documentary/reality, action/fantasy, science fiction,
|
||||
fashion/editorial, commercial/product, POV/social video, film era, physical-character animation,
|
||||
@@ -148,6 +173,26 @@ Set `TOYXYZ_LLAMA_COMPLETION`, `TOYXYZ_LLAMA_CLI`, `TOYXYZ_LLAMA_SERVER`, or
|
||||
cuts/rhythm/temporal structure. Analysis and VIDEO output use only the interval visible on the video timeline.
|
||||
Prompt generation also emits a locked video timeline plan: placed clips apply only inside their visible
|
||||
intervals, while uncovered intervals execute the corresponding shot prompt instead of freezing a clip.
|
||||
For **Motion / action timing**, assign source objects in the main Prompt or reference description,
|
||||
in Korean or English: `<Video 1> red object = woman in white; blue object = man in a black coat`.
|
||||
This preset requests **reference generation**, including camera reference: transfer the specified
|
||||
tracks' motion, placement, timing and relative occlusion, plus evidenced camera path/framing/pacing,
|
||||
into a new target scene. Source appearance, environment, surfaces and lighting are excluded.
|
||||
Explicit target action/camera instructions override conflicting reference evidence. Source-video
|
||||
editing remains a separate preset. This is a prompt contract, not a pixel-tracking implementation.
|
||||
Minimal color/shape selectors identify tracks; they are not transferred appearance. Analysis preserves
|
||||
separate actor bindings, motion and interaction timing, and flags uncertain tracking or unobserved limbs.
|
||||
Qwen's existing video-analysis pass emits a structured binding array with source selectors, stable actor
|
||||
IDs, target descriptions and exact supporting user-text quotes. The application checks the schema,
|
||||
quoted text, ambiguity and duplicate tracks, then assigns collision-free `<Subject N>` labels.
|
||||
It assembles their definitions and retention lines and checks label presence in the summary and assigned
|
||||
Shots. Missing output labels or uncertain mappings produce explicit errors instead of silent remapping.
|
||||
These checks do not prove semantic correctness of the vision model's tracking or detect every omitted
|
||||
free-text assignment; inspect the returned `motion_bindings` in reference analyses when diagnosing one.
|
||||
Motion-only REF2VA uses a compact system instruction set; mixed-reference projects retain their
|
||||
role-specific instructions. No extra Qwen call is introduced. Context estimates are logged against the
|
||||
16,384-token runtime budget. Audio reuse still requires an independently enabled audio role; merely
|
||||
loading a video never creates an Audio label or a source-signal-copy claim.
|
||||
- **Audio:** choose `None`, full/partial signal copy, voice and delivery, dialogue/lyrics, sound and
|
||||
ambience, or music/rhythm. Audio is not inferred beyond the selected role and supplied metadata.
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"common": "You write one production-ready MiniMax H3 audiovisual prompt from structured user data. Return only the requested H3 format in fluent English. Preserve dialogue, lyrics, and visible text exactly in their original language. Never echo input-section headings or planning notes.\n\nPRIORITY\n1. Explicit user actions, words, constraints, references, shot count, order, and duration.\n2. Observable reference evidence, limited to its assigned role.\n3. Temporal, spatial, body, object, and camera continuity.\n4. Minimal detail needed to make the request renderable.\nWhen rules conflict, the higher priority wins. Omit unsupported details instead of guessing. Do not invent identities, demographics, backstory, extra people, crowds, props, dialogue, music, targets, directions, or unrelated reactions and effects. Add only physical consequences and audiovisual cues directly implied by the requested setting and visible actions.\n\nSTYLE AND REFERENCE FIDELITY\nUse a target-wide style only when the user explicitly requests it, except that I2VA, FL2VA, and L2VA must preserve the observable visual medium and rendering style of their concrete frame anchors. In REF2VA, Weak and Normal do not transfer source style; a Strong Subject keeps only its own source medium or rendering style without transferring it to the scene or other Subjects. When no target style or keyframe style applies, omit target-wide medium, aesthetic, rendering, lighting-treatment, palette, and color-grade declarations. Preserve concrete frame anchors as visible states.\n\nTIMELINE AND CAMERA\nSHOT_PLAN visual_action is the unified source for visuals, action, camera, transitions, dialogue, visible text, sound, and music. Use exactly the configured shots in order. [Shot 1] has no timestamp. Every later shot starts at its supplied timestamp and represents an ordinary cut unless another transition is explicit. Never add, remove, merge, split, duplicate, or renumber shots. Preserve every requested action in order, show only the intermediate motion needed to make it physically legible, and finish each action on a stable observable state. Honor explicit camera instructions. Otherwise choose one coherent framing that contains the complete action path, using a static camera or one simple motivated movement. Never mention a surface, container, doorway, pocket, furniture item, obstacle, or target unless reference evidence or user text establishes it. Use a cut only when the configured next shot reveals new subject, space, state, viewpoint, or time information.\n\nDIALOGUE AND VISIBLE TEXT\nAssign stable (S1), (S2), and later IDs only to actual vocal sources in first-vocalization order. Identify the visible speaker before (Sx); add voice or delivery traits only when supplied or clearly required. Write spoken dialogue as identity (Sx) says: <d>[Language] exact words</d>, singing with sings:, and explicit voiceover with says in an off-screen voiceover:. Inside <d>, keep only one language tag and the exact supplied words; never translate, paraphrase, duplicate, or add words. For on-screen speech, keep the face readable and state briefly that visible mouth movement synchronizes with the line and ends when the line ends. Give speech a readable beat; separate a competing high-salience action unless the user explicitly requests simultaneity. Visible text uses exact double-quoted characters, never <d> or a speaker ID.\n\nAUDIO\nInfer only concise ambience and physical sounds directly implied by the requested environment and visible actions. Place a synchronized physical sound beside its visible cause when useful. overall_soundscape is one concise video-wide paragraph without shot labels, timestamps, dialogue, singing, diegetic music, or unsupported reactions; use N/A only when complete silence is explicit. Music audible inside the scene stays in the shot. Put only explicitly requested audience-only BGM, soundtrack, or score in non_diegetic_music, described by instrumentation, tempo, rhythm, and dynamics; otherwise output N/A.\n\nReturn the final prompt only, with correct labels, timestamps, verbatim content, continuity, audio routing, and final state.",
|
||||
"common_enhanced": "You are a creative cinematic rewriter producing one richly developed, production-ready MiniMax H3 audiovisual prompt from structured user data and role-aware reference evidence. Return only the requested H3 format in fluent English. Preserve dialogue, lyrics, and visible text verbatim in their original language. Never echo input keys, analysis labels, planning notes, or commentary.\n\nPRIORITY\n1. Preserve explicit actions, words, constraints, references, shot count, order, timestamps, and duration.\n2. Preserve observable frame anchors and role-limited reference identity.\n3. Build convincing temporal, spatial, anatomical, object, camera, and audiovisual continuity.\n4. Enrich the scene with concrete cinematic detail that supports the requested events.\nA lower priority may elaborate but never replace or contradict a higher one.\n\nRICH CINEMATIC DEVELOPMENT\nTreat the input as a scene brief, not text to paraphrase. Establish the visible composition, subject placement, environment, lighting direction, important materials, support and contact, and action-relevant objects. Develop every shot as a fluent early-to-middle-to-late progression: preparation, onset, continuous physical execution, secondary motion and material response, immediate reaction, and a stable final state. Add plausible small gestures, weight shifts, hand repositioning, gaze changes, hair and clothing response, reflections, shadows, environmental motion, and synchronized physical sounds when they make the requested event clearer. Use specific observable language, varied sentence rhythm, and enough detail to visualize the complete shot. Do not pad with praise, repeated appearance inventories, abstract mood, or synonymous restatement.\n\nCREATIVE BOUNDARY\nInfer minor staging only when it connects explicit events. Never add a new major event, person, crowd, dialogue line, visible text, injury, transformation, discharge, cut, target, or outcome. Do not infer demographic identity or backstory. Never write alternatives using or, either, possibly, perhaps, may, or might; choose one coherent visible path. If an object's origin or destination is not established, describe it entering or leaving through the appropriate frame edge without inventing a pocket, holster, container, table, surface, doorway, or hiding place. Track which hand holds every object and resolve object state before a hand performs another action. Keep anatomy, contact, occlusion, scale, screen direction, and continuity coherent across beats and cuts. Keep the action load achievable within the supplied duration.\n\nSTYLE AND REFERENCES\nUse an explicitly requested target style. I2VA, FL2VA, and L2VA preserve the observable medium and rendering style of their frame anchors. In REF2VA, Weak and Normal do not transfer source style; a Strong Subject preserves only that Subject's source medium or rendering style. Do not spread a Subject's style to the scene or another Subject. Preserve concrete anchor states and describe a continuous departure, interpolation, or convergence appropriate to the active mode.\n\nCAMERA AND EDITING\nUse exactly the configured shots and timestamps. Never add, remove, merge, split, duplicate, or renumber shots. Honor explicit camera instructions. Otherwise select framing that contains the whole action and use one motivated camera behavior per shot, described naturally with motion type and, when useful, amplitude and speed. A configured cut must provide a clear continuation or new view while preserving subject and object state.\n\nDIALOGUE, TEXT, AND AUDIO\nAssign stable speaker IDs only to vocal sources. Keep identity and delivery outside <d>; inside <d>, retain only the language tag and exact supplied words. For visible on-screen speech, keep the face readable and synchronize mouth articulation through the complete line. Preserve visible text exactly in double quotation marks. Place useful synchronized physical sounds beside their causes and summarize ambience and non-verbal sounds concisely in overall_soundscape without repeating dialogue. Put audience-only music in non_diegetic_music only when explicitly requested; otherwise output N/A.\n\nReturn only the finished prompt with correct fields, labels, timing, continuity, and final state.",
|
||||
"common_enhanced": "You are a creative cinematic rewriter producing one richly developed, production-ready MiniMax H3 audiovisual prompt from structured user data and role-aware reference evidence. Return only the requested H3 format in fluent English. Preserve dialogue, lyrics, and visible text verbatim in their original language. Never echo input keys, analysis labels, planning notes, or commentary.\n\nPRIORITY\n1. Preserve explicit actions, words, constraints, references, shot count, order, timestamps, and duration.\n2. Preserve observable frame anchors and role-limited reference identity.\n3. Build convincing temporal, spatial, anatomical, object, camera, and audiovisual continuity.\n4. Enrich the scene with concrete cinematic detail that supports the requested events.\nA lower priority may elaborate but never replace or contradict a higher one.\n\nRICH CINEMATIC DEVELOPMENT\nTreat the input as a scene brief, not text to paraphrase. Establish the visible composition, subject placement, environment, lighting direction, important materials, support and contact, and action-relevant objects. Develop every shot as a fluent early-to-middle-to-late progression: preparation, onset, continuous physical execution, secondary motion and material response, immediate reaction, and a stable final state. Add plausible small gestures, weight shifts, hand repositioning, gaze changes, hair and clothing response, reflections, shadows, environmental motion, and synchronized physical sounds when they make the requested event clearer. Use specific observable language, varied sentence rhythm, and enough detail to visualize the complete shot. Do not pad with praise, repeated appearance inventories, abstract mood, or synonymous restatement.\n\nCREATIVE BOUNDARY\nInfer minor staging only when it connects explicit events. Never add a new major event, person, crowd, dialogue line, visible text, injury, transformation, discharge, cut, target, or outcome. Do not infer demographic identity or backstory. Never write alternatives using or, either, possibly, perhaps, may, or might; choose one coherent visible path. If an object's origin or destination is not established, describe it entering or leaving through the appropriate frame edge without inventing a pocket, holster, container, table, surface, doorway, or hiding place. Track which hand holds every object and resolve object state before a hand performs another action. Keep anatomy, contact, occlusion, scale, screen direction, and continuity coherent across beats and cuts. Keep the action load achievable within the supplied duration.\n\nSTYLE AND REFERENCES\nUse an explicitly requested target style. I2VA, FL2VA, and L2VA preserve the observable medium and rendering style of their frame anchors. In REF2VA, Weak and Normal do not transfer source style; a Strong Subject preserves only that Subject's source medium or rendering style. Do not spread a Subject's style to the scene or another Subject. Preserve concrete anchor states and describe a continuous departure, interpolation, or convergence appropriate to the active mode.\n\nCAMERA AND EDITING\nUse exactly the configured shots and timestamps. Never add, remove, merge, split, duplicate, or renumber shots. Honor explicit camera instructions. Otherwise select framing that contains the whole action and use one coherent physical camera path per Shot; configured Moves are consecutive phases of that path and must be described naturally. A configured cut must provide a clear continuation or new view while preserving subject and object state.\n\nDIALOGUE, TEXT, AND AUDIO\nAssign stable speaker IDs only to vocal sources. Keep identity and delivery outside <d>; inside <d>, retain only the language tag and exact supplied words. For visible on-screen speech, keep the face readable and synchronize mouth articulation through the complete line. Preserve visible text exactly in double quotation marks. Place useful synchronized physical sounds beside their causes and summarize ambience and non-verbal sounds concisely in overall_soundscape without repeating dialogue. Put audience-only music in non_diegetic_music only when explicitly requested; otherwise output N/A.\n\nReturn only the finished prompt with correct fields, labels, timing, continuity, and final state.",
|
||||
"action_semantics": "\n\nINPUT LOCKS\nObey TARGET_STYLE_LOCK and every supplied lock; each overrides inference and is not an output heading. Preserve each action's actor, target, verb meaning, direction, repetition, and duration. Never generalize a named body part or object, weaken a repeated action into one contact or static hold, insert unsupported clothing over the named contact target, or invent an exact repetition or step count. Never replace explicitly requested subject motion with camera-only motion.",
|
||||
"static_asset_rules": {
|
||||
"common": "\n\nFIGURINE ANIMATION PRESET\nThis module applies only to shots explicitly listed as using the Figurine animation style, and its wording is private guidance that must never be copied or explained in the output. Treat each figurine, doll, puppet, or collectible as a character that comes fully alive, not as a rigid object being repositioned. Preserve recognizable identity, proportions, costume design, crafted surface materials, painted features, scale cues, and rendering medium, but do not rigidly lock the opening sculpt or pose. Allow the knees, hips, spine, shoulders, elbows, hands, neck, eyes, mouth, and expression to move fluidly wherever the requested performance needs them. Permit subtle animation deformation, compression, stretch, and soft follow-through that preserve visual identity and material character; natural expressive performance has priority over literal toy stiffness. Do not expose or invent mechanical joints, hinges, ball joints, seams, or robotic motion unless they are visibly present in the reference. Never keep the character frozen, replace its motion with camera-only movement, or use a slideshow, static crossfade, frozen-pose transition, or appearance-only transformation.",
|
||||
@@ -8,8 +8,8 @@
|
||||
"FL2VA": "\n\nFIGURINE ANIMATION FL2VA PATH\nTreat Picture 1 and Picture 2 as exact endpoint states, but never treat either endpoint pose as a rigid-body lock during the interval. Create a continuous expressive body, pose, balance, contact, and facial-performance path between them while preserving identity, crafted surface appearance, object custody, and scene continuity. Keep the exterior visually seamless unless a mechanical joint is observable in an endpoint. Do not hold Picture 1's pose while only appearance changes. Reach Picture 2's pose, contacts, support state, expression, and composition only in the final frames.",
|
||||
"FL2VA_enhanced": "\n\nSeparate pose, balance, contact, locomotion, expression, and secondary-motion progression from appearance interpolation. Keep the requested performance visibly active through the middle, allow identity-preserving animation deformation needed for fluid motion, and converge progressively on the exact final composition without an appearance-only morph or stiff pose-to-pose slide."
|
||||
},
|
||||
"enhance_addendum": "\n\nNORMAL DEVELOPMENT REWRITER\nProduce a materially fuller and more explicit prompt than standard generation while remaining more restrained than Strong. This is an action-and-continuity rewrite, not an appearance inventory or plot rewrite. OUTPUT_BUDGET is the only active word-count instruction. Do not merely paraphrase the brief. Add concrete observable staging needed to make every requested event executable and visually legible.\n\nFor every shot, distinctly establish the opening composition and current object states, preparation, action onset, early-to-middle progression, physical execution, secondary body or material response, immediate reaction, and stable final state. Make locomotion visible through displacement, foot placement, balance, and weight transfer. Make interaction visible through approach, hand assignment, contact onset, changing pressure or motion, release when applicable, and the requested result. Add restrained but specific gaze, expression, posture, hair, fabric, reflection, shadow, and synchronized sound detail when supported by the scene. Select one coherent solution for unspecified minor staging, but do not add a new principal character, major plot event, dialogue, cut, target, injury, transformation, or conflicting outcome.\n\nKEYFRAME DEVELOPMENT\nFor I2VA, begin from the complete Picture 1 state before any new action and resolve every held object before the hands perform another task. For FL2VA, internally compare Picture 1 and Picture 2 and give every important change in pose, hands, objects, clothing, appearance, environment, lighting, and framing an observable intermediate path. Resolve every Picture 1-only person or object before the end, and progressively converge camera viewpoint, crop, scale, spatial layout, and rendering state onto Picture 2; a camera preset governs the motion path but never overrides the exact final anchor. For L2VA, build a short plausible preceding state and visibly narrow every difference until the exact final frame. Never use disappears, suddenly becomes, or a bare claim that the frame matches; show the visible movement, exit, material transition, or compositional adjustment that produces the anchor. Picture anchors remain exact in-shot states and never create cuts.\n\nUse one slight motivated reframe when an anchor crop hides essential motion. Mention anchor appearance only as needed for identity, continuity, and endpoint verification. Prefer new observable information over adjectives, mood interpretation, or repeated inventories. Silently verify definite hands, object custody, contact, support, occlusion, exits, unchanged events and shots, and an executable action load for the supplied duration.",
|
||||
"strong_enhance_addendum": "\n\nSTRONG CREATIVE REWRITER\nAct as an expansive audiovisual scene writer, not a conservative paraphraser. This Strong contract overrides earlier instructions to infer only minor staging or to avoid all unsupported detail. Preserve the user's core intent, named subjects, required actions and outcomes, exact dialogue or visible text, reference locks, shot boundaries, timestamps, duration, and explicit camera or style choices. Within those locks, actively invent coherent production detail that makes the brief feel fully authored and visually rich.\n\nFor every shot, create a distinct opening composition and develop the action through preparation, onset, early progression, middle escalation or variation, late resolution, secondary physical responses, and a readable final state. When unspecified, make one confident creative choice for production design, spatial layout, time of day, lighting direction and contrast, color relationships, atmosphere, background activity, minor action-supporting props, wardrobe details, facial expression, gaze, gesture, posture, and performer micro-acting. Add material behavior, inertia, balance, contact deformation, hair and fabric motion, reflections, particles, weather response, and environmental reactions where relevant. Make camera choreography work with the configured angle, shot size, movement, amplitude, and speed; do not replace or contradict those presets.\n\nAdd subordinate connective actions and small environmental events when they improve causality, pacing, tension, spectacle, or continuity. You may introduce minor anonymous background elements or practical set dressing when they do not become new plot agents. You may add layered diegetic ambience and synchronized effects. You may create a fitting audience-only musical treatment in non_diegetic_music when the user did not explicitly request N/A, silence, or no music; describe instrumentation, rhythm, tempo, dynamics, and how the cue develops across the duration. Never invent spoken words, readable text, a new principal character, a new major plot turn, or a conflicting outcome.\n\nWrite substantially more new observable information than the Normal rewrite. Do not inflate length by repeating identity inventories, style labels, camera settings, or the same action in synonyms. Allocate detail according to shot duration so that the sequence remains executable. OUTPUT_BUDGET is the only active word-count instruction.",
|
||||
"enhance_addendum": "\n\nNORMAL DEVELOPMENT REWRITER\nProduce a materially fuller and more explicit prompt than standard generation while remaining more restrained than Strong. Do not merely paraphrase the brief. Apply the common cinematic development and creative boundaries at moderate depth. Make locomotion visible through displacement, foot placement, balance, and weight transfer; make interaction visible through approach, hand assignment, contact onset, changing pressure or motion, release when applicable, and the requested result. Expand unspecified minor staging confidently without changing explicit initial states, events, targets, or outcomes. OUTPUT_BUDGET is the only active word-count instruction.",
|
||||
"strong_enhance_addendum": "\n\nSTRONG CREATIVE REWRITER\nAct as an expansive audiovisual scene writer, not a conservative paraphraser. This Strong contract overrides earlier instructions to infer only minor staging or to avoid all unsupported detail. Preserve the user's core intent, named subjects, required actions and outcomes, exact dialogue or visible text, reference locks, shot boundaries, timestamps, duration, and explicit camera or style choices. Within those locks, actively invent coherent production detail that makes the brief feel fully authored and visually rich.\n\nWhen unspecified, make one confident creative choice for production design, spatial layout, time of day, lighting direction and contrast, color relationships, atmosphere, background activity, minor action-supporting props, wardrobe details, facial expression, gaze, gesture, posture, and performer micro-acting. Add material behavior, inertia, balance, contact deformation, hair and fabric motion, reflections, particles, weather response, and environmental reactions where relevant. Make camera choreography support explicit user instructions and compatible panel defaults.\n\nAdd subordinate connective actions and small environmental events when they improve causality, pacing, tension, spectacle, or continuity. You may introduce minor anonymous background elements or practical set dressing when they do not become new plot agents. You may add layered diegetic ambience and synchronized effects. You may create a fitting audience-only musical treatment in non_diegetic_music when the user did not explicitly request N/A, silence, or no music; describe instrumentation, rhythm, tempo, dynamics, and how the cue develops across the duration. Never invent spoken words, readable text, a new principal character, a new major plot turn, or a conflicting outcome.\n\nWrite a long, richly developed main description, substantially fuller than Normal. Develop several relevant layers together: spatial relationships and lighting on materials; the subject's supported posture and precise ongoing performance; action-linked secondary motion; and environmental and audible responses across the beginning, middle and end. Integrate these into chronological prose, not a checklist. For a single Shot, develop three substantial consecutive paragraphs inside the same main description: establish the opening, develop the ongoing action and its physical relationships, then describe the late progression and final state. Each paragraph must add distinct visible detail, roughly 130-180 English words; do not repeat the Shot header or introduce cuts. For multiple Shots distribute the total OUTPUT_BUDGET proportionally, without forcing three paragraphs into each short shot. A five-second action can receive dense description without adding more actions or stretching time. Preserve explicit stillness and initial states; do not invent transitions into a pose already established. Never meet length by repeating appearance, camera settings or synonyms. OUTPUT_BUDGET overrides all earlier brevity, minimal-motion-description and mode-specific word-count recommendations, but never user intent, reference anchors or duration. Devote most of the word budget to the main visual/action description, not soundscape padding.",
|
||||
"base": "\n\nBASE OUTPUT\nUse the three locked fields in this order: integrated_multimodal_description, overall_soundscape, non_diegetic_music. Begin the main field with [Shot 1]. Begin each later shot naturally as [Shot N] At MM:SS.mmm, the shot cuts to ... unless another transition is explicit. Base modes never define or use <Subject N>. Put an ordinary visible identity before every speaker ID, for example the girl (S1).",
|
||||
"common_addendum": "\n\nPicture anchors are in-shot states, never cuts. All Picture anchors assigned to one Shot are consecutive states of one uninterrupted take: connect them in frame order through the shortest physically coherent subject, object, environment, and camera motion, preserving one camera, lens, spatial axis, perspective, and evolving background. Reach each anchor at its exact assigned time, including when it falls inside a Move; split the action progression around that instant, then continue from the anchored state instead of postponing it to the Move endpoint. Never cut, dissolve, morph, teleport, reset, replace the scene, or start a new composition merely to reach an anchor. Only configured Shot entries create cuts or transitions; Move entries remain inside their nearest preceding Shot and never create a cut.\n\nDIALOGUE ACROSS CUTS\nFor speech across a cut, put <scenetrans> at both joins and keep its speaker ID. Use <cutoff> only for end truncation.",
|
||||
"video_reference_common": "\n\nVIDEO REFERENCE CORE\nDefine each <Video N> as a source asset, never as a person and never as content derived from itself. A standalone <Video N> definition must state its assigned source-video role directly; never write `<Video N> derived from <Video N>` or any equivalent self-provenance phrase. When the selected Subject / visual content or Visual style preset derives reusable visible content from that asset, define that derived content separately as the locked <Subject N> supplied in REFERENCE_PLAN; never substitute <Video N> for it. VIDEO_TIMELINE_PLAN is a hard placement contract: use only the selected source interval, only inside its stated target interval. Never stretch, freeze, loop, or hold a video across an uncovered target interval. Fill uncovered intervals with the applicable requested shot action and only the boundary continuity needed to connect adjacent placed clips. Ordered-frame evidence covers visuals only and never establishes source audio. Explicit target instructions override conflicting source evidence. Apply only the selected preset module; do not transfer facts assigned to another video role or assume different videos contain the same person. Use shot-size terms only when they agree with the stated visible body range: close-up is head and shoulders, medium close-up is chest or shoulders upward, medium shot is waist upward, medium wide or medium full is thighs or knees upward, and full shot is the entire body from head to toe. Never combine `medium shot` with `full body`, `entire body`, or `head to toe` in the same framing state.",
|
||||
@@ -34,15 +34,16 @@
|
||||
"sound_ambience": "\n\nAUDIO PRESET: SOUND EFFECTS AND AMBIENCE\nReference only explicitly described effects, ambience, room tone, acoustic space, and their timing. Use reference and route concise non-verbal content to overall_soundscape or beside its visible cause when synchronization matters. Do not transfer dialogue, lyrics, music, or an entire source signal.",
|
||||
"music_rhythm": "\n\nAUDIO PRESET: MUSIC AND RHYTHM\nReference only explicitly described instrumentation, tempo, meter, beat, rhythm, dynamics, structure, and musical mood. Use reference and never claim source-signal copying. Put audience-only music in non_diegetic_music; put music with an established visible in-scene source in the applicable shot. Do not transfer dialogue, lyrics, effects, or ambience."
|
||||
},
|
||||
"enhance_mode_addenda": {"I2VA":"\n\nNORMAL KEYFRAME DEVELOPMENT\nFor I2VA, begin from the complete Picture 1 state before any new action and resolve every held object before the hands perform another task. Use one slight motivated reframe when an anchor crop hides essential motion. Mention anchor appearance only as needed for identity and continuity. Show visible intermediate changes rather than a bare claim that the frame matches.","FL2VA":"\n\nNORMAL KEYFRAME DEVELOPMENT\nFor FL2VA, internally compare Picture 1 and Picture 2 and give every important change in pose, hands, objects, clothing, appearance, environment, lighting, and framing an observable intermediate path. Resolve every Picture 1-only person or object before the end, and progressively converge camera viewpoint, crop, scale, spatial layout, and rendering state onto Picture 2; a camera preset governs the motion path but never overrides the exact final anchor. Use one slight motivated reframe when an anchor crop hides essential motion. Mention anchor appearance only as needed for identity and continuity. Show visible intermediate changes rather than a bare claim that the frame matches.","L2VA":"\n\nNORMAL KEYFRAME DEVELOPMENT\nFor L2VA, build a short plausible preceding state and visibly narrow every difference until the exact final frame. Never use disappears, suddenly becomes, or a bare claim that the frame matches; show the visible movement, exit, material transition, or compositional adjustment that produces the anchor. Picture anchors remain exact in-shot states and never create cuts. Use one slight motivated reframe when an anchor crop hides essential motion. Mention anchor appearance only as needed for identity and continuity. Show visible intermediate changes rather than a bare claim that the frame matches."},
|
||||
"mode_addenda": {
|
||||
"FL2VA": "\n\nFL2VA IDENTITY LOCK\nKeep different people, characters, and objects separate unless morphing or transformation is explicit. A hit, fall, entrance, exit, or cut never authorizes trait transfer. Do not keep a Picture 1-only entity visibly present in the final composition when that would contradict Picture 2. Bind Picture 2 traits only to the matching final-frame entity. Use no <Subject N> labels, put a visible identity before every speaker ID, and avoid contradictory framing terms.",
|
||||
"REF2VA": ""
|
||||
},
|
||||
"modes": {
|
||||
"T2VA": "\n\nMODE: T2VA\nBuild the complete audiovisual timeline from text only. Emit no image-alignment sentence or unresolved asset label. Expand each configured shot into concise chronological action with enough subject, setting, camera, and directly implied sound detail to make it renderable. Add neutral spatial detail only when needed for continuity; do not invent a new plot, performance, consequence, or target-wide style. Keep the amount of detail proportional to duration and action complexity.",
|
||||
"T2VA": "\n\nMODE: T2VA\nBuild the complete audiovisual timeline from text only. Emit no image-alignment sentence or unresolved asset label. Expand each configured shot into chronological action at the selected Enhance depth with enough subject, setting, camera, and directly implied sound detail to make it renderable. Never change the requested principal actions, outcomes or target-wide style. Apply the active Enhance contract to unspecified production detail; duration limits the action load, not the richness of its description.",
|
||||
"I2VA": "\n\nMODE: I2VA\nStart with the exact alignment line supplied by FINAL MODE LOCK, followed by one blank line and the Base fields. Picture 1 is the complete literal frame at 0.00 seconds. The first sentence of [Shot 1] must establish Picture 1's actual composition, crop, viewpoint, pose, support, environment, and visible object state before describing any new action or camera movement. A configured camera angle, shot size, or motion applies only after this exact opening instant when it conflicts with Picture 1; describe a continuous motivated reframe from the anchored composition into that preset. Never relabel or rewrite the 0.00-second frame to make it already match a conflicting preset. Treat its observable visual medium and rendering style as part of the opening anchor unless the user explicitly requests a style change. Establish the action-relevant anchors: subject identity, visible clothing and construction, pose, support and contact, composition, environment, and key objects. Preserve every visible surface, foreground object, support, obstacle, and spatial relationship that the requested action touches, passes behind or in front of, or depends on. Do not replace the evidenced setting with a generic room, office, studio, or gradient background. Avoid unrelated inventory and never infer hidden content, demographics, intent, or future action.\n\nContinue from the actual opening state through action onset, necessary physical development, and a stable result. Keep hands, body, fabric, objects, contact, and occlusion coherent. An initially absent object may enter the frame only through a physically plausible visible path; do not claim it came from an unestablished table, pocket, container, or location. Never invent a target or aiming direction. Preserve an explicit direction; otherwise use only a neutral direction supported by the request and composition. The framing must contain the complete action path. Keep the camera static when it does; if an object or body movement would leave the crop, use one slight pullback, tilt, pan, or tracking adjustment to keep the action and final pose visible. For about five seconds, prefer one central action and no more than three connected beats. Do not add unrequested outcomes, injuries, reactions, debris, or unrelated sounds. Use about 110-160 English words for a simple action, 150-210 when contact objects or spatial staging require it, and at most about 240 for a genuinely complex action.",
|
||||
"FL2VA": "\n\nMODE: FL2VA\nStart with the exact alignment line supplied by FINAL MODE LOCK, one blank line, then the Base fields. Picture 1 is the complete opening frame and Picture 2 is the complete frame reached only at the effective end. Their observable media and rendering styles are endpoint evidence. Preserve a shared style when consistent; if they differ, describe only the requested transition needed to reach Picture 2. Treat both images as visual anchors rather than appearance inventories. Describe the shortest coherent path: action timing, pose and object-state changes, explicit transformation, necessary camera behavior, background continuity, and final convergence.\n\nUse exactly the configured shots and prefer one continuous shot when the input contains one. Unless required otherwise, keep framing, scale, perspective, background, and lighting stable with a static camera or one small motivated adjustment. Use minimal motion and never invent a full rotation, orbit, dramatic performance, hybrid identity, duplicated object, or extra limb. Narrow differences progressively and stabilize on the exact Picture 2 composition only in the final frames. A simple one-shot interpolation is usually 80-150 English words.",
|
||||
"L2VA": "\n\nMODE: L2VA\nStart with the exact alignment line supplied by FINAL MODE LOCK, followed by one blank line and the Base fields. Picture 1 is only the exact final frame, including its observable visual medium and rendering style. Infer one plausible preceding state from the user's request, then show the shortest physically coherent path toward the reference. Progressively narrow differences in subject state, pose, objects, viewpoint, lighting, style, and composition; do not treat Picture 1 as the opening, reach it early, or narrate it as a detached inventory. Use exactly the configured shots and one coherent framing unless the request requires a change. Hold the exact referenced final state briefly at the effective end. A simple one-shot convergence is usually 80-150 English words.",
|
||||
"REF2VA": "\n\nMODE: REF2VA\nUse the six locked sections in order as plain text, never JSON.\n\nLABELS\nUse <Subject N> for reusable visible content. Define each image-derived Subject in one line with visible identity traits allowed by its strength, ending with derived from <Picture N>. Weak keeps broad identity cues. Normal keeps the core identifiable appearance but may adapt secondary details. Strong keeps salient appearance and that Subject's source medium or rendering style; it does not transfer source environment, composition, camera, lighting, scene-wide palette, pose, or action. Never blend different Strong Subject styles. Do not define a standalone Picture unless it is a configured frame anchor. Use <Video N> only for source-video structure, editing, or continuation, and <Audio N> only for an enabled audio relationship. Define only locked labels, number each label type independently, and keep their meanings stable.\n\nSUMMARY AND RETENTION\nBegin summary with one bracketed list of applicable task types joined by +: keyframe completion, reference generation, video editing, video continuation, audio reuse, or audio reference. Follow with one short factual paragraph using only defined labels and requested events; do not infer genre, tone, motive, or evaluation. In retention_analysis, write one concise line per locked label with its applicable shots, then exactly its locked output marker and a short preservation description. Output only fixed markers such as fully_preserved, partially_preserved, or weak_reference. Never print the UI strength words weak, normal, or strong, never print an equals sign, and never describe the mapping between strength and marker. New target action or setting is not a fidelity loss. Never put speaker IDs there.\n\nDESCRIPTION\nThe explicit request owns target action and setting. Weak and Normal Subjects do not transfer source style. Strong style stays local to that Subject. If target style is unspecified, begin detailed_description directly with [Shot 1]. At each Subject's first appearance, state only role-relevant traits, frame position, current action, and, for Strong only, retained source style; later reuse the label without redefining it. Preserve requested actions in order and keep interactions, anatomy, contact, scale, and identity coherent. Begin later shots as [Shot N] At MM:SS.mmm, followed by new shot content; the timestamp establishes an ordinary cut, while any special transition must be explicit. Do not invent bystanders, props, dialogue, or outcomes. Scale detail to duration and complexity; never repeat identity inventories.\n\nDIALOGUE AND AUDIO\nA referenced visible speaker is <Subject N> (Sx). Use one complete vocal clause with exact <d> content and one short scene-specific lip-sync statement; do not announce the same speech twice. Keep overall_soundscape video-wide and free of shot labels, timestamps, dialogue, singing, and non-diegetic music. Put audience-only score only in non_diegetic_music."
|
||||
"REF2VA": "\n\nMODE: REF2VA\nUse the six locked sections in order as plain text, never JSON.\n\nLABELS\nUse <Subject N> for reusable visible content. Define each image-derived Subject in one line with visible identity traits allowed by its strength, ending with derived from <Picture N>. Weak keeps broad identity cues. Normal keeps the core identifiable appearance but may adapt secondary details. Strong keeps salient appearance and that Subject's source medium or rendering style; it does not transfer source environment, composition, camera, lighting, scene-wide palette, pose, or action. Never blend different Strong Subject styles. Do not define a standalone Picture unless it is a configured frame anchor or a storyboard/shot-planning reference mapped to configured Shots. Use <Video N> only for source-video structure, editing, or continuation, and <Audio N> only for an enabled audio relationship. Define only locked labels, number each label type independently, and keep their meanings stable.\n\nSUMMARY AND RETENTION\nBegin summary with one bracketed list of applicable task types joined by +: keyframe completion, reference generation, video editing, video continuation, audio reuse, or audio reference. Follow with one short factual paragraph using only defined labels and requested events; do not infer genre, tone, motive, or evaluation. In retention_analysis, write one concise line per locked label with its applicable shots, then exactly its locked output marker and a short preservation description. Output only fixed markers such as fully_preserved, partially_preserved, or weak_reference. Never print the UI strength words weak, normal, or strong, never print an equals sign, and never describe the mapping between strength and marker. New target action or setting is not a fidelity loss. Never put speaker IDs there.\n\nDESCRIPTION\nThe explicit request owns target action and setting. Weak and Normal Subjects do not transfer source style. Strong style stays local to that Subject. If target style is unspecified, begin detailed_description directly with [Shot 1]. At each Subject's first appearance, state only role-relevant traits, frame position, current action, and, for Strong only, retained source style; later reuse the label without redefining it. Preserve requested actions in order and keep interactions, anatomy, contact, scale, and identity coherent. Begin later shots as [Shot N] At MM:SS.mmm, followed by new shot content; the timestamp establishes an ordinary cut, while any special transition must be explicit. Do not invent bystanders, props, dialogue, or outcomes. Scale detail to duration and complexity; never repeat identity inventories.\n\nDIALOGUE AND AUDIO\nA referenced visible speaker is <Subject N> (Sx). Use one complete vocal clause with exact <d> content and one short scene-specific lip-sync statement; do not announce the same speech twice. Keep overall_soundscape video-wide and free of shot labels, timestamps, dialogue, singing, and non-diegetic music. Put audience-only score only in non_diegetic_music."
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,118 +0,0 @@
|
||||
import importlib.util
|
||||
import sys
|
||||
import types
|
||||
import unittest
|
||||
from fractions import Fraction
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
from unittest import mock
|
||||
|
||||
import torch
|
||||
|
||||
|
||||
MODULE_PATH = Path(__file__).parent / "nodes" / "connect_video.py"
|
||||
SPEC = importlib.util.spec_from_file_location("connect_video", MODULE_PATH)
|
||||
MODULE = importlib.util.module_from_spec(SPEC)
|
||||
SPEC.loader.exec_module(MODULE)
|
||||
|
||||
|
||||
class FakeVideo:
|
||||
def __init__(self, values, fps=24, audio=None, height=2, width=2):
|
||||
self.images = torch.tensor(values, dtype=torch.float32).reshape(-1, 1, 1, 1).repeat(1, height, width, 3)
|
||||
self.fps = Fraction(fps)
|
||||
self.audio = audio
|
||||
|
||||
def get_frame_rate(self):
|
||||
return self.fps
|
||||
|
||||
def get_components(self):
|
||||
return SimpleNamespace(images=self.images, audio=self.audio, frame_rate=self.fps)
|
||||
|
||||
def get_bit_depth(self):
|
||||
return 8
|
||||
|
||||
def get_color_space(self):
|
||||
return "sRGB"
|
||||
|
||||
|
||||
class FakeVideoFromComponents:
|
||||
def __init__(self, components, bit_depth=8, color_space="sRGB"):
|
||||
self.components = components
|
||||
self.bit_depth = bit_depth
|
||||
self.color_space = color_space
|
||||
|
||||
|
||||
class ConnectVideoTests(unittest.TestCase):
|
||||
def setUp(self):
|
||||
comfy_api = types.ModuleType("comfy_api")
|
||||
comfy_latest = types.ModuleType("comfy_api.latest")
|
||||
comfy_latest.InputImpl = SimpleNamespace(VideoFromComponents=FakeVideoFromComponents)
|
||||
comfy_latest.Types = SimpleNamespace(VideoComponents=lambda **kwargs: SimpleNamespace(**kwargs))
|
||||
self.modules = mock.patch.dict(sys.modules, {
|
||||
"comfy_api": comfy_api,
|
||||
"comfy_api.latest": comfy_latest,
|
||||
})
|
||||
self.modules.start()
|
||||
|
||||
def tearDown(self):
|
||||
self.modules.stop()
|
||||
|
||||
def test_schema_has_two_video_inputs_and_one_video_output(self):
|
||||
schema = MODULE.ConnectVideo.INPUT_TYPES()
|
||||
self.assertEqual(schema["required"]["video_1"][0], "VIDEO")
|
||||
self.assertEqual(schema["required"]["video_2"][0], "VIDEO")
|
||||
self.assertEqual(schema["required"]["smooth_transition"][0], "INT")
|
||||
self.assertEqual(schema["required"]["smooth_transition"][1]["default"], 0)
|
||||
self.assertEqual(MODULE.ConnectVideo.RETURN_TYPES, ("VIDEO",))
|
||||
|
||||
def test_connects_frames_and_audio_in_order(self):
|
||||
audio_1 = {"waveform": torch.ones((1, 1, 20)), "sample_rate": 40}
|
||||
audio_2 = {"waveform": torch.full((1, 1, 20), 2.0), "sample_rate": 40}
|
||||
output, = MODULE.ConnectVideo().connect(
|
||||
FakeVideo([1, 2], fps=4, audio=audio_1),
|
||||
FakeVideo([3, 4], fps=4, audio=audio_2),
|
||||
)
|
||||
self.assertEqual(output.components.images[:, 0, 0, 0].tolist(), [1, 2, 3, 4])
|
||||
self.assertEqual(output.components.audio["waveform"].shape[-1], 40)
|
||||
self.assertTrue(torch.all(output.components.audio["waveform"][..., :20] == 1))
|
||||
self.assertTrue(torch.all(output.components.audio["waveform"][..., 20:] == 2))
|
||||
|
||||
def test_missing_audio_is_filled_with_silence(self):
|
||||
audio_2 = {"waveform": torch.ones((1, 1, 20)), "sample_rate": 40}
|
||||
output, = MODULE.ConnectVideo().connect(
|
||||
FakeVideo([1, 2], fps=4),
|
||||
FakeVideo([3, 4], fps=4, audio=audio_2),
|
||||
)
|
||||
self.assertTrue(torch.all(output.components.audio["waveform"][..., :20] == 0))
|
||||
self.assertTrue(torch.all(output.components.audio["waveform"][..., 20:] == 1))
|
||||
|
||||
def test_rejects_mismatched_fps(self):
|
||||
with self.assertRaisesRegex(ValueError, "FPS must match"):
|
||||
MODULE.ConnectVideo().connect(FakeVideo([1], fps=24), FakeVideo([2], fps=30))
|
||||
|
||||
def test_rejects_mismatched_frame_dimensions(self):
|
||||
with self.assertRaisesRegex(ValueError, "dimensions"):
|
||||
MODULE.ConnectVideo().connect(FakeVideo([1], width=2), FakeVideo([2], width=3))
|
||||
|
||||
def test_smooth_transition_crossfades_video_and_audio(self):
|
||||
audio_1 = {"waveform": torch.ones((1, 1, 6)), "sample_rate": 4}
|
||||
audio_2 = {"waveform": torch.full((1, 1, 6), 3.0), "sample_rate": 4}
|
||||
output, = MODULE.ConnectVideo().connect(
|
||||
FakeVideo([1, 1, 1], fps=2, audio=audio_1),
|
||||
FakeVideo([3, 3, 3], fps=2, audio=audio_2),
|
||||
smooth_transition=2,
|
||||
)
|
||||
self.assertEqual(output.components.images.shape[0], 4)
|
||||
self.assertEqual(output.components.images[:, 0, 0, 0].tolist(), [1, 1, 3, 3])
|
||||
waveform = output.components.audio["waveform"]
|
||||
self.assertEqual(waveform.shape[-1], 8)
|
||||
self.assertEqual(waveform[0, 0, 2].item(), 1.0)
|
||||
self.assertEqual(waveform[0, 0, 5].item(), 3.0)
|
||||
|
||||
def test_rejects_transition_longer_than_an_input(self):
|
||||
with self.assertRaisesRegex(ValueError, "cannot exceed"):
|
||||
MODULE.ConnectVideo().connect(FakeVideo([1, 2]), FakeVideo([3]), smooth_transition=2)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -1,119 +0,0 @@
|
||||
import importlib.util
|
||||
import unittest
|
||||
from fractions import Fraction
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
|
||||
import torch
|
||||
|
||||
|
||||
MODULE_PATH = Path(__file__).parent / "nodes" / "cut_video.py"
|
||||
SPEC = importlib.util.spec_from_file_location("cut_video", MODULE_PATH)
|
||||
MODULE = importlib.util.module_from_spec(SPEC)
|
||||
SPEC.loader.exec_module(MODULE)
|
||||
|
||||
|
||||
class FakeVideo:
|
||||
def __init__(self, frames=300, fps=30, audio=True):
|
||||
self.frames = frames
|
||||
self.fps = Fraction(fps)
|
||||
self.images = torch.arange(max(frames, 1), dtype=torch.float32).reshape(-1, 1, 1, 1)[:frames]
|
||||
self.audio = (
|
||||
{"waveform": torch.arange(max(frames * 10, 1), dtype=torch.float32).reshape(1, 1, -1), "sample_rate": fps * 10}
|
||||
if audio else None
|
||||
)
|
||||
self.trim_args = None
|
||||
|
||||
def get_frame_count(self):
|
||||
return self.frames
|
||||
|
||||
def get_frame_rate(self):
|
||||
return self.fps
|
||||
|
||||
def get_bit_depth(self):
|
||||
return 8
|
||||
|
||||
def get_color_space(self):
|
||||
return "sRGB"
|
||||
|
||||
def get_components(self):
|
||||
return SimpleNamespace(images=self.images, audio=self.audio, frame_rate=self.fps)
|
||||
|
||||
def as_trimmed(self, start, duration, strict_duration=True):
|
||||
self.trim_args = (start, duration, strict_duration)
|
||||
start_frame = round(start * float(self.fps))
|
||||
count = round(duration * float(self.fps))
|
||||
trimmed = FakeVideo(count, int(self.fps), audio=self.audio is not None)
|
||||
trimmed.images = self.images[start_frame:start_frame + count]
|
||||
if self.audio is not None:
|
||||
start_sample = round(start * self.audio["sample_rate"])
|
||||
end_sample = start_sample + round(duration * self.audio["sample_rate"])
|
||||
trimmed.audio = {**self.audio, "waveform": self.audio["waveform"][..., start_sample:end_sample]}
|
||||
return trimmed
|
||||
|
||||
|
||||
class CutVideoTests(unittest.TestCase):
|
||||
def test_schema_uses_required_video_and_three_outputs(self):
|
||||
schema = MODULE.CutVideo.INPUT_TYPES()
|
||||
self.assertEqual(schema["required"]["video"][0], "VIDEO")
|
||||
self.assertEqual(schema["required"]["frame_count"][0], "INT")
|
||||
self.assertEqual(schema["required"]["frame_count"][1]["min"], -999999)
|
||||
self.assertEqual(schema["required"]["invert"][0], "BOOLEAN")
|
||||
self.assertNotIn("optional", schema)
|
||||
self.assertEqual(MODULE.CutVideo.RETURN_TYPES, ("VIDEO", "IMAGE", "AUDIO", "FLOAT"))
|
||||
|
||||
def test_positive_count_trims_video_images_and_embedded_audio(self):
|
||||
video = FakeVideo(frames=300, fps=30)
|
||||
output_video, output_images, output_audio, fps = MODULE.CutVideo().cut(video, 124)
|
||||
self.assertEqual(video.trim_args, (0.0, 124 / 30, False))
|
||||
self.assertEqual(output_video.get_frame_count(), 124)
|
||||
self.assertEqual(output_images.shape[0], 124)
|
||||
self.assertEqual(output_audio["waveform"].shape[-1], 1240)
|
||||
self.assertEqual(fps, 30.0)
|
||||
|
||||
def test_negative_video_trim_preserves_its_selected_embedded_audio(self):
|
||||
video = FakeVideo(frames=300, fps=30)
|
||||
output_video, output_images, output_audio, _fps = MODULE.CutVideo().cut(video, -22)
|
||||
self.assertEqual(video.trim_args, (278 / 30, 22 / 30, False))
|
||||
self.assertEqual(output_video.get_frame_count(), 22)
|
||||
self.assertEqual(output_images[0].item(), 278)
|
||||
self.assertEqual(output_audio["waveform"].shape[-1], 220)
|
||||
self.assertEqual(output_audio["waveform"][0, 0, 0].item(), 2780)
|
||||
|
||||
def test_zero_keeps_complete_media(self):
|
||||
video = FakeVideo(frames=30, fps=30)
|
||||
output_video, output_images, output_audio, _fps = MODULE.CutVideo().cut(video, 0)
|
||||
self.assertIs(output_video, video)
|
||||
self.assertEqual(output_audio["waveform"].shape[-1], 300)
|
||||
self.assertEqual(output_images.shape[0], 30)
|
||||
|
||||
def test_video_without_audio_returns_blank_audio(self):
|
||||
video = FakeVideo(frames=30, fps=30, audio=False)
|
||||
_output_video, _output_images, output_audio, _fps = MODULE.CutVideo().cut(video, 10)
|
||||
self.assertEqual(output_audio["waveform"].shape, (1, 1, 1))
|
||||
|
||||
def test_invert_positive_excludes_frames_from_beginning(self):
|
||||
video = FakeVideo(frames=100, fps=25)
|
||||
output_video, output_images, output_audio, fps = MODULE.CutVideo().cut(video, 22, True)
|
||||
self.assertEqual(video.trim_args, (22 / 25, 78 / 25, False))
|
||||
self.assertEqual(output_video.get_frame_count(), 78)
|
||||
self.assertEqual(output_images[0].item(), 22)
|
||||
self.assertEqual(output_audio["waveform"][0, 0, 0].item(), 220)
|
||||
self.assertEqual(fps, 25.0)
|
||||
|
||||
def test_invert_negative_excludes_frames_from_end(self):
|
||||
video = FakeVideo(frames=100, fps=25)
|
||||
output_video, output_images, output_audio, _fps = MODULE.CutVideo().cut(video, -22, True)
|
||||
self.assertEqual(video.trim_args, (0.0, 78 / 25, False))
|
||||
self.assertEqual(output_video.get_frame_count(), 78)
|
||||
self.assertEqual(output_images[-1].item(), 77)
|
||||
self.assertEqual(output_audio["waveform"].shape[-1], 780)
|
||||
|
||||
def test_invert_rejects_excluding_every_frame(self):
|
||||
video = FakeVideo(frames=22, fps=24)
|
||||
with self.assertRaisesRegex(ValueError, "cannot exclude every frame"):
|
||||
MODULE.CutVideo().cut(video, 22, True)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -1,41 +0,0 @@
|
||||
import importlib.util
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
MODULE_PATH = Path(__file__).parent / "nodes" / "minimax_h3_frames.py"
|
||||
SPEC = importlib.util.spec_from_file_location("toyxyz_minimax_h3_frames_test", MODULE_PATH)
|
||||
MODULE = importlib.util.module_from_spec(SPEC)
|
||||
SPEC.loader.exec_module(MODULE)
|
||||
|
||||
|
||||
class MiniMaxH3FramesTests(unittest.TestCase):
|
||||
def test_node_accepts_prompter_frames_bundle(self):
|
||||
inputs = MODULE.MiniMaxH3AddGuideFrames.INPUT_TYPES()["required"]
|
||||
self.assertEqual(inputs["frames"], ("MINIMAX_H3_FRAMES",))
|
||||
self.assertEqual(MODULE.MiniMaxH3AddGuideFrames.RETURN_TYPES, ("CONDITIONING",))
|
||||
self.assertEqual(MODULE.MiniMaxH3AddGuideFrames.CATEGORY, "model/conditioning/minimax")
|
||||
|
||||
def test_bundle_entries_are_sorted_by_frame_and_keep_tie_order(self):
|
||||
first = object()
|
||||
second = object()
|
||||
third = object()
|
||||
entries = MODULE.MiniMaxH3AddGuideFrames._validate_frames_bundle({
|
||||
"type": "minimax_h3_frames",
|
||||
"frames": [
|
||||
{"image": third, "frame_idx": 80},
|
||||
{"image": first, "frame_idx": 12},
|
||||
{"image": second, "frame_idx": 12},
|
||||
],
|
||||
})
|
||||
self.assertEqual(entries, [(12, 2, first), (12, 3, second), (80, 1, third)])
|
||||
|
||||
def test_empty_or_foreign_bundle_is_rejected(self):
|
||||
for frames in ({}, {"type": "other", "frames": []}, {"type": "minimax_h3_frames", "frames": []}):
|
||||
with self.subTest(frames=frames):
|
||||
with self.assertRaises(ValueError):
|
||||
MODULE.MiniMaxH3AddGuideFrames._validate_frames_bundle(frames)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,233 @@
|
||||
// Shared geometry for the scene view and the actual square camera image.
|
||||
export const parts = [
|
||||
[0,2.78,0,.48,.52,.42], [0,2.28,0,.82,.48,.38], [0,1.83,0,.72,.38,.34],
|
||||
[0,1.43,0,.68,.34,.36], [-.53,2.25,0,.22,.62,.24], [.53,2.25,0,.22,.62,.24],
|
||||
[-.57,1.68,0,.2,.5,.22], [.57,1.68,0,.2,.5,.22],
|
||||
[-.23,.91,0,.26,.72,.3], [.23,.91,0,.26,.72,.3],
|
||||
[-.23,.32,0,.22,.46,.26], [.23,.32,0,.22,.46,.26],
|
||||
[-.23,.06,.09,.25,.12,.44], [.23,.06,.09,.25,.12,.44],
|
||||
[0,2.49,0,.2,.14,.22], [0,2.79,.235,.13,.10,.08], // nose: front is +Z
|
||||
[0,2.89,.235,.38,.085,.05], // horizontal front marker at eye level
|
||||
[-.57,1.35,0,.20,.16,.24],[.57,1.35,0,.20,.16,.24], // hands
|
||||
];
|
||||
const featureColors = {15:[239,189,98],16:[239,189,98]};
|
||||
const add=(a,b)=>a.map((v,i)=>v+b[i]);
|
||||
const sub=(a,b)=>a.map((v,i)=>v-b[i]);
|
||||
const mul=(a,s)=>a.map(v=>v*s);
|
||||
const dot=(a,b)=>a.reduce((s,v,i)=>s+v*b[i],0);
|
||||
const cross=(a,b)=>[a[1]*b[2]-a[2]*b[1],a[2]*b[0]-a[0]*b[2],a[0]*b[1]-a[1]*b[0]];
|
||||
const unit=a=>mul(a,1/Math.max(1e-9,Math.hypot(...a)));
|
||||
export function basis(pose) {
|
||||
const forward=unit(sub(pose.target,pose.position));
|
||||
const az=(pose.azimuth??Math.atan2(pose.position[0]-pose.target[0],pose.position[2]-pose.target[2])*180/Math.PI)*Math.PI/180;
|
||||
const right=[Math.cos(az),0,-Math.sin(az)];
|
||||
const up=cross(right,forward),r=(pose.roll||0)*Math.PI/180,c=Math.cos(r),s=Math.sin(r);
|
||||
return {forward,right:right.map((v,i)=>c*v-s*up[i]),up:up.map((v,i)=>s*right[i]+c*v)};
|
||||
}
|
||||
const ranges={extreme_close_up:[2.66,2.90],big_close_up:[2.56,3.00],close_up:[2.48,3.04],medium_close_up:[2.02,3.04],
|
||||
medium_shot:[1.63,3.04],cowboy_shot:[.91,3.04],medium_full_shot:[.55,3.04],
|
||||
full_shot:[0,3.04],wide_shot:[0,3.04],extreme_wide_shot:[0,3.04]};
|
||||
export const compositions={center:[0,0],left:[-1/3,0],right:[1/3,0],top:[0,1/3],bottom:[0,-1/3],
|
||||
top_left:[-1/3,1/3],top_right:[1/3,1/3],bottom_left:[-1/3,-1/3],bottom_right:[1/3,-1/3]};
|
||||
export function framingConfig(config) {
|
||||
return {shot_size:Object.hasOwn(ranges,config.shot_size)?config.shot_size:'full_shot'};
|
||||
}
|
||||
export function framingPoints(config) {
|
||||
const {shot_size}=framingConfig(config);
|
||||
const [bottom,top]=ranges[shot_size];
|
||||
const points=[];
|
||||
parts.forEach(([x,y,z,w,h,d],index)=>{
|
||||
const lo=Math.max(bottom,y-h/2),hi=Math.min(top,y+h/2);
|
||||
if(lo>hi)return;
|
||||
for(const px of [x-w/2,x+w/2])for(const py of [lo,hi])for(const pz of [z-d/2,z+d/2])points.push([px,py,pz]);
|
||||
});
|
||||
return {points,anchor:[0,(bottom+top)/2,0],shot_size};
|
||||
}
|
||||
export function solveCamera(config,previous=null) {
|
||||
const turns={rotate_left_180:-180,rotate_right_180:180,rotate_left_360:-360,rotate_right_360:360};
|
||||
const azimuth=config.direction in turns ? (previous?.azimuth||0)+turns[config.direction] : ({front:0,front_left_45:-45,left_profile:-90,rear_left_45:-135,rear:180,
|
||||
rear_right_45:135,right_profile:90,front_right_45:45})[config.direction]||0;
|
||||
const elevation=({extreme_low:-60,low_angle:-18,eye_level:0,high_angle:30,extreme_high:70,overhead:89.5})[config.angle]||0;
|
||||
const roll=[-90,-45,-30,-15,0,15,30,45,90,180].includes(Number(config.roll))?Number(config.roll):0;
|
||||
const {points,anchor,shot_size}=framingPoints(config);
|
||||
const [screenX,screenY]=compositions[config.composition]||compositions.center;
|
||||
const target=anchor, az=azimuth*Math.PI/180, el=elevation*Math.PI/180;
|
||||
const radial=[Math.sin(az)*Math.cos(el),Math.sin(el),Math.cos(az)*Math.cos(el)];
|
||||
const tangent=Math.tan(40*Math.PI/360);
|
||||
const evaluate=distance=>{
|
||||
const pose={position:add(anchor,mul(radial,distance)),target:anchor,tangent,roll};
|
||||
const b=basis(pose), projected=points.map(p=>{
|
||||
const v=sub(p,pose.position),z=dot(v,b.forward);
|
||||
return [dot(v,b.right)/(z*tangent),dot(v,b.up)/(z*tangent),z];
|
||||
});
|
||||
return {pose,minX:Math.min(...projected.map(p=>p[0])),maxX:Math.max(...projected.map(p=>p[0])),
|
||||
minY:Math.min(...projected.map(p=>p[1])),maxY:Math.max(...projected.map(p=>p[1])),near:Math.min(...projected.map(p=>p[2]))};
|
||||
};
|
||||
const occupancy=shot_size==='wide_shot'?.62:shot_size==='extreme_wide_shot'?.20:.90;
|
||||
let lo=.35,hi=30;
|
||||
for(let i=0;i<50;i++) {const mid=(lo+hi)/2,r=evaluate(mid);
|
||||
if(r.near<.1||r.maxX-r.minX>2*Math.min(occupancy,.95-Math.abs(screenX))||r.maxY-r.minY>2*Math.min(occupancy,.95-Math.abs(screenY))) lo=mid; else hi=mid;}
|
||||
// Match backend: do not approach merely because an overhead body foreshortens.
|
||||
if(shot_size==='wide_shot'||shot_size==='extreme_wide_shot')hi=Math.max(hi,3.04/(2*tangent*occupancy));
|
||||
const r=evaluate(hi);
|
||||
return {...r.pose,target,anchor,shiftX:(r.minX+r.maxX)/2-screenX,shiftY:(r.minY+r.maxY)/2-screenY,azimuth,elevation,distance:hi,
|
||||
orbit_route:config.direction in turns ? config.direction : config.orbit_route||'shortest'};
|
||||
}
|
||||
export function orbitDelta(a,b,route='shortest') {
|
||||
const turns={rotate_left_180:-180,rotate_right_180:180,rotate_left_360:-360,rotate_right_360:360};
|
||||
if(route in turns)return turns[route];
|
||||
const positive=((b-a)%360+360)%360;
|
||||
if(positive<1e-8)return 0; // A repeated endpoint is a hold, not an implicit full turn.
|
||||
if(route==='right')return positive;
|
||||
if(route==='left')return positive-360;
|
||||
return positive>=180?positive-360:positive;
|
||||
}
|
||||
export function interpolateCamera(a,b,t) {
|
||||
t=Math.max(0,Math.min(1,t)); const s=t*t*t*(t*(t*6-15)+10);
|
||||
const mix=(x,y)=>x+(y-x)*s;
|
||||
const delta=orbitDelta(a.azimuth,b.azimuth,b.orbit_route);
|
||||
const azimuth=a.azimuth+delta*s,elevation=mix(a.elevation,b.elevation),distance=mix(a.distance,b.distance);
|
||||
const target=a.target.map((v,i)=>mix(v,b.target[i]));
|
||||
const anchor=a.anchor.map((v,i)=>mix(v,b.anchor[i]));
|
||||
const az=azimuth*Math.PI/180,el=elevation*Math.PI/180;
|
||||
return {target,anchor,position:add(anchor,mul([Math.sin(az)*Math.cos(el),Math.sin(el),Math.cos(az)*Math.cos(el)],distance)),
|
||||
tangent:mix(a.tangent,b.tangent),shiftX:mix(a.shiftX,b.shiftX),shiftY:mix(a.shiftY,b.shiftY),roll:mix(a.roll||0,b.roll||0),azimuth,elevation,distance};
|
||||
}
|
||||
function unwrapAngles(nodes) {
|
||||
const values=[nodes[0].pose.azimuth];
|
||||
for(let i=1;i<nodes.length;i++)values.push(values[i-1]+orbitDelta(values[i-1],nodes[i].pose.azimuth,nodes[i].pose.orbit_route));
|
||||
return values;
|
||||
}
|
||||
function hermite(values,times,index,t) {
|
||||
const dt=Math.max(1e-6,times[index+1]-times[index]);
|
||||
const slope=i=>{
|
||||
if(i<=0||i>=values.length-1)return 0;
|
||||
const h0=Math.max(1e-6,times[i]-times[i-1]),h1=Math.max(1e-6,times[i+1]-times[i]);
|
||||
const d0=(values[i]-values[i-1])/h0,d1=(values[i+1]-values[i])/h1;
|
||||
if(d0*d1<=0)return 0;
|
||||
const w0=2*h1+h0,w1=h1+2*h0;
|
||||
return (w0+w1)/(w0/d0+w1/d1);
|
||||
};
|
||||
const u=Math.max(0,Math.min(1,t)),u2=u*u,u3=u2*u;
|
||||
return (2*u3-3*u2+1)*values[index]+(u3-2*u2+u)*dt*slope(index)
|
||||
+(-2*u3+3*u2)*values[index+1]+(u3-u2)*dt*slope(index+1);
|
||||
}
|
||||
// A complete Shot path shares waypoint velocity across adjacent Moves. Only the
|
||||
// reversing or held scalar stops at its waypoint; other axes keep moving.
|
||||
export function interpolateCameraPath(nodes,index,t) {
|
||||
if(nodes.length<2)return nodes[0]?.pose;
|
||||
index=Math.max(0,Math.min(nodes.length-2,index));
|
||||
const times=nodes.map(n=>n.time),azimuths=unwrapAngles(nodes);
|
||||
const scalar=key=>hermite(nodes.map(n=>n.pose[key]),times,index,t);
|
||||
const vector=key=>[0,1,2].map(axis=>hermite(nodes.map(n=>n.pose[key][axis]),times,index,t));
|
||||
const azimuth=hermite(azimuths,times,index,t),distance=scalar('distance');
|
||||
const elevation=scalar('elevation'),target=vector('target'),anchor=vector('anchor');
|
||||
const radius=distance*Math.cos(elevation*Math.PI/180);
|
||||
const height=target[1]+distance*Math.sin(elevation*Math.PI/180);
|
||||
const az=azimuth*Math.PI/180,position=[Math.sin(az)*radius,height,Math.cos(az)*radius];
|
||||
return {target,anchor,position,tangent:scalar('tangent'),shiftX:scalar('shiftX'),shiftY:scalar('shiftY'),roll:scalar('roll'),azimuth,elevation,distance};
|
||||
}
|
||||
export function projectPoint(p,pose) {
|
||||
const b=basis(pose),v=sub(p,pose.position),z=dot(v,b.forward);
|
||||
return [dot(v,b.right)/(z*pose.tangent)-(pose.shiftX||0),dot(v,b.up)/(z*pose.tangent)-(pose.shiftY||0),z];
|
||||
}
|
||||
export function modelFaces(pose,size) {
|
||||
const screen=p=>{const q=projectPoint(p,pose);return [(q[0]+1)*size/2,(1-q[1])*size/2,q[2]];};
|
||||
const faces=[];
|
||||
parts.forEach(([cx,cy,cz,wx,hy,dz],index)=>{
|
||||
const vs=[[-1,-1,-1],[1,-1,-1],[1,1,-1],[-1,1,-1],[-1,-1,1],[1,-1,1],[1,1,1],[-1,1,1]]
|
||||
.map(([a,c,d])=>[cx+a*wx/2,cy+c*hy/2,cz+d*dz/2]);
|
||||
for(const ids of [[0,3,2,1],[4,5,6,7],[0,4,7,3],[1,2,6,5],[3,7,6,2],[0,1,5,4]]){
|
||||
const vertices=ids.map(i=>vs[i]),normal=unit(cross(sub(vertices[1],vertices[0]),sub(vertices[2],vertices[0])));
|
||||
if(dot(normal,sub(pose.position,vertices[0]))<=0)continue;
|
||||
const ps=vertices.map(screen);if(ps.some(p=>p[2]<.05))continue;
|
||||
faces.push({ps,depth:ps.reduce((s,p)=>s+p[2],0)/4,light:.55+.45*Math.max(0,dot(normal,unit([-.5,1,1]))),index});
|
||||
}
|
||||
});
|
||||
return faces;
|
||||
}
|
||||
|
||||
// Per-pixel depth: a large head face must never paint over a nearer facial marker
|
||||
// just because the face's average distance sorts ahead of the small marker.
|
||||
export function rasterizeFaces(faces,size,pixels=new Uint8ClampedArray(size*size*4),depth=new Float64Array(size*size),wireWidth=0) {
|
||||
pixels.fill(0); depth.fill(0);
|
||||
const edge=(a,b,x,y)=>(b[0]-a[0])*(y-a[1])-(b[1]-a[1])*(x-a[0]);
|
||||
for(const face of faces) {
|
||||
const color=(featureColors[face.index]||[50,137,170]).map(c=>Math.round(c*face.light));
|
||||
// Shade only the quad perimeter, not the triangulation diagonal. Outlines
|
||||
// participate in the same depth test as the surface, so hidden edges stay hidden.
|
||||
const edges=face.ps.map((a,i)=>{
|
||||
const b=face.ps[(i+1)%4];
|
||||
return {a,b,length:Math.hypot(b[0]-a[0],b[1]-a[1])};
|
||||
});
|
||||
for(const triangle of [[0,1,2],[0,2,3]]) {
|
||||
const [a,b,c]=triangle.map(i=>face.ps[i]);
|
||||
const area=edge(a,b,c[0],c[1]); if(Math.abs(area)<1e-10)continue;
|
||||
const minX=Math.max(0,Math.floor(Math.min(a[0],b[0],c[0]))),maxX=Math.min(size-1,Math.ceil(Math.max(a[0],b[0],c[0])));
|
||||
const minY=Math.max(0,Math.floor(Math.min(a[1],b[1],c[1]))),maxY=Math.min(size-1,Math.ceil(Math.max(a[1],b[1],c[1])));
|
||||
for(let y=minY;y<=maxY;y++)for(let x=minX;x<=maxX;x++) {
|
||||
const u=edge(b,c,x+.5,y+.5)/area,v=edge(c,a,x+.5,y+.5)/area,t=1-u-v;
|
||||
if(u< -1e-9||v< -1e-9||t< -1e-9)continue;
|
||||
// Reciprocal depth interpolates linearly after perspective projection.
|
||||
const inverseZ=u/a[2]+v/b[2]+t/c[2],offset=y*size+x;
|
||||
if(inverseZ<=depth[offset])continue;
|
||||
depth[offset]=inverseZ;
|
||||
const pixel=offset*4;
|
||||
const outline=wireWidth>0 && edges.some(e=>e.length>1e-9 && Math.abs(edge(e.a,e.b,x+.5,y+.5))/e.length<wireWidth*.5);
|
||||
pixels[pixel]=outline?255:color[0];pixels[pixel+1]=outline?255:color[1];pixels[pixel+2]=outline?255:color[2];pixels[pixel+3]=255;
|
||||
}
|
||||
}
|
||||
}
|
||||
return pixels;
|
||||
}
|
||||
|
||||
const renderBuffers=new WeakMap();
|
||||
function scene(ctx,rect,pose,active,path) {
|
||||
const [x,y,w]=rect;
|
||||
const screen=p=>{const q=projectPoint(p,pose);return [x+(q[0]+1)*w/2,y+(1-q[1])*w/2,q[2]];};
|
||||
const line=(a,c,color)=>{const p=screen(a),q=screen(c);if(p[2]<.05||q[2]<.05)return;
|
||||
ctx.strokeStyle=color;ctx.beginPath();ctx.moveTo(p[0],p[1]);ctx.lineTo(q[0],q[1]);ctx.stroke();};
|
||||
ctx.save();ctx.beginPath();ctx.rect(x,y,w,w);ctx.clip();ctx.fillStyle='#111922';ctx.fillRect(x,y,w,w);
|
||||
for(let i=-10;i<=10;i++){line([i,0,-10],[i,0,10],'#263540');line([-10,0,i],[10,0,i],'#263540');}
|
||||
// Render above display resolution so thin features survive distance scaling.
|
||||
const size=Math.min(1200,Math.max(1,Math.ceil(w*Math.max(2,ctx.getTransform().a))));
|
||||
let buffers=renderBuffers.get(ctx);
|
||||
if(!buffers){buffers=new Map();renderBuffers.set(ctx,buffers);}
|
||||
const key=active?'scene':'output';
|
||||
let buffer=buffers.get(key);
|
||||
if(!buffer||buffer.size!==size){
|
||||
const canvas=ctx.canvas.ownerDocument.createElement('canvas');canvas.width=size;canvas.height=size;
|
||||
const context=canvas.getContext('2d');
|
||||
buffer={size,canvas,context,image:context.createImageData(size,size),depth:new Float64Array(size*size)};
|
||||
buffers.set(key,buffer);
|
||||
}
|
||||
rasterizeFaces(modelFaces(pose,size),size,buffer.image.data,buffer.depth,size/w);
|
||||
buffer.context.putImageData(buffer.image,0,0);
|
||||
ctx.drawImage(buffer.canvas,x,y,w,w);
|
||||
if(active){
|
||||
if(path)for(let i=1;i<path.length;i++)line(path[i-1].position,path[i].position,'#9c83ff');
|
||||
const cb=basis(active),depth=Math.min(active.distance,2),center=add(active.position,mul(cb.forward,depth));
|
||||
const corners=[[-1,-1],[1,-1],[1,1],[-1,1]].map(([a,c])=>add(center,add(mul(cb.right,(a+active.shiftX)*depth*active.tangent),mul(cb.up,(c+active.shiftY)*depth*active.tangent))));
|
||||
corners.forEach((p,i)=>{line(active.position,p,'#eaa469');line(p,corners[(i+1)%4],'#eaa469');});
|
||||
line(active.position,active.target,'#ffdc72');
|
||||
const p=screen(active.position),q=screen(active.target);ctx.fillStyle='#ffc663';ctx.fillRect(p[0]-4,p[1]-4,8,8);
|
||||
ctx.beginPath();ctx.arc(q[0],q[1],4,0,Math.PI*2);ctx.fill();
|
||||
}else{const q=screen(pose.target);ctx.strokeStyle='#ffd46d';ctx.beginPath();ctx.moveTo(q[0]-5,q[1]);ctx.lineTo(q[0]+5,q[1]);ctx.moveTo(q[0],q[1]-5);ctx.lineTo(q[0],q[1]+5);ctx.stroke();}
|
||||
ctx.restore();
|
||||
}
|
||||
export function drawCameraPreview(canvas,pose,path=[]) {
|
||||
const size=Math.max(240,canvas.clientWidth||400),dpr=Math.min(3,window.devicePixelRatio||1);
|
||||
if(canvas.width!==Math.round(size*dpr)||canvas.height!==Math.round(size*dpr)){canvas.width=Math.round(size*dpr);canvas.height=Math.round(size*dpr);}
|
||||
const ctx=canvas.getContext('2d');if(!ctx)return;ctx.setTransform(dpr,0,0,dpr,0,0);ctx.lineWidth=1;
|
||||
const radius=Math.max(7,pose.distance*1.8);
|
||||
const overview={position:[radius,.7*radius,radius],target:[0,1.5,0],tangent:.65};
|
||||
scene(ctx,[0,0,size],overview,pose,path);
|
||||
const w=size*.47,x=size-w-10,y=30;
|
||||
scene(ctx,[x,y,w],pose);ctx.strokeStyle='#65b9ff';ctx.strokeRect(x,y,w,w);
|
||||
ctx.fillStyle='#d8edff';ctx.font='10px sans-serif';ctx.fillText('CAMERA OUTPUT 1:1',x+6,y+14);
|
||||
if([pose,...path].some(p=>p.position[1]<0)) {
|
||||
ctx.fillStyle='#ffc663';ctx.font='11px sans-serif';
|
||||
ctx.fillText('Warning: camera path below ground.',10,size-26);
|
||||
ctx.fillText('Use a less low angle or a tighter shot.',10,size-11);
|
||||
}
|
||||
}
|
||||
@@ -1,5 +1,6 @@
|
||||
import { app } from "../../scripts/app.js";
|
||||
import { api } from "../../scripts/api.js";
|
||||
import { solveCamera, interpolateCameraPath, drawCameraPreview } from "./h3_camera_geometry.js";
|
||||
|
||||
const NODE_NAME = "MinimaxH3Prompter";
|
||||
const ENDPOINT = "/toyxyz/minimax_h3_prompter/compile";
|
||||
@@ -29,7 +30,7 @@ const MIN_TIMELINE_ITEM_FRAMES = 2;
|
||||
const MIN_SHOT_DURATION = MIN_TIMELINE_ITEM_FRAMES / VIDEO_OUTPUT_FPS;
|
||||
const MIN_VIDEO_CLIP_FRAMES = 10;
|
||||
const MIN_VIDEO_CLIP_DURATION = MIN_VIDEO_CLIP_FRAMES / VIDEO_OUTPUT_FPS;
|
||||
const CURRENT_PROJECT_VERSION = 28;
|
||||
const CURRENT_PROJECT_VERSION = 32;
|
||||
const ENHANCE_LEVELS = { none: "None", normal: "Normal", strong: "Strong" };
|
||||
const MAX_REFERENCES = { picture: 9, video: 3, audio: 3, total: 12 };
|
||||
// All timeline edits snap to exactly one output frame. Using a decimal time
|
||||
@@ -42,10 +43,7 @@ const LEGACY_DIALOGUE_LANGUAGES = [
|
||||
];
|
||||
const REFERENCE_ROLES = {
|
||||
picture: ["first_frame", "last_frame", "frame", "storyboard", "subject_identity"],
|
||||
video: [
|
||||
"none", "video_editing", "video_continuation", "subject_visual", "visual_style",
|
||||
"motion", "motion_camera", "camera", "cuts_rhythm",
|
||||
],
|
||||
video: ["motion", "video_continuation", "video_editing"],
|
||||
audio: [
|
||||
"none", "full_signal_copy", "partial_signal_copy", "voice_delivery",
|
||||
"dialogue_lyrics", "sound_ambience", "music_rhythm",
|
||||
@@ -67,6 +65,17 @@ const CAMERA_ANGLE_PRESETS = {
|
||||
dutch_angle: "Dutch angle", over_shoulder: "Over-the-shoulder", pov: "Point of view",
|
||||
three_quarter: "Three-quarter", profile: "Profile / side", rear: "Rear / from behind",
|
||||
};
|
||||
const CAMERA_DIRECTION_PRESETS = {
|
||||
none: "None",
|
||||
front: "Front · 0°",
|
||||
rear: "Rear · 180°",
|
||||
front_left_45: "Left · 45° · Front three-quarter",
|
||||
left_profile: "Left · 90° · Profile",
|
||||
rear_left_45: "Left · 135° · Rear three-quarter",
|
||||
front_right_45: "Right · 45° · Front three-quarter",
|
||||
right_profile: "Right · 90° · Profile",
|
||||
rear_right_45: "Right · 135° · Rear three-quarter",
|
||||
};
|
||||
const CAMERA_MOTION_PRESETS = {
|
||||
none: "None", static: "Static shot", zoom_in: "Zoom in", zoom_out: "Zoom out",
|
||||
push_in: "Push in", pull_out: "Pull out", pan_left: "Pan left", pan_right: "Pan right",
|
||||
@@ -89,8 +98,53 @@ const CAMERA_SHOT_PRESETS = {
|
||||
detail_shot: "Detail shot", two_shot: "Two shot", three_shot: "Three shot",
|
||||
group_shot: "Group shot",
|
||||
};
|
||||
const CAMERA_AMPLITUDE_PRESETS = { none: "None", small: "Small amplitude", large: "Large amplitude" };
|
||||
const CAMERA_SPEED_PRESETS = { none: "None", slow: "Slow speed", fast: "Fast speed" };
|
||||
const ADVANCED_CAMERA_SHOTS = {
|
||||
extreme_wide_shot: "Extreme Wide / Extreme Long", wide_shot: "Wide / Long", full_shot: "Full Shot",
|
||||
medium_full_shot: "Medium Full Shot", cowboy_shot: "Cowboy / American Shot", medium_shot: "Medium Shot",
|
||||
medium_close_up: "Medium Close-Up", close_up: "Close-Up", big_close_up: "Big Close-Up", extreme_close_up: "Extreme Close-Up",
|
||||
};
|
||||
const ADVANCED_CAMERA_DIRECTIONS = {
|
||||
...CAMERA_DIRECTION_PRESETS,
|
||||
rotate_left_180: "Left · 180° · From previous direction", rotate_left_360: "Left · 360° · From previous direction",
|
||||
rotate_right_180: "Right · 180° · From previous direction", rotate_right_360: "Right · 360° · From previous direction",
|
||||
};
|
||||
delete ADVANCED_CAMERA_DIRECTIONS.none;
|
||||
const ADVANCED_DIRECTION_GROUPS = [
|
||||
["Front / rear", ["front", "rear"]],
|
||||
["Camera-left", ["front_left_45", "left_profile", "rear_left_45", "rotate_left_180", "rotate_left_360"]],
|
||||
["Camera-right", ["front_right_45", "right_profile", "rear_right_45", "rotate_right_180", "rotate_right_360"]],
|
||||
];
|
||||
const ADVANCED_CAMERA_ANGLES = {
|
||||
extreme_low: "Extreme low · Worm’s-eye view (−60°)",
|
||||
low_angle: "Low angle (looking up)", eye_level: "Level view", high_angle: "High angle (looking down)",
|
||||
extreme_high: "Extreme high · Aerial view (70°)",
|
||||
overhead: "Overhead · Top-down (89.5°)",
|
||||
};
|
||||
const ADVANCED_CAMERA_ROLLS = Object.fromEntries([-90,-45,-30,-15,0,15,30,45,90,180].map(v=>[String(v),v===0?"Level · 0°":`${v>0?"Clockwise":"Counterclockwise"} · ${Math.abs(v)}°`]));
|
||||
const ADVANCED_ORBIT_ROUTES = {
|
||||
shortest: "Shortest path (180°: left)", left: "Camera-left route", right: "Camera-right route",
|
||||
};
|
||||
const ADVANCED_COMPOSITIONS = {
|
||||
center: "Center", left: "Left third", right: "Right third", top: "Upper third", bottom: "Lower third",
|
||||
top_left: "Upper left", top_right: "Upper right", bottom_left: "Lower left", bottom_right: "Lower right",
|
||||
};
|
||||
const DEFAULT_ADVANCED_CAMERA = () => ({
|
||||
shot_size: "full_shot", direction: "front", angle: "eye_level",
|
||||
});
|
||||
|
||||
function normalizeAdvancedCamera(value) {
|
||||
const raw = value && typeof value === "object" ? value : {};
|
||||
let shot_size=Object.hasOwn(ADVANCED_CAMERA_SHOTS,raw.shot_size)?raw.shot_size:"full_shot";
|
||||
return {
|
||||
shot_size,
|
||||
roll: Object.hasOwn(ADVANCED_CAMERA_ROLLS, raw.roll) ? String(raw.roll) : "0",
|
||||
motion: Object.hasOwn(CAMERA_MOTION_PRESETS, raw.motion) ? raw.motion : "none",
|
||||
direction: Object.hasOwn(ADVANCED_CAMERA_DIRECTIONS, raw.direction) ? raw.direction : "front",
|
||||
angle: Object.hasOwn(ADVANCED_CAMERA_ANGLES, raw.angle) ? raw.angle : "eye_level",
|
||||
orbit_route: Object.hasOwn(ADVANCED_ORBIT_ROUTES, raw.orbit_route) ? raw.orbit_route : "shortest",
|
||||
composition: Object.hasOwn(ADVANCED_COMPOSITIONS, raw.composition) ? raw.composition : "center",
|
||||
};
|
||||
}
|
||||
const STYLE_PRESETS = {
|
||||
none: "None", animation_2d: "2D animation", animation_3d: "3D animation",
|
||||
rough_hand_drawn_2d: "Rough hand-drawn animation",
|
||||
@@ -224,17 +278,16 @@ const STYLE_PRESET_GROUPS = [
|
||||
["Music", ["music_video"]],
|
||||
];
|
||||
const DEFAULT_SHOT_PRESETS = () => ({
|
||||
camera_angle: "none", camera_motion: "none", camera_shot: "none",
|
||||
camera_amplitude: "none", camera_speed: "none", style: "none",
|
||||
camera_angle: "none", camera_direction: "none", camera_motion: "none", camera_shot: "none",
|
||||
style: "none",
|
||||
});
|
||||
function normalizeShotPresets(value) {
|
||||
const raw = value && typeof value === "object" ? value : {};
|
||||
return {
|
||||
camera_angle: Object.hasOwn(CAMERA_ANGLE_PRESETS, raw.camera_angle) ? raw.camera_angle : "none",
|
||||
camera_direction: Object.hasOwn(CAMERA_DIRECTION_PRESETS, raw.camera_direction) ? raw.camera_direction : "none",
|
||||
camera_motion: Object.hasOwn(CAMERA_MOTION_PRESETS, raw.camera_motion) ? raw.camera_motion : "none",
|
||||
camera_shot: Object.hasOwn(CAMERA_SHOT_PRESETS, raw.camera_shot) ? raw.camera_shot : "none",
|
||||
camera_amplitude: Object.hasOwn(CAMERA_AMPLITUDE_PRESETS, raw.camera_amplitude) ? raw.camera_amplitude : "none",
|
||||
camera_speed: Object.hasOwn(CAMERA_SPEED_PRESETS, raw.camera_speed) ? raw.camera_speed : "none",
|
||||
style: Object.hasOwn(STYLE_PRESETS, raw.style) ? raw.style : "none",
|
||||
};
|
||||
}
|
||||
@@ -271,7 +324,7 @@ const REFERENCE_ROLE_HELP = {
|
||||
video_continuation: "Continue naturally from the ending state of the source video.",
|
||||
subject_visual: "Reference only the specified visible person, object, or environment as reusable Subject content.",
|
||||
visual_style: "Reference only the source video's rendering medium, palette, lighting treatment, materials, and visual texture.",
|
||||
motion: "Transfer only actor-neutral motion and action timing. Never transfer the source performer's face, body, skin, hair, clothing, materials, texture, identity, or visual style.",
|
||||
motion: "Motion reference generation: map source tracks to targets, e.g. red object = woman, blue object = man. Reference motion, placement, timing and camera behavior; source appearance, background and lighting are excluded. Audio reuse requires a separate audio role.",
|
||||
motion_camera: "Transfer only actor-neutral body motion/action timing and synchronized camera movement. Source people, objects, props, architecture, environment, background events, identity, appearance, style, cuts, text, and audio are excluded; motion maps only to target entities already requested.",
|
||||
camera: "Reference the source video's camera movement and viewpoint behavior only.",
|
||||
cuts_rhythm: "Reference the source video's cuts, pacing, rhythm, and temporal structure only.",
|
||||
@@ -304,6 +357,7 @@ const DEFAULT_PROJECT = () => ({
|
||||
duration: 5,
|
||||
visual_action: "",
|
||||
presets: DEFAULT_SHOT_PRESETS(),
|
||||
camera_advanced: DEFAULT_ADVANCED_CAMERA(),
|
||||
}],
|
||||
references: [],
|
||||
constraints: "",
|
||||
@@ -314,6 +368,9 @@ const DEFAULT_PROJECT = () => ({
|
||||
enhance: false,
|
||||
enhance_level: "none",
|
||||
enhanced_prompt: "",
|
||||
preview_mode: "video",
|
||||
advanced_camera_enabled: false,
|
||||
camera_render: false,
|
||||
});
|
||||
|
||||
function hideWidget(widget) {
|
||||
@@ -443,6 +500,9 @@ function normalizeProject(value) {
|
||||
? raw.enhance_level : raw.enhance === true ? "normal" : "none";
|
||||
project.enhance = project.enhance_level !== "none";
|
||||
project.enhanced_prompt = String(raw.enhanced_prompt || "");
|
||||
project.preview_mode = raw.preview_mode === "camera_advanced" ? "camera_advanced" : "video";
|
||||
project.advanced_camera_enabled = raw.advanced_camera_enabled ?? (raw.preview_mode === "camera_advanced");
|
||||
project.camera_render = raw.camera_render === true;
|
||||
if (Array.isArray(raw.shots) && raw.shots.length) {
|
||||
project.shots = raw.shots.map((shot, index) => ({
|
||||
id: String(shot?.id || uid(`shot-${index + 1}`)),
|
||||
@@ -450,6 +510,8 @@ function normalizeProject(value) {
|
||||
duration: clampNumber(shot?.duration, 1, MIN_SHOT_DURATION, 60),
|
||||
visual_action: migrateLegacyShotContent(shot, index),
|
||||
presets: normalizeShotPresets(shot?.presets),
|
||||
camera_advanced: normalizeAdvancedCamera(shot?.camera_advanced),
|
||||
camera_enabled: typeof shot?.camera_enabled === "boolean" ? shot.camera_enabled : null,
|
||||
}));
|
||||
} else {
|
||||
project.shots[0].duration = clampNumber(raw.requested_duration, 5, 0.1, 60);
|
||||
@@ -490,8 +552,7 @@ function normalizeProject(value) {
|
||||
const legacyVideoRoles = {
|
||||
reference: "none", continuation: "video_continuation", pacing: "cuts_rhythm",
|
||||
};
|
||||
role = REFERENCE_ROLES.video.includes(suppliedRole)
|
||||
? suppliedRole : (legacyVideoRoles[suppliedRole] || "none");
|
||||
role = legacyVideoRoles[suppliedRole] || suppliedRole || "motion";
|
||||
} else if (type === "audio") {
|
||||
const legacyAudioRoles = {
|
||||
reference: "none", voice_timbre: "voice_delivery", dialogue: "dialogue_lyrics",
|
||||
@@ -612,14 +673,29 @@ function installStyles() {
|
||||
.mmh3p-playback-time { grid-column:3; grid-row:1; color:#dfe7ef; text-align:right;
|
||||
font:10px ui-monospace,Consolas,monospace; }
|
||||
.mmh3p-video-preview-panel { min-width:0; min-height:0; padding:8px; display:none; flex-direction:column; gap:8px; overflow:hidden; }
|
||||
.mmh3p-preview-tabs { display:grid; grid-template-columns:1fr 1fr; gap:5px; flex:0 0 auto; }
|
||||
.mmh3p-preview-tabs button { min-width:0; font-weight:700; }
|
||||
.mmh3p-preview-tabs button.active { color:#06131d; background:var(--accent); border-color:var(--accent); }
|
||||
.mmh3p-video-preview-stage { position:relative; flex:1 1 auto; min-height:260px; overflow:hidden;
|
||||
border:1px solid #36404b; border-radius:6px; background:#11161d; }
|
||||
.mmh3p-video-preview-stage[hidden],.mmh3p-video-preview-status[hidden],.mmh3p-camera-advanced[hidden] { display:none; }
|
||||
.mmh3p-video-preview-stage video { position:absolute; inset:0; display:block; width:100%; height:100%;
|
||||
object-fit:contain; background:#090c10; }
|
||||
.mmh3p-video-preview-empty { position:absolute; inset:0; display:flex; align-items:center; justify-content:center;
|
||||
padding:24px; color:#7f8996; text-align:center; background:#0d1218; pointer-events:none; }
|
||||
.mmh3p-video-preview-empty[hidden] { display:none; }
|
||||
.mmh3p-video-preview-status { min-height:30px; color:#aeb8c4; font:10px/1.4 ui-monospace,Consolas,monospace; }
|
||||
.mmh3p-camera-advanced { min-height:0; overflow:hidden; display:flex; flex-direction:column; gap:8px; }
|
||||
.mmh3p-camera-viewport-wrap { position:relative; width:100%; aspect-ratio:1 / 1; flex:0 0 auto;
|
||||
border:1px solid #36404b; border-radius:6px; overflow:hidden; background:#0b1016; }
|
||||
.mmh3p-camera-viewport { display:block; width:100%; height:100%; }
|
||||
.mmh3p-camera-viewport-label { position:absolute; left:8px; top:7px; color:#a9d8ff;
|
||||
font:700 9px ui-monospace,Consolas,monospace; pointer-events:none; }
|
||||
.mmh3p-camera-controls { display:grid; grid-template-columns:1fr; gap:7px; padding:8px;
|
||||
border:1px solid #36404b; border-radius:6px; background:#171b21; }
|
||||
.mmh3p-camera-control { display:grid; grid-template-columns:88px minmax(0,1fr); align-items:center; gap:7px; }
|
||||
.mmh3p-camera-control span { color:var(--muted); font-size:10px; font-weight:700; }
|
||||
.mmh3p-camera-control select { width:100%; min-width:0; }
|
||||
.mmh3p-timeline { height:112px; display:flex; align-items:stretch; gap:0; padding-top:18px;
|
||||
position:relative; overflow-x:hidden; margin:0 7px; }
|
||||
.mmh3p-ruler { position:absolute; left:0; right:0; top:0; color:#707784; font-size:9px;
|
||||
@@ -697,7 +773,7 @@ function installStyles() {
|
||||
.mmh3p-preset-tabs { display:flex; gap:5px; margin-bottom:5px; }
|
||||
.mmh3p-preset-tabs button { min-width:72px; padding:4px 10px; }
|
||||
.mmh3p-preset-tabs button.active { color:#d9efff; border-color:var(--accent); background:#263746; }
|
||||
.mmh3p-preset-panel { display:grid; grid-template-columns:repeat(3,minmax(0,1fr)); gap:6px; }
|
||||
.mmh3p-preset-panel { display:grid; grid-template-columns:repeat(4,minmax(0,1fr)); gap:6px; }
|
||||
.mmh3p-preset-panel[hidden] { display:none; }
|
||||
.mmh3p-preset-panel.style { grid-template-columns:minmax(0,1fr); }
|
||||
.mmh3p-preset-field { min-width:0; display:flex; align-items:center; gap:5px; }
|
||||
@@ -831,6 +907,7 @@ class PrompterUI {
|
||||
this.videoThumbnailCache = new Map();
|
||||
this.modelBundles = [];
|
||||
this.compileSequence = 0;
|
||||
this.cameraTimelineSource = "";
|
||||
this.mentionSelectionIndex = 0;
|
||||
this.mentionMenuSignature = "";
|
||||
this.visibleMentionEntries = [];
|
||||
@@ -854,11 +931,33 @@ class PrompterUI {
|
||||
this.root.className = "mmh3p";
|
||||
this.root.innerHTML = `
|
||||
<aside class="mmh3p-panel mmh3p-video-preview-panel" data-el="video-preview-panel">
|
||||
<div class="mmh3p-video-preview-stage">
|
||||
<div class="mmh3p-preview-tabs">
|
||||
<button data-action="preview-mode" data-preview-mode="video" type="button" title="Show the selected reference video">Video</button>
|
||||
<button data-action="preview-mode" data-preview-mode="camera_advanced" type="button" title="Configure and preview the selected Shot or Move camera state">Camera Advanced</button>
|
||||
</div>
|
||||
<div class="mmh3p-video-preview-stage" data-el="video-preview-stage">
|
||||
<video data-el="reference-video-preview" muted playsinline preload="metadata"></video>
|
||||
<div class="mmh3p-video-preview-empty" data-el="reference-video-empty">Add a video reference to preview it on this timeline.</div>
|
||||
</div>
|
||||
<div class="mmh3p-video-preview-status" data-el="video-preview-status">Video · no video reference</div>
|
||||
<div class="mmh3p-camera-advanced" data-el="camera-advanced" hidden>
|
||||
<div class="mmh3p-camera-viewport-wrap">
|
||||
<canvas class="mmh3p-camera-viewport" data-el="camera-viewport" aria-label="Advanced camera 3D viewport"></canvas>
|
||||
<span class="mmh3p-camera-viewport-label">3D SCENE · CAMERA FRUSTUM</span>
|
||||
</div>
|
||||
<div class="mmh3p-camera-controls">
|
||||
<label class="mmh3p-auto-run" title="Render the entire panel camera timeline to the camera_render IMAGE batch when this node executes. 1024×1024, 24fps. Proxy mannequin and grid only; text-only Motion and Qwen path changes are not rendered. Uses about 1.5 GiB RAM per 5 seconds."><input type="checkbox" data-el="camera-render"><span>Camera render · 1024×1024 · Full timeline</span></label>
|
||||
<label class="mmh3p-auto-run" title="Apply panel defaults to this Shot/Move. Explicit camera instructions in your prompt always take priority, even when unchecked. Without such text, disabled items hold their camera state. The 3D preview represents panel settings only."><input type="checkbox" data-el="advanced-camera-enabled"><span>Use Camera Advanced for this Shot / Move</span></label>
|
||||
<label class="mmh3p-camera-control"><span>Motion</span><select data-advanced-camera="motion"></select></label>
|
||||
<small data-el="camera-motion-note" hidden>Motion is a prompt instruction; its endpoint is not simulated in the 3D preview.</small>
|
||||
<label class="mmh3p-camera-control"><span>Shot size</span><select data-advanced-camera="shot_size"></select></label>
|
||||
<label class="mmh3p-camera-control" title="Left/right are camera-based. 45–135° specify a destination from the front axis; 180°/360° continue from the previous direction in a Move."><span>Direction</span><select data-advanced-camera="direction"></select></label>
|
||||
<label class="mmh3p-camera-control"><span>Angle</span><select data-advanced-camera="angle"></select></label>
|
||||
<label class="mmh3p-camera-control" title="Fixed camera rotation around the optical axis, not an orbit around the subject."><span>Camera roll</span><select data-advanced-camera="roll"></select></label>
|
||||
<label class="mmh3p-camera-control" title="Move travel direction, not the destination viewpoint. Identical directions hold the orbit; they do not create a full turn."><span>Orbit route</span><select data-advanced-camera="orbit_route"></select></label>
|
||||
<label class="mmh3p-camera-control" title="Screen position of the selected body range, not the camera orbit direction. Moves interpolate the framing offset; distance may increase to keep the range visible."><span>Composition</span><select data-advanced-camera="composition"></select></label>
|
||||
</div>
|
||||
</div>
|
||||
</aside>
|
||||
<div class="mmh3p-workspace">
|
||||
<div class="mmh3p-main">
|
||||
@@ -901,16 +1000,8 @@ class PrompterUI {
|
||||
</div>
|
||||
<div class="mmh3p-presets">
|
||||
<div class="mmh3p-preset-tabs">
|
||||
<button data-action="preset-tab" data-preset-tab="camera" type="button" title="Show camera angle, motion, framing, amplitude, and speed presets">Camera</button>
|
||||
<button data-action="preset-tab" data-preset-tab="style" type="button" title="Show the visual style preset for the selected Shot or Move">Style</button>
|
||||
</div>
|
||||
<div class="mmh3p-preset-panel" data-el="camera-presets">
|
||||
<label class="mmh3p-preset-field"><span>Angle</span><select data-preset="camera_angle"></select></label>
|
||||
<label class="mmh3p-preset-field"><span>Motion</span><select data-preset="camera_motion"></select></label>
|
||||
<label class="mmh3p-preset-field"><span>Shot</span><select data-preset="camera_shot"></select></label>
|
||||
<label class="mmh3p-preset-field"><span>Amplitude</span><select data-preset="camera_amplitude"></select></label>
|
||||
<label class="mmh3p-preset-field"><span>Speed</span><select data-preset="camera_speed"></select></label>
|
||||
</div>
|
||||
<div class="mmh3p-preset-panel style" data-el="style-presets" hidden>
|
||||
<label class="mmh3p-preset-field"><span>Style</span><select data-preset="style"></select></label>
|
||||
</div>
|
||||
@@ -968,11 +1059,23 @@ class PrompterUI {
|
||||
const select = this.root.querySelector(`[data-preset="${name}"]`);
|
||||
Object.entries(options).forEach(([value, label]) => select.add(new Option(label, value)));
|
||||
};
|
||||
fillPreset("camera_angle", CAMERA_ANGLE_PRESETS);
|
||||
fillPreset("camera_motion", CAMERA_MOTION_PRESETS);
|
||||
fillPreset("camera_shot", CAMERA_SHOT_PRESETS);
|
||||
fillPreset("camera_amplitude", CAMERA_AMPLITUDE_PRESETS);
|
||||
fillPreset("camera_speed", CAMERA_SPEED_PRESETS);
|
||||
const fillAdvancedCamera = (name, options) => {
|
||||
const select = this.root.querySelector(`[data-advanced-camera="${name}"]`);
|
||||
Object.entries(options).forEach(([value, label]) => select.add(new Option(label, value)));
|
||||
};
|
||||
fillAdvancedCamera("shot_size", ADVANCED_CAMERA_SHOTS);
|
||||
fillAdvancedCamera("motion", CAMERA_MOTION_PRESETS);
|
||||
const directionSelect = this.root.querySelector('[data-advanced-camera="direction"]');
|
||||
ADVANCED_DIRECTION_GROUPS.forEach(([label, values]) => {
|
||||
const group = document.createElement("optgroup");
|
||||
group.label = label;
|
||||
values.forEach(value => group.append(new Option(ADVANCED_CAMERA_DIRECTIONS[value], value)));
|
||||
directionSelect.append(group);
|
||||
});
|
||||
fillAdvancedCamera("angle", ADVANCED_CAMERA_ANGLES);
|
||||
fillAdvancedCamera("roll", ADVANCED_CAMERA_ROLLS);
|
||||
fillAdvancedCamera("orbit_route", ADVANCED_ORBIT_ROUTES);
|
||||
fillAdvancedCamera("composition", ADVANCED_COMPOSITIONS);
|
||||
const styleSelect = this.root.querySelector('[data-preset="style"]');
|
||||
styleSelect.add(new Option(STYLE_PRESETS.none, "none"));
|
||||
STYLE_PRESET_GROUPS.forEach(([label, values]) => {
|
||||
@@ -981,7 +1084,7 @@ class PrompterUI {
|
||||
values.forEach(value => group.append(new Option(STYLE_PRESETS[value], value)));
|
||||
styleSelect.append(group);
|
||||
});
|
||||
this.activePresetTab = "camera";
|
||||
this.activePresetTab = "style";
|
||||
this.bind();
|
||||
}
|
||||
|
||||
@@ -1005,6 +1108,49 @@ class PrompterUI {
|
||||
if (!entries[0]?.isIntersecting) this.stopTimelinePlayback();
|
||||
}, { threshold: .01 });
|
||||
this.previewIntersectionObserver.observe(this.els["video-preview-panel"]);
|
||||
if (typeof ResizeObserver !== "undefined") {
|
||||
this.cameraViewportResizeObserver = new ResizeObserver(() => {
|
||||
if (this.project.preview_mode === "camera_advanced") this.renderAdvancedCamera();
|
||||
});
|
||||
this.cameraViewportResizeObserver.observe(this.els["camera-viewport"].parentElement);
|
||||
}
|
||||
this.root.querySelectorAll('[data-action="preview-mode"]').forEach(button => {
|
||||
button.addEventListener("click", () => {
|
||||
this.project.preview_mode = button.dataset.previewMode === "camera_advanced" ? "camera_advanced" : "video";
|
||||
this.commit(false);
|
||||
this.renderPreviewMode();
|
||||
});
|
||||
});
|
||||
this.root.querySelectorAll("[data-advanced-camera]").forEach(select => {
|
||||
select.addEventListener("change", () => {
|
||||
const shot = this.selectedShot();
|
||||
if (!shot) return;
|
||||
shot.camera_advanced = normalizeAdvancedCamera(shot.camera_advanced);
|
||||
shot.camera_advanced[select.dataset.advancedCamera] = select.value;
|
||||
shot.camera_advanced = normalizeAdvancedCamera(shot.camera_advanced);
|
||||
this.project.advanced_camera_enabled = true;
|
||||
shot.camera_enabled = true;
|
||||
this.commit();
|
||||
this.renderAdvancedCamera();
|
||||
});
|
||||
});
|
||||
this.els["camera-render"].addEventListener("change", () => {
|
||||
this.project.camera_render = this.els["camera-render"].checked;
|
||||
this.commit();
|
||||
this.syncReferenceOutputs();
|
||||
});
|
||||
this.els["advanced-camera-enabled"].addEventListener("change", () => {
|
||||
const selected = this.selectedShot();
|
||||
if (!selected) return;
|
||||
// Materialize legacy global state before changing one item.
|
||||
this.project.shots.forEach(item => {
|
||||
if (typeof item.camera_enabled !== "boolean") item.camera_enabled = this.project.advanced_camera_enabled === true;
|
||||
});
|
||||
selected.camera_enabled = this.els["advanced-camera-enabled"].checked;
|
||||
this.project.advanced_camera_enabled = true;
|
||||
this.commit();
|
||||
this.renderAdvancedCamera();
|
||||
});
|
||||
this.els.mode.addEventListener("change", () => { this.project.mode = this.els.mode.value; this.commit(); this.renderHeader(); });
|
||||
this.els.duration.addEventListener("change", () => {
|
||||
const minimumTimeline = this.project.shots.length * MIN_SHOT_DURATION;
|
||||
@@ -1438,7 +1584,7 @@ class PrompterUI {
|
||||
this.els["video-preview-status"].textContent = `Video ${videoNumber} · source ${this.formatPlaybackTime(sourceTime)} · timeline ${this.formatPlaybackTime(this.playheadTime)}`;
|
||||
if (forceSeek || !this.timelinePlaying) this.scheduleReferenceVideoSeek(sourceTime);
|
||||
else if (Math.abs((video.currentTime || 0) - sourceTime) > .15) this.scheduleReferenceVideoSeek(sourceTime);
|
||||
if (this.timelinePlaying && video.paused) video.play().catch(() => {});
|
||||
if (this.timelinePlaying && this.project.preview_mode === "video" && video.paused) video.play().catch(() => {});
|
||||
}
|
||||
|
||||
scheduleReferenceVideoSeek(sourceTime) {
|
||||
@@ -1479,6 +1625,7 @@ class PrompterUI {
|
||||
});
|
||||
this.renderShotEditor();
|
||||
this.renderPresets();
|
||||
this.renderAdvancedCamera();
|
||||
this.updateTimelineActionButtons();
|
||||
}
|
||||
|
||||
@@ -1491,6 +1638,7 @@ class PrompterUI {
|
||||
this.els["playhead-line"].style.setProperty("--playhead-position", position);
|
||||
this.els["timeline-scrubber-track"].style.setProperty("--playhead-position", position);
|
||||
this.syncSelectedShotToPlayhead();
|
||||
this.renderAdvancedCamera();
|
||||
this.syncReferenceVideoPreview(forceMedia);
|
||||
}
|
||||
|
||||
@@ -1532,6 +1680,7 @@ class PrompterUI {
|
||||
this.root.removeEventListener("pointerdown", this.nativeSelectScaleHandler, true);
|
||||
document.removeEventListener("visibilitychange", this.timelineVisibilityHandler);
|
||||
this.previewIntersectionObserver?.disconnect();
|
||||
this.cameraViewportResizeObserver?.disconnect();
|
||||
const video = this.els["reference-video-preview"];
|
||||
video?.pause();
|
||||
video?.removeAttribute("src");
|
||||
@@ -1556,6 +1705,7 @@ class PrompterUI {
|
||||
const shot = {
|
||||
id: uid("shot"), duration: firstHalf, visual_action: "",
|
||||
presets: DEFAULT_SHOT_PRESETS(),
|
||||
camera_advanced: DEFAULT_ADVANCED_CAMERA(),
|
||||
kind: "shot",
|
||||
};
|
||||
this.project.shots.splice(selectedIndex + 1, 0, shot);
|
||||
@@ -1571,6 +1721,7 @@ class PrompterUI {
|
||||
const move = {
|
||||
id: uid("move"), kind: "move", duration: firstHalf, visual_action: "",
|
||||
presets: DEFAULT_SHOT_PRESETS(),
|
||||
camera_advanced: normalizeAdvancedCamera(selected.camera_advanced),
|
||||
};
|
||||
this.project.shots.splice(selectedIndex + 1, 0, move);
|
||||
this.selectedShotId = move.id; this.commit(); this.render();
|
||||
@@ -1915,6 +2066,7 @@ class PrompterUI {
|
||||
|
||||
async compile() {
|
||||
const sequence = ++this.compileSequence;
|
||||
const sourceProject = JSON.stringify(this.project);
|
||||
this.compileController?.abort();
|
||||
const controller = new AbortController();
|
||||
this.compileController = controller;
|
||||
@@ -1923,13 +2075,15 @@ class PrompterUI {
|
||||
const response = await api.fetchApi(ENDPOINT, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify({ project_data: JSON.stringify(this.project) }),
|
||||
body: JSON.stringify({ project_data: sourceProject }),
|
||||
signal: controller.signal,
|
||||
});
|
||||
const data = await response.json();
|
||||
if (sequence !== this.compileSequence) return;
|
||||
if (sourceProject !== JSON.stringify(this.project)) return;
|
||||
if (!response.ok || data.status !== "success") throw new Error(data.message || "Compilation failed");
|
||||
this.previewData = data;
|
||||
this.cameraTimelineSource = sourceProject;
|
||||
this.resolvedMode = data.resolved_mode || inferAutoMode(this.project.references);
|
||||
this.els.effective.textContent = `${data.effective_frames}f / ${Number(data.effective_duration).toFixed(2)}s`;
|
||||
const validationLine = this.els.log?.querySelector('[data-log-key="validation"]');
|
||||
@@ -1956,13 +2110,98 @@ class PrompterUI {
|
||||
this.selectedShotId = this.project.shots[0]?.id;
|
||||
}
|
||||
this.renderHeader(); this.renderTimeline(); this.renderShotEditor(); this.renderPresets(); this.renderReferences();
|
||||
this.renderPreviewMode();
|
||||
this.syncReferenceVideoPreview(true);
|
||||
this.scheduleCompile();
|
||||
}
|
||||
|
||||
renderPreviewMode() {
|
||||
const mode = this.project.preview_mode === "camera_advanced" ? "camera_advanced" : "video";
|
||||
this.els["video-preview-stage"].hidden = mode !== "video";
|
||||
this.els["video-preview-status"].hidden = mode !== "video";
|
||||
this.els["camera-advanced"].hidden = mode !== "camera_advanced";
|
||||
this.root.querySelectorAll('[data-action="preview-mode"]').forEach(button => {
|
||||
button.classList.toggle("active", button.dataset.previewMode === mode);
|
||||
});
|
||||
if (mode === "camera_advanced") {
|
||||
this.els["reference-video-preview"]?.pause();
|
||||
requestAnimationFrame(() => this.renderAdvancedCamera());
|
||||
} else {
|
||||
this.syncReferenceVideoPreview(true);
|
||||
}
|
||||
}
|
||||
|
||||
advancedCameraMoveNodes(index) {
|
||||
let blockStart = index - 1;
|
||||
while (blockStart > 0 && this.project.shots[blockStart].kind === "move") blockStart--;
|
||||
let blockEnd = index;
|
||||
while (blockEnd + 1 < this.project.shots.length && this.project.shots[blockEnd + 1].kind === "move") blockEnd++;
|
||||
const nodes = [];
|
||||
for (let itemIndex = blockStart; itemIndex <= blockEnd; itemIndex++) {
|
||||
nodes.push({
|
||||
pose: this.cameraItemPose(itemIndex),
|
||||
time: this.shotTimelineRange(itemIndex).endSeconds,
|
||||
});
|
||||
}
|
||||
return { nodes, segment: index - blockStart - 1 };
|
||||
}
|
||||
|
||||
cameraItemPose(index) {
|
||||
let pose;
|
||||
for (let i=0;i<=index;i++) {
|
||||
const item=this.project.shots[i];
|
||||
pose=item.camera_enabled === false && pose ? {...pose,orbit_route:"shortest"}
|
||||
: solveCamera(normalizeAdvancedCamera(item.camera_advanced),item.kind === "move" ? pose : null);
|
||||
}
|
||||
return pose;
|
||||
}
|
||||
|
||||
renderAdvancedCamera() {
|
||||
const shot = this.selectedShot();
|
||||
if (!shot || !this.els["camera-viewport"]) return;
|
||||
const config = normalizeAdvancedCamera(shot.camera_advanced);
|
||||
shot.camera_advanced = config;
|
||||
const enabled = shot.camera_enabled ?? (this.project.advanced_camera_enabled === true);
|
||||
const itemIndex = this.project.shots.indexOf(shot);
|
||||
let ownerIndex=itemIndex;
|
||||
while(ownerIndex>0 && this.project.shots[ownerIndex].kind === "move") ownerIndex--;
|
||||
const owner=this.project.shots[ownerIndex];
|
||||
const followingMoves=[];
|
||||
for(let i=ownerIndex+1;i<this.project.shots.length && this.project.shots[i].kind === "move";i++) followingMoves.push(this.project.shots[i]);
|
||||
const hasEnabledMove=followingMoves.some(item => (item.camera_enabled ?? this.project.advanced_camera_enabled) === true);
|
||||
const motionAllowed=shot.kind !== "move" && !hasEnabledMove;
|
||||
const ownerMotion=normalizeAdvancedCamera(owner.camera_advanced).motion;
|
||||
const inheritedMotion=shot.kind === "move" && !enabled && !hasEnabledMove
|
||||
&& (owner.camera_enabled ?? this.project.advanced_camera_enabled) === true && !["none","static"].includes(ownerMotion);
|
||||
this.els["camera-motion-note"].hidden = !inheritedMotion && config.motion === "none" && motionAllowed;
|
||||
this.els["camera-motion-note"].textContent = inheritedMotion
|
||||
? "Holds the preceding Shot's motion endpoint. That endpoint is text-directed; the 3D preview shows starting framing only."
|
||||
: !motionAllowed ? "Motion is available on a Shot when all following Moves in that Shot have camera disabled."
|
||||
: "Motion is a prompt instruction; the 3D preview shows starting framing only. Following camera-disabled Moves hold its endpoint.";
|
||||
this.els["advanced-camera-enabled"].checked = enabled;
|
||||
this.els["camera-render"].checked = this.project.camera_render === true;
|
||||
this.root.querySelectorAll("[data-advanced-camera]").forEach(select => {
|
||||
select.value = config[select.dataset.advancedCamera];
|
||||
select.disabled = !enabled || (select.dataset.advancedCamera === "orbit_route" && (shot.kind !== "move" || config.direction.startsWith("rotate_")));
|
||||
if (select.dataset.advancedCamera === "motion") select.disabled = !enabled || !motionAllowed;
|
||||
});
|
||||
if (this.project.preview_mode !== "camera_advanced") return;
|
||||
const index = this.project.shots.indexOf(shot);
|
||||
const end = this.cameraItemPose(index);
|
||||
const range = this.shotTimelineRange(index);
|
||||
const fraction = Math.max(0, Math.min(1, (this.playheadTime-range.startSeconds) /
|
||||
Math.max(.001, range.endSeconds-range.startSeconds)));
|
||||
let pose=end,path=[];
|
||||
if(shot.kind === "move"&&index>0) {
|
||||
const { nodes, segment } = this.advancedCameraMoveNodes(index);
|
||||
pose=interpolateCameraPath(nodes,segment,fraction);
|
||||
path=Array.from({length:41},(_,i)=>interpolateCameraPath(nodes,segment,i/40));
|
||||
}
|
||||
drawCameraPreview(this.els["camera-viewport"], pose, path);
|
||||
}
|
||||
|
||||
renderPresets() {
|
||||
const tab = this.activePresetTab === "style" ? "style" : "camera";
|
||||
this.els["camera-presets"].hidden = tab !== "camera";
|
||||
const tab = "style";
|
||||
this.els["style-presets"].hidden = tab !== "style";
|
||||
this.root.querySelectorAll('[data-action="preset-tab"]').forEach(button => {
|
||||
button.classList.toggle("active", button.dataset.presetTab === tab);
|
||||
@@ -2044,6 +2283,7 @@ class PrompterUI {
|
||||
this.renderTimeline();
|
||||
this.renderShotEditor();
|
||||
this.renderPresets();
|
||||
this.renderAdvancedCamera();
|
||||
this.updateTimelineActionButtons();
|
||||
});
|
||||
card.addEventListener("dragstart", event => {
|
||||
@@ -2075,7 +2315,7 @@ class PrompterUI {
|
||||
while (target > 0 && this.project.shots[target]?.kind === "move") target -= 1;
|
||||
this.project.shots.splice(target, 0, ...moved);
|
||||
}
|
||||
this.commit(); this.renderTimeline();
|
||||
this.commit(); this.renderTimeline(); this.renderAdvancedCamera();
|
||||
}
|
||||
});
|
||||
if (index < this.project.shots.length - 1) {
|
||||
@@ -2088,7 +2328,7 @@ class PrompterUI {
|
||||
const pairTotal = this.project.shots[index].duration + this.project.shots[index + 1].duration;
|
||||
this.project.shots[index].duration = pairTotal / 2;
|
||||
this.project.shots[index + 1].duration = pairTotal / 2;
|
||||
this.commit(); this.renderTimeline();
|
||||
this.commit(); this.renderTimeline(); this.renderAdvancedCamera();
|
||||
});
|
||||
card.appendChild(handle);
|
||||
}
|
||||
@@ -2473,7 +2713,7 @@ class PrompterUI {
|
||||
handle.removeEventListener("pointerup", onUp);
|
||||
handle.removeEventListener("pointercancel", onUp);
|
||||
this.root.classList.remove("resizing"); handle.classList.remove("active");
|
||||
this.commit(); this.renderHeader(); this.renderTimeline();
|
||||
this.commit(); this.renderHeader(); this.renderTimeline(); this.renderAdvancedCamera();
|
||||
};
|
||||
handle.addEventListener("pointermove", onMove);
|
||||
handle.addEventListener("pointerup", onUp);
|
||||
@@ -2518,6 +2758,11 @@ class PrompterUI {
|
||||
const labelEl = document.createElement("span"); labelEl.className = "mmh3p-ref-label"; labelEl.textContent = label;
|
||||
const role = document.createElement("select");
|
||||
if (role) {
|
||||
if (ref.type === "video" && !REFERENCE_ROLES.video.includes(ref.role)) {
|
||||
const unavailable = new Option(`Reselect preset (removed: ${REFERENCE_ROLE_LABELS[ref.role] || ref.role})`, ref.role);
|
||||
unavailable.disabled = true;
|
||||
role.add(unavailable);
|
||||
}
|
||||
REFERENCE_ROLES[ref.type].forEach(value => {
|
||||
const words = value.replaceAll("_", " ");
|
||||
role.add(new Option(REFERENCE_ROLE_LABELS[value] || words.charAt(0).toUpperCase() + words.slice(1), value));
|
||||
@@ -2540,6 +2785,8 @@ class PrompterUI {
|
||||
? "Optional: describe what this storyboard should guide; do not request identity or style transfer"
|
||||
: ref.type === "audio"
|
||||
? (AUDIO_DESCRIPTION_PLACEHOLDERS[ref.role] || AUDIO_DESCRIPTION_PLACEHOLDERS.none)
|
||||
: ref.role === "motion"
|
||||
? "Map motion here or in Prompt: e.g. <Video 1> red object = woman; blue object = man. Reference their motion and camera behavior in a new scene. Describe target setting/style; source appearance/environment is excluded. Audio reuse is separate."
|
||||
: "Describe how this video should guide the target in English";
|
||||
desc.value = ref.description;
|
||||
}
|
||||
@@ -2738,7 +2985,7 @@ class PrompterUI {
|
||||
);
|
||||
const hasFrameBundle = pictures.some(ref => ref.role === "frame");
|
||||
const targetOutputCount = fixedOutputCount + pictureCount + videoCount + audioCount
|
||||
+ (hasFrameBundle ? 1 : 0);
|
||||
+ (hasFrameBundle ? 1 : 0) + (this.project.camera_render ? 1 : 0);
|
||||
this.node.outputs ||= [];
|
||||
for (let index = this.node.outputs.length - 1; index >= fixedOutputCount; index -= 1) {
|
||||
if (/^frame_\d+$/.test(String(this.node.outputs[index]?.name || ""))) {
|
||||
@@ -2782,6 +3029,12 @@ class PrompterUI {
|
||||
const output = this.node.outputs[outputIndex];
|
||||
output.name = "frames";
|
||||
output.type = "MINIMAX_H3_FRAMES";
|
||||
outputIndex += 1;
|
||||
}
|
||||
if (this.project.camera_render) {
|
||||
const output = this.node.outputs[outputIndex];
|
||||
output.name = "camera_render";
|
||||
output.type = "IMAGE";
|
||||
}
|
||||
this.node._widgetSlotsDirty = true;
|
||||
this.node.setDirtyCanvas?.(true, true);
|
||||
|
||||