updated requirements and reconfigured

This commit is contained in:
Dag Thomas Olsen
2026-08-13 14:02:14 +02:00
parent a0cad027fe
commit 25578a7ebe
19 changed files with 2824 additions and 193 deletions
+121 -12
View File
@@ -488,6 +488,113 @@ Optimized latent generation for Stable Diffusion 3 pipelines.
---
### 📏 Resolution Planning
#### H3 Resolution Planner (Crop Only)
**Display Name:** `APNext H3 Resolution Planner (Crop Only) - by gabbo`
> Original node and algorithm by **gabbo**. Ported into this pack with the planning logic unchanged.
Plans a two-stage *generate → upscale* resolution pair and center-crops the input image to the **exact** aspect ratio of that plan, so nothing in the chain has to resample or pad. Step sizes are chosen so both stages always land on clean multiples of 32:
| Upscale | Stage 1 steps | Stage 2 steps |
|---------|---------------|---------------|
| `2x` | 32 | 64 |
| `1.5x` | 64 | 96 |
| Input | Description |
|-------|-------------|
| `image` | Source image; only its dimensions drive the plan |
| `resolution_mode` | `target_megapixels`, `max_stage1_from_input`, `max_final_from_input` |
| `stage1_megapixels` | Target stage 1 size in MP (0.05–4.00). `target_megapixels` mode only |
| `upscale_mode` | `2x` or `1.5x` |
| `max_crop_percent` | Max share of input area croppable (0–25%). The two `max_*` modes only; falls back to the least-lossy candidate if nothing fits |
**Modes**
- `target_megapixels` — hits the requested stage 1 megapixels while staying as close as possible to the input aspect ratio.
- `max_stage1_from_input` — largest stage 1 the input can feed natively within the crop budget.
- `max_final_from_input` — largest stage 2 (final) the input can feed natively within the crop budget.
**Returns:** `(cropped_image, stage1_width, stage1_height, stage2_width, stage2_height, upscale_factor, plan_info)`
`plan_info` is a human-readable summary of the chosen plan:
```
mode: target_megapixels @ 2x
input: 1920x1080
crop: 1917x1065 at (1,7) - 1.54% of area removed
aspect: 9:5
stage 1: 864x480 (0.40 MP)
stage 2: 1728x960 (1.58 MP)
```
---
### 🎥 MiniMax-H3 Prompt Nodes
Both nodes take a **short idea, an image, or both** and expand it into a complete, spec-compliant MiniMax-H3 video prompt. The official MiniMax writing guides ship verbatim in `data/h3/` and are used as the system prompt, so the model follows the real spec rather than a paraphrase — edit those files to tune behaviour globally.
Any provider works: `auto-detect` picks the first of Claude → GPT → Gemini → Grok → Groq that has an API key set. When an image is connected it is sent as vision input, so the model describes the frame itself instead of you writing the description.
#### APNext H3 Prompt Writer
**Display Name:** `APNext H3 Prompt Writer`
Writes the base format — `integrated_multimodal_description`, `overall_soundscape`, `non_diegetic_music` — per [VIDEO_PROMPT_WRITING_GUIDE_base_en.md](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md).
| Input | Description |
|-------|-------------|
| `idea` | Your short prompt or image description — the thing being expanded |
| `task_type` | `T2VA` (text only), `I2VA` (first frame), `FL2VA` (first + last), `L2VA` (last frame). Non-T2VA emits the exact reference-alignment instruction line |
| `duration_seconds` | Drives cut times and the `S.SS` value in the alignment line |
| `shot_plan` | Auto, or force 1–4 shots |
| `visual_style` | Auto, or one of the guide's styles (`Cinematic`, `live-action`, `2D-animated`, `3D CG`, `claymation`, `watercolor`, `vintage film`) |
| `wildness` | **0 = literal, 100 = fully unhinged.** See below |
| `camera_motion` / `camera_amplitude` / `camera_speed` | The guide's full camera vocabulary. Medium amplitude and normal speed are omitted from the output, as the spec requires |
| `include_dialogue` | Off ⇒ no `(Sx)` IDs and no `<d>` blocks at all |
| `dialogue_language` | Language tag written inside `<d>[...]</d>` |
| `include_on_screen_text` | Whether readable signs/banners/subtitles appear |
| `include_soundscape` / `include_non_diegetic_music` | Off writes `N/A` into that field |
| `model`, `temperature`, `seed` | Provider selection and sampling |
| `image` *(optional)* | Reference frame(s), sent as vision input |
| `extra_instructions` *(optional)* | Free-form extra direction |
**Returns:** `(h3_prompt, integrated_multimodal_description, overall_soundscape, non_diegetic_music, model_used)` — the full prompt plus each field split out for separate wiring.
---
#### APNext H3 Reference Prompt Writer
**Display Name:** `APNext H3 Reference Prompt Writer`
Writes the six-section full-reference format per [VIDEO_PROMPT_WRITING_GUIDE_ref_en.md](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md). Shares every option above, plus:
| Input | Description |
|-------|-------------|
| `task_type` | The `[bracketed]` summary prefix: `keyframe completion`, `reference generation`, `video editing`, `video continuation`, `audio reuse`, `audio reference`. Auto lets the model combine them with ` + ` |
| `reference_role` | How attached images get labelled: auto, `<Subject N>`, standalone `<Picture N>`, style-only, or storyboard |
| `word_target` | Target length of `detailed_description` (guide recommends 350–500) |
| `image_1` … `image_4` *(optional)* | Up to four reference images, in label order |
| `reference_notes` *(optional)* | Per-reference notes, one per line — also how you describe video/audio references you can't attach |
**Returns:** `(h3_prompt, subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music, model_used)`
---
#### The `wildness` slider
One dial from conservative to unhinged. Above 40 it also injects concrete surreal **events** (not mood words) drawn from a 40-entry pool — selection is driven by `seed`, so the same seed gives the same weirdness.
| Range | Band | Behaviour | Random elements |
|-------|------|-----------|-----------------|
| 0–15 | Conservative | Strictly literal, no invented events | 0 |
| 16–40 | Grounded | Believable, well-directed embellishment | 0 |
| 41–65 | Bold | Strong authorial choices, physics still holds | 1 |
| 66–85 | Wild | Surreal juxtapositions, dreamlike logic | 2 |
| 86–100 | Unhinged | Scale, gravity and continuity all negotiable | 3 |
Injected elements are filmable, e.g. *"the subject's shadow moves a beat out of sync"*, *"a doorway opens onto a completely different biome"*, *"rain falls upward into the sky"*.
---
### 🎲 Prompt Generators
#### Auto Prompter
@@ -815,21 +922,23 @@ Example workflows are available in the `examples/` directory:
## 📋 Requirements
```
Pillow>=10.4.0
requests>=2.32.5
openai>=1.44.0
blend-modes>=2.1.0
huggingface_hub>=0.34.0
color_matcher>=0.5.0
openai>=2.54.0,<3.0.0
anthropic>=0.121.0
google-genai>=2.18.0
httpx>=0.28.1
huggingface_hub[hf_xet]>=0.34.0
chardet>=5.2.0
google-generativeai>=0.7.2
anthropic
transformers>=4.40.0
decord>=0.6.0
scipy>=1.10.0
tqdm>=4.67.1
```
Anything ComfyUI already ships in its own `requirements.txt` — `Pillow`, `requests`, `transformers`, `scipy`, `tqdm`, `numpy`, `torch` — is deliberately **not** repeated, since re-pinning it only risks downgrading the base install.
Two constraints worth knowing about:
- **`openai` is capped below 3.0.** v3 switched to HTTPX2 and stopped shipping `httpx`; the GPT/Grok/Groq nodes pass an `httpx.Client` as `http_client=`, which v3 rejects.
- **Gemini uses `google-genai`, not `google-generativeai`.** The legacy SDK hard-pinned `google-ai-generativelanguage==0.6.15`, which forced `protobuf<6` and dragged grpcio into the ComfyUI environment. The current SDK needs neither.
`decord` is listed but commented out: it is unmaintained and not numpy-2 safe, and the QwenVL/MiniCPM video nodes fall back to OpenCV automatically. Uncomment it in `requirements.txt` if you specifically want decord-based frame decoding.
---
## 🔄 Model Support Matrix
+24
View File
@@ -381,6 +381,30 @@ try:
except Exception:
pass
# Resolution Planning Nodes
try:
# H3 Resolution Planner - original node and algorithm by gabbo
from .nodes.resolution.h3_resolution_planner import H3ResolutionPlannerCropOnly
NEW_MAPPINGS["H3ResolutionPlannerCropOnly"] = H3ResolutionPlannerCropOnly
NEW_DISPLAY_MAPPINGS["H3ResolutionPlannerCropOnly"] = "APNext H3 Resolution Planner (Crop Only) - by gabbo"
except Exception:
pass
# MiniMax-H3 Prompt Nodes
try:
from .nodes.h3.base_prompt_writer import H3BasePromptWriter
NEW_MAPPINGS["H3BasePromptWriter"] = H3BasePromptWriter
NEW_DISPLAY_MAPPINGS["H3BasePromptWriter"] = "APNext H3 Prompt Writer"
except Exception:
pass
try:
from .nodes.h3.ref_prompt_writer import H3RefPromptWriter
NEW_MAPPINGS["H3RefPromptWriter"] = H3RefPromptWriter
NEW_DISPLAY_MAPPINGS["H3RefPromptWriter"] = "APNext H3 Reference Prompt Writer"
except Exception:
pass
# Combine mappings (modular nodes + dynamic nodes)
NODE_CLASS_MAPPINGS = {**NEW_MAPPINGS, **DYNAMIC_MAPPINGS}
NODE_DISPLAY_NAME_MAPPINGS = {**NEW_DISPLAY_MAPPINGS, **DYNAMIC_DISPLAY_MAPPINGS}
+222
View File
@@ -0,0 +1,222 @@
# Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
## 1. Task Overview
- **T2VA**: Builds a complete audiovisual timeline from text.
- **I2VA**: T2VA body + first-frame instruction + a visual path that develops forward from the first frame.
- **FL2VA**: T2VA body + first-and-last-frame instruction + a continuous path from the first frame to the last frame.
- **L2VA**: T2VA body + last-frame instruction + a path that converges from a plausible preceding state to the last frame.
## 2. Final Prompt Structure
### 2.1 Part One Is the Instruction
**T2VA** has no image-alignment instruction and begins directly with the three core fields.
**I2VA** always uses:
```text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
```
**FL2VA** always uses:
```text
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.
```
**L2VA** always uses:
```text
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
```
Here, `N` is the index of the actual final shot, and `S.SS` is the effective video duration formatted to exactly two decimal places. The instruction must be the first line of the final prompt, followed by one blank line before the core fields.
### 2.2 Part Two Contains the Three Core Fields
```text
integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...
```
- **integrated_multimodal_description**: Describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline.
- **overall_soundscape**: Summarizes ambient sound, physical action sounds, and non-verbal human sounds across the entire video.
- **non_diegetic_music**: Describes background music that the characters cannot hear and only the audience can hear.
## 3. How to Incorporate Keyframes into the Multimodal Description
### 3.1 I2VA: Begin from the Image and Develop Forward
`<Picture 1>` is the actual first frame of the video at 0.00 seconds and belongs to `[Shot 1]`. The description should first establish the style, subjects, composition, and scene anchors in the image, then describe the next action. Character identity, clothing, colors, key objects, and spatial relationships should remain consistent.
Recommended structure: **first-frame anchor → action onset → continuous development → result or reaction**.
### 3.2 FL2VA: Describe the Path Between the First and Last Frames
Picture 1 is the opening, and Picture 2 is the ending. Focus on how the subject moves, how poses change, how objects are manipulated, how the composition evolves, and how the scene or lighting transitions.
FL2VA generally favors a single shot so the model can interpolate continuously from the first frame to the last frame. Use multiple shots only when they are explicitly specified. The last frame must be reached by the final `[Shot N]` at the end of the video.
Recommended structure: **first-frame state → observable intermediate changes → progressively narrowing differences → last-frame state**.
### 3.3 L2VA: Infer the Opening and Land on the Image at the End
`<Picture 1>` is the final frame of the video and belongs to the last `[Shot N]`; it does not inherently belong to Shot 1. Infer a plausible earlier state from the user's intent and the last frame, then describe how the characters, objects, camera, and scene gradually approach the reference image.
Recommended structure: **plausible preceding state → explicit action and transition path → gradual convergence in the final shot → last-frame landing**.
## 4. How to Write the Three Shared Core Sections
### 4.1 Develop the Multimodal Description Along the Timeline
`integrated_multimodal_description` is the main body of the rewritten prompt. Every detail should correspond to something visible or audible: visual style, initial composition, subject appearance and position, scene and key props, actions and reactions, shot changes, spoken language, and synchronized diegetic sound.
At the beginning of `[Shot 1]`, state the overall style and initial composition. Common styles include `Cinematic`, `live-action`, `2D-animated`, `3D CG`, `claymation`, `watercolor`, and `vintage film`. For keyframe tasks, derive the style from the reference image; for T2VA, select it from the user's text.
```text
[Shot 1] Live-action, cinematic, a medium-wide shot frames...
```
### 4.2 Shots and Cuts
Do not add a timestamp to the first shot. Use sequential shot numbers for later shots, and begin each one with a strictly increasing cut time that falls within the video duration:
```text
[Shot 2] At 00:03.500, the camera cuts to...
```
For ordinary cuts, use `the camera cuts to`, `the shot cuts to`, `the shot transitions to`, `the shot changes to`, or `the shot switches to`. When explicitly requested by the user, cross-dissolve, fade, or wipe may also be used. A cut should introduce new information about the subject, space, state, viewpoint, or time. If only the distance or a slight angle needs to change, prefer camera motion.
### 4.3 Camera Motion: Motion Type + Amplitude + Speed
A complete camera-motion expression has three dimensions: the **motion type** defines how the camera moves, **amplitude** defines the range of compositional change, and **speed** defines the pacing of that change. Add amplitude and speed only when they are meaningful; medium amplitude and normal speed are usually omitted.
| Dimension | Available Expression | Description |
|-|-|-|
| Motion type | `Zoom In / Zoom Out` | The focal length changes while the camera body remains stationary |
| Motion type | `Push In / Pull Out` | The camera moves forward / backward |
| Motion type | `Pan Left / Pan Right` | The camera remains in place while the lens pivots horizontally |
| Motion type | `Truck Left / Truck Right` | The camera translates horizontally |
| Motion type | `Tilt Up / Tilt Down` | The camera remains in place while the lens pivots vertically |
| Motion type | `Pedestal Up / Pedestal Down` | The entire camera moves upward / downward |
| Motion type | `Arc Shot` | The camera moves in an arc around the subject |
| Motion type | `Tracking Shot` | The camera follows a moving subject |
| Motion type | `Static Shot` | The camera position and lens remain still |
| Motion type | `Shake Slightly / Shake Strongly` | Slight / strong camera shake |
| Motion type | `POV` | The subject's point of view |
| Motion type | `Roll Clockwise / Roll Counterclockwise` | The camera rolls clockwise / counterclockwise around the lens axis |
| Amplitude | `with small amplitude` | Small-range change |
| Amplitude | `with large amplitude` | Large-range change |
| Speed | `at slow speed` | Slow movement |
| Speed | `at fast speed` | Fast movement |
Camera motion should be written as a natural English action within the shot, rather than stacked as separate labels at the end of a sentence:
```text
The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
The camera pans right with large amplitude at fast speed, revealing the open doorway.
The camera holds a static shot as the runner exits the frame.
```
### 4.4 Speakers, Dialogue, and Singing
Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as `(S1)` and `(S2)`. When multiple already-numbered speakers speak or sing together, use a compound ID such as `(S1,S2)`. A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID.
When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent. Place the speaker's identifying phrase, ID, action, and delivery outside `<d>`. Inside `<d>`, include only the language tag and the actual user-provided spoken content. Preserve every original word and punctuation mark verbatim; do not translate or rewrite them.
```text
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>
```
For voiceover, use the exact phrase `says in an off-screen voiceover`. Immediately after every voiceover `<d>` block, state that the corresponding on-screen character's lips remain closed:
```text
The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.
```
When the same line of dialogue or lyrics crosses a cut, use `<scenetrans>` at the connecting points in both parts and explicitly state that the audio continues across the cut. Use `<cutoff>` when speech is truncated by the end of the video. Continuity may be expressed with `continues seamlessly across the cut`, `continues uninterrupted into the next shot`, `carries over from the previous shot`, or `remains audible across the transition`.
### 4.5 On-Screen Text
Place any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks. Preserve the original text and punctuation verbatim, without translation.
```text
A red neon sign reading "营业中" glows above the doorway.
```
### 4.6 overall_soundscape
Use 1–4 English sentences in one continuous paragraph to summarize the ambient sound, physical action sounds, and non-verbal human sounds across the full video, such as wind, rain, traffic, footsteps, fabric movement, impacts, breathing, laughter, or panting. Dialogue, singing, and diegetic music already belong in the multimodal description and should not be repeated here. Use `N/A` only when the user explicitly requests complete silence throughout the video.
```text
overall_soundscape: Steady rain taps against the café windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.
```
### 4.7 non_diegetic_music
Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear. Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score. Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description. Use `N/A` when there is no non-diegetic music.
```text
non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.
```
## 5. Cases
### Case 1: T2VA
With no reference image, construct the complete timeline directly from the text. You may add scene, character, action, and sound details that remain consistent with the user's intent.
```text
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.
overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.
non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.
```
### Case 2: I2VA
Write the first-frame instruction first, then use the subject, composition, and scene in Picture 1 as the starting point of Shot 1 before describing how the scene continues to develop.
```text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.
overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.
non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.
```
### Case 3: FL2VA
The two images anchor the opening and ending respectively. The body should not repeat two static image descriptions; instead, it should supply the motion path that connects them. The following example is an eight-second single shot.
```text
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.
overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.
non_diegetic_music: N/A
```
### Case 4: L2VA
The image anchors only the final moment. First establish a compatible earlier state, then let the actions, object states, and composition gradually land on Picture 1 in the final shot. The following example is a six-second single shot.
```text
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>.
overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.
non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.
```
+341
View File
@@ -0,0 +1,341 @@
# Full-Reference Mode Rewrite Output Format Guide
This guide explains how rewrite outputs are organized and written in full-reference mode.
Write all six rewrite sections in English. Preserve the original language only for dialogue and lyrics inside `<d>` and for text visibly present in the scene.
**Description detail:** Make `detailed_description` as detailed and explicit as possible. For each shot, clearly establish the current composition, subject appearance and position, environment and lighting, actions and state changes, camera movement, current sound, and the points where referenced content actually appears or takes effect. Avoid reducing the description to a plot summary or a list of reference relationships.
> The basic formats for shots, camera movement, speakers, dialogue, and ordinary sound are shared with the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA). This guide focuses on the reference labels, analysis sections, and format differences specific to full-reference mode.
## 1. Overall Structure
A complete rewrite output consists of six sections in the following order:
| Section | Purpose |
| --- | --- |
| `subject_definitions` | Defines referenced content and its reference labels |
| `summary` | Summarizes the task type, target video, and main reference relationships |
| `retention_analysis` | Describes how referenced content is preserved, transferred, or reused |
| `detailed_description` | Describes visuals, actions, shots, sound, and dialogue in playback order |
| `overall_soundscape` | Summarizes ambience and physical sounds |
| `non_diegetic_music` | Describes background music audible only to the audience |
## 2. Reference Labels and Definitions (`subject_definitions`)
Full-reference rewrites use four types of labels to identify the source and role of referenced content:
| Label | Meaning |
| --- | --- |
| `<Subject N>` | Visible content abstracted from reference assets that can be reused or modified in the target video |
| `<Picture N>` | A reference image used as a concrete target frame or shot-planning anchor |
| `<Video N>` | A reference video that provides an editing source, continuation starting point, or whole-video temporal structure |
| `<Audio N>` | An audio signal that is copied or referenced |
> Once a reference label is assigned to a piece of content, it keeps the same meaning across `subject_definitions`, `summary`, `retention_analysis`, `detailed_description`, and the audio sections.
`subject_definitions` defines each piece of referenced content that must be tracked separately later, such as a person, an environment, a source video's structure, or an audio track. Give each item its own line and explain what its label denotes, its reference role, and the main features to follow; name the corresponding source asset when its provenance needs to be made explicit. If `<Picture N>` or `<Video N>` only identifies the source of another referenced item and will not be analyzed or used separately later, cite it inside that item's definition without adding a separate line. `retention_analysis` records where each referenced item appears and whether it is fully preserved, partially preserved, transferred, or reused.
### 2.1 `<Subject N>`
`<Subject N>` is used for reusable visible content, including:
- People, animals, or objects
- Scenes, backgrounds, or environments
- Clothing, props, interfaces, or visual effects
- Styles, actions, expressions, or poses
It represents a content unit that will actually be used in the target video, rather than the source file itself. One subject may be defined by multiple reference assets, and one reference asset may provide multiple subjects.
```text
<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.
```
When the same subject comes from multiple assets, combine the sources and state what each asset provides:
```text
<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.
```
### 2.2 `<Picture N>`
Use a standalone `<Picture N>` when the reference image itself serves as a shot's first frame, keyframe, last frame, edited keyframe, or composition anchor:
```text
<Picture 2> is the first frame of [Shot 1], showing a woman seated beside a café window.
```
If an image is used only to define a character, scene, costume, or style, do not create a standalone picture entry. Instead, cite the image source inside the corresponding `<Subject N>` definition.
When an image acts as a storyboard or shot-planning reference, state which shots it maps to and what planning information it provides:
```text
<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.
```
### 2.3 `<Video N>`
`<Video N>` is reserved for whole-video relationships, such as:
- Editing an original video
- Continuing from the end of an original video
- Referencing the original video's camera movement, cuts, rhythm, or temporal structure
```text
<Video 1> is the source video for the target video edit.
```
If a person, object, scene, action, or effect from a reference video is reused as visible content, it still belongs under `<Subject N>`. `<Video N>` identifies the asset or structural source and does not replace subject labels.
### 2.4 `<Audio N>`
`<Audio N>` represents a standalone audio asset or an enabled synchronized audio track from a reference video. Common uses include:
- Copying all or part of an audio signal
- Referencing a background-music style
- Referencing a speaker's voice timbre and delivery
- Using dialogue, lyrics, or sound effects from the original audio
- Referencing beat, rhythm, or audio continuity
When an `<Audio N>` explicitly corresponds to a target speaker, reuse that speaker's global ID in the definition: write `<Subject N> (Sx)` when the speaker maps to a defined subject, or use a stable voice description followed by `(Sx)` otherwise. The ID comes from the target video's global speaker order and is not independently assigned or renumbered in the audio definition. See Section 5.4 for the speaker-numbering rules:
```text
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
```
When one audio asset serves multiple roles, describe those roles in one natural sentence rather than creating additional subsections.
### 2.5 Visual and Audio Tracks from the Same Reference Video
`<Video N>` and `<Audio N>` are numbered independently. Each index indicates only the label's order within its own category and does not encode a pairing between the two categories. The same reference video may therefore correspond to `<Video 1>` and `<Audio 2>`; different indices do not prevent them from coming from the same source asset.
An ordinary reference video does not create `<Audio N>` merely because the file contains sound.
An `<Audio N>` definition primarily states the audio's role and does not have to name the `<Video N>` it comes from. State the shared source only when needed to remove provenance ambiguity, for example:
```text
<Video 1> is the source video for the target video edit.
<Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.
```
## 3. `summary`
This section uses one short English paragraph to summarize the target video and its reference relationships. It begins with a square-bracketed task-type prefix:
```text
[reference generation] ...
[video editing + reference generation + audio reuse] ...
```
Choose task types according to the actual role each reference asset plays in the target video:
| Task type | When to use it |
| --- | --- |
| `keyframe completion` | An image serves as the target video's first frame, keyframe, last frame, edited keyframe, or another concrete frame anchor |
| `reference generation` | An image, video, or audio asset provides generation guidance for a character, scene, style, action, camera movement, storyboard, and so on, without serving as a concrete frame or as the source video being edited or continued |
| `video editing` | An existing source video is directly modified; editing an image or generating between still keyframes does not belong to this type |
| `video continuation` | New content continues, extends, resumes, or transitions from an existing source video |
| `audio reuse` | The same audio signal is reused in full or in part |
| `audio reference` | The audio signal is not copied directly; only its music style, timbre, dialogue or lyric content, sound-effect texture, beat, or continuity is referenced |
When a task satisfies multiple relationships, combine the task types with ` + ` and do not repeat a type. For example, continuing from a source video while using an image as the last frame is written as `[video continuation + keyframe completion]`. Editing a source video while retaining its original audio may be written as `[video editing + audio reuse]`.
The mere presence of video or audio does not automatically create a corresponding task type. If a reference video provides only camera movement, cuts, or rhythm, it normally belongs to `reference generation`. Use `video editing` or `video continuation` only when that video is directly edited or continued.
When editing a source video, use `audio reuse` as well if its original audio remains audible. When continuing a source video without directly copying the audio signal, use `audio reference` if the new audio only continues the original track's audible characteristics.
The summary uses the previously defined `<Subject N>`, `<Picture N>`, `<Video N>`, and `<Audio N>` labels to describe the main subjects, shot flow, and roles of the reference assets. Do not introduce new reference labels in this section.
For video-editing tasks, begin the summary after the task-type prefix with:
```text
The target video is an edited version of <Video 1>.
```
## 4. `retention_analysis`
This section describes how each piece of referenced content is preserved, transferred, copied, or referenced in the target video. Use one line for each reference label and preserve the meaning established in `subject_definitions`.
### 4.1 Visible Content
`<Subject N>`, `<Picture N>`, and `<Video N>` use the following relationship markers. These markers are fixed English values in the output format:
| Relationship marker | Meaning |
| --- | --- |
| `fully_preserved` | The defined role of the referenced content is fully preserved |
| `partially_preserved` | The referenced content is still used, but some defined characteristics are changed or only partially retained |
| `attribute_transfer` | Referenced characteristics are transferred to a different identifiable target subject |
| `weak_reference` | Only broad similarity in style, category, composition, or atmosphere is retained |
Subject entry:
```text
<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...
```
Picture entry:
```text
<Picture 2> ([Shot 1] first frame): fully_preserved - ...
```
Video-structure entry:
```text
<Video 1> (cut and pacing structure): weak_reference - ...
```
### 4.2 Audio
`<Audio N>` uses the following relationship markers:
| Relationship marker | Meaning |
| --- | --- |
| `fully_copy` | The complete source audio serves as the target video's complete final audio track |
| `partially_copy` | Only part of the timeline or selected audio layers are copied, or other sounds are added, removed, or replaced after copying |
| `reference` | The signal is not copied directly; only timbre, rhythm, music style, dialogue content, or sound texture is referenced |
| `weak_reference` | Only broad similarity in category or atmosphere is retained |
```text
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
```
```text
<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.
```
Choose each relationship marker only within the reference role already defined for that label in `subject_definitions`. Do not treat newly added actions, backgrounds, or plot events in the target video as losses of reference fidelity.
## 5. `detailed_description`
This is the main body of a full-reference rewrite. It describes visuals, actions, sound, and dialogue shot by shot in target-video playback order and inserts reference labels where they apply.
### 5.1 Basic Format
The basic format follows the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA):
- Write the body in English. Preserve the original language of dialogue, lyrics, and visible text.
- `[Shot 1]` marks the opening shot and has no timestamp. Later shots use `[Shot N] At MM:SS.mmm, ...` to mark cut times.
- Write camera movement as natural English within the current shot, including movement type, amplitude, and speed when they need to be expressed.
- Give vocal sources stable `(S1)`, `(S2)`, and subsequent IDs. Write dialogue and lyrics as `<d>[Language] ...</d>`.
- Use `<scenetrans>`, `<cutoff>`, and the corresponding continuity descriptions for dialogue crossing a cut, speech truncated by the video ending, and continuous audio across shots.
For complete rules and examples covering camera vocabulary, group speech, voice-over, dialogue across cuts, and visible text, see the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).
### 5.2 Full-Reference Mode Differences
| Dimension | T2VA | Full-reference mode |
| --- | --- | --- |
| Main field | `integrated_multimodal_description` | `detailed_description` |
| Style opening | Written after `[Shot 1]` | Established in one or two English sentences before `[Shot 1]` |
| Reference information | Does not use full-reference labels | Inserts `<Subject N>`, `<Picture N>`, `<Video N>`, and `<Audio N>` at their first appearance and where their roles apply |
| Audio relationships | Describes the target video's own sound | Cites `<Audio N>` in the corresponding shot or audio phase and states whether the signal is copied or referenced |
Opening example:
```text
The target video is in a cinematic, literary music-video style with soft lighting and a slightly desaturated color palette.
[Shot 1] The scene opens in a crowded urban street...
[Shot 2] At 00:09.000, the shot cuts to an extreme close-up...
```
For generation tasks, `detailed_description` is normally 350-500 English words. Dialogue-dense content prioritizes fitting the complete spoken timeline rather than mechanically reaching a word count. Video-editing descriptions scale with the complexity of the source video and do not have to follow the generation-task range. A single shot does not automatically justify a shorter description; distribute detail across multiple shots according to their information load.
### 5.3 Using Reference Labels in Shots
At the first clear appearance of an important `<Subject N>`, describe its referenced characteristics, position in the frame, and current action within what is actually visible in the shot. Continue using the same label in later shots without redefining what the label represents.
Use natural phrasing for concrete frame anchors:
```text
the shot begins from <Picture 1>
the shot's keyframe corresponds to <Picture 2>
the shot ends on <Picture 3>
```
When editing or continuing an original video, cite `<Video N>` naturally where its source state, structure, or continuation relationship applies. Cite `<Audio N>` in the shot or semantic phase where the audio relationship is active.
### 5.4 Speakers, Audio Sources, and Dialogue
The basic speaker-ID and `<d>` formats follow T2VA. When a referenced subject physically speaks, retain both the visual reference label and the speaker ID:
```text
<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>
```
`<Subject N>` identifies the referenced subject, while `(Sx)` identifies the actual speaker. When the subject speaks, write `<Subject N> (Sx)`. If the same subject speaks off-screen, keep the same form and mark it as `off-screen`. When the speaker does not correspond to a defined subject, use a stable voice description followed by `(Sx)`.
When verbal content is only a cue within a directly reused BGM or complete soundtrack, and no person, character, narrator, or other independent vocal source physically produces it, use `<Audio N>` as the audible source and do not invent an additional `(Sx)`. If a concrete person, character, narrator, or other independent vocal source produces the voice, assign and reuse `(Sx)` for that source:
```text
When <Audio 1> reaches the phrase <d>[English] I'm lonely lonely lonely lonely lonely I'm lonely</d>, <Subject 1> performs the corresponding hand gesture without becoming a separate speaker source.
```
When dialogue, narration, or lyrics from reference audio are directly reused, or when the input prompt explicitly requests their reperformance, preserve the exact source words and original language inside `<d>`. Write `[unclear]` for unintelligible spans instead of guessing or paraphrasing them. Standardize punctuation to the basic written marks needed to express the sentence, such as `,`, `.`, `?`, and `!`; remove repeated tildes, emoji, bullets, and repeated or decorative punctuation. End complete statements, questions, and exclamations with `.`, `?`, or `!` respectively before `</d>`.
When only timbre, rhythm, emotion, or delivery is referenced, do not carry the original dialogue from the reference audio into the target video.
Assign `(Sx)` once according to the order of actual vocal events in the target video. Reuse the corresponding ID at every actual vocal event in `detailed_description`; an `<Audio N>` definition bound to a target speaker in `subject_definitions` also reuses the same `(Sx)` but never assigns a new one independently. Do not write `(Sx)` in `retention_analysis`. Verbal cues that exist only within a directly reused BGM or complete soundtrack use `<Audio N>`; voices physically produced by a concrete person, character, narrator, or other independent vocal source use `(Sx)`.
## 6. `overall_soundscape` and `non_diegetic_music`
The definitions of these two sound categories follow the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).
`overall_soundscape` summarizes ambience and physical sounds across the full video. Dialogue, singing, and sound events synchronized to a particular shot remain in `detailed_description`:
```text
overall_soundscape: Quiet indoor room tone and a low ventilation hum continue throughout the video.
```
`non_diegetic_music` describes background music that the characters cannot hear and that is audible only to the audience. When music is present, state its instrumentation, tempo, and dynamic development:
```text
non_diegetic_music: A restrained solo-piano score at a slow tempo, with sustained low cello underneath and no swell.
```
When reference audio is used, state its copy or reference relationship only in the section that matches the audible layer: ambience and sound effects belong in `overall_soundscape`, while audience-only score belongs in `non_diegetic_music`. If the same audio provides both kinds of content, describe the corresponding relationship in each section:
```text
overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.
non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.
```
Write complete dialogue and lyrics only inside `<d>` in `detailed_description`; do not repeat them in these two sections.
## 7. Complete Example
<details>
<summary>Show the complete example</summary>
```text
subject_definitions:
<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.
<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.
<Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.
<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
The target video uses a realistic multi-camera sitcom style with warm indoor lighting.
[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.
overall_soundscape:
Soft indoor coffee-shop room tone continues throughout the scene.
non_diegetic_music:
N/A
```
</details>
+5 -24
View File
@@ -6,9 +6,9 @@ import random
import torch
import numpy as np
from PIL import Image
import google.generativeai as genai
from ...utils.constants import CUSTOM_CATEGORY, gemini_models
from ...utils.gemini_client import get_gemini_client, gemini_generate
from ...utils.image_utils import tensor2pil, pil2tensor
@@ -17,7 +17,7 @@ class GeminiCustomVision:
self.gemini_api_key = os.environ.get("GEMINI_API_KEY")
if not self.gemini_api_key:
raise ValueError("GEMINI_API_KEY environment variable is not set")
genai.configure(api_key=self.gemini_api_key)
self.client = get_gemini_client(self.gemini_api_key)
@classmethod
def INPUT_TYPES(s):
@@ -140,28 +140,9 @@ class GeminiCustomVision:
combined_image = self.fade_images(pil_images, fade_percentage)
safety_settings = [
{
"category": "HARM_CATEGORY_HARASSMENT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_HATE_SPEECH",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_DANGEROUS_CONTENT",
"threshold": "BLOCK_NONE",
},
]
model = genai.GenerativeModel(gemini_model, safety_settings=safety_settings)
response = model.generate_content([full_prompt, combined_image])
response = gemini_generate(
self.client, gemini_model, [full_prompt, combined_image]
)
result = response.text
+13 -27
View File
@@ -6,9 +6,13 @@ import random
import torch
import numpy as np
from PIL import Image
import google.generativeai as genai
from ...utils.constants import CUSTOM_CATEGORY, gemini_models
from ...utils.gemini_client import (
get_gemini_client,
gemini_generate,
gemini_finished_normally,
)
class GeminiNextScene:
@@ -26,8 +30,8 @@ class GeminiNextScene:
self.gemini_api_key = os.environ.get("GEMINI_API_KEY")
if not self.gemini_api_key:
raise ValueError("GEMINI_API_KEY environment variable is not set")
genai.configure(api_key=self.gemini_api_key)
self.client = get_gemini_client(self.gemini_api_key)
# Load the custom prompt template
prompt_file = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(__file__))), "data", "custom_prompts", "next_scene.txt")
try:
@@ -188,30 +192,12 @@ CRITICAL: OUTPUT ONLY THE FINISHED "NEXT SCENE:" PROMPT - NO explanations or pre
}
full_prompt += f"\n\nINTENSITY: {intensity_map[transition_intensity]}"
# Configure safety settings to allow creative content
safety_settings = [
{
"category": "HARM_CATEGORY_HARASSMENT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_HATE_SPEECH",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_DANGEROUS_CONTENT",
"threshold": "BLOCK_NONE",
},
]
# Safety filters are set to BLOCK_NONE by the shared helper so
# creative content is allowed through.
response = gemini_generate(
self.client, gemini_model, [full_prompt, pil_image]
)
model = genai.GenerativeModel(gemini_model, safety_settings=safety_settings)
response = model.generate_content([full_prompt, pil_image])
# Check if response was blocked
if not response.candidates:
print("⚠️ WARNING: Gemini returned no candidates!")
@@ -219,7 +205,7 @@ CRITICAL: OUTPUT ONLY THE FINISHED "NEXT SCENE:" PROMPT - NO explanations or pre
fallback_text = "The camera pulls back to reveal more of the surrounding environment, as lighting shifts to create a different mood and atmosphere."
result = f"{scene_prefix_text} {fallback_text}" if add_scene_prefix else fallback_text
short_description = result
elif response.candidates[0].finish_reason != 1: # 1 = STOP (normal completion)
elif not gemini_finished_normally(response.candidates[0]):
print(f"⚠️ WARNING: Response did not complete normally. Finish reason: {response.candidates[0].finish_reason}")
print("Checking safety ratings...")
+5 -44
View File
@@ -2,16 +2,17 @@
import random
import os
import google.generativeai as genai
from ...utils.constants import CUSTOM_CATEGORY, CINEMATIC_TERMS, gemini_models
from ...utils.gemini_client import get_gemini_client, gemini_generate
class GeminiPromptEnhancer:
def __init__(self):
self.gemini_api_key = os.environ.get("GEMINI_API_KEY")
self.client = None
if self.gemini_api_key:
genai.configure(api_key=self.gemini_api_key)
self.client = get_gemini_client(self.gemini_api_key)
@classmethod
def INPUT_TYPES(s):
@@ -193,27 +194,7 @@ class GeminiPromptEnhancer:
full_prompt = f"{system_prompt}\n\nOriginal prompt: {prompt}\n\nEnhanced prompt:"
safety_settings = [
{
"category": "HARM_CATEGORY_HARASSMENT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_HATE_SPEECH",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_DANGEROUS_CONTENT",
"threshold": "BLOCK_NONE",
},
]
model = genai.GenerativeModel(gemini_model, safety_settings=safety_settings)
response = model.generate_content(full_prompt)
response = gemini_generate(self.client, gemini_model, full_prompt)
enhanced_result = response.text.strip()
@@ -252,27 +233,7 @@ Incorporate these elements naturally into your enhancement where relevant."""
full_prompt = f"{system_prompt}\n\nOriginal prompt: {prompt}\n\nEnhanced prompt:"
safety_settings = [
{
"category": "HARM_CATEGORY_HARASSMENT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_HATE_SPEECH",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_DANGEROUS_CONTENT",
"threshold": "BLOCK_NONE",
},
]
model = genai.GenerativeModel(gemini_model, safety_settings=safety_settings)
response = model.generate_content(full_prompt)
response = gemini_generate(self.client, gemini_model, full_prompt)
enhanced_result = response.text.strip()
+3 -24
View File
@@ -3,9 +3,9 @@
import os
import re
import random
import google.generativeai as genai
from ...utils.constants import CUSTOM_CATEGORY, gemini_models
from ...utils.gemini_client import get_gemini_client, gemini_generate
class GeminiTextOnly:
@@ -13,7 +13,7 @@ class GeminiTextOnly:
self.gemini_api_key = os.environ.get("GEMINI_API_KEY")
if not self.gemini_api_key:
raise ValueError("GEMINI_API_KEY environment variable is not set")
genai.configure(api_key=self.gemini_api_key)
self.client = get_gemini_client(self.gemini_api_key)
@classmethod
def INPUT_TYPES(s):
@@ -80,28 +80,7 @@ class GeminiTextOnly:
full_prompt = f"{additive_prompt} {custom_prompt}".strip() if additive_prompt else custom_prompt
safety_settings = [
{
"category": "HARM_CATEGORY_HARASSMENT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_HATE_SPEECH",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_DANGEROUS_CONTENT",
"threshold": "BLOCK_NONE",
},
]
model = genai.GenerativeModel(gemini_model, safety_settings=safety_settings)
response = model.generate_content(full_prompt)
response = gemini_generate(self.client, gemini_model, full_prompt)
result = response.text
+14
View File
@@ -0,0 +1,14 @@
# APNext MiniMax-H3 Nodes
from .base_prompt_writer import H3BasePromptWriter
from .ref_prompt_writer import H3RefPromptWriter
NODE_CLASS_MAPPINGS = {
"H3BasePromptWriter": H3BasePromptWriter,
"H3RefPromptWriter": H3RefPromptWriter,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"H3BasePromptWriter": "APNext H3 Prompt Writer",
"H3RefPromptWriter": "APNext H3 Reference Prompt Writer",
}
+353
View File
@@ -0,0 +1,353 @@
# APNext H3 Prompt Writer
#
# Turns a short idea (and optionally a reference image) into a complete
# MiniMax-H3 video prompt in the T2VA / I2VA / FL2VA / L2VA format, using the
# official writing guide as the system prompt.
import random
from ...utils.constants import CUSTOM_CATEGORY
from ...utils.image_utils import tensor2pil
from ...utils.llm_router import AUTO_DETECT, call_llm, list_all_models
from .common import (
AUTO,
CAMERA_AMPLITUDES,
CAMERA_MOTIONS,
CAMERA_SPEEDS,
SHOT_PLANS,
VISUAL_STYLES,
camera_directive,
extract_section,
load_guide,
shot_directive,
strip_code_fence,
toggle_directives,
wildness_directive,
)
TASK_TYPES = [
"T2VA (text only)",
"I2VA (first frame)",
"FL2VA (first + last frame)",
"L2VA (last frame)",
]
_FIELDS = (
"integrated_multimodal_description",
"overall_soundscape",
"non_diegetic_music",
)
class H3BasePromptWriter:
"""
APNext H3 Prompt Writer
Writes a MiniMax-H3 video prompt from a short idea or an image description,
following the official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).
Attach images to describe them directly instead of writing the description
yourself; leave them unconnected for pure text-to-video.
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"idea": ("STRING", {
"multiline": True,
"default": "",
"tooltip": "Your short prompt or image description. This is what gets expanded into a full H3 prompt.",
}),
"task_type": (TASK_TYPES, {
"default": "T2VA (text only)",
"tooltip": "Which H3 task the prompt targets. Anything other than T2VA emits the matching reference-alignment instruction line.",
}),
"duration_seconds": ("FLOAT", {
"default": 6.0, "min": 1.0, "max": 60.0, "step": 0.5,
"tooltip": "Effective video duration. Drives the cut times and the S.SS value in the alignment instruction.",
}),
"shot_plan": (SHOT_PLANS, {"default": AUTO}),
"visual_style": (VISUAL_STYLES, {
"default": AUTO,
"tooltip": "Style stated at the start of [Shot 1]. Auto derives it from the idea or the attached image.",
}),
"wildness": ("INT", {
"default": 25, "min": 0, "max": 100, "step": 1,
"tooltip": "0 = literal and conservative, 100 = fully unhinged. Above 40 the node also injects concrete surreal elements picked from the seed.",
}),
"camera_motion": (CAMERA_MOTIONS, {
"default": AUTO,
"tooltip": "Primary camera movement, using the guide's vocabulary.",
}),
"camera_amplitude": (CAMERA_AMPLITUDES, {"default": AUTO}),
"camera_speed": (CAMERA_SPEEDS, {"default": AUTO}),
"include_dialogue": ("BOOLEAN", {
"default": True,
"tooltip": "Off means no (Sx) speaker IDs and no <d> blocks at all.",
}),
"dialogue_language": ("STRING", {
"default": "English",
"tooltip": "Language tag written inside <d>[...]</d>.",
}),
"include_on_screen_text": ("BOOLEAN", {"default": False}),
"include_soundscape": ("BOOLEAN", {
"default": True,
"tooltip": "Off writes N/A into overall_soundscape.",
}),
"include_non_diegetic_music": ("BOOLEAN", {
"default": True,
"tooltip": "Off writes N/A into non_diegetic_music.",
}),
"model": (list_all_models(), {
"default": AUTO_DETECT,
"tooltip": "Which LLM writes the prompt. auto-detect picks the first provider with an API key set.",
}),
"temperature": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 2.0, "step": 0.05}),
"seed": ("INT", {"default": -1, "min": -1, "max": 0xffffffffffffffff}),
},
"optional": {
"image": ("IMAGE", {
"tooltip": "Reference frame(s). Sent to the model as vision input so it can describe them itself.",
}),
"extra_instructions": ("STRING", {
"multiline": True,
"default": "",
"tooltip": "Free-form extra direction appended to the request.",
}),
},
}
RETURN_TYPES = ("STRING", "STRING", "STRING", "STRING", "STRING")
RETURN_NAMES = (
"h3_prompt",
"integrated_multimodal_description",
"overall_soundscape",
"non_diegetic_music",
"model_used",
)
FUNCTION = "write"
CATEGORY = f"{CUSTOM_CATEGORY}/H3"
DESCRIPTION = (
"Expands a short idea or image description into a full MiniMax-H3 video prompt "
"(T2VA / I2VA / FL2VA / L2VA) using the official writing guide."
)
def _instruction_rule(self, task_type, duration_seconds):
duration = f"{duration_seconds:.2f}"
if task_type.startswith("T2VA"):
return (
"Task: T2VA. There is no image-alignment instruction. Begin the output "
"directly with `integrated_multimodal_description:`."
)
if task_type.startswith("I2VA"):
return (
"Task: I2VA. The first line of the output must be exactly:\n"
"For the target video, at 0.00 seconds into the target video, <Picture 1> "
"(from [Shot 1]) is fully referenced.\n"
"Follow it with one blank line, then the core fields. <Picture 1> is the "
"real first frame at 0.00s and belongs to [Shot 1]: anchor its style, "
"subjects, composition and scene, then develop forward "
"(anchor -> action onset -> continuous development -> result)."
)
if task_type.startswith("FL2VA"):
return (
"Task: FL2VA. The first line of the output must be exactly:\n"
f"How the reference pictures align with the target video - Picture 1 (from "
f"Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 "
f"(from Shot N) aligns with the {duration}-second mark of the target video.\n"
"Replace N with the index of the actual final shot. Follow it with one blank "
"line, then the core fields. Favour a single shot so the model can interpolate, "
"and supply the motion path between the two frames rather than two static "
"descriptions (first-frame state -> intermediate changes -> narrowing "
"differences -> last-frame state)."
)
return (
"Task: L2VA. The first line of the output must be exactly:\n"
f"How the reference pictures align with the target video - <Picture 1> (from "
f"[Shot N]) aligns with the {duration}-second mark of the target video.\n"
"Replace N with the index of the actual final shot. Follow it with one blank "
"line, then the core fields. <Picture 1> is the final frame and belongs to the "
"last shot, not Shot 1: infer a plausible earlier state and converge onto the "
"image (preceding state -> transition path -> gradual convergence -> landing)."
)
def _build_system_prompt(self):
guide = load_guide("guide_base_en.md")
return (
"You are a MiniMax-H3 video prompt engineer. You rewrite a user's short idea "
"into a complete, production-ready H3 prompt.\n\n"
"The authoritative specification follows. Obey it exactly - field names, shot "
"labels, timestamp formats, camera vocabulary, speaker IDs and <d> blocks all "
"follow this document.\n\n"
"=== BEGIN MINIMAX-H3 VIDEO PROMPT WRITING GUIDE ===\n"
f"{guide}\n"
"=== END MINIMAX-H3 VIDEO PROMPT WRITING GUIDE ===\n\n"
"Output rules:\n"
"- Emit only the finished prompt. No preamble, no commentary, no markdown "
"fences, no headings of your own.\n"
"- Keep the exact field names `integrated_multimodal_description:`, "
"`overall_soundscape:` and `non_diegetic_music:`, each separated by one blank line.\n"
"- Write everything in English except dialogue, lyrics and on-screen text, which "
"keep their original language.\n"
"- Never invent reference labels for a task type that does not use them."
)
def _build_user_prompt(
self,
idea,
task_type,
duration_seconds,
shot_plan,
visual_style,
wildness,
camera_motion,
camera_amplitude,
camera_speed,
include_dialogue,
dialogue_language,
include_on_screen_text,
include_soundscape,
include_non_diegetic_music,
extra_instructions,
has_image,
rng,
):
directives = [self._instruction_rule(task_type, duration_seconds)]
directives.append(shot_directive(shot_plan, duration_seconds))
if visual_style == AUTO:
directives.append(
"Visual style: choose one that fits the idea"
+ (" and the attached image" if has_image else "")
+ ", and state it at the start of [Shot 1]."
)
else:
directives.append(
f"Visual style: open [Shot 1] with `{visual_style}` as the stated style."
)
directives.append(camera_directive(camera_motion, camera_amplitude, camera_speed))
directives.extend(
toggle_directives(
include_dialogue,
include_on_screen_text,
include_soundscape,
include_non_diegetic_music,
dialogue_language,
)
)
wild_lines, wild_label = wildness_directive(wildness, rng)
directives.extend(wild_lines)
if extra_instructions.strip():
directives.append(f"Additional direction from the user: {extra_instructions.strip()}")
numbered = "\n".join(f"{i}. {line}" for i, line in enumerate(directives, 1))
if has_image:
source = (
"The attached image(s) are the reference frames. Read them directly: derive "
"style, subjects, clothing, colours, key objects and spatial relationships "
"from what you actually see, and keep them consistent."
)
if idea.strip():
source += f"\n\nThe user also wrote:\n{idea.strip()}"
else:
source = f"The user's idea:\n{idea.strip()}"
return (
f"{source}\n\n"
f"Write the H3 prompt under these constraints:\n{numbered}\n\n"
"Return only the finished prompt."
), wild_label
def write(
self,
idea,
task_type,
duration_seconds,
shot_plan,
visual_style,
wildness,
camera_motion,
camera_amplitude,
camera_speed,
include_dialogue,
dialogue_language,
include_on_screen_text,
include_soundscape,
include_non_diegetic_music,
model,
temperature,
seed,
image=None,
extra_instructions="",
):
try:
if not idea.strip() and image is None:
raise ValueError("Provide an idea, an image, or both.")
current_seed = seed if seed != -1 else random.randint(0, 0xffffffffffffffff)
rng = random.Random(current_seed)
images = None
if image is not None:
images = [tensor2pil(frame) for frame in image]
user_prompt, wild_label = self._build_user_prompt(
idea,
task_type,
duration_seconds,
shot_plan,
visual_style,
wildness,
camera_motion,
camera_amplitude,
camera_speed,
include_dialogue,
dialogue_language,
include_on_screen_text,
include_soundscape,
include_non_diegetic_music,
extra_instructions,
images is not None,
rng,
)
print(
f"🎬 H3 Prompt Writer | {task_type} | {duration_seconds:.2f}s | "
f"wildness {wildness} ({wild_label}) | seed {current_seed}"
)
text, resolved_model = call_llm(
model,
user_prompt,
system_prompt=self._build_system_prompt(),
images=images,
temperature=temperature,
seed=current_seed,
max_tokens=4000,
)
prompt = strip_code_fence(text)
return (
prompt,
extract_section(prompt, "integrated_multimodal_description", _FIELDS),
extract_section(prompt, "overall_soundscape", _FIELDS),
extract_section(prompt, "non_diegetic_music", _FIELDS),
resolved_model,
)
except Exception as exc:
print(f"❌ H3 Prompt Writer error: {exc}")
import traceback
print(traceback.format_exc())
error_message = f"Error occurred while writing the H3 prompt: {exc}"
return (error_message, error_message, "", "", "error")
+378
View File
@@ -0,0 +1,378 @@
# Shared helpers for the MiniMax-H3 prompt writer nodes
#
# The system prompts are the official MiniMax guides shipped verbatim under
# data/h3/, so the model is steered by the real spec rather than a paraphrase:
# data/h3/guide_base_en.md - VIDEO_PROMPT_WRITING_GUIDE_base_en.md (T2VA/I2VA/FL2VA/L2VA)
# data/h3/guide_ref_en.md - VIDEO_PROMPT_WRITING_GUIDE_ref_en.md (full-reference mode)
#
# Source: https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main/docs
import os
import re
_DATA_DIR = os.path.join(
os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))),
"data",
"h3",
)
_guide_cache = {}
def load_guide(name):
"""
Read (and cache) one of the shipped H3 guides from data/h3/.
Only a bare `.md` filename is accepted, so the lookup can never escape the
guide directory even if a caller later wires this to a node widget.
"""
safe_name = os.path.basename(name)
if safe_name != name or not safe_name.endswith(".md"):
raise ValueError(
f"H3 guide name must be a bare .md filename inside data/h3, got: {name!r}"
)
if safe_name in _guide_cache:
return _guide_cache[safe_name]
path = os.path.join(_DATA_DIR, safe_name)
try:
with open(path, "r", encoding="utf-8") as handle:
text = handle.read()
except Exception as exc:
raise FileNotFoundError(
f"Could not read the H3 prompt guide at {path}: {exc}"
)
_guide_cache[safe_name] = text
return text
# ----------------------------------------------------------------------
# Vocabulary lifted from section 4.3 of the base guide
# ----------------------------------------------------------------------
AUTO = "Auto"
CAMERA_MOTIONS = [
AUTO,
"Static Shot",
"Zoom In",
"Zoom Out",
"Push In",
"Pull Out",
"Pan Left",
"Pan Right",
"Truck Left",
"Truck Right",
"Tilt Up",
"Tilt Down",
"Pedestal Up",
"Pedestal Down",
"Arc Shot",
"Tracking Shot",
"Shake Slightly",
"Shake Strongly",
"POV",
"Roll Clockwise",
"Roll Counterclockwise",
]
# "medium amplitude and normal speed are usually omitted" - guide 4.3
CAMERA_AMPLITUDES = [AUTO, "with small amplitude", "medium (omit)", "with large amplitude"]
CAMERA_SPEEDS = [AUTO, "at slow speed", "normal (omit)", "at fast speed"]
VISUAL_STYLES = [
AUTO,
"Cinematic",
"live-action",
"2D-animated",
"3D CG",
"claymation",
"watercolor",
"vintage film",
]
SHOT_PLANS = [AUTO, "Single shot", "Two shots", "Three shots", "Four shots"]
_SHOT_COUNTS = {
"Single shot": 1,
"Two shots": 2,
"Three shots": 3,
"Four shots": 4,
}
CUT_STYLES = [
AUTO,
"the camera cuts to",
"the shot cuts to",
"the shot transitions to",
"the shot changes to",
"the shot switches to",
]
def camera_directive(motion, amplitude, speed):
"""
Turn the three camera widgets into one instruction line, honouring the
guide's rule that medium amplitude and normal speed are left unwritten.
"""
if motion == AUTO:
return (
"Camera motion: choose motion types that suit the action, and write them as "
"natural English inside the shot (motion type, plus amplitude and speed only "
"when meaningful). Do not stack them as labels at the end of a sentence."
)
parts = [motion]
if amplitude != AUTO and not amplitude.startswith("medium"):
parts.append(amplitude)
if speed != AUTO and not speed.startswith("normal"):
parts.append(speed)
phrase = " ".join(parts)
return (
f"Camera motion: the primary camera movement is `{phrase}`. Express it as natural "
"English action inside the shot rather than as a trailing label. Additional shots "
"may use other motion types when the action calls for it."
)
def shot_directive(shot_plan, duration_seconds):
"""Instruction covering shot count and cut-time formatting."""
duration = f"{duration_seconds:.2f}"
if shot_plan == AUTO:
count_rule = (
"Choose the shot count that fits the action. Prefer a single shot unless a cut "
"genuinely introduces new information about the subject, space, state, viewpoint or time."
)
else:
count = _SHOT_COUNTS[shot_plan]
count_rule = (
f"Use exactly {count} shot{'s' if count > 1 else ''}."
if count > 1
else "Use exactly 1 shot."
)
return (
f"{count_rule} The effective video duration is {duration} seconds. [Shot 1] carries no "
f"timestamp; every later shot opens with a strictly increasing cut time in `[Shot N] At "
f"MM:SS.mmm, ...` form that falls inside {duration} seconds."
)
def toggle_directives(
include_dialogue,
include_on_screen_text,
include_soundscape,
include_non_diegetic_music,
dialogue_language,
):
"""Feature switches shared by both writer nodes."""
lines = []
if include_dialogue:
lines.append(
f"Dialogue: include spoken lines. Give each vocal source a stable (S1)/(S2) ID and "
f"wrap only the spoken words in <d>[{dialogue_language}] ...</d>, keeping the "
f"identifying phrase, action and delivery outside the <d> block."
)
else:
lines.append(
"Dialogue: no character speaks, sings, or delivers a voiceover. Do not emit any "
"(Sx) speaker IDs or <d> blocks."
)
if include_on_screen_text:
lines.append(
'On-screen text: signs, banners, labels or subtitles that are actually visible go in '
'English double quotation marks, verbatim and untranslated.'
)
else:
lines.append("On-screen text: keep the frame free of readable signs, banners, labels or subtitles.")
if include_soundscape:
lines.append(
"overall_soundscape: 1-4 English sentences in one paragraph covering ambience, "
"physical action sounds and non-verbal human sounds. Do not repeat dialogue or "
"singing here."
)
else:
lines.append("overall_soundscape: output exactly `N/A`.")
if include_non_diegetic_music:
lines.append(
"non_diegetic_music: 1-3 English sentences on instrumentation, tempo, rhythm and "
"dynamic change. No abstract mood words and no explanation of emotional function."
)
else:
lines.append("non_diegetic_music: output exactly `N/A`.")
return lines
# ----------------------------------------------------------------------
# Wildness
# ----------------------------------------------------------------------
# Bands are (upper_bound_inclusive, label, directive, number_of_random_elements).
_WILDNESS_BANDS = (
(
15,
"Conservative",
"Stay literal. Render only what the input actually implies, adding just enough "
"concrete detail to make the timeline filmable. No invented events, no surreal "
"flourishes, no unmotivated camera tricks.",
0,
),
(
40,
"Grounded",
"Stay believable, but direct it properly. Choose expressive framing, motivated "
"lighting and small human behaviour that enrich the input without changing what "
"it is about.",
0,
),
(
65,
"Bold",
"Make strong authorial choices. Heightened lighting, striking compositions, "
"expressive camera work and one memorable visual idea are welcome, as long as the "
"scene still obeys ordinary physics.",
1,
),
(
85,
"Wild",
"Break realism on purpose. Surreal juxtapositions, impossible transitions and "
"dreamlike logic are encouraged. The result must still be a coherent, shootable "
"timeline rather than a list of random images.",
2,
),
(
100,
"Unhinged",
"Go fully unhinged. Reality is negotiable: scale, gravity, continuity and material "
"behaviour can all misbehave. Commit hard to the strangeness, and still deliver a "
"timeline a video model can actually follow, shot by shot.",
3,
),
)
# Concrete, filmable weirdness. Each is a visual event, not a mood word.
UNHINGED_ELEMENTS = [
"gravity reverses for a single object while everything else stays put",
"the subject's shadow moves a beat out of sync with the subject",
"a mirror or reflective surface shows something that is not in the room",
"one material behaves like another - stone flows, cloth turns molten, water holds an edge",
"an impossible scale shift: something small becomes enormous, or the reverse",
"the environment breathes, expanding and contracting like a slow lung",
"weather that belongs outdoors happens indoors",
"a doorway or window opens onto a completely different biome",
"the subject multiplies into synchronized copies that share one motion",
"everyone in the background freezes while the subject keeps moving",
"the floor turns to water and nobody reacts to it",
"practical lights pulse in time with a rhythm no one on screen can hear",
"the camera passes straight through a solid surface",
"time stutters: one action repeats half a beat before continuing",
"the scene briefly rewinds, then resumes forward",
"ordinary objects swarm and move as a single organism",
"the horizon tilts past vertical while the subject stays upright",
"an out-of-place animal crosses frame with complete confidence",
"the set reveals itself as a miniature, then becomes full scale again",
"colour drains from everything except one object",
"the colour palette inverts for a single beat",
"a texture spreads across the frame like frost, converting whatever it touches",
"the subject walks and the background scrolls the wrong way",
"objects rearrange themselves the instant the camera looks away",
"a second, older version of the scene bleeds through as a double exposure",
"the light source is physically present and can be picked up",
"rain falls upward into the sky",
"the subject's clothing changes between one cut and the next without comment",
"a hallway extends further the longer the camera pushes down it",
"sound arrives visibly, distorting the air before it is heard",
"the frame edge becomes a physical wall the subject can lean on",
"one object stays perfectly sharp while everything else smears into motion",
"the ground tessellates into moving tiles",
"a crowd moves in perfect unison as if choreographed by accident",
"the subject steps out of frame and immediately re-enters from the opposite side",
"smoke or steam holds a solid shape long after it should disperse",
"the scene is briefly lit as if from underwater",
"an object falls upward off the table and settles on the ceiling",
"the subject's reflection stays behind when they walk away",
"a season changes across a single continuous shot",
]
def wildness_directive(wildness, rng):
"""
Map the 0-100 wildness slider onto a creative-latitude instruction, plus a
seeded selection of concrete unhinged elements once the slider is high enough.
"""
wildness = max(0, min(100, int(wildness)))
for upper, label, directive, element_count in _WILDNESS_BANDS:
if wildness <= upper:
break
lines = [f"Creative latitude ({label}, wildness {wildness}/100): {directive}"]
if element_count and UNHINGED_ELEMENTS:
picks = rng.sample(UNHINGED_ELEMENTS, min(element_count, len(UNHINGED_ELEMENTS)))
joined = "; ".join(picks)
lines.append(
f"Weave in {'this element' if len(picks) == 1 else 'these elements'} and make "
f"{'it' if len(picks) == 1 else 'them'} land as real, visible events on the "
f"timeline: {joined}."
)
return lines, label
# ----------------------------------------------------------------------
# Output parsing
# ----------------------------------------------------------------------
def strip_code_fence(text):
"""Unwrap a ```...``` block if the model wrapped its whole answer in one."""
stripped = text.strip()
if not stripped.startswith("```"):
return stripped
lines = stripped.splitlines()
if len(lines) < 2:
return stripped
lines = lines[1:]
if lines and lines[-1].strip().startswith("```"):
lines = lines[:-1]
return "\n".join(lines).strip()
def extract_section(text, field, all_fields):
"""
Pull one `field:` section out of an H3 prompt, stopping at the next known
field label or at the end of the text. Returns "" when the field is absent.
Handles both layouts the guides use: a value on the same line as the label
(base mode) and a value starting on the next line (full-reference mode).
"""
others = [re.escape(f) for f in all_fields if f != field]
# The end-of-text alternative matters for the final section, which has no
# following label to stop at.
stop = (
r"(?=^[ \t]*(?:" + "|".join(others) + r")[ \t]*:|\Z)"
if others
else r"(?=\Z)"
)
pattern = re.compile(
r"^[ \t]*" + re.escape(field) + r"[ \t]*:(.*?)" + stop,
re.DOTALL | re.MULTILINE,
)
match = pattern.search(text)
return match.group(1).strip() if match else ""
+402
View File
@@ -0,0 +1,402 @@
# APNext H3 Reference Prompt Writer
#
# Writes a MiniMax-H3 full-reference rewrite (six sections) from a short idea
# plus up to four reference images, using the official full-reference guide.
import random
from ...utils.constants import CUSTOM_CATEGORY
from ...utils.image_utils import tensor2pil
from ...utils.llm_router import AUTO_DETECT, call_llm, list_all_models
from .common import (
AUTO,
CAMERA_AMPLITUDES,
CAMERA_MOTIONS,
CAMERA_SPEEDS,
SHOT_PLANS,
VISUAL_STYLES,
camera_directive,
extract_section,
load_guide,
shot_directive,
strip_code_fence,
toggle_directives,
wildness_directive,
)
TASK_TYPES = [
"Auto (decide from the references)",
"keyframe completion",
"reference generation",
"video editing",
"video continuation",
"audio reuse",
"audio reference",
]
# How each attached image should be labelled in subject_definitions.
REFERENCE_ROLES = [
"Auto (decide per image)",
"Subject (character / object / scene to reuse)",
"Picture (concrete frame anchor)",
"Style reference only",
"Storyboard / shot-planning reference",
]
_FIELDS = (
"subject_definitions",
"summary",
"retention_analysis",
"detailed_description",
"overall_soundscape",
"non_diegetic_music",
)
class H3RefPromptWriter:
"""
APNext H3 Reference Prompt Writer
Writes a MiniMax-H3 full-reference rewrite - subject_definitions, summary,
retention_analysis, detailed_description, overall_soundscape and
non_diegetic_music - from a short idea plus reference images, following the
official Full-Reference Mode Rewrite Output Format Guide.
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"idea": ("STRING", {
"multiline": True,
"default": "",
"tooltip": "Your short prompt: what should happen in the target video.",
}),
"task_type": (TASK_TYPES, {
"default": "Auto (decide from the references)",
"tooltip": "Square-bracketed prefix of the summary section. Auto lets the model combine types with ' + '.",
}),
"reference_role": (REFERENCE_ROLES, {
"default": "Auto (decide per image)",
"tooltip": "How the attached images should be labelled in subject_definitions.",
}),
"duration_seconds": ("FLOAT", {
"default": 8.0, "min": 1.0, "max": 60.0, "step": 0.5,
}),
"shot_plan": (SHOT_PLANS, {"default": AUTO}),
"visual_style": (VISUAL_STYLES, {
"default": AUTO,
"tooltip": "In full-reference mode the style is stated in one or two sentences BEFORE [Shot 1].",
}),
"wildness": ("INT", {
"default": 25, "min": 0, "max": 100, "step": 1,
"tooltip": "0 = literal and conservative, 100 = fully unhinged. Above 40 the node also injects concrete surreal elements picked from the seed.",
}),
"word_target": ("INT", {
"default": 425, "min": 150, "max": 1200, "step": 25,
"tooltip": "Target length of detailed_description. The guide recommends 350-500 words for generation tasks.",
}),
"camera_motion": (CAMERA_MOTIONS, {"default": AUTO}),
"camera_amplitude": (CAMERA_AMPLITUDES, {"default": AUTO}),
"camera_speed": (CAMERA_SPEEDS, {"default": AUTO}),
"include_dialogue": ("BOOLEAN", {"default": True}),
"dialogue_language": ("STRING", {"default": "English"}),
"include_on_screen_text": ("BOOLEAN", {"default": False}),
"include_soundscape": ("BOOLEAN", {"default": True}),
"include_non_diegetic_music": ("BOOLEAN", {"default": True}),
"model": (list_all_models(), {"default": AUTO_DETECT}),
"temperature": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 2.0, "step": 0.05}),
"seed": ("INT", {"default": -1, "min": -1, "max": 0xffffffffffffffff}),
},
"optional": {
"image_1": ("IMAGE",),
"image_2": ("IMAGE",),
"image_3": ("IMAGE",),
"image_4": ("IMAGE",),
"reference_notes": ("STRING", {
"multiline": True,
"default": "",
"tooltip": "Optional per-reference notes, one per line, e.g. 'Image 1: the woman, keep her cardigan'. Also use this to describe video or audio references you cannot attach.",
}),
"extra_instructions": ("STRING", {"multiline": True, "default": ""}),
},
}
RETURN_TYPES = ("STRING",) * 8
RETURN_NAMES = (
"h3_prompt",
"subject_definitions",
"summary",
"retention_analysis",
"detailed_description",
"overall_soundscape",
"non_diegetic_music",
"model_used",
)
FUNCTION = "write"
CATEGORY = f"{CUSTOM_CATEGORY}/H3"
DESCRIPTION = (
"Writes a MiniMax-H3 full-reference rewrite (six sections) from a short idea plus "
"reference images, following the official full-reference format guide."
)
def _build_system_prompt(self):
guide = load_guide("guide_ref_en.md")
base_guide = load_guide("guide_base_en.md")
return (
"You are a MiniMax-H3 video prompt engineer working in full-reference mode. You "
"rewrite a user's short idea plus their reference assets into a complete "
"six-section H3 rewrite.\n\n"
"Two authoritative specifications follow. The full-reference guide governs the "
"output structure and reference labels; the base guide governs shots, camera "
"vocabulary, speakers, dialogue and sound. Obey both exactly.\n\n"
"=== BEGIN FULL-REFERENCE MODE GUIDE ===\n"
f"{guide}\n"
"=== END FULL-REFERENCE MODE GUIDE ===\n\n"
"=== BEGIN BASE VIDEO PROMPT WRITING GUIDE ===\n"
f"{base_guide}\n"
"=== END BASE VIDEO PROMPT WRITING GUIDE ===\n\n"
"Output rules:\n"
"- Emit only the finished rewrite. No preamble, no commentary, no markdown "
"fences, no headings of your own.\n"
"- Emit all six sections, in this order, with these exact labels: "
"`subject_definitions:`, `summary:`, `retention_analysis:`, "
"`detailed_description:`, `overall_soundscape:`, `non_diegetic_music:`.\n"
"- A reference label keeps one meaning across every section. Do not introduce "
"new labels in `summary`.\n"
"- Do not write (Sx) speaker IDs in `retention_analysis`.\n"
"- Write everything in English except dialogue, lyrics and on-screen text."
)
def _reference_directive(self, reference_role, image_count, reference_notes):
if image_count == 0:
base = (
"No images are attached. Build the reference labels from the user's notes "
"below; if there are no usable references either, define the minimum set of "
"<Subject N> entries the idea implies and say so honestly in the summary."
)
else:
listing = ", ".join(f"image {i}" for i in range(1, image_count + 1))
base = (
f"{image_count} reference image(s) are attached in order ({listing}). Read them "
"directly and define them in subject_definitions."
)
if reference_role.startswith("Auto"):
base += (
" Decide per image whether it is a <Subject N> (reusable visible content), "
"a standalone <Picture N> (a concrete frame or composition anchor), or only "
"a source cited inside another item's definition."
)
elif reference_role.startswith("Subject"):
base += (
" Treat each one as reusable visible content: define a <Subject N> per image "
"and cite the image inside that definition rather than creating standalone "
"<Picture N> entries."
)
elif reference_role.startswith("Picture"):
base += (
" Treat each one as a concrete frame anchor: define a standalone <Picture N> "
"per image and state which shot and which position (first frame, keyframe, "
"last frame) it anchors."
)
elif reference_role.startswith("Style"):
base += (
" Treat them as style references only: do not create standalone <Picture N> "
"entries, and fold the style provenance into the relevant <Subject N> "
"definitions and the style sentence before [Shot 1]."
)
else:
base += (
" Treat them as storyboard / shot-planning references: define standalone "
"<Picture N> entries stating which shots they map to and what planning "
"information they provide."
)
if reference_notes.strip():
base += f"\n User notes on the references:\n {reference_notes.strip()}"
return base
def _build_user_prompt(
self,
idea,
task_type,
reference_role,
duration_seconds,
shot_plan,
visual_style,
wildness,
word_target,
camera_motion,
camera_amplitude,
camera_speed,
include_dialogue,
dialogue_language,
include_on_screen_text,
include_non_diegetic_music,
include_soundscape,
reference_notes,
extra_instructions,
image_count,
rng,
):
directives = [self._reference_directive(reference_role, image_count, reference_notes)]
if task_type.startswith("Auto"):
directives.append(
"Task type: choose the prefix that matches the actual role each reference "
"plays, combining several with ' + ' when they apply. Do not add a type just "
"because an asset exists."
)
else:
directives.append(
f"Task type: the summary begins with `[{task_type}]`, extended with ' + ' only "
"if another relationship genuinely applies."
)
directives.append(shot_directive(shot_plan, duration_seconds))
if visual_style == AUTO:
directives.append(
"Visual style: choose one that fits the idea and the references, and state it "
"in one or two English sentences before [Shot 1]."
)
else:
directives.append(
f"Visual style: state `{visual_style}` in one or two English sentences before "
"[Shot 1], not inside the shot label."
)
directives.append(camera_directive(camera_motion, camera_amplitude, camera_speed))
directives.append(
f"Length: aim for roughly {word_target} English words in detailed_description, "
"distributing detail across the shots by information load. Fitting a complete "
"spoken timeline matters more than hitting the number exactly."
)
directives.extend(
toggle_directives(
include_dialogue,
include_on_screen_text,
include_soundscape,
include_non_diegetic_music,
dialogue_language,
)
)
wild_lines, wild_label = wildness_directive(wildness, rng)
directives.extend(wild_lines)
if extra_instructions.strip():
directives.append(f"Additional direction from the user: {extra_instructions.strip()}")
numbered = "\n".join(f"{i}. {line}" for i, line in enumerate(directives, 1))
idea_text = idea.strip() or "(no written idea - build the target video from the references)"
return (
f"The user's idea:\n{idea_text}\n\n"
f"Write the full-reference H3 rewrite under these constraints:\n{numbered}\n\n"
"Return only the six finished sections."
), wild_label
def write(
self,
idea,
task_type,
reference_role,
duration_seconds,
shot_plan,
visual_style,
wildness,
word_target,
camera_motion,
camera_amplitude,
camera_speed,
include_dialogue,
dialogue_language,
include_on_screen_text,
include_soundscape,
include_non_diegetic_music,
model,
temperature,
seed,
image_1=None,
image_2=None,
image_3=None,
image_4=None,
reference_notes="",
extra_instructions="",
):
try:
images = []
for tensor in (image_1, image_2, image_3, image_4):
if tensor is not None:
# Only the first frame of each input acts as that reference.
images.append(tensor2pil(tensor[0]))
if not idea.strip() and not images and not reference_notes.strip():
raise ValueError("Provide an idea, at least one reference image, or reference notes.")
current_seed = seed if seed != -1 else random.randint(0, 0xffffffffffffffff)
rng = random.Random(current_seed)
user_prompt, wild_label = self._build_user_prompt(
idea,
task_type,
reference_role,
duration_seconds,
shot_plan,
visual_style,
wildness,
word_target,
camera_motion,
camera_amplitude,
camera_speed,
include_dialogue,
dialogue_language,
include_on_screen_text,
include_non_diegetic_music,
include_soundscape,
reference_notes,
extra_instructions,
len(images),
rng,
)
print(
f"🎬 H3 Reference Prompt Writer | {len(images)} image(s) | "
f"{duration_seconds:.2f}s | wildness {wildness} ({wild_label}) | seed {current_seed}"
)
text, resolved_model = call_llm(
model,
user_prompt,
system_prompt=self._build_system_prompt(),
images=images or None,
temperature=temperature,
seed=current_seed,
max_tokens=6000,
)
prompt = strip_code_fence(text)
return (
prompt,
extract_section(prompt, "subject_definitions", _FIELDS),
extract_section(prompt, "summary", _FIELDS),
extract_section(prompt, "retention_analysis", _FIELDS),
extract_section(prompt, "detailed_description", _FIELDS),
extract_section(prompt, "overall_soundscape", _FIELDS),
extract_section(prompt, "non_diegetic_music", _FIELDS),
resolved_model,
)
except Exception as exc:
print(f"❌ H3 Reference Prompt Writer error: {exc}")
import traceback
print(traceback.format_exc())
error_message = f"Error occurred while writing the H3 reference prompt: {exc}"
return (error_message, error_message, "", "", "", "", "", "error")
+6 -26
View File
@@ -2,12 +2,12 @@
import os
import random
import google.generativeai as genai
from openai import OpenAI
import anthropic
import httpx
from ...utils.constants import CUSTOM_CATEGORY, gpt_models, gemini_models, grok_models, claude_models, groq_models
from ...utils.gemini_client import get_gemini_client, gemini_generate
class APNextGenerator:
@@ -22,6 +22,7 @@ class APNextGenerator:
self.grok_client = None
self.claude_client = None
self.groq_client = None
self.gemini_client = None
self.gemini_configured = False
@classmethod
@@ -177,37 +178,16 @@ Generate a prompt that would create compelling, high-quality images. Be specific
api_key = os.environ.get("GEMINI_API_KEY")
if not api_key:
raise ValueError("GEMINI_API_KEY environment variable not set")
genai.configure(api_key=api_key)
self.gemini_client = get_gemini_client(api_key)
self.gemini_configured = True
# Extract model name (remove "gemini:" prefix)
gemini_model = model_name.replace("gemini:", "")
safety_settings = [
{
"category": "HARM_CATEGORY_HARASSMENT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_HATE_SPEECH",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT",
"threshold": "BLOCK_NONE",
},
{
"category": "HARM_CATEGORY_DANGEROUS_CONTENT",
"threshold": "BLOCK_NONE",
},
]
# Use EXACTLY the same pattern as your working nodes - no generation_config
model = genai.GenerativeModel(gemini_model, safety_settings=safety_settings)
print(f"🔄 Sending to Gemini model: {gemini_model}")
response = model.generate_content(prompt)
response = gemini_generate(self.gemini_client, gemini_model, prompt)
print("📥 Gemini response received!")
# EXACT same pattern as your working GeminiTextOnly node
@@ -19,12 +19,11 @@ except ImportError:
OpenAI = None
# Gemini imports (optional)
try:
import google.generativeai as genai
GEMINI_AVAILABLE = True
except ImportError:
GEMINI_AVAILABLE = False
genai = None
from ...utils.gemini_client import (
GEMINI_AVAILABLE,
get_gemini_client,
gemini_generate,
)
# Anthropic imports (optional)
try:
@@ -49,6 +48,7 @@ class UniversalVisionCloner:
self.openai_client = None
self.grok_client = None
self.claude_client = None
self.gemini_client = None
self.gemini_configured = False
@classmethod
@@ -127,7 +127,7 @@ class UniversalVisionCloner:
if not OPENAI_AVAILABLE:
missing_deps.append("openai library")
if not GEMINI_AVAILABLE:
missing_deps.append("google-generativeai library")
missing_deps.append("google-genai library")
if not ANTHROPIC_AVAILABLE:
missing_deps.append("anthropic library")
@@ -401,31 +401,24 @@ class UniversalVisionCloner:
"""Analyze image using Google Gemini"""
try:
if not GEMINI_AVAILABLE:
raise ValueError("Google Generative AI library not available. Install with: pip install google-generativeai")
raise ValueError("google-genai library not available. Install with: pip install google-genai")
# Lazy initialization of Gemini
if not self.gemini_configured:
api_key = os.environ.get("GEMINI_API_KEY")
if not api_key:
raise ValueError("GEMINI_API_KEY environment variable not set")
genai.configure(api_key=api_key)
self.gemini_client = get_gemini_client(api_key)
self.gemini_configured = True
# Extract model name (remove "gemini:" prefix)
gemini_model = model_name.replace("gemini:", "")
safety_settings = [
{"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_NONE"},
{"category": "HARM_CATEGORY_HATE_SPEECH", "threshold": "BLOCK_NONE"},
{"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT", "threshold": "BLOCK_NONE"},
{"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_NONE"},
]
model = genai.GenerativeModel(gemini_model, safety_settings=safety_settings)
print(f"🔄 Sending to Gemini vision model: {gemini_model}")
response = model.generate_content([prompt, image])
response = gemini_generate(
self.gemini_client, gemini_model, [prompt, image]
)
print("📥 Gemini vision response received!")
result = response.text
print(f"📝 Response length: {len(result)} characters")
+11
View File
@@ -0,0 +1,11 @@
# APNext Resolution Planning Nodes
from .h3_resolution_planner import H3ResolutionPlannerCropOnly
NODE_CLASS_MAPPINGS = {
"H3ResolutionPlannerCropOnly": H3ResolutionPlannerCropOnly,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"H3ResolutionPlannerCropOnly": "APNext H3 Resolution Planner (Crop Only) - by gabbo",
}
+475
View File
@@ -0,0 +1,475 @@
# H3 Resolution Planner (Crop Only)
#
# Original node and algorithm by **gabbo**.
# Ported into comfyui_dagthomas with the planning logic kept intact; only the
# ComfyUI plumbing (category, tooltips, English messages, plan_info output)
# was adapted to match this pack's conventions.
import math
from ...utils.constants import CUSTOM_CATEGORY
class H3ResolutionPlannerCropOnly:
"""
H3 Resolution Planner (Crop Only) - by gabbo
Plans a two-stage generate-then-upscale resolution pair and center-crops the
input image to the exact aspect ratio of that plan, so no resampling or
padding is needed anywhere in the chain.
Stage 1 is the generation size, stage 2 is stage 1 multiplied by the chosen
upscale factor. Step sizes are picked so that both stages land on clean
multiples of 32:
- 2x -> stage 1 steps of 32, stage 2 steps of 64
- 1.5x -> stage 1 steps of 64, stage 2 steps of 96
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"image": ("IMAGE",),
"resolution_mode": (
[
"target_megapixels",
"max_stage1_from_input",
"max_final_from_input",
],
{
"default": "target_megapixels",
"tooltip": (
"target_megapixels: hit stage1_megapixels while "
"staying close to the input aspect ratio.\n"
"max_stage1_from_input: largest stage 1 the input "
"can feed within max_crop_percent.\n"
"max_final_from_input: largest stage 2 (final) the "
"input can feed within max_crop_percent."
),
}
),
"stage1_megapixels": ("FLOAT", {
"default": 0.40,
"min": 0.05,
"max": 4.00,
"step": 0.01,
"tooltip": "Target stage 1 size in megapixels. Only used by target_megapixels mode.",
}),
"upscale_mode": (
["2x", "1.5x"],
{
"default": "2x",
"tooltip": "Stage 1 -> stage 2 factor. 1.5x forces stage 1 onto multiples of 64.",
}
),
"max_crop_percent": ("FLOAT", {
"default": 2.0,
"min": 0.0,
"max": 25.0,
"step": 0.1,
"tooltip": (
"Maximum share of the input area allowed to be cropped away. "
"Only used by the two max_* modes; if nothing fits, the "
"least-lossy candidate is used instead."
),
}),
}
}
RETURN_TYPES = (
"IMAGE",
"INT",
"INT",
"INT",
"INT",
"FLOAT",
"STRING",
)
RETURN_NAMES = (
"cropped_image",
"stage1_width",
"stage1_height",
"stage2_width",
"stage2_height",
"upscale_factor",
"plan_info",
)
FUNCTION = "plan"
CATEGORY = f"{CUSTOM_CATEGORY}/Resolution"
DESCRIPTION = (
"Plans stage 1 / stage 2 resolutions and center-crops the input to that "
"exact aspect ratio, so the whole chain stays on clean multiples of 32. "
"Node and algorithm by gabbo."
)
# ------------------------------------------------------------
# Helpers
# ------------------------------------------------------------
def _get_stage1_step_and_scale(self, upscale_mode):
# For 1.5x, stage1 must be multiple of 64 so stage2 remains
# a multiple of 32.
if upscale_mode == "1.5x":
return 64, 1.5
return 32, 2.0
def _get_stage2_step(self, upscale_mode):
# If stage1 is multiple of 32 and scale is 2x, stage2 is multiple of 64.
# If stage1 is multiple of 64 and scale is 1.5x, stage2 is multiple of 96.
if upscale_mode == "1.5x":
return 96
return 64
def _stage2_size(self, w1, h1, upscale_mode):
if upscale_mode == "1.5x":
return int(w1 * 3 / 2), int(h1 * 3 / 2)
return w1 * 2, h1 * 2
def _stage1_from_stage2(self, w2, h2, upscale_mode):
if upscale_mode == "1.5x":
return int(w2 * 2 / 3), int(h2 * 2 / 3)
return int(w2 / 2), int(h2 / 2)
def _largest_exact_crop(self, in_w, in_h, target_w, target_h):
g = math.gcd(int(target_w), int(target_h))
rw = int(target_w) // g
rh = int(target_h) // g
k = min(in_w // rw, in_h // rh)
if k < 1:
return None
crop_w = rw * k
crop_h = rh * k
crop_x = max(0, (in_w - crop_w) // 2)
crop_y = max(0, (in_h - crop_h) // 2)
return crop_x, crop_y, crop_w, crop_h
def _crop_loss(self, in_w, in_h, crop_w, crop_h):
input_area = max(1, in_w * in_h)
crop_area = crop_w * crop_h
return 1.0 - (crop_area / input_area)
# ------------------------------------------------------------
# Mode 1: target MP
# ------------------------------------------------------------
def _find_target_mp_resolution(
self,
in_w,
in_h,
stage1_megapixels,
upscale_mode,
search_radius_steps=12
):
target_ratio = in_w / in_h
target_pixels = stage1_megapixels * 1024 * 1024
step1, _ = self._get_stage1_step_and_scale(upscale_mode)
ideal_w = math.sqrt(target_pixels * target_ratio)
ideal_h = math.sqrt(target_pixels / target_ratio)
center_w = max(step1, round(ideal_w / step1) * step1)
center_h = max(step1, round(ideal_h / step1) * step1)
best = None
best_score = float("inf")
for wi in range(-search_radius_steps, search_radius_steps + 1):
for hi in range(-search_radius_steps, search_radius_steps + 1):
w1 = center_w + wi * step1
h1 = center_h + hi * step1
if w1 < step1 or h1 < step1:
continue
w2, h2 = self._stage2_size(w1, h1, upscale_mode)
if (w2 % 32) != 0 or (h2 % 32) != 0:
continue
crop = self._largest_exact_crop(in_w, in_h, w1, h1)
if crop is None:
continue
ratio = w1 / h1
aspect_err = abs(math.log(ratio / target_ratio))
pixel_err = abs((w1 * h1) - target_pixels) / target_pixels
crop_loss = self._crop_loss(in_w, in_h, crop[2], crop[3])
score = aspect_err * 5.0 + pixel_err + crop_loss * 0.25
if score < best_score:
best_score = score
best = (w1, h1, w2, h2, crop)
if best is None:
raise ValueError(
f"No valid resolution found for a {in_w}x{in_h} input at "
f"{stage1_megapixels:.2f} MP ({upscale_mode}). "
"Try a different target or upscale mode."
)
return best
# ------------------------------------------------------------
# Mode 2: max Stage 1 from input
# ------------------------------------------------------------
def _find_max_stage1_from_input(
self,
in_w,
in_h,
upscale_mode,
max_crop_percent
):
step1, _ = self._get_stage1_step_and_scale(upscale_mode)
max_crop_loss = max_crop_percent / 100.0
valid = []
fallback = []
for w1 in range(step1, in_w + 1, step1):
for h1 in range(step1, in_h + 1, step1):
w2, h2 = self._stage2_size(w1, h1, upscale_mode)
if (w2 % 32) != 0 or (h2 % 32) != 0:
continue
crop = self._largest_exact_crop(in_w, in_h, w1, h1)
if crop is None:
continue
_, _, crop_w, crop_h = crop
# Stage 1 canvas must fit within the cropped source.
if w1 > crop_w or h1 > crop_h:
continue
loss = self._crop_loss(in_w, in_h, crop_w, crop_h)
area = w1 * h1
item = (area, -loss, w1, h1, w2, h2, crop)
fallback.append(item)
if loss <= max_crop_loss + 1e-12:
valid.append(item)
pool = valid if valid else fallback
if not pool:
raise ValueError(
f"No valid resolution found for a {in_w}x{in_h} input "
f"({upscale_mode}). The image is likely smaller than one "
f"{step1}px step."
)
if valid:
best = max(pool, key=lambda x: (x[0], x[1]))
else:
best = max(pool, key=lambda x: (x[1], x[0]))
_, _, w1, h1, w2, h2, crop = best
return w1, h1, w2, h2, crop, not valid
# ------------------------------------------------------------
# Mode 3: max FINAL from input
# ------------------------------------------------------------
def _find_max_final_from_input(
self,
in_w,
in_h,
upscale_mode,
max_crop_percent
):
step2 = self._get_stage2_step(upscale_mode)
step1, _ = self._get_stage1_step_and_scale(upscale_mode)
max_crop_loss = max_crop_percent / 100.0
valid = []
fallback = []
for w2 in range(step2, in_w + 1, step2):
for h2 in range(step2, in_h + 1, step2):
w1, h1 = self._stage1_from_stage2(w2, h2, upscale_mode)
if w1 < step1 or h1 < step1:
continue
if (w1 % step1) != 0 or (h1 % step1) != 0:
continue
crop = self._largest_exact_crop(in_w, in_h, w2, h2)
if crop is None:
continue
_, _, crop_w, crop_h = crop
# Final canvas must fit within the cropped source.
if w2 > crop_w or h2 > crop_h:
continue
loss = self._crop_loss(in_w, in_h, crop_w, crop_h)
area = w2 * h2
item = (area, -loss, w1, h1, w2, h2, crop)
fallback.append(item)
if loss <= max_crop_loss + 1e-12:
valid.append(item)
pool = valid if valid else fallback
if not pool:
raise ValueError(
f"No valid resolution found for a {in_w}x{in_h} input "
f"({upscale_mode}). The image is likely smaller than one "
f"{step2}px final step."
)
if valid:
best = max(pool, key=lambda x: (x[0], x[1]))
else:
best = max(pool, key=lambda x: (x[1], x[0]))
_, _, w1, h1, w2, h2, crop = best
return w1, h1, w2, h2, crop, not valid
# ------------------------------------------------------------
# Main
# ------------------------------------------------------------
def plan(
self,
image,
resolution_mode,
stage1_megapixels,
upscale_mode,
max_crop_percent
):
if len(image.shape) != 4:
raise ValueError(f"Unexpected image format: {tuple(image.shape)}")
_, in_h, in_w, _ = image.shape
in_w = int(in_w)
in_h = int(in_h)
# True when a max_* mode could not satisfy max_crop_percent and had to
# settle for the least-lossy candidate instead.
over_budget = False
if resolution_mode == "max_stage1_from_input":
stage1_w, stage1_h, stage2_w, stage2_h, crop, over_budget = (
self._find_max_stage1_from_input(
in_w=in_w,
in_h=in_h,
upscale_mode=upscale_mode,
max_crop_percent=max_crop_percent
)
)
elif resolution_mode == "max_final_from_input":
stage1_w, stage1_h, stage2_w, stage2_h, crop, over_budget = (
self._find_max_final_from_input(
in_w=in_w,
in_h=in_h,
upscale_mode=upscale_mode,
max_crop_percent=max_crop_percent
)
)
else:
stage1_w, stage1_h, stage2_w, stage2_h, crop = (
self._find_target_mp_resolution(
in_w=in_w,
in_h=in_h,
stage1_megapixels=stage1_megapixels,
upscale_mode=upscale_mode,
search_radius_steps=12
)
)
crop_x, crop_y, crop_w, crop_h = crop
cropped = image[
:,
crop_y:crop_y + crop_h,
crop_x:crop_x + crop_w,
:
]
_, scale = self._get_stage1_step_and_scale(upscale_mode)
info = self._format_plan_info(
in_w=in_w,
in_h=in_h,
crop=crop,
stage1_w=stage1_w,
stage1_h=stage1_h,
stage2_w=stage2_w,
stage2_h=stage2_h,
resolution_mode=resolution_mode,
upscale_mode=upscale_mode,
max_crop_percent=max_crop_percent,
over_budget=over_budget,
)
return (
cropped,
int(stage1_w),
int(stage1_h),
int(stage2_w),
int(stage2_h),
float(scale),
info,
)
def _format_plan_info(
self,
in_w,
in_h,
crop,
stage1_w,
stage1_h,
stage2_w,
stage2_h,
resolution_mode,
upscale_mode,
max_crop_percent,
over_budget,
):
crop_x, crop_y, crop_w, crop_h = crop
loss_pct = self._crop_loss(in_w, in_h, crop_w, crop_h) * 100.0
g = math.gcd(int(stage1_w), int(stage1_h))
lines = [
f"mode: {resolution_mode} @ {upscale_mode}",
f"input: {in_w}x{in_h}",
f"crop: {crop_w}x{crop_h} at ({crop_x},{crop_y}) - {loss_pct:.2f}% of area removed",
f"aspect: {int(stage1_w) // g}:{int(stage1_h) // g}",
f"stage 1: {stage1_w}x{stage1_h} ({(stage1_w * stage1_h) / 1048576.0:.2f} MP)",
f"stage 2: {stage2_w}x{stage2_h} ({(stage2_w * stage2_h) / 1048576.0:.2f} MP)",
]
# The max_* modes silently fall back to the least-lossy candidate when
# nothing fits the budget, so say when that happened.
if over_budget:
lines.append(
f"WARNING: no plan fit within max_crop_percent ({max_crop_percent:.1f}%); "
f"used the least-lossy option at {loss_pct:.2f}%."
)
return "\n".join(lines)
NODE_CLASS_MAPPINGS = {
"H3ResolutionPlannerCropOnly": H3ResolutionPlannerCropOnly,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"H3ResolutionPlannerCropOnly": "APNext H3 Resolution Planner (Crop Only) - by gabbo",
}
+29 -14
View File
@@ -1,14 +1,29 @@
Pillow==10.4.0
requests==2.32.5
openai==1.44.0
blend-modes==2.1.0
huggingface_hub>=0.34.0
color_matcher==0.5.0
chardet==5.2.0
google-generativeai==0.7.2
anthropic
transformers>=4.40.0
decord>=0.6.0
scipy>=1.10.0
tqdm>=4.67.1
huggingface_hub[hf_xet]
# comfyui_dagthomas
#
# Anything ComfyUI already declares in its own requirements.txt is deliberately
# NOT repeated here: Pillow, requests, transformers, scipy, tqdm, numpy, torch.
# Re-pinning them only risks downgrading the versions the base install ships.
# --- LLM / API SDKs ---
# openai 3.x switched to httpx2 and no longer accepts the httpx.Client we pass
# as http_client=, so stay on 2.x until the nodes are migrated.
openai>=2.54.0,<3.0.0
anthropic>=0.121.0
# Current Gemini SDK. The legacy `google-generativeai` package was dropped
# because it hard-pinned google-ai-generativelanguage==0.6.15, which forced
# protobuf<6 and pulled grpcio into the ComfyUI env. This one needs neither.
google-genai>=2.18.0
# Imported directly by the GPT / Grok / Groq nodes, not just via the SDKs.
httpx>=0.28.1
# --- model hub ---
huggingface_hub[hf_xet]>=0.34.0
# --- misc ---
chardet>=5.2.0
# --- optional ---
# decord (Qwen-VL / MiniCPM video nodes) is unmaintained since 2022 and is not
# numpy-2 safe; those nodes fall back to opencv automatically. Uncomment only if
# you specifically want decord-based frame decoding.
# decord>=0.6.0
+106
View File
@@ -0,0 +1,106 @@
# Shared Gemini helper (google-genai)
#
# These nodes used to talk to the legacy `google-generativeai` SDK, which hard
# pinned google-ai-generativelanguage==0.6.15 and therefore dragged protobuf<6
# plus the whole grpcio / google-api stack into the ComfyUI environment.
# `google-genai` is the current SDK and needs none of that.
#
# Legacy -> current mapping used across the pack:
# genai.configure(api_key=...) -> genai.Client(api_key=...)
# genai.GenerativeModel(m, safety=...) -> config=types.GenerateContentConfig(...)
# model.generate_content(contents) -> client.models.generate_content(
# model=m, contents=..., config=...)
import os
try:
from google import genai
from google.genai import types as genai_types
GEMINI_AVAILABLE = True
except ImportError:
genai = None
genai_types = None
GEMINI_AVAILABLE = False
GEMINI_INSTALL_HINT = (
"google-genai is not installed. Install it with: pip install google-genai"
)
# The permissive safety posture these nodes have always shipped with.
_SAFETY_CATEGORIES = (
"HARM_CATEGORY_HARASSMENT",
"HARM_CATEGORY_HATE_SPEECH",
"HARM_CATEGORY_SEXUALLY_EXPLICIT",
"HARM_CATEGORY_DANGEROUS_CONTENT",
)
def get_gemini_client(api_key=None):
"""
Build a google-genai client.
Falls back to the GEMINI_API_KEY environment variable and raises the same
errors the nodes raised before the migration.
"""
if not GEMINI_AVAILABLE:
raise ImportError(GEMINI_INSTALL_HINT)
key = api_key or os.environ.get("GEMINI_API_KEY")
if not key:
raise ValueError("GEMINI_API_KEY environment variable is not set")
return genai.Client(api_key=key)
def build_gemini_config(**kwargs):
"""
GenerateContentConfig with every safety filter set to BLOCK_NONE.
Any extra keyword (temperature, max_output_tokens, system_instruction, ...)
is forwarded; None values are dropped so callers can pass optionals
unconditionally.
"""
if not GEMINI_AVAILABLE:
raise ImportError(GEMINI_INSTALL_HINT)
return genai_types.GenerateContentConfig(
safety_settings=[
genai_types.SafetySetting(category=category, threshold="BLOCK_NONE")
for category in _SAFETY_CATEGORIES
],
**{k: v for k, v in kwargs.items() if v is not None},
)
def gemini_generate(client, model, contents, **config_kwargs):
"""
client.models.generate_content with the shared safety config applied.
`contents` may be a string, a PIL image, or a list mixing both - google-genai
converts PIL images to inline image parts on its own.
"""
if not isinstance(contents, list):
contents = [contents]
return client.models.generate_content(
model=model,
contents=contents,
config=build_gemini_config(**config_kwargs),
)
def gemini_finished_normally(candidate):
"""
True when a candidate completed with FinishReason.STOP.
The legacy SDK exposed finish_reason as the int 1; google-genai returns a
FinishReason string enum, so compare by name instead.
"""
reason = getattr(candidate, "finish_reason", None)
if reason is None:
return True
name = getattr(reason, "name", None) or str(reason).rsplit(".", 1)[-1]
return str(name).upper() == "STOP"
+301
View File
@@ -0,0 +1,301 @@
# Shared multi-provider LLM router
#
# One entry point (`call_llm`) that talks to GPT, Gemini, Claude, Grok and Groq
# with an optional system prompt and optional images. Providers are selected by
# a "provider:model" string, the same convention the universal nodes already use.
import base64
import io
import os
from .constants import (
gpt_models,
gemini_models,
grok_models,
claude_models,
groq_models,
)
try:
from openai import OpenAI
OPENAI_AVAILABLE = True
except ImportError:
OpenAI = None
OPENAI_AVAILABLE = False
try:
import anthropic
ANTHROPIC_AVAILABLE = True
except ImportError:
anthropic = None
ANTHROPIC_AVAILABLE = False
try:
import httpx
except ImportError:
httpx = None
from .gemini_client import GEMINI_AVAILABLE, get_gemini_client, gemini_generate
# Provider prefix -> (env var names, OpenAI-compatible base url or None)
_OPENAI_COMPATIBLE = {
"gpt": (("OPENAI_API_KEY",), None),
"grok": (("XAI_API_KEY", "GROK_API_KEY"), "https://api.x.ai/v1"),
"groq": (("GROQ_API_KEY",), "https://api.groq.com/openai/v1"),
}
# Preference order used by auto-detect, with the fallback model for each.
_AUTO_DETECT_ORDER = (
(("ANTHROPIC_API_KEY", "CLAUDE_API_KEY"), "claude:claude-sonnet-4.5"),
(("OPENAI_API_KEY",), "gpt:gpt-4o"),
(("GEMINI_API_KEY",), "gemini:gemini-2.5-flash"),
(("XAI_API_KEY", "GROK_API_KEY"), "grok:grok-beta"),
(("GROQ_API_KEY",), "groq:llama-3.3-70b-versatile"),
)
AUTO_DETECT = "auto-detect"
# Clients are cached per process so repeated node runs reuse connections.
_client_cache = {}
def list_all_models():
"""Every selectable model string, auto-detect first."""
return (
[AUTO_DETECT]
+ [f"claude:{m}" for m in claude_models]
+ [f"gpt:{m}" for m in gpt_models]
+ [f"gemini:{m}" for m in gemini_models]
+ [f"grok:{m}" for m in grok_models]
+ [f"groq:{m}" for m in groq_models]
)
def _first_env(names):
for name in names:
value = os.environ.get(name)
if value:
return value
return None
def auto_detect_model():
"""Pick the best model whose API key is actually present."""
for env_names, model in _AUTO_DETECT_ORDER:
if _first_env(env_names):
return model
raise ValueError(
"No API keys found. Set one of: ANTHROPIC_API_KEY, OPENAI_API_KEY, "
"GEMINI_API_KEY, XAI_API_KEY, GROQ_API_KEY"
)
def resolve_model(model_name):
"""Turn 'auto-detect' into a concrete 'provider:model' string."""
if not model_name or model_name == AUTO_DETECT:
return auto_detect_model()
return model_name
def split_model(model_name):
"""'gpt:gpt-4o' -> ('gpt', 'gpt-4o'). Unprefixed names default to gpt."""
if ":" in model_name:
provider, model = model_name.split(":", 1)
return provider.strip().lower(), model.strip()
return "gpt", model_name.strip()
def _http_client():
if httpx is None:
return None
try:
return httpx.Client(timeout=180.0)
except TypeError:
return httpx.Client()
def _get_openai_compatible_client(provider):
cached = _client_cache.get(provider)
if cached is not None:
return cached
if not OPENAI_AVAILABLE:
raise ImportError("openai is not installed. Install it with: pip install 'openai<3'")
env_names, base_url = _OPENAI_COMPATIBLE[provider]
api_key = _first_env(env_names)
if not api_key:
raise ValueError(f"{' or '.join(env_names)} environment variable not set")
kwargs = {"api_key": api_key}
http_client = _http_client()
if http_client is not None:
kwargs["http_client"] = http_client
if base_url:
kwargs["base_url"] = base_url
client = OpenAI(**kwargs)
_client_cache[provider] = client
return client
def _get_claude_client():
cached = _client_cache.get("claude")
if cached is not None:
return cached
if not ANTHROPIC_AVAILABLE:
raise ImportError("anthropic is not installed. Install it with: pip install anthropic")
api_key = _first_env(("ANTHROPIC_API_KEY", "CLAUDE_API_KEY"))
if not api_key:
raise ValueError("ANTHROPIC_API_KEY or CLAUDE_API_KEY environment variable not set")
client = anthropic.Anthropic(api_key=api_key)
_client_cache["claude"] = client
return client
def _get_gemini_client():
cached = _client_cache.get("gemini")
if cached is not None:
return cached
client = get_gemini_client()
_client_cache["gemini"] = client
return client
def _encode_png(image):
"""PIL image -> raw PNG bytes."""
buffer = io.BytesIO()
image.convert("RGB").save(buffer, format="PNG")
return buffer.getvalue()
def _call_openai_compatible(provider, model, user_prompt, system_prompt, images, temperature, seed, max_tokens):
client = _get_openai_compatible_client(provider)
if images:
content = [{"type": "text", "text": user_prompt}]
for image in images:
encoded = base64.b64encode(_encode_png(image)).decode("utf-8")
content.append(
{
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{encoded}"},
}
)
else:
content = user_prompt
messages = []
if system_prompt:
messages.append({"role": "system", "content": system_prompt})
messages.append({"role": "user", "content": content})
kwargs = {
"model": model,
"messages": messages,
"max_tokens": max_tokens,
"temperature": temperature,
}
# Groq rejects the seed parameter on several models, so only GPT and Grok get it.
if seed is not None and seed != -1 and provider in ("gpt", "grok"):
kwargs["seed"] = seed
response = client.chat.completions.create(**kwargs)
return (response.choices[0].message.content or "").strip()
def _call_claude(model, user_prompt, system_prompt, images, temperature, max_tokens):
client = _get_claude_client()
if images:
content = []
for image in images:
encoded = base64.b64encode(_encode_png(image)).decode("utf-8")
content.append(
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png",
"data": encoded,
},
}
)
content.append({"type": "text", "text": user_prompt})
else:
content = user_prompt
kwargs = {
"model": model,
"max_tokens": max_tokens,
"temperature": temperature,
"messages": [{"role": "user", "content": content}],
}
if system_prompt:
kwargs["system"] = system_prompt
response = client.messages.create(**kwargs)
# Concatenate every text block; thinking-capable models can emit several.
parts = [block.text for block in response.content if getattr(block, "type", None) == "text"]
return "".join(parts).strip()
def _call_gemini(model, user_prompt, system_prompt, images, temperature, max_tokens):
client = _get_gemini_client()
contents = [user_prompt]
if images:
contents.extend(images)
response = gemini_generate(
client,
model,
contents,
system_instruction=system_prompt,
temperature=temperature,
max_output_tokens=max_tokens,
)
return (response.text or "").strip()
def call_llm(
model_name,
user_prompt,
system_prompt=None,
images=None,
temperature=1.0,
seed=-1,
max_tokens=4000,
):
"""
Send a prompt to whichever provider `model_name` selects.
`images` is a list of PIL images; providers that support vision receive them
inline. Raises on failure so callers can decide how to surface the error.
"""
resolved = resolve_model(model_name)
provider, model = split_model(resolved)
if provider in _OPENAI_COMPATIBLE:
text = _call_openai_compatible(
provider, model, user_prompt, system_prompt, images, temperature, seed, max_tokens
)
elif provider == "claude":
text = _call_claude(model, user_prompt, system_prompt, images, temperature, max_tokens)
elif provider == "gemini":
if not GEMINI_AVAILABLE:
raise ImportError("google-genai is not installed. Install it with: pip install google-genai")
text = _call_gemini(model, user_prompt, system_prompt, images, temperature, max_tokens)
else:
raise ValueError(f"Unknown provider '{provider}' in model '{model_name}'")
return text, resolved