diff --git a/README.md b/README.md index 2dde570..9092413 100644 --- a/README.md +++ b/README.md @@ -488,6 +488,113 @@ Optimized latent generation for Stable Diffusion 3 pipelines. --- +### πŸ“ Resolution Planning + +#### H3 Resolution Planner (Crop Only) +**Display Name:** `APNext H3 Resolution Planner (Crop Only) - by gabbo` + +> Original node and algorithm by **gabbo**. Ported into this pack with the planning logic unchanged. + +Plans a two-stage *generate β†’ upscale* resolution pair and center-crops the input image to the **exact** aspect ratio of that plan, so nothing in the chain has to resample or pad. Step sizes are chosen so both stages always land on clean multiples of 32: + +| Upscale | Stage 1 steps | Stage 2 steps | +|---------|---------------|---------------| +| `2x` | 32 | 64 | +| `1.5x` | 64 | 96 | + +| Input | Description | +|-------|-------------| +| `image` | Source image; only its dimensions drive the plan | +| `resolution_mode` | `target_megapixels`, `max_stage1_from_input`, `max_final_from_input` | +| `stage1_megapixels` | Target stage 1 size in MP (0.05–4.00). `target_megapixels` mode only | +| `upscale_mode` | `2x` or `1.5x` | +| `max_crop_percent` | Max share of input area croppable (0–25%). The two `max_*` modes only; falls back to the least-lossy candidate if nothing fits | + +**Modes** +- `target_megapixels` β€” hits the requested stage 1 megapixels while staying as close as possible to the input aspect ratio. +- `max_stage1_from_input` β€” largest stage 1 the input can feed natively within the crop budget. +- `max_final_from_input` β€” largest stage 2 (final) the input can feed natively within the crop budget. + +**Returns:** `(cropped_image, stage1_width, stage1_height, stage2_width, stage2_height, upscale_factor, plan_info)` + +`plan_info` is a human-readable summary of the chosen plan: + +``` +mode: target_megapixels @ 2x +input: 1920x1080 +crop: 1917x1065 at (1,7) - 1.54% of area removed +aspect: 9:5 +stage 1: 864x480 (0.40 MP) +stage 2: 1728x960 (1.58 MP) +``` + +--- + +### πŸŽ₯ MiniMax-H3 Prompt Nodes + +Both nodes take a **short idea, an image, or both** and expand it into a complete, spec-compliant MiniMax-H3 video prompt. The official MiniMax writing guides ship verbatim in `data/h3/` and are used as the system prompt, so the model follows the real spec rather than a paraphrase β€” edit those files to tune behaviour globally. + +Any provider works: `auto-detect` picks the first of Claude β†’ GPT β†’ Gemini β†’ Grok β†’ Groq that has an API key set. When an image is connected it is sent as vision input, so the model describes the frame itself instead of you writing the description. + +#### APNext H3 Prompt Writer +**Display Name:** `APNext H3 Prompt Writer` + +Writes the base format β€” `integrated_multimodal_description`, `overall_soundscape`, `non_diegetic_music` β€” per [VIDEO_PROMPT_WRITING_GUIDE_base_en.md](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md). + +| Input | Description | +|-------|-------------| +| `idea` | Your short prompt or image description β€” the thing being expanded | +| `task_type` | `T2VA` (text only), `I2VA` (first frame), `FL2VA` (first + last), `L2VA` (last frame). Non-T2VA emits the exact reference-alignment instruction line | +| `duration_seconds` | Drives cut times and the `S.SS` value in the alignment line | +| `shot_plan` | Auto, or force 1–4 shots | +| `visual_style` | Auto, or one of the guide's styles (`Cinematic`, `live-action`, `2D-animated`, `3D CG`, `claymation`, `watercolor`, `vintage film`) | +| `wildness` | **0 = literal, 100 = fully unhinged.** See below | +| `camera_motion` / `camera_amplitude` / `camera_speed` | The guide's full camera vocabulary. Medium amplitude and normal speed are omitted from the output, as the spec requires | +| `include_dialogue` | Off β‡’ no `(Sx)` IDs and no `` blocks at all | +| `dialogue_language` | Language tag written inside `[...]` | +| `include_on_screen_text` | Whether readable signs/banners/subtitles appear | +| `include_soundscape` / `include_non_diegetic_music` | Off writes `N/A` into that field | +| `model`, `temperature`, `seed` | Provider selection and sampling | +| `image` *(optional)* | Reference frame(s), sent as vision input | +| `extra_instructions` *(optional)* | Free-form extra direction | + +**Returns:** `(h3_prompt, integrated_multimodal_description, overall_soundscape, non_diegetic_music, model_used)` β€” the full prompt plus each field split out for separate wiring. + +--- + +#### APNext H3 Reference Prompt Writer +**Display Name:** `APNext H3 Reference Prompt Writer` + +Writes the six-section full-reference format per [VIDEO_PROMPT_WRITING_GUIDE_ref_en.md](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md). Shares every option above, plus: + +| Input | Description | +|-------|-------------| +| `task_type` | The `[bracketed]` summary prefix: `keyframe completion`, `reference generation`, `video editing`, `video continuation`, `audio reuse`, `audio reference`. Auto lets the model combine them with ` + ` | +| `reference_role` | How attached images get labelled: auto, ``, standalone ``, style-only, or storyboard | +| `word_target` | Target length of `detailed_description` (guide recommends 350–500) | +| `image_1` … `image_4` *(optional)* | Up to four reference images, in label order | +| `reference_notes` *(optional)* | Per-reference notes, one per line β€” also how you describe video/audio references you can't attach | + +**Returns:** `(h3_prompt, subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music, model_used)` + +--- + +#### The `wildness` slider + +One dial from conservative to unhinged. Above 40 it also injects concrete surreal **events** (not mood words) drawn from a 40-entry pool β€” selection is driven by `seed`, so the same seed gives the same weirdness. + +| Range | Band | Behaviour | Random elements | +|-------|------|-----------|-----------------| +| 0–15 | Conservative | Strictly literal, no invented events | 0 | +| 16–40 | Grounded | Believable, well-directed embellishment | 0 | +| 41–65 | Bold | Strong authorial choices, physics still holds | 1 | +| 66–85 | Wild | Surreal juxtapositions, dreamlike logic | 2 | +| 86–100 | Unhinged | Scale, gravity and continuity all negotiable | 3 | + +Injected elements are filmable, e.g. *"the subject's shadow moves a beat out of sync"*, *"a doorway opens onto a completely different biome"*, *"rain falls upward into the sky"*. + +--- + ### 🎲 Prompt Generators #### Auto Prompter @@ -815,21 +922,23 @@ Example workflows are available in the `examples/` directory: ## πŸ“‹ Requirements ``` -Pillow>=10.4.0 -requests>=2.32.5 -openai>=1.44.0 -blend-modes>=2.1.0 -huggingface_hub>=0.34.0 -color_matcher>=0.5.0 +openai>=2.54.0,<3.0.0 +anthropic>=0.121.0 +google-genai>=2.18.0 +httpx>=0.28.1 +huggingface_hub[hf_xet]>=0.34.0 chardet>=5.2.0 -google-generativeai>=0.7.2 -anthropic -transformers>=4.40.0 -decord>=0.6.0 -scipy>=1.10.0 -tqdm>=4.67.1 ``` +Anything ComfyUI already ships in its own `requirements.txt` β€” `Pillow`, `requests`, `transformers`, `scipy`, `tqdm`, `numpy`, `torch` β€” is deliberately **not** repeated, since re-pinning it only risks downgrading the base install. + +Two constraints worth knowing about: + +- **`openai` is capped below 3.0.** v3 switched to HTTPX2 and stopped shipping `httpx`; the GPT/Grok/Groq nodes pass an `httpx.Client` as `http_client=`, which v3 rejects. +- **Gemini uses `google-genai`, not `google-generativeai`.** The legacy SDK hard-pinned `google-ai-generativelanguage==0.6.15`, which forced `protobuf<6` and dragged grpcio into the ComfyUI environment. The current SDK needs neither. + +`decord` is listed but commented out: it is unmaintained and not numpy-2 safe, and the QwenVL/MiniCPM video nodes fall back to OpenCV automatically. Uncomment it in `requirements.txt` if you specifically want decord-based frame decoding. + --- ## πŸ”„ Model Support Matrix diff --git a/__init__.py b/__init__.py index bdca84e..9b1eb10 100644 --- a/__init__.py +++ b/__init__.py @@ -381,6 +381,30 @@ try: except Exception: pass +# Resolution Planning Nodes +try: + # H3 Resolution Planner - original node and algorithm by gabbo + from .nodes.resolution.h3_resolution_planner import H3ResolutionPlannerCropOnly + NEW_MAPPINGS["H3ResolutionPlannerCropOnly"] = H3ResolutionPlannerCropOnly + NEW_DISPLAY_MAPPINGS["H3ResolutionPlannerCropOnly"] = "APNext H3 Resolution Planner (Crop Only) - by gabbo" +except Exception: + pass + +# MiniMax-H3 Prompt Nodes +try: + from .nodes.h3.base_prompt_writer import H3BasePromptWriter + NEW_MAPPINGS["H3BasePromptWriter"] = H3BasePromptWriter + NEW_DISPLAY_MAPPINGS["H3BasePromptWriter"] = "APNext H3 Prompt Writer" +except Exception: + pass + +try: + from .nodes.h3.ref_prompt_writer import H3RefPromptWriter + NEW_MAPPINGS["H3RefPromptWriter"] = H3RefPromptWriter + NEW_DISPLAY_MAPPINGS["H3RefPromptWriter"] = "APNext H3 Reference Prompt Writer" +except Exception: + pass + # Combine mappings (modular nodes + dynamic nodes) NODE_CLASS_MAPPINGS = {**NEW_MAPPINGS, **DYNAMIC_MAPPINGS} NODE_DISPLAY_NAME_MAPPINGS = {**NEW_DISPLAY_MAPPINGS, **DYNAMIC_DISPLAY_MAPPINGS} diff --git a/data/h3/guide_base_en.md b/data/h3/guide_base_en.md new file mode 100644 index 0000000..40cf586 --- /dev/null +++ b/data/h3/guide_base_en.md @@ -0,0 +1,222 @@ +# Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA) + +## 1. Task Overview + +- **T2VA**: Builds a complete audiovisual timeline from text. +- **I2VA**: T2VA body + first-frame instruction + a visual path that develops forward from the first frame. +- **FL2VA**: T2VA body + first-and-last-frame instruction + a continuous path from the first frame to the last frame. +- **L2VA**: T2VA body + last-frame instruction + a path that converges from a plausible preceding state to the last frame. + +## 2. Final Prompt Structure + +### 2.1 Part One Is the Instruction + +**T2VA** has no image-alignment instruction and begins directly with the three core fields. + +**I2VA** always uses: + +```text +For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced. +``` + +**FL2VA** always uses: + +```text +How the reference pictures align with the target video β€” Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video. +``` + +**L2VA** always uses: + +```text +How the reference pictures align with the target video β€” (from [Shot N]) aligns with the S.SS-second mark of the target video. +``` + +Here, `N` is the index of the actual final shot, and `S.SS` is the effective video duration formatted to exactly two decimal places. The instruction must be the first line of the final prompt, followed by one blank line before the core fields. + +### 2.2 Part Two Contains the Three Core Fields + +```text +integrated_multimodal_description: [Shot 1] ... + +overall_soundscape: ... + +non_diegetic_music: ... +``` + +- **integrated_multimodal_description**: Describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline. +- **overall_soundscape**: Summarizes ambient sound, physical action sounds, and non-verbal human sounds across the entire video. +- **non_diegetic_music**: Describes background music that the characters cannot hear and only the audience can hear. + +## 3. How to Incorporate Keyframes into the Multimodal Description + +### 3.1 I2VA: Begin from the Image and Develop Forward + +`` is the actual first frame of the video at 0.00 seconds and belongs to `[Shot 1]`. The description should first establish the style, subjects, composition, and scene anchors in the image, then describe the next action. Character identity, clothing, colors, key objects, and spatial relationships should remain consistent. + +Recommended structure: **first-frame anchor β†’ action onset β†’ continuous development β†’ result or reaction**. + +### 3.2 FL2VA: Describe the Path Between the First and Last Frames + +Picture 1 is the opening, and Picture 2 is the ending. Focus on how the subject moves, how poses change, how objects are manipulated, how the composition evolves, and how the scene or lighting transitions. + +FL2VA generally favors a single shot so the model can interpolate continuously from the first frame to the last frame. Use multiple shots only when they are explicitly specified. The last frame must be reached by the final `[Shot N]` at the end of the video. + +Recommended structure: **first-frame state β†’ observable intermediate changes β†’ progressively narrowing differences β†’ last-frame state**. + +### 3.3 L2VA: Infer the Opening and Land on the Image at the End + +`` is the final frame of the video and belongs to the last `[Shot N]`; it does not inherently belong to Shot 1. Infer a plausible earlier state from the user's intent and the last frame, then describe how the characters, objects, camera, and scene gradually approach the reference image. + +Recommended structure: **plausible preceding state β†’ explicit action and transition path β†’ gradual convergence in the final shot β†’ last-frame landing**. + +## 4. How to Write the Three Shared Core Sections + +### 4.1 Develop the Multimodal Description Along the Timeline + +`integrated_multimodal_description` is the main body of the rewritten prompt. Every detail should correspond to something visible or audible: visual style, initial composition, subject appearance and position, scene and key props, actions and reactions, shot changes, spoken language, and synchronized diegetic sound. + +At the beginning of `[Shot 1]`, state the overall style and initial composition. Common styles include `Cinematic`, `live-action`, `2D-animated`, `3D CG`, `claymation`, `watercolor`, and `vintage film`. For keyframe tasks, derive the style from the reference image; for T2VA, select it from the user's text. + +```text +[Shot 1] Live-action, cinematic, a medium-wide shot frames... +``` + +### 4.2 Shots and Cuts + +Do not add a timestamp to the first shot. Use sequential shot numbers for later shots, and begin each one with a strictly increasing cut time that falls within the video duration: + +```text +[Shot 2] At 00:03.500, the camera cuts to... +``` + +For ordinary cuts, use `the camera cuts to`, `the shot cuts to`, `the shot transitions to`, `the shot changes to`, or `the shot switches to`. When explicitly requested by the user, cross-dissolve, fade, or wipe may also be used. A cut should introduce new information about the subject, space, state, viewpoint, or time. If only the distance or a slight angle needs to change, prefer camera motion. + +### 4.3 Camera Motion: Motion Type + Amplitude + Speed + +A complete camera-motion expression has three dimensions: the **motion type** defines how the camera moves, **amplitude** defines the range of compositional change, and **speed** defines the pacing of that change. Add amplitude and speed only when they are meaningful; medium amplitude and normal speed are usually omitted. + +| Dimension | Available Expression | Description | +|-|-|-| +| Motion type | `Zoom In / Zoom Out` | The focal length changes while the camera body remains stationary | +| Motion type | `Push In / Pull Out` | The camera moves forward / backward | +| Motion type | `Pan Left / Pan Right` | The camera remains in place while the lens pivots horizontally | +| Motion type | `Truck Left / Truck Right` | The camera translates horizontally | +| Motion type | `Tilt Up / Tilt Down` | The camera remains in place while the lens pivots vertically | +| Motion type | `Pedestal Up / Pedestal Down` | The entire camera moves upward / downward | +| Motion type | `Arc Shot` | The camera moves in an arc around the subject | +| Motion type | `Tracking Shot` | The camera follows a moving subject | +| Motion type | `Static Shot` | The camera position and lens remain still | +| Motion type | `Shake Slightly / Shake Strongly` | Slight / strong camera shake | +| Motion type | `POV` | The subject's point of view | +| Motion type | `Roll Clockwise / Roll Counterclockwise` | The camera rolls clockwise / counterclockwise around the lens axis | +| Amplitude | `with small amplitude` | Small-range change | +| Amplitude | `with large amplitude` | Large-range change | +| Speed | `at slow speed` | Slow movement | +| Speed | `at fast speed` | Fast movement | + +Camera motion should be written as a natural English action within the shot, rather than stacked as separate labels at the end of a sentence: + +```text +The camera pushes in with small amplitude at slow speed toward the folded letter in her hands. +The camera pans right with large amplitude at fast speed, revealing the open doorway. +The camera holds a static shot as the runner exits the frame. +``` + +### 4.4 Speakers, Dialogue, and Singing + +Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as `(S1)` and `(S2)`. When multiple already-numbered speakers speak or sing together, use a compound ID such as `(S1,S2)`. A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID. + +When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent. Place the speaker's identifying phrase, ID, action, and delivery outside ``. Inside ``, include only the language tag and the actual user-provided spoken content. Preserve every original word and punctuation mark verbatim; do not translate or rewrite them. + +```text +The young woman with a quiet, breathy voice (S1) says: [English] I get off at the next station. +The two children (S1,S2) shout together, [English] Wait for us! +``` + +For voiceover, use the exact phrase `says in an off-screen voiceover`. Immediately after every voiceover `` block, state that the corresponding on-screen character's lips remain closed: + +```text +The man (S1) says in an off-screen voiceover: [English] I still remember that road. while his lips remain completely closed. +``` + +When the same line of dialogue or lyrics crosses a cut, use `` at the connecting points in both parts and explicitly state that the audio continues across the cut. Use `` when speech is truncated by the end of the video. Continuity may be expressed with `continues seamlessly across the cut`, `continues uninterrupted into the next shot`, `carries over from the previous shot`, or `remains audible across the transition`. + +### 4.5 On-Screen Text + +Place any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks. Preserve the original text and punctuation verbatim, without translation. + +```text +A red neon sign reading "θ₯业中" glows above the doorway. +``` + +### 4.6 overall_soundscape + +Use 1–4 English sentences in one continuous paragraph to summarize the ambient sound, physical action sounds, and non-verbal human sounds across the full video, such as wind, rain, traffic, footsteps, fabric movement, impacts, breathing, laughter, or panting. Dialogue, singing, and diegetic music already belong in the multimodal description and should not be repeated here. Use `N/A` only when the user explicitly requests complete silence throughout the video. + +```text +overall_soundscape: Steady rain taps against the cafΓ© windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair. +``` + +### 4.7 non_diegetic_music + +Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear. Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score. Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description. Use `N/A` when there is no non-diegetic music. + +```text +non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out. +``` + +## 5. Cases + +### Case 1: T2VA + +With no reference image, construct the complete timeline directly from the text. You may add scene, character, action, and sound details that remain consistent with the user's intent. + +```text +integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: [English] First batch of the morning. [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot. + +overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced. + +non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end. +``` + +### Case 2: I2VA + +Write the first-frame instruction first, then use the subject, composition, and scene in Picture 1 as the starting point of Shot 1 before describing how the scene continues to develop. + +```text +For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced. + +integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: [English] I get off at the next station. She folds the letter along its existing crease. + +overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands. + +non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume. +``` + +### Case 3: FL2VA + +The two images anchor the opening and ending respectively. The body should not repeat two static image descriptions; instead, it should supply the motion path that connects them. The following example is an eight-second single shot. + +```text +How the reference pictures align with the target video β€” Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video. + +integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot. + +overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes. + +non_diegetic_music: N/A +``` + +### Case 4: L2VA + +The image anchors only the final moment. First establish a compatible earlier state, then let the actions, object states, and composition gradually land on Picture 1 in the final shot. The following example is a six-second single shot. + +```text +How the reference pictures align with the target video β€” (from [Shot 1]) aligns with the 6.00-second mark of the target video. + +integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by . + +overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor. + +non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks. +``` diff --git a/data/h3/guide_ref_en.md b/data/h3/guide_ref_en.md new file mode 100644 index 0000000..7ae1b2d --- /dev/null +++ b/data/h3/guide_ref_en.md @@ -0,0 +1,341 @@ +# Full-Reference Mode Rewrite Output Format Guide + +This guide explains how rewrite outputs are organized and written in full-reference mode. + +Write all six rewrite sections in English. Preserve the original language only for dialogue and lyrics inside `` and for text visibly present in the scene. + +**Description detail:** Make `detailed_description` as detailed and explicit as possible. For each shot, clearly establish the current composition, subject appearance and position, environment and lighting, actions and state changes, camera movement, current sound, and the points where referenced content actually appears or takes effect. Avoid reducing the description to a plot summary or a list of reference relationships. + +> The basic formats for shots, camera movement, speakers, dialogue, and ordinary sound are shared with the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA). This guide focuses on the reference labels, analysis sections, and format differences specific to full-reference mode. + +## 1. Overall Structure + +A complete rewrite output consists of six sections in the following order: + +| Section | Purpose | +| --- | --- | +| `subject_definitions` | Defines referenced content and its reference labels | +| `summary` | Summarizes the task type, target video, and main reference relationships | +| `retention_analysis` | Describes how referenced content is preserved, transferred, or reused | +| `detailed_description` | Describes visuals, actions, shots, sound, and dialogue in playback order | +| `overall_soundscape` | Summarizes ambience and physical sounds | +| `non_diegetic_music` | Describes background music audible only to the audience | + +## 2. Reference Labels and Definitions (`subject_definitions`) + +Full-reference rewrites use four types of labels to identify the source and role of referenced content: + +| Label | Meaning | +| --- | --- | +| `` | Visible content abstracted from reference assets that can be reused or modified in the target video | +| `` | A reference image used as a concrete target frame or shot-planning anchor | +| `