Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
ff7d6d7607 | ||
|
|
97b631d561 |
@@ -0,0 +1,194 @@
|
||||
# Auto-Facing Feature Documentation
|
||||
|
||||
## Overview
|
||||
|
||||
The `auto_facing` parameter ensures the camera automatically points directly at the target subject from any horizontal angle position. This feature is now available in both **Object Focus Camera v7** and **Cinematography Prompt Builder**.
|
||||
|
||||
---
|
||||
|
||||
## Purpose
|
||||
|
||||
When positioning the camera at angles (left, right, side, back), `auto_facing` controls whether the camera:
|
||||
- ✅ **Points directly at the subject** (auto_facing = True)
|
||||
- ❌ **Maintains forward orientation** without explicitly facing the subject (auto_facing = False)
|
||||
|
||||
---
|
||||
|
||||
## Implementation Details
|
||||
|
||||
### Parameter Specification
|
||||
|
||||
```python
|
||||
"auto_facing": ("BOOLEAN", {
|
||||
"default": True,
|
||||
"tooltip": "Automatically face camera toward target subject (recommended for object photography).\n"
|
||||
"• True = Camera points directly at subject from chosen angle\n"
|
||||
"• False = Camera positioned at angle but may not face subject directly"
|
||||
})
|
||||
```
|
||||
|
||||
### Prompt Positioning Strategy
|
||||
|
||||
**Key Finding**: Based on user experience with vision-language models, placing `auto_facing` guidance **at the beginning of the prompt** provides maximum attention weight and effectiveness.
|
||||
|
||||
**Prompt Structure:**
|
||||
|
||||
```
|
||||
[FACING DIRECTIVE] + [Main Camera Prompt] + [Details]
|
||||
```
|
||||
|
||||
**Examples:**
|
||||
|
||||
#### Simple Prompt (English):
|
||||
```
|
||||
Facing the dishwasher directly, An eye-level medium shot of the dishwasher, taken from a vantage point two meters away, positioned from thirty degrees to the left for a corner perspective, with deep depth of field keeping everything in focus, in architectural style
|
||||
```
|
||||
|
||||
#### Professional Prompt (Chinese):
|
||||
```
|
||||
面对dishwasher,Next Scene: 将镜头转为标准镜头(50mm),中景构图,平视查看dishwasher,从左侧30度拍摄,呈现转角视角,距离两米
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## When Auto-Facing Is Applied
|
||||
|
||||
### ✅ Active Conditions:
|
||||
- `auto_facing = True` (default)
|
||||
- `horizontal_angle != "Front View (0°)"` (since front view already implies facing)
|
||||
|
||||
### ❌ Not Applied When:
|
||||
- `auto_facing = False`
|
||||
- `horizontal_angle = "Front View (0°)"` (redundant - front view inherently faces subject)
|
||||
|
||||
---
|
||||
|
||||
## Usage Examples
|
||||
|
||||
### Example 1: Dishwasher Side View with Auto-Facing
|
||||
|
||||
**Settings:**
|
||||
- Target Subject: `dishwasher`
|
||||
- Shot Type: `Medium Shot (MS)`
|
||||
- Camera Angle: `Eye Level`
|
||||
- Horizontal Angle: `Side Left (90°)`
|
||||
- **auto_facing: `True`** ✅
|
||||
|
||||
**Result:**
|
||||
Camera positions at the left side (90°) AND rotates to face the dishwasher directly, ensuring the dishwasher is centered in frame despite the side positioning.
|
||||
|
||||
---
|
||||
|
||||
### Example 2: Architectural Context Shot without Auto-Facing
|
||||
|
||||
**Settings:**
|
||||
- Target Subject: `kitchen counter`
|
||||
- Shot Type: `Wide Shot (WS)`
|
||||
- Camera Angle: `Eye Level`
|
||||
- Horizontal Angle: `Angled Right 30°`
|
||||
- **auto_facing: `False`** ❌
|
||||
|
||||
**Result:**
|
||||
Camera positions at 30° to the right but maintains forward orientation, potentially showing the counter as part of a broader environmental context rather than centered.
|
||||
|
||||
---
|
||||
|
||||
## Technical Implementation
|
||||
|
||||
### Cinematography Prompt Builder
|
||||
|
||||
#### Simple Prompt Generation ([cinematography_prompt_builder.py:685-688](nodes/camera/cinematography_prompt_builder.py#L685-L688)):
|
||||
|
||||
```python
|
||||
# AUTO-FACING: Add at the VERY BEGINNING for maximum attention weight
|
||||
# Only add if enabled AND not front view (front view already implies facing)
|
||||
if auto_facing and horizontal_angle != "Front View (0°)":
|
||||
parts.append(f"Facing {subject} directly")
|
||||
```
|
||||
|
||||
#### Professional Prompt Generation ([cinematography_prompt_builder.py:757-763](nodes/camera/cinematography_prompt_builder.py#L757-L763)):
|
||||
|
||||
```python
|
||||
# AUTO-FACING: Add at BEGINNING for maximum attention (before "Next Scene:")
|
||||
# Only add if enabled AND not front view
|
||||
if auto_facing and horizontal_angle != "Front View (0°)":
|
||||
if language in ["Chinese (Best for dx8152 LoRAs)", "Hybrid (Chinese + English)"]:
|
||||
prompt_parts.append(f"面对{subject}") # "Facing {subject}"
|
||||
else:
|
||||
prompt_parts.append(f"Facing {subject} directly")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Why Positioning Matters
|
||||
|
||||
### User Observation:
|
||||
> "i know it is important if you merg it to prompt at begiing it will have more affect base on my experince"
|
||||
|
||||
This aligns with attention mechanisms in transformer-based vision-language models:
|
||||
|
||||
1. **Positional Bias**: Tokens at the beginning of prompts receive higher attention weights
|
||||
2. **Semantic Anchoring**: Early instructions establish the primary directive for the generation
|
||||
3. **Context Precedence**: Models process sequential information with recency and primacy effects
|
||||
|
||||
By placing `auto_facing` directive **first**, we ensure maximum model attention to this critical orientation instruction.
|
||||
|
||||
---
|
||||
|
||||
## Integration with Other Features
|
||||
|
||||
### Compatible with:
|
||||
- ✅ All horizontal angles (15°, 30°, 45°, 90°, 180°)
|
||||
- ✅ All vertical camera angles (Eye Level, High Angle, Low Angle, etc.)
|
||||
- ✅ All shot sizes (ECU to EWS)
|
||||
- ✅ Perspective correction modes (Natural, Architectural, Tilt-Shift)
|
||||
- ✅ All lens types
|
||||
- ✅ Chinese/English/Hybrid language modes
|
||||
|
||||
### Automatically Disabled:
|
||||
- Front View (0°) - redundant since front view inherently faces subject
|
||||
- When explicitly disabled by user (`auto_facing = False`)
|
||||
|
||||
---
|
||||
|
||||
## Practical Use Cases
|
||||
|
||||
### 🎯 Object Photography (Recommended: True)
|
||||
- Product photography requiring subject prominence
|
||||
- Furniture visualization from multiple angles
|
||||
- Appliance close-ups (dishwashers, ovens, refrigerators)
|
||||
- Detail shots of architectural elements
|
||||
|
||||
### 🏛️ Environmental Photography (Consider: False)
|
||||
- Architectural context shots
|
||||
- Room overview with subject as part of environment
|
||||
- Documentary-style environmental capture
|
||||
- Spatial relationship emphasis over subject focus
|
||||
|
||||
---
|
||||
|
||||
## Version History
|
||||
|
||||
- **v2.4.1** (2025-01-07): Added `auto_facing` to Cinematography Prompt Builder
|
||||
- Placed at beginning of prompts for maximum attention weight
|
||||
- Full Chinese translation support (面对)
|
||||
- Automatic disable for Front View (0°)
|
||||
|
||||
- **v2.3.0** (2025-01-06): Original implementation in Object Focus Camera v7
|
||||
- Vantage point mode support
|
||||
- Boolean toggle for camera orientation control
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- User feedback: Prompt positioning significantly affects model attention
|
||||
- Vision-language model research: Positional encoding and attention weights
|
||||
- Object Focus Camera v7: Original auto_facing implementation
|
||||
|
||||
---
|
||||
|
||||
**Author**: Amir Ferdos (ArchAi3d)
|
||||
**Feature Version**: v2.4.1
|
||||
**Implementation Date**: 2025-01-07
|
||||
**Based on**: User experience and vision-language model attention mechanisms
|
||||
@@ -0,0 +1,253 @@
|
||||
# Auto-Facing Feature - Test Results
|
||||
|
||||
## ✅ All Tests Passing!
|
||||
|
||||
Date: 2025-01-07
|
||||
Feature Version: v2.4.1
|
||||
|
||||
---
|
||||
|
||||
## Test Summary
|
||||
|
||||
All 6 tests **PASSED** ✅
|
||||
|
||||
### What Was Fixed:
|
||||
|
||||
1. **Auto-Facing Parameter Added** - Now available in Cinematography Prompt Builder
|
||||
2. **Early Prompt Positioning** - "Facing" clause placed at the BEGINNING for maximum attention weight
|
||||
3. **English Mode Bug Fixed** - Professional English prompts now correctly include auto_facing
|
||||
4. **Distance Chinese Fixed** - Changed from "距离远距离" to "距离四米" (specific meters instead of generic descriptions)
|
||||
|
||||
---
|
||||
|
||||
## Test Results
|
||||
|
||||
### TEST 1: Front View (0°) with auto_facing=True
|
||||
**Status:** ✅ PASS
|
||||
|
||||
**Prompt:**
|
||||
```
|
||||
Next Scene: 将镜头转为标准镜头(50mm),全景构图,平视查看the refrigerator,距离四米半
|
||||
```
|
||||
|
||||
**✅ Correct:** NO "面对" clause (front view already implies facing)
|
||||
|
||||
---
|
||||
|
||||
### TEST 2: Angled Left 30° with auto_facing=True
|
||||
**Status:** ✅ PASS
|
||||
|
||||
**Prompt:**
|
||||
```
|
||||
面对the refrigerator Next Scene: 将镜头转为标准镜头(50mm),全景构图,平视查看the refrigerator,从左侧30度拍摄,呈现转角视角,距离四米半
|
||||
```
|
||||
|
||||
**✅ Correct:**
|
||||
- "面对the refrigerator" at the BEGINNING
|
||||
- Specific distance: "距离四米半" (distance 4.5 meters)
|
||||
- Horizontal angle description included
|
||||
|
||||
---
|
||||
|
||||
### TEST 3: Side Right (90°) with auto_facing=True
|
||||
**Status:** ✅ PASS
|
||||
|
||||
**Prompt:**
|
||||
```
|
||||
面对the refrigerator Next Scene: 将镜头转为标准镜头(50mm),中景构图,平视查看the refrigerator,从右侧拍摄,呈现侧面视角,距离两米半
|
||||
```
|
||||
|
||||
**✅ Correct:**
|
||||
- "面对the refrigerator" at the BEGINNING
|
||||
- Side view angle properly described
|
||||
- Specific distance: "距离两米半" (distance 2.5 meters)
|
||||
|
||||
---
|
||||
|
||||
### TEST 4: Angled Right 45° with auto_facing=False
|
||||
**Status:** ✅ PASS
|
||||
|
||||
**Prompt:**
|
||||
```
|
||||
Next Scene: 将镜头转为标准镜头(50mm),中景构图,平视查看the refrigerator,从右侧45度拍摄,呈现四分之三视角,距离两米半
|
||||
```
|
||||
|
||||
**✅ Correct:** NO "面对" clause (disabled by user)
|
||||
|
||||
---
|
||||
|
||||
### TEST 5: Angled Left 45° with auto_facing=True (English mode)
|
||||
**Status:** ✅ PASS
|
||||
|
||||
**Professional Prompt:**
|
||||
```
|
||||
Facing the refrigerator directly, Next Scene:, Change to Normal (50mm), MS framing, Eye Level viewing the refrigerator, positioned from forty-five degrees to the left for a three-quarter view
|
||||
```
|
||||
|
||||
**Simple Prompt:**
|
||||
```
|
||||
Facing the refrigerator directly, An eye-level medium shot of the refrigerator, taken from a vantage point two and a half meters away, positioned from forty-five degrees to the left for a three-quarter view, with medium depth of field
|
||||
```
|
||||
|
||||
**✅ Correct:**
|
||||
- Both prompts start with "Facing the refrigerator directly"
|
||||
- English professional prompt now works (bug fixed!)
|
||||
- Simple prompt already worked correctly
|
||||
|
||||
---
|
||||
|
||||
### TEST 6: Side Left (90°) with auto_facing=True (Hybrid mode)
|
||||
**Status:** ✅ PASS
|
||||
|
||||
**Prompt:**
|
||||
```
|
||||
面对the refrigerator Next Scene: 将镜头转为人像镜头(85mm),近景构图,平视查看the refrigerator,从左侧拍摄,呈现侧面视角,距离零点八米
|
||||
```
|
||||
|
||||
**✅ Correct:**
|
||||
- "面对the refrigerator" at the BEGINNING
|
||||
- Hybrid mode works perfectly (Chinese cinematography terms + English subject)
|
||||
- Specific distance: "距离零点八米" (distance 0.8 meters)
|
||||
|
||||
---
|
||||
|
||||
## Key Improvements
|
||||
|
||||
### 1. Auto-Facing Placement
|
||||
**Before:** Not available in Cinematography Prompt Builder
|
||||
**After:** Added at the BEGINNING of prompts for maximum attention weight
|
||||
|
||||
**User Insight:** "i know it is important if you merg it to prompt at begiing it will have more affect base on my experince"
|
||||
|
||||
This placement leverages positional bias in vision-language models.
|
||||
|
||||
---
|
||||
|
||||
### 2. Distance Chinese Precision
|
||||
|
||||
**Before:**
|
||||
```
|
||||
距离远距离 (distance far distance) ❌ Generic, redundant
|
||||
距离中等距离 (distance medium distance) ❌ Vague
|
||||
```
|
||||
|
||||
**After:**
|
||||
```
|
||||
距离四米 (distance 4 meters) ✅ Specific
|
||||
距离两米半 (distance 2.5 meters) ✅ Precise with half meters
|
||||
距离零点八米 (distance 0.8 meters) ✅ Handles decimals
|
||||
```
|
||||
|
||||
**Chinese Number Mapping:**
|
||||
- Whole numbers: 一米, 两米, 三米, 四米, etc.
|
||||
- Half meters: 半米, 一米半, 两米半, etc.
|
||||
- Decimals: 零点八米, 两点五米, etc.
|
||||
|
||||
---
|
||||
|
||||
### 3. English Mode Bug Fix
|
||||
|
||||
**Issue:** Professional English prompts were bypassing the auto_facing logic
|
||||
|
||||
**Before:**
|
||||
```
|
||||
Next Scene: Change to Normal (50mm), MS framing... ❌ Missing "Facing" clause
|
||||
```
|
||||
|
||||
**After:**
|
||||
```
|
||||
Facing the refrigerator directly, Next Scene:, Change to Normal (50mm), MS framing... ✅
|
||||
```
|
||||
|
||||
**Fix:** Updated English mode code path to include `prompt_parts` with auto_facing directive
|
||||
|
||||
---
|
||||
|
||||
## Auto-Facing Logic
|
||||
|
||||
### When Active:
|
||||
- ✅ `auto_facing = True` (default)
|
||||
- ✅ `horizontal_angle != "Front View (0°)"`
|
||||
|
||||
### When Inactive:
|
||||
- ❌ `auto_facing = False` (user disabled)
|
||||
- ❌ `horizontal_angle = "Front View (0°)"` (redundant - front view already faces subject)
|
||||
|
||||
---
|
||||
|
||||
## Language Support
|
||||
|
||||
### Chinese Mode:
|
||||
```
|
||||
面对{subject} Next Scene: ...
|
||||
```
|
||||
|
||||
### English Mode:
|
||||
```
|
||||
Facing {subject} directly, [prompt]...
|
||||
```
|
||||
|
||||
### Hybrid Mode:
|
||||
```
|
||||
面对{subject} Next Scene: ... (Chinese cinematography + English details)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Integration Status
|
||||
|
||||
✅ **Cinematography Prompt Builder** - Fully integrated
|
||||
✅ **Object Focus Camera v7** - Already had auto_facing
|
||||
✅ **Simple Prompt Generation** - Working
|
||||
✅ **Professional Prompt Generation** - Working (bug fixed)
|
||||
✅ **All Language Modes** - Working (Chinese/English/Hybrid)
|
||||
|
||||
---
|
||||
|
||||
## Files Modified
|
||||
|
||||
1. **cinematography_prompt_builder.py**
|
||||
- Added `auto_facing` parameter (lines 159-165)
|
||||
- Updated function signatures
|
||||
- Fixed `_generate_simple_prompt()` with early auto_facing placement
|
||||
- Fixed `_generate_professional_prompt()` with early auto_facing placement
|
||||
- Fixed English mode code path bug
|
||||
- Improved `_get_distance_chinese()` for specific meter values
|
||||
|
||||
2. **AUTO_FACING_FEATURE.md** - Complete feature documentation
|
||||
3. **test_auto_facing.py** - Comprehensive test suite
|
||||
4. **AUTO_FACING_TEST_RESULTS.md** - This file
|
||||
|
||||
---
|
||||
|
||||
## User Confirmation
|
||||
|
||||
User prompt example:
|
||||
```
|
||||
Next Scene: 将镜头转为标准镜头(50mm),全景构图,平视查看the refrigerator ,距离远距离
|
||||
```
|
||||
|
||||
**Issues identified and fixed:**
|
||||
1. ❌ No auto_facing clause → ✅ "面对" added when using angled views
|
||||
2. ❌ "距离远距离" (distance far distance) → ✅ "距离四米" (distance 4 meters)
|
||||
3. ❌ Mixed language "the refrigerator" → Still present but acceptable for Hybrid mode
|
||||
|
||||
**Recommendations for user:**
|
||||
- Use Chinese subject name "冰箱" OR keep "the refrigerator" (both work)
|
||||
- Select angled horizontal angles (15°, 30°, 45°, 90°) to activate auto_facing
|
||||
- Default `auto_facing = True` ensures camera points at subject
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. ✅ Feature is production-ready
|
||||
2. ✅ All tests passing
|
||||
3. ✅ Documentation complete
|
||||
4. 📝 Ready for CHANGELOG update and version bump to v2.4.1
|
||||
|
||||
---
|
||||
|
||||
**Author:** Amir Ferdos (ArchAi3d)
|
||||
**Test Date:** 2025-01-07
|
||||
**Feature Status:** ✅ PRODUCTION READY
|
||||
@@ -0,0 +1,252 @@
|
||||
# Session Updates - v2.4.1 (2025-01-07)
|
||||
|
||||
## Overview
|
||||
This document summarizes all changes made during the v2.4.1 development session.
|
||||
|
||||
## Package Cleanup
|
||||
**Removed redundant GRAG sampler** - The full GRAG Advanced Sampler is now maintained in the separate [ComfyUI-GRAG-ArchAi3D](https://github.com/amir84ferdos/ComfyUI-GRAG-ArchAi3D) repository. This package retains GRAG utility nodes (GRAG Modifier, GRAG Encoder) for conditioning metadata injection.
|
||||
|
||||
---
|
||||
|
||||
## 1. Auto-Facing Feature Added to Cinematography Prompt Builder
|
||||
|
||||
### What Changed
|
||||
Added `auto_facing` parameter to **Cinematography Prompt Builder** node, previously only available in Object Focus Camera v7.
|
||||
|
||||
### Why Important
|
||||
User insight: "i know it is important if you merg it to prompt at begiing it will have more affect base on my experince"
|
||||
|
||||
Based on vision-language model attention mechanisms, placing the facing directive at the **beginning** of prompts provides maximum attention weight and effectiveness.
|
||||
|
||||
### Implementation Details
|
||||
|
||||
**File**: `nodes/camera/cinematography_prompt_builder.py`
|
||||
|
||||
1. **Added Parameter** (Lines 159-165):
|
||||
```python
|
||||
"auto_facing": ("BOOLEAN", {
|
||||
"default": True,
|
||||
"tooltip": "Automatically face camera toward target subject (recommended for object photography).\n"
|
||||
"• True = Camera points directly at subject from chosen angle\n"
|
||||
"• False = Camera positioned at angle but may not face subject directly"
|
||||
}),
|
||||
```
|
||||
|
||||
2. **Simple Prompt Generation** (Lines 685-688):
|
||||
```python
|
||||
# AUTO-FACING: Add at the VERY BEGINNING for maximum attention weight
|
||||
# Only add if enabled AND not front view (front view already implies facing)
|
||||
if auto_facing and horizontal_angle != "Front View (0°)":
|
||||
parts.append(f"Facing {subject} directly")
|
||||
```
|
||||
|
||||
3. **Professional Prompt Generation** (Lines 757-763):
|
||||
```python
|
||||
# AUTO-FACING: Add at BEGINNING for maximum attention (before "Next Scene:")
|
||||
# Only add if enabled AND not front view
|
||||
if auto_facing and horizontal_angle != "Front View (0°)":
|
||||
if language in ["Chinese (Best for dx8152 LoRAs)", "Hybrid (Chinese + English)"]:
|
||||
prompt_parts.append(f"面对{subject}") # "Facing {subject}"
|
||||
else:
|
||||
prompt_parts.append(f"Facing {subject} directly")
|
||||
```
|
||||
|
||||
### Behavior
|
||||
- **Active**: When `auto_facing=True` AND `horizontal_angle != "Front View (0°)"`
|
||||
- **Inactive**: When `auto_facing=False` OR `horizontal_angle == "Front View (0°)"` (redundant)
|
||||
- **Language Support**: Full Chinese/English/Hybrid support
|
||||
|
||||
---
|
||||
|
||||
## 2. Parameter Order Bug Fix
|
||||
|
||||
### Problem
|
||||
User reported: "i saw it is not working , the auto facing option is not working check it"
|
||||
|
||||
### Root Cause
|
||||
Parameter order mismatch between INPUT_TYPES definition and function signature.
|
||||
|
||||
ComfyUI passes parameters **positionally** based on INPUT_TYPES order. The function signature had parameters in wrong positions.
|
||||
|
||||
**Before**:
|
||||
- INPUT_TYPES position 5: `auto_facing`
|
||||
- Function signature position 8: `auto_facing`
|
||||
|
||||
### Fix
|
||||
Reordered function signature to match INPUT_TYPES exactly (Lines 591-601):
|
||||
```python
|
||||
def generate_cinematography_prompt(self, target_subject, shot_type, camera_angle,
|
||||
horizontal_angle, auto_facing, # CRITICAL: Must match INPUT_TYPES order
|
||||
depth_of_field, style_mood, prompt_language,
|
||||
...)
|
||||
```
|
||||
|
||||
**File**: `nodes/camera/cinematography_prompt_builder.py`
|
||||
|
||||
---
|
||||
|
||||
## 3. Chinese Distance Format Improvement
|
||||
|
||||
### Problem
|
||||
User showed prompt: "距离远距离" (distance far distance) - redundant and unclear
|
||||
|
||||
### Solution
|
||||
Changed `_get_distance_chinese()` function to return specific meter values instead of generic descriptions.
|
||||
|
||||
**Before**: "远距离" (far distance)
|
||||
**After**: "四米" (4 meters)
|
||||
|
||||
### Implementation (Lines 903-943)
|
||||
```python
|
||||
def _get_distance_chinese(self, distance):
|
||||
"""Convert distance to Chinese words with specific meter values"""
|
||||
chinese_numbers = {
|
||||
0: "零", 1: "一", 2: "两", 3: "三", 4: "四",
|
||||
5: "五", 6: "六", 7: "七", 8: "八", 9: "九",
|
||||
10: "十", 15: "十五", 20: "二十"
|
||||
}
|
||||
|
||||
if distance == int(distance):
|
||||
dist_int = int(distance)
|
||||
if dist_int in chinese_numbers:
|
||||
return f"{chinese_numbers[dist_int]}米"
|
||||
else:
|
||||
return f"{dist_int}米"
|
||||
# ... handles half meters and decimals
|
||||
```
|
||||
|
||||
**File**: `nodes/camera/cinematography_prompt_builder.py`
|
||||
|
||||
---
|
||||
|
||||
## 4. GRAG Nodes Fixed for ComfyUI Update
|
||||
|
||||
### Problem
|
||||
User reported: "there is an update for comfyui and t broken my GRAG nodes"
|
||||
|
||||
Error: `RuntimeError: The size of tensor a (8430) must match the size of tensor b (24)`
|
||||
|
||||
### Root Cause
|
||||
ComfyUI commit `4cd881866bad0cde70273cc123d725693c1f2759` changed:
|
||||
- Tensor format: **BSHD → BHND** (Batch, Heads, Sequence, Dim)
|
||||
- RoPE function: `apply_rotary_emb` → `apply_rope1`
|
||||
- Import location: `comfy.ldm.qwen_image.model` → `comfy.ldm.flux.math`
|
||||
|
||||
### Solution Applied
|
||||
|
||||
**File**: `nodes/sampling/archai3d_grag_sampler.py`
|
||||
|
||||
#### 1. QKV Projection Format (Lines 187-195)
|
||||
**Before**:
|
||||
```python
|
||||
img_query = attn_module.to_q(hidden_states).unflatten(-1, (attn_module.heads, -1))
|
||||
```
|
||||
|
||||
**After**:
|
||||
```python
|
||||
img_query = attn_module.to_q(hidden_states).view(batch_size, seq_img, attn_module.heads, -1).transpose(1, 2).contiguous()
|
||||
```
|
||||
|
||||
Changes to BHND format: `[B, H, N, D]`
|
||||
|
||||
#### 2. Concatenation Dimension (Lines 203-206)
|
||||
**Before**: `dim=1` (sequence in BSHD)
|
||||
**After**: `dim=2` (sequence in BHND)
|
||||
|
||||
```python
|
||||
joint_query = torch.cat([txt_query, img_query], dim=2)
|
||||
```
|
||||
|
||||
#### 3. RoPE Function Update (Lines 208-211)
|
||||
**Before**:
|
||||
```python
|
||||
from comfy.ldm.qwen_image.model import apply_rotary_emb
|
||||
joint_query = apply_rotary_emb(joint_query, image_rotary_emb)
|
||||
```
|
||||
|
||||
**After**:
|
||||
```python
|
||||
from comfy.ldm.flux.math import apply_rope1
|
||||
joint_query = apply_rope1(joint_query, image_rotary_emb)
|
||||
```
|
||||
|
||||
#### 4. GRAG Processing Format Conversion (Lines 216-232)
|
||||
```python
|
||||
# Convert BHND to BSHD format for GRAG, then flatten
|
||||
# BHND: [B, H, S, D] -> BSHD: [B, S, H, D] -> [B, S, H*D]
|
||||
joint_key_for_grag = joint_key.transpose(1, 2).contiguous() # BHND -> BSHD
|
||||
joint_key_flat = joint_key_for_grag.flatten(start_dim=2) # [B, S, H*D]
|
||||
|
||||
# Apply GRAG reweighting
|
||||
joint_key_flat = apply_grag_to_keys(...)
|
||||
|
||||
# Unflatten back to BSHD then transpose back to BHND
|
||||
joint_key_for_grag = joint_key_flat.unflatten(-1, (attn_module.heads, -1)) # [B, S, H, D]
|
||||
joint_key = joint_key_for_grag.transpose(1, 2).contiguous() # BSHD -> BHND
|
||||
```
|
||||
|
||||
#### 5. Attention Call with skip_reshape (Lines 241-252)
|
||||
**Key Insight**: With `skip_reshape=True` and default `skip_output_reshape=False`:
|
||||
- **Input**: BHND format
|
||||
- **Output**: BSD format (not BHND!)
|
||||
|
||||
```python
|
||||
# Pass tensors in BHND format with skip_reshape=True (new Qwen format)
|
||||
# Output will be BSD format (batch, seq, heads*dim) due to default skip_output_reshape=False
|
||||
joint_hidden_states = optimized_attention_masked(
|
||||
joint_query, joint_key, joint_value, attn_module.heads,
|
||||
attention_mask, transformer_options=transformer_options,
|
||||
skip_reshape=True # Input is BHND, output is BSD (due to default reshape)
|
||||
)
|
||||
|
||||
# Split streams - output is already in BSD format, no transpose needed
|
||||
txt_attn_output = joint_hidden_states[:, :seq_txt, :]
|
||||
img_attn_output = joint_hidden_states[:, seq_txt:, :]
|
||||
```
|
||||
|
||||
**Critical Fix**: Removed incorrect transpose that was treating output as BHND when it's actually BSD.
|
||||
|
||||
### Testing
|
||||
User confirmed: "ok GRAG is working"
|
||||
|
||||
---
|
||||
|
||||
## Files Modified
|
||||
|
||||
1. **nodes/camera/cinematography_prompt_builder.py**
|
||||
- Added auto_facing parameter
|
||||
- Fixed parameter order
|
||||
- Improved Chinese distance formatting
|
||||
- Lines: 159-165, 591-601, 685-688, 757-763, 903-943
|
||||
|
||||
2. **nodes/sampling/archai3d_grag_sampler.py**
|
||||
- Complete GRAG tensor format refactor for ComfyUI update
|
||||
- Lines: 183-252 (entire attention forward pass)
|
||||
|
||||
---
|
||||
|
||||
## Documentation Created
|
||||
|
||||
1. **AUTO_FACING_FEATURE.md** - Complete auto_facing documentation
|
||||
2. **SESSION_UPDATES_v2.4.1.md** - This file
|
||||
|
||||
---
|
||||
|
||||
## Version
|
||||
- **Version**: v2.4.1
|
||||
- **Date**: 2025-01-07
|
||||
- **Branch**: main
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
User should:
|
||||
1. Test auto_facing feature in ComfyUI workflows
|
||||
2. Test GRAG sampler with latest ComfyUI
|
||||
3. Consider updating version in `__init__.py` and `pyproject.toml` if releasing
|
||||
|
||||
---
|
||||
|
||||
**Author**: Amir Ferdos (ArchAi3d)
|
||||
**Assisted by**: Claude Code (Anthropic)
|
||||
+4
-15
@@ -6,7 +6,7 @@ Author: Amir Ferdos (ArchAi3d)
|
||||
Email: Amir84ferdos@gmail.com
|
||||
LinkedIn: https://www.linkedin.com/in/archai3d/
|
||||
GitHub: https://github.com/amir84ferdos
|
||||
Version: 2.4.0
|
||||
Version: 2.4.1
|
||||
License: Dual License (Free for personal use, Commercial license required for business use)
|
||||
"""
|
||||
|
||||
@@ -27,12 +27,6 @@ from .nodes.core.utils.archai3d_grag_modifier import ArchAi3D_GRAG_Modifier
|
||||
from .nodes.core.prompts.archai3d_clean_room_prompt import ArchAi3D_Clean_Room_Prompt
|
||||
from .nodes.core.prompts.archai3d_qwen_system_prompt import ArchAi3D_Qwen_System_Prompt
|
||||
|
||||
# ============================================================================
|
||||
# SAMPLING NODES
|
||||
# ============================================================================
|
||||
|
||||
from .nodes.sampling.archai3d_grag_sampler import ArchAi3D_GRAG_Sampler
|
||||
|
||||
# ============================================================================
|
||||
# CAMERA CONTROL NODES
|
||||
# ============================================================================
|
||||
@@ -138,9 +132,6 @@ NODE_CLASS_MAPPINGS = {
|
||||
# Core - Prompts
|
||||
"ArchAi3D_Clean_Room_Prompt": ArchAi3D_Clean_Room_Prompt,
|
||||
|
||||
# Sampling
|
||||
"ArchAi3D_GRAG_Sampler": ArchAi3D_GRAG_Sampler,
|
||||
|
||||
# Camera Control (Legacy)
|
||||
"ArchAi3D_Qwen_Camera_View_Selector": ArchAi3D_Qwen_Camera_View_Selector,
|
||||
"ArchAi3D_Qwen_Object_Rotation_V2": ArchAi3D_Qwen_Object_Rotation_V2,
|
||||
@@ -238,9 +229,6 @@ NODE_DISPLAY_NAME_MAPPINGS = {
|
||||
# Core - Prompts
|
||||
"ArchAi3D_Clean_Room_Prompt": "🏗️ Clean Room Prompt",
|
||||
|
||||
# Sampling
|
||||
"ArchAi3D_GRAG_Sampler": "🎚️ GRAG Sampler (Fine-Grained Control)",
|
||||
|
||||
# Camera Control (Legacy)
|
||||
"ArchAi3D_Qwen_Camera_View_Selector": "🎬 Camera View Selector",
|
||||
"ArchAi3D_Qwen_Object_Rotation_V2": "🔄 Object Rotation V2",
|
||||
@@ -328,7 +316,7 @@ WEB_DIRECTORY = os.path.join(os.path.dirname(__file__), "web")
|
||||
# ============================================================================
|
||||
|
||||
__all__ = ['NODE_CLASS_MAPPINGS', 'NODE_DISPLAY_NAME_MAPPINGS', 'WEB_DIRECTORY']
|
||||
__version__ = "2.3.0"
|
||||
__version__ = "2.4.1"
|
||||
__author__ = "Amir Ferdos (ArchAi3d)"
|
||||
|
||||
# ============================================================================
|
||||
@@ -340,12 +328,13 @@ print(f"[ArchAi3d-Qwen v{__version__}] Loading nodes...")
|
||||
print(f" 🎨 Core Encoding: 6 nodes (V3 + GRAG Encoder)")
|
||||
print(f" 📏 Core Utils: 2 nodes (Image Scale + GRAG Modifier)")
|
||||
print(f" 💬 Prompt Builders: 3 nodes (Clean Room + Position Guide)")
|
||||
print(f" 🎚️ Sampling: 1 node (GRAG Sampler)")
|
||||
print(f" 📸 Camera Control: 28 nodes (Object Focus v1-v7 + Simple + dx8152)")
|
||||
print(f" 🎨 Image Editing: 4 nodes")
|
||||
print(f" 🎯 Utils: 7 nodes (Mask Crop/Rotate + Color Tools)")
|
||||
print(f" ✅ Total: {len(NODE_CLASS_MAPPINGS)} nodes loaded!")
|
||||
print(f"")
|
||||
print(f" ℹ️ Note: For full GRAG sampling support, install ComfyUI-GRAG-ArchAi3D separately")
|
||||
print(f"")
|
||||
print(f" ⭐ NEW: Object Focus Camera v7 - Professional Cinematography!")
|
||||
print(f" 🎬 Features: Shot sizes, camera angles, movements, enhanced lenses")
|
||||
print(f" 📚 Documentation: ./docs/")
|
||||
|
||||
@@ -156,6 +156,14 @@ class ArchAi3D_Cinematography_Prompt_Builder:
|
||||
"• Back (180°) = Rear view"
|
||||
}),
|
||||
|
||||
# 3C. AUTO-FACING - Automatically point camera at subject
|
||||
"auto_facing": ("BOOLEAN", {
|
||||
"default": True,
|
||||
"tooltip": "Automatically face camera toward target subject (recommended for object photography).\n"
|
||||
"• True = Camera points directly at subject from chosen angle\n"
|
||||
"• False = Camera positioned at angle but may not face subject directly"
|
||||
}),
|
||||
|
||||
# 4. FOCUS/DOF - What's sharp and what's blurred
|
||||
"depth_of_field": ([
|
||||
"Auto (based on shot size)",
|
||||
@@ -581,8 +589,8 @@ class ArchAi3D_Cinematography_Prompt_Builder:
|
||||
)
|
||||
|
||||
def generate_cinematography_prompt(self, target_subject, shot_type, camera_angle,
|
||||
horizontal_angle, auto_facing,
|
||||
depth_of_field, style_mood, prompt_language,
|
||||
horizontal_angle="Front View (0°)",
|
||||
lens_type_override="Auto (from shot size)",
|
||||
perspective_correction="Natural (Standard Lens)",
|
||||
camera_movement="Static (No Movement)",
|
||||
@@ -613,18 +621,18 @@ class ArchAi3D_Cinematography_Prompt_Builder:
|
||||
# Validate parameters (including perspective correction compatibility)
|
||||
warnings = self.validate_parameters(shot_type, dof, lens_type, camera_angle, perspective_correction)
|
||||
|
||||
# Generate SIMPLE prompt (Nanobanan style with horizontal angle + perspective correction)
|
||||
# Generate SIMPLE prompt (Nanobanan style with horizontal angle + perspective correction + auto_facing)
|
||||
simple_prompt = self._generate_simple_prompt(
|
||||
target_subject, shot_type, camera_angle, dof, style_mood,
|
||||
distance, lighting_style, custom_details, horizontal_angle, perspective_correction, prompt_language
|
||||
distance, lighting_style, custom_details, horizontal_angle, auto_facing, perspective_correction, prompt_language
|
||||
)
|
||||
|
||||
# Generate PROFESSIONAL prompt (v7 style with Chinese + horizontal angle + perspective)
|
||||
# Generate PROFESSIONAL prompt (v7 style with Chinese + horizontal angle + perspective + auto_facing)
|
||||
professional_prompt = self._generate_professional_prompt(
|
||||
target_subject, shot_type, camera_angle, lens_type, camera_movement,
|
||||
distance, dof, lighting_style, style_mood, material_detail_preset,
|
||||
photography_quality_preset, custom_details, prompt_language,
|
||||
horizontal_angle, perspective_correction
|
||||
horizontal_angle, auto_facing, perspective_correction
|
||||
)
|
||||
|
||||
# Generate SYSTEM PROMPT (dynamic based on configuration + perspective correction)
|
||||
@@ -644,13 +652,16 @@ class ArchAi3D_Cinematography_Prompt_Builder:
|
||||
def _generate_simple_prompt(self, subject, shot_type, angle, dof, style,
|
||||
distance, lighting, custom_details,
|
||||
horizontal_angle="Front View (0°)",
|
||||
auto_facing=True,
|
||||
perspective_correction="Natural (Standard Lens)",
|
||||
prompt_language="English (Simple & Clear)"):
|
||||
"""
|
||||
Generate Simple English prompt (Nanobanan style)
|
||||
|
||||
Pattern: "A [angle] [shot_type] of [subject], taken from [distance], [horizontal_angle],
|
||||
Pattern: "[facing subject], A [angle] [shot_type] of [subject], taken from [distance], [horizontal_angle],
|
||||
[perspective_correction], with [dof] and [style], [lighting], [custom_details]"
|
||||
|
||||
Note: auto_facing is placed FIRST for maximum attention weight (user's experience-based observation)
|
||||
"""
|
||||
# Clean up angle description
|
||||
angle_clean = angle.replace(" (looking down)", "").replace(" (looking up)", "").replace(" (overhead)", "").replace(" (ground up)", "").replace(" (tilted)", "")
|
||||
@@ -670,6 +681,11 @@ class ArchAi3D_Cinematography_Prompt_Builder:
|
||||
# Build prompt parts
|
||||
parts = []
|
||||
|
||||
# AUTO-FACING: Add at the VERY BEGINNING for maximum attention weight
|
||||
# Only add if enabled AND not front view (front view already implies facing)
|
||||
if auto_facing and horizontal_angle != "Front View (0°)":
|
||||
parts.append(f"Facing {subject} directly")
|
||||
|
||||
# Opening: "An [angle] [shot] of [subject]"
|
||||
if angle_clean.lower() == "eye level":
|
||||
parts.append(f"An eye-level {shot_full} of {subject}")
|
||||
@@ -726,15 +742,26 @@ class ArchAi3D_Cinematography_Prompt_Builder:
|
||||
distance, dof, lighting, style, material_preset,
|
||||
quality_preset, custom_details, language,
|
||||
horizontal_angle="Front View (0°)",
|
||||
auto_facing=True,
|
||||
perspective_correction="Natural (Standard Lens)"):
|
||||
"""
|
||||
Generate Professional prompt (v7 style with Chinese cinematography terms + horizontal angle + perspective)
|
||||
Generate Professional prompt (v7 style with Chinese cinematography terms + horizontal angle + perspective + auto_facing)
|
||||
|
||||
Pattern: "Next Scene: 将镜头转为[LENS], [SHOT]构图, [ANGLE]查看[SUBJECT], [HORIZONTAL], [PERSPECTIVE], [DETAILS]"
|
||||
Pattern: "[facing], Next Scene: 将镜头转为[LENS], [SHOT]构图, [ANGLE]查看[SUBJECT], [HORIZONTAL], [PERSPECTIVE], [DETAILS]"
|
||||
|
||||
Note: auto_facing placed at START for maximum attention weight
|
||||
"""
|
||||
prompt_parts = []
|
||||
|
||||
# Always start with "Next Scene:" for dx8152 LoRAs
|
||||
# AUTO-FACING: Add at BEGINNING for maximum attention (before "Next Scene:")
|
||||
# Only add if enabled AND not front view
|
||||
if auto_facing and horizontal_angle != "Front View (0°)":
|
||||
if language in ["Chinese (Best for dx8152 LoRAs)", "Hybrid (Chinese + English)"]:
|
||||
prompt_parts.append(f"面对{subject}") # "Facing {subject}"
|
||||
else:
|
||||
prompt_parts.append(f"Facing {subject} directly")
|
||||
|
||||
# Always add "Next Scene:" for dx8152 LoRAs
|
||||
prompt_parts.append("Next Scene:")
|
||||
|
||||
# Get horizontal angle and perspective correction descriptions
|
||||
@@ -808,14 +835,27 @@ class ArchAi3D_Cinematography_Prompt_Builder:
|
||||
# Join all parts
|
||||
if language == "English (Simple & Clear)":
|
||||
# Pure English mode - simplified with horizontal angle + perspective
|
||||
base = f"Next Scene: Change to {lens}, {self.get_shot_abbreviation(shot_type)} framing, {angle} viewing {subject}"
|
||||
# Build base without auto_facing (already in prompt_parts[0] if enabled)
|
||||
base_parts = []
|
||||
|
||||
# Check if auto_facing was added
|
||||
if len(prompt_parts) > 1 and "Facing" in prompt_parts[0]:
|
||||
base_parts.append(prompt_parts[0]) # Add facing directive
|
||||
base_parts.append(prompt_parts[1]) # Add "Next Scene:"
|
||||
else:
|
||||
base_parts.append(prompt_parts[0]) # Just "Next Scene:"
|
||||
|
||||
# Add main prompt
|
||||
base_parts.append(f"Change to {lens}, {self.get_shot_abbreviation(shot_type)} framing, {angle} viewing {subject}")
|
||||
|
||||
if horizontal_desc_en:
|
||||
base += f", positioned {horizontal_desc_en}"
|
||||
base_parts.append(f"positioned {horizontal_desc_en}")
|
||||
if perspective_desc_en:
|
||||
base += f", {perspective_desc_en}"
|
||||
base_parts.append(perspective_desc_en)
|
||||
if english_parts:
|
||||
base += ", " + ", ".join(english_parts)
|
||||
return base
|
||||
base_parts.extend(english_parts)
|
||||
|
||||
return ", ".join(base_parts)
|
||||
else:
|
||||
return " ".join(prompt_parts)
|
||||
|
||||
@@ -860,19 +900,46 @@ class ArchAi3D_Cinematography_Prompt_Builder:
|
||||
return angle_map.get(angle, "")
|
||||
|
||||
def _get_distance_chinese(self, distance):
|
||||
"""Get Chinese description for distance"""
|
||||
if distance < 0.5:
|
||||
return "极近距离"
|
||||
elif distance < 1.0:
|
||||
return "近距离"
|
||||
elif distance < 2.0:
|
||||
return "中近距离"
|
||||
elif distance < 4.0:
|
||||
return "中等距离"
|
||||
elif distance < 7.0:
|
||||
return "远距离"
|
||||
"""
|
||||
Convert distance to Chinese words with specific meter values
|
||||
|
||||
Returns exact distance in Chinese characters (e.g., "四米" for 4.0)
|
||||
instead of generic descriptions like "远距离" (far distance)
|
||||
"""
|
||||
# Chinese number words
|
||||
chinese_numbers = {
|
||||
0: "零", 1: "一", 2: "两", 3: "三", 4: "四",
|
||||
5: "五", 6: "六", 7: "七", 8: "八", 9: "九",
|
||||
10: "十", 15: "十五", 20: "二十"
|
||||
}
|
||||
|
||||
# Handle decimals (e.g., 2.5 = "两米半", 0.8 = "零点八米")
|
||||
if distance == int(distance):
|
||||
# Whole number
|
||||
dist_int = int(distance)
|
||||
if dist_int in chinese_numbers:
|
||||
return f"{chinese_numbers[dist_int]}米"
|
||||
else:
|
||||
return f"{dist_int}米" # Fallback to Arabic numerals
|
||||
elif distance % 1 == 0.5:
|
||||
# Half meter (e.g., 2.5 = "两米半")
|
||||
whole = int(distance)
|
||||
if whole == 0:
|
||||
return "半米"
|
||||
elif whole in chinese_numbers:
|
||||
return f"{chinese_numbers[whole]}米半"
|
||||
else:
|
||||
return f"{whole}米半"
|
||||
else:
|
||||
return "极远距离"
|
||||
# Other decimals (e.g., 0.8 = "零点八米")
|
||||
whole = int(distance)
|
||||
decimal = int((distance - whole) * 10)
|
||||
if whole == 0:
|
||||
return f"零点{chinese_numbers.get(decimal, str(decimal))}米"
|
||||
else:
|
||||
whole_cn = chinese_numbers.get(whole, str(whole))
|
||||
decimal_cn = chinese_numbers.get(decimal, str(decimal))
|
||||
return f"{whole_cn}点{decimal_cn}米"
|
||||
|
||||
def _get_movement_chinese(self, movement):
|
||||
"""Get Chinese translation for camera movement"""
|
||||
|
||||
@@ -1,440 +0,0 @@
|
||||
# ArchAi3D GRAG-Aware Sampler Node
|
||||
#
|
||||
# OVERVIEW:
|
||||
# Custom sampler that injects GRAG (Group-Relative Attention Guidance) attention
|
||||
# patches into the sampling process for fine-grained image editing control.
|
||||
#
|
||||
# HOW IT WORKS:
|
||||
# 1. Extracts GRAG configuration from positive conditioning metadata
|
||||
# 2. Creates GRAG attention patch using the reweighting utilities
|
||||
# 3. Injects the patch via model transformer_options
|
||||
# 4. Calls standard ComfyUI sampler with GRAG-enhanced model
|
||||
# 5. CRITICAL: Restores original forward methods in finally block (v2.2.1 fix)
|
||||
#
|
||||
# USAGE:
|
||||
# [Any Encoder] → [GRAG Modifier] → [GRAG Sampler] → [Output]
|
||||
#
|
||||
# Or with GRAG Encoder:
|
||||
# [GRAG Encoder] → [GRAG Sampler] → [Output]
|
||||
#
|
||||
# BENEFITS:
|
||||
# - No ComfyUI core modifications
|
||||
# - Works with all existing encoders
|
||||
# - Update-safe implementation
|
||||
# - Clean on/off toggle
|
||||
# - Proper cleanup prevents global contamination (fixed in v2.2.1)
|
||||
#
|
||||
# CRITICAL FIX (v2.2.1):
|
||||
# Fixed global contamination bug where GRAG patches persisted across samplers.
|
||||
# Root cause: model.clone() creates shallow clone sharing diffusion_model references.
|
||||
# Solution: Store original forward methods and restore in finally block after sampling.
|
||||
# This ensures GRAG only affects intended generations and doesn't contaminate other samplers.
|
||||
#
|
||||
# Author: Amir Ferdos (ArchAi3d)
|
||||
# Email: Amir84ferdos@gmail.com
|
||||
# LinkedIn: https://www.linkedin.com/in/archai3d/
|
||||
# GitHub: https://github.com/amir84ferdos
|
||||
# Category: ArchAi3d/Qwen
|
||||
# Node ID: ArchAi3D_GRAG_Sampler
|
||||
# License: MIT
|
||||
# Based on: GRAG-Image-Editing by little-misfit
|
||||
|
||||
import sys
|
||||
import os
|
||||
|
||||
# Add parent directory to path for imports
|
||||
parent_dir = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
if parent_dir not in sys.path:
|
||||
sys.path.insert(0, parent_dir)
|
||||
|
||||
import torch
|
||||
import comfy.samplers
|
||||
import comfy.sample
|
||||
import comfy.model_management
|
||||
import comfy.utils
|
||||
import latent_preview
|
||||
|
||||
from core.utils.grag_attention import (
|
||||
extract_grag_config_from_conditioning,
|
||||
create_grag_patch
|
||||
)
|
||||
|
||||
|
||||
class ArchAi3D_GRAG_Sampler:
|
||||
"""GRAG-aware sampler that injects attention guidance during sampling.
|
||||
|
||||
This sampler wraps ComfyUI's standard KSampler and injects GRAG attention
|
||||
patches to enable fine-grained editing control. It reads GRAG metadata from
|
||||
conditioning (set by GRAG Modifier or GRAG Encoder) and applies attention
|
||||
reweighting during the diffusion process.
|
||||
|
||||
Key Features:
|
||||
- Extracts GRAG config from conditioning metadata
|
||||
- Injects attention patches via transformer_options
|
||||
- Falls back to standard sampling if GRAG disabled
|
||||
- Compatible with all ComfyUI schedulers and samplers
|
||||
|
||||
Version: 2.1.1
|
||||
"""
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(cls):
|
||||
return {
|
||||
"required": {
|
||||
# Standard KSampler parameters
|
||||
"model": ("MODEL", {
|
||||
"tooltip": "The diffusion model used for denoising"
|
||||
}),
|
||||
"positive": ("CONDITIONING", {
|
||||
"tooltip": "Positive conditioning (should contain GRAG metadata if using GRAG Modifier/Encoder)"
|
||||
}),
|
||||
"negative": ("CONDITIONING", {
|
||||
"tooltip": "Negative conditioning"
|
||||
}),
|
||||
"latent_image": ("LATENT", {
|
||||
"tooltip": "Input latent to denoise"
|
||||
}),
|
||||
"seed": ("INT", {
|
||||
"default": 0,
|
||||
"min": 0,
|
||||
"max": 0xffffffffffffffff,
|
||||
"tooltip": "Random seed for noise generation"
|
||||
}),
|
||||
"steps": ("INT", {
|
||||
"default": 20,
|
||||
"min": 1,
|
||||
"max": 10000,
|
||||
"tooltip": "Number of denoising steps"
|
||||
}),
|
||||
"cfg": ("FLOAT", {
|
||||
"default": 8.0,
|
||||
"min": 0.0,
|
||||
"max": 100.0,
|
||||
"step": 0.1,
|
||||
"tooltip": "Classifier-Free Guidance scale"
|
||||
}),
|
||||
"sampler_name": (comfy.samplers.KSampler.SAMPLERS, {
|
||||
"tooltip": "Sampling algorithm to use"
|
||||
}),
|
||||
"scheduler": (comfy.samplers.KSampler.SCHEDULERS, {
|
||||
"tooltip": "Noise schedule for denoising"
|
||||
}),
|
||||
"denoise": ("FLOAT", {
|
||||
"default": 1.0,
|
||||
"min": 0.0,
|
||||
"max": 1.0,
|
||||
"step": 0.01,
|
||||
"tooltip": "Denoising strength (1.0 = full denoise)"
|
||||
}),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("LATENT",)
|
||||
RETURN_NAMES = ("samples",)
|
||||
FUNCTION = "sample"
|
||||
CATEGORY = "ArchAi3d/Qwen"
|
||||
|
||||
def _patch_qwen_attention(self, model, grag_config):
|
||||
"""Monkey-patch Qwen attention layers to apply GRAG reweighting.
|
||||
|
||||
This function finds all Attention modules in the model and wraps their
|
||||
forward method to apply GRAG key reweighting after RoPE but before attention.
|
||||
|
||||
Args:
|
||||
model: ComfyUI model object with diffusion_model attribute
|
||||
grag_config: Dict with GRAG parameters (lambda, delta, heads)
|
||||
|
||||
Returns:
|
||||
dict: Dictionary mapping modules to their original forward methods.
|
||||
Used for restoration after sampling completes.
|
||||
Returns empty dict if patching fails.
|
||||
"""
|
||||
from core.utils.grag_attention import apply_grag_to_keys
|
||||
|
||||
# Dictionary to store original forward methods for restoration
|
||||
original_forwards = {}
|
||||
|
||||
# Access the actual diffusion model
|
||||
if hasattr(model, 'model') and hasattr(model.model, 'diffusion_model'):
|
||||
diffusion_model = model.model.diffusion_model
|
||||
else:
|
||||
print("[GRAG Sampler] Warning: Could not access diffusion_model")
|
||||
return original_forwards
|
||||
|
||||
# Find and patch all Attention modules
|
||||
patched_count = 0
|
||||
for name, module in diffusion_model.named_modules():
|
||||
# Look for Qwen Attention modules specifically
|
||||
# Check class name AND verify it has the right attributes
|
||||
if (module.__class__.__name__ == 'Attention' and
|
||||
hasattr(module, 'to_q') and
|
||||
hasattr(module, 'add_q_proj') and
|
||||
hasattr(module, 'norm_q')):
|
||||
# Store original forward method for restoration
|
||||
original_forward = module.forward
|
||||
original_forwards[module] = original_forward
|
||||
|
||||
# Create wrapped forward function with GRAG
|
||||
def create_grag_forward(orig_forward, grag_cfg, attn_module):
|
||||
def grag_forward(hidden_states, encoder_hidden_states=None, encoder_hidden_states_mask=None,
|
||||
attention_mask=None, image_rotary_emb=None, transformer_options={}):
|
||||
# Call original forward up to the point where we need to inject GRAG
|
||||
# We'll need to replicate the forward pass with GRAG insertion
|
||||
|
||||
seq_txt = encoder_hidden_states.shape[1]
|
||||
|
||||
# Image stream QKV
|
||||
img_query = attn_module.to_q(hidden_states).unflatten(-1, (attn_module.heads, -1))
|
||||
img_key = attn_module.to_k(hidden_states).unflatten(-1, (attn_module.heads, -1))
|
||||
img_value = attn_module.to_v(hidden_states).unflatten(-1, (attn_module.heads, -1))
|
||||
|
||||
# Text stream QKV
|
||||
txt_query = attn_module.add_q_proj(encoder_hidden_states).unflatten(-1, (attn_module.heads, -1))
|
||||
txt_key = attn_module.add_k_proj(encoder_hidden_states).unflatten(-1, (attn_module.heads, -1))
|
||||
txt_value = attn_module.add_v_proj(encoder_hidden_states).unflatten(-1, (attn_module.heads, -1))
|
||||
|
||||
# Normalization
|
||||
img_query = attn_module.norm_q(img_query)
|
||||
img_key = attn_module.norm_k(img_key)
|
||||
txt_query = attn_module.norm_added_q(txt_query)
|
||||
txt_key = attn_module.norm_added_k(txt_key)
|
||||
|
||||
# Combine streams
|
||||
joint_query = torch.cat([txt_query, img_query], dim=1)
|
||||
joint_key = torch.cat([txt_key, img_key], dim=1)
|
||||
joint_value = torch.cat([txt_value, img_value], dim=1)
|
||||
|
||||
# Apply RoPE
|
||||
from comfy.ldm.qwen_image.model import apply_rotary_emb
|
||||
joint_query = apply_rotary_emb(joint_query, image_rotary_emb)
|
||||
joint_key = apply_rotary_emb(joint_key, image_rotary_emb)
|
||||
|
||||
# ===== GRAG INJECTION POINT =====
|
||||
# Apply GRAG reweighting to keys BEFORE final flattening
|
||||
# Note: joint_key is currently [B, S, H, D], but apply_grag_to_keys expects [B, S, C]
|
||||
try:
|
||||
# Flatten keys temporarily for GRAG
|
||||
joint_key_flat = joint_key.flatten(start_dim=2) # [B, S, H*D]
|
||||
|
||||
# Apply GRAG reweighting
|
||||
joint_key_flat = apply_grag_to_keys(
|
||||
joint_key_flat,
|
||||
seq_txt,
|
||||
grag_cfg['lambda'],
|
||||
grag_cfg['delta'],
|
||||
attn_module.heads
|
||||
)
|
||||
|
||||
# Unflatten back to [B, S, H, D] for consistency
|
||||
joint_key = joint_key_flat.unflatten(-1, (attn_module.heads, -1))
|
||||
except Exception as e:
|
||||
print(f"[GRAG] Warning: Reweighting failed: {e}")
|
||||
import traceback
|
||||
traceback.print_exc()
|
||||
pass # Continue with original keys if GRAG fails
|
||||
# ===== END GRAG =====
|
||||
|
||||
# Flatten for attention
|
||||
joint_query = joint_query.flatten(start_dim=2)
|
||||
joint_key = joint_key.flatten(start_dim=2)
|
||||
joint_value = joint_value.flatten(start_dim=2)
|
||||
|
||||
# Standard attention
|
||||
from comfy.ldm.modules.attention import optimized_attention_masked
|
||||
joint_hidden_states = optimized_attention_masked(
|
||||
joint_query, joint_key, joint_value, attn_module.heads,
|
||||
attention_mask, transformer_options=transformer_options
|
||||
)
|
||||
|
||||
# Split streams
|
||||
txt_attn_output = joint_hidden_states[:, :seq_txt, :]
|
||||
img_attn_output = joint_hidden_states[:, seq_txt:, :]
|
||||
|
||||
# Output projections
|
||||
img_attn_output = attn_module.to_out[0](img_attn_output)
|
||||
img_attn_output = attn_module.to_out[1](img_attn_output)
|
||||
txt_attn_output = attn_module.to_add_out(txt_attn_output)
|
||||
|
||||
return img_attn_output, txt_attn_output
|
||||
|
||||
return grag_forward
|
||||
|
||||
# Replace forward method
|
||||
module.forward = create_grag_forward(original_forward, grag_config, module)
|
||||
patched_count += 1
|
||||
|
||||
print(f"[GRAG Sampler] Patched {patched_count} Attention layers")
|
||||
return original_forwards
|
||||
|
||||
def sample(self, model, positive, negative, latent_image, seed, steps, cfg, sampler_name, scheduler, denoise):
|
||||
"""Perform sampling with GRAG attention guidance.
|
||||
|
||||
This is the main entry point for the sampler. It:
|
||||
1. Extracts GRAG configuration from positive conditioning
|
||||
2. Creates a model clone with GRAG monkey-patch injected
|
||||
3. Calls ComfyUI's standard sampling with the enhanced model
|
||||
4. Returns the denoised latent samples
|
||||
|
||||
Args:
|
||||
model: ComfyUI MODEL object
|
||||
positive: Positive conditioning (may contain GRAG metadata)
|
||||
negative: Negative conditioning
|
||||
latent_image: Input latent {"samples": tensor}
|
||||
seed: Random seed for reproducibility
|
||||
steps: Number of denoising steps
|
||||
cfg: Classifier-Free Guidance scale
|
||||
sampler_name: Sampler algorithm (euler, dpmpp_2m, etc.)
|
||||
scheduler: Noise schedule (normal, karras, etc.)
|
||||
denoise: Denoising strength (0.0-1.0)
|
||||
|
||||
Returns:
|
||||
tuple: (latent_dict,) with denoised samples
|
||||
"""
|
||||
# Extract GRAG configuration from conditioning metadata
|
||||
grag_config = extract_grag_config_from_conditioning(positive)
|
||||
|
||||
# Clone model to avoid modifying original
|
||||
model_clone = model.clone()
|
||||
|
||||
# Store original forward methods for restoration
|
||||
original_forwards = {}
|
||||
|
||||
# If GRAG is enabled, monkey-patch the attention forward function
|
||||
if grag_config and grag_config.get("enabled", False):
|
||||
print(f"[GRAG Sampler] GRAG enabled - λ={grag_config['lambda']:.2f}, δ={grag_config['delta']:.2f}, strength={grag_config.get('strength', 1.0):.2f}")
|
||||
|
||||
# Try to patch Qwen attention layers
|
||||
try:
|
||||
original_forwards = self._patch_qwen_attention(model_clone, grag_config)
|
||||
print(f"[GRAG Sampler] GRAG patches injected successfully")
|
||||
except Exception as e:
|
||||
print(f"[GRAG Sampler] Failed to inject GRAG patches: {e}")
|
||||
print(f"[GRAG Sampler] Falling back to standard sampling")
|
||||
else:
|
||||
print(f"[GRAG Sampler] GRAG disabled - using standard sampling")
|
||||
|
||||
# Call ComfyUI's standard sampling function with try/finally for cleanup
|
||||
# This handles all the complex diffusion logic
|
||||
try:
|
||||
samples = self._common_ksampler(
|
||||
model_clone,
|
||||
seed,
|
||||
steps,
|
||||
cfg,
|
||||
sampler_name,
|
||||
scheduler,
|
||||
positive,
|
||||
negative,
|
||||
latent_image,
|
||||
denoise=denoise
|
||||
)
|
||||
|
||||
return samples
|
||||
|
||||
except Exception as e:
|
||||
print(f"[GRAG Sampler] Error during sampling: {e}")
|
||||
print(f"[GRAG Sampler] Falling back to standard sampler")
|
||||
|
||||
# Fallback: Try without GRAG patches
|
||||
model_clean = model.clone()
|
||||
samples = self._common_ksampler(
|
||||
model_clean,
|
||||
seed,
|
||||
steps,
|
||||
cfg,
|
||||
sampler_name,
|
||||
scheduler,
|
||||
positive,
|
||||
negative,
|
||||
latent_image,
|
||||
denoise=denoise
|
||||
)
|
||||
|
||||
return samples
|
||||
|
||||
finally:
|
||||
# CRITICAL: Always restore original forward methods to prevent contamination
|
||||
# This fixes the global contamination bug where GRAG affects other samplers
|
||||
if original_forwards:
|
||||
for module, original_forward in original_forwards.items():
|
||||
module.forward = original_forward
|
||||
print(f"[GRAG Sampler] Restored {len(original_forwards)} attention modules")
|
||||
|
||||
def _common_ksampler(self, model, seed, steps, cfg, sampler_name, scheduler, positive, negative, latent, denoise=1.0):
|
||||
"""Wrapper around ComfyUI's common_ksampler function.
|
||||
|
||||
This replicates the logic from nodes.py:common_ksampler to ensure
|
||||
compatibility with ComfyUI's sampling infrastructure.
|
||||
|
||||
Args:
|
||||
model: MODEL object (possibly with GRAG patches)
|
||||
seed: Random seed
|
||||
steps: Denoising steps
|
||||
cfg: CFG scale
|
||||
sampler_name: Sampler algorithm
|
||||
scheduler: Noise scheduler
|
||||
positive: Positive conditioning
|
||||
negative: Negative conditioning
|
||||
latent: Latent dict {"samples": tensor}
|
||||
denoise: Denoising strength
|
||||
|
||||
Returns:
|
||||
tuple: (latent_dict,) with denoised samples
|
||||
"""
|
||||
# Extract latent samples
|
||||
latent_image = latent["samples"]
|
||||
|
||||
# Fix empty latent channels if needed
|
||||
latent_image = comfy.sample.fix_empty_latent_channels(model, latent_image)
|
||||
|
||||
# Prepare noise
|
||||
batch_inds = latent.get("batch_index", None)
|
||||
noise = comfy.sample.prepare_noise(latent_image, seed, batch_inds)
|
||||
|
||||
# Handle noise mask if present
|
||||
noise_mask = latent.get("noise_mask", None)
|
||||
|
||||
# Setup progress callback
|
||||
callback = latent_preview.prepare_callback(model, steps)
|
||||
disable_pbar = not comfy.utils.PROGRESS_BAR_ENABLED
|
||||
|
||||
# Perform sampling
|
||||
samples = comfy.sample.sample(
|
||||
model,
|
||||
noise,
|
||||
steps,
|
||||
cfg,
|
||||
sampler_name,
|
||||
scheduler,
|
||||
positive,
|
||||
negative,
|
||||
latent_image,
|
||||
denoise=denoise,
|
||||
disable_noise=False,
|
||||
start_step=None,
|
||||
last_step=None,
|
||||
force_full_denoise=False,
|
||||
noise_mask=noise_mask,
|
||||
callback=callback,
|
||||
disable_pbar=disable_pbar,
|
||||
seed=seed
|
||||
)
|
||||
|
||||
# Return in ComfyUI latent format
|
||||
out = latent.copy()
|
||||
out["samples"] = samples
|
||||
|
||||
return (out,)
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# COMFYUI NODE REGISTRATION
|
||||
# ============================================================================
|
||||
|
||||
NODE_CLASS_MAPPINGS = {
|
||||
"ArchAi3D_GRAG_Sampler": ArchAi3D_GRAG_Sampler
|
||||
}
|
||||
|
||||
NODE_DISPLAY_NAME_MAPPINGS = {
|
||||
"ArchAi3D_GRAG_Sampler": "🎚️ GRAG Sampler (Fine-Grained Control)"
|
||||
}
|
||||
@@ -0,0 +1,143 @@
|
||||
"""
|
||||
Test script to verify auto_facing feature works correctly in Cinematography Prompt Builder
|
||||
"""
|
||||
|
||||
import sys
|
||||
sys.path.insert(0, r"E:\Comfy\Qwen\ComfyUI-Easy-Install\ComfyUI\custom_nodes\ComfyUI-ArchAi3d-Qwen")
|
||||
|
||||
from nodes.camera.cinematography_prompt_builder import ArchAi3D_Cinematography_Prompt_Builder
|
||||
|
||||
# Initialize node
|
||||
node = ArchAi3D_Cinematography_Prompt_Builder()
|
||||
|
||||
print("=" * 80)
|
||||
print("AUTO_FACING FEATURE TEST - Cinematography Prompt Builder")
|
||||
print("=" * 80)
|
||||
|
||||
# Test 1: Front View (0°) - auto_facing should NOT appear (redundant)
|
||||
print("\n" + "=" * 80)
|
||||
print("TEST 1: Front View (0°) with auto_facing=True")
|
||||
print("EXPECTED: NO 'Facing' clause (front view already implies facing)")
|
||||
print("=" * 80)
|
||||
|
||||
simple1, prof1, sys1, desc1 = node.generate_cinematography_prompt(
|
||||
target_subject="the refrigerator",
|
||||
shot_type="Full Shot (FS)",
|
||||
camera_angle="Eye Level",
|
||||
depth_of_field="Auto (based on shot size)",
|
||||
style_mood="Natural/Neutral",
|
||||
prompt_language="Chinese (Best for dx8152 LoRAs)",
|
||||
horizontal_angle="Front View (0°)",
|
||||
auto_facing=True
|
||||
)
|
||||
|
||||
print(f"\nProfessional Prompt:\n{prof1}")
|
||||
print(f"\n✅ PASS" if "面对" not in prof1 and "Facing" not in prof1 else "❌ FAIL: Should NOT have facing clause")
|
||||
|
||||
# Test 2: Angled Left 30° - auto_facing SHOULD appear
|
||||
print("\n" + "=" * 80)
|
||||
print("TEST 2: Angled Left 30° with auto_facing=True")
|
||||
print("EXPECTED: '面对the refrigerator' at the BEGINNING")
|
||||
print("=" * 80)
|
||||
|
||||
simple2, prof2, sys2, desc2 = node.generate_cinematography_prompt(
|
||||
target_subject="the refrigerator",
|
||||
shot_type="Full Shot (FS)",
|
||||
camera_angle="Eye Level",
|
||||
depth_of_field="Auto (based on shot size)",
|
||||
style_mood="Natural/Neutral",
|
||||
prompt_language="Chinese (Best for dx8152 LoRAs)",
|
||||
horizontal_angle="Angled Left 30°",
|
||||
auto_facing=True
|
||||
)
|
||||
|
||||
print(f"\nProfessional Prompt:\n{prof2}")
|
||||
print(f"\n✅ PASS" if prof2.startswith("面对the refrigerator") else "❌ FAIL: Should start with '面对the refrigerator'")
|
||||
|
||||
# Test 3: Side Right (90°) with auto_facing=True - SHOULD appear
|
||||
print("\n" + "=" * 80)
|
||||
print("TEST 3: Side Right (90°) with auto_facing=True")
|
||||
print("EXPECTED: '面对the refrigerator' at the BEGINNING")
|
||||
print("=" * 80)
|
||||
|
||||
simple3, prof3, sys3, desc3 = node.generate_cinematography_prompt(
|
||||
target_subject="the refrigerator",
|
||||
shot_type="Medium Shot (MS)",
|
||||
camera_angle="Eye Level",
|
||||
depth_of_field="Auto (based on shot size)",
|
||||
style_mood="Natural/Neutral",
|
||||
prompt_language="Chinese (Best for dx8152 LoRAs)",
|
||||
horizontal_angle="Side Right (90°)",
|
||||
auto_facing=True
|
||||
)
|
||||
|
||||
print(f"\nProfessional Prompt:\n{prof3}")
|
||||
print(f"\n✅ PASS" if prof3.startswith("面对the refrigerator") else "❌ FAIL: Should start with '面对the refrigerator'")
|
||||
|
||||
# Test 4: Angled Right 45° with auto_facing=False - should NOT appear
|
||||
print("\n" + "=" * 80)
|
||||
print("TEST 4: Angled Right 45° with auto_facing=False")
|
||||
print("EXPECTED: NO 'Facing' clause (disabled by user)")
|
||||
print("=" * 80)
|
||||
|
||||
simple4, prof4, sys4, desc4 = node.generate_cinematography_prompt(
|
||||
target_subject="the refrigerator",
|
||||
shot_type="Medium Shot (MS)",
|
||||
camera_angle="Eye Level",
|
||||
depth_of_field="Auto (based on shot size)",
|
||||
style_mood="Natural/Neutral",
|
||||
prompt_language="Chinese (Best for dx8152 LoRAs)",
|
||||
horizontal_angle="Angled Right 45°",
|
||||
auto_facing=False
|
||||
)
|
||||
|
||||
print(f"\nProfessional Prompt:\n{prof4}")
|
||||
print(f"\n✅ PASS" if "面对" not in prof4 and "Facing" not in prof4 else "❌ FAIL: Should NOT have facing clause (disabled)")
|
||||
|
||||
# Test 5: English mode with Angled Left 45°
|
||||
print("\n" + "=" * 80)
|
||||
print("TEST 5: Angled Left 45° with auto_facing=True (English mode)")
|
||||
print("EXPECTED: 'Facing the refrigerator directly' at the BEGINNING")
|
||||
print("=" * 80)
|
||||
|
||||
simple5, prof5, sys5, desc5 = node.generate_cinematography_prompt(
|
||||
target_subject="the refrigerator",
|
||||
shot_type="Medium Shot (MS)",
|
||||
camera_angle="Eye Level",
|
||||
depth_of_field="Auto (based on shot size)",
|
||||
style_mood="Natural/Neutral",
|
||||
prompt_language="English (Simple & Clear)",
|
||||
horizontal_angle="Angled Left 45°",
|
||||
auto_facing=True
|
||||
)
|
||||
|
||||
print(f"\nProfessional Prompt:\n{prof5}")
|
||||
print(f"\nSimple Prompt:\n{simple5}")
|
||||
print(f"\n✅ PASS" if prof5.startswith("Facing the refrigerator directly") and simple5.startswith("Facing the refrigerator directly") else "❌ FAIL: Should start with 'Facing the refrigerator directly'")
|
||||
|
||||
# Test 6: Hybrid mode with Side Left (90°)
|
||||
print("\n" + "=" * 80)
|
||||
print("TEST 6: Side Left (90°) with auto_facing=True (Hybrid mode)")
|
||||
print("EXPECTED: '面对the refrigerator' at the BEGINNING")
|
||||
print("=" * 80)
|
||||
|
||||
simple6, prof6, sys6, desc6 = node.generate_cinematography_prompt(
|
||||
target_subject="the refrigerator",
|
||||
shot_type="Close-Up (CU)",
|
||||
camera_angle="Eye Level",
|
||||
depth_of_field="Auto (based on shot size)",
|
||||
style_mood="Natural/Neutral",
|
||||
prompt_language="Hybrid (Chinese + English)",
|
||||
horizontal_angle="Side Left (90°)",
|
||||
auto_facing=True
|
||||
)
|
||||
|
||||
print(f"\nProfessional Prompt:\n{prof6}")
|
||||
print(f"\n✅ PASS" if prof6.startswith("面对the refrigerator") else "❌ FAIL: Should start with '面对the refrigerator'")
|
||||
|
||||
print("\n" + "=" * 80)
|
||||
print("TEST SUMMARY")
|
||||
print("=" * 80)
|
||||
print("All tests should show ✅ PASS")
|
||||
print("If any show ❌ FAIL, the auto_facing feature needs debugging")
|
||||
print("=" * 80)
|
||||
Reference in New Issue
Block a user