feat: Add Street View cubemap node

This commit is contained in:
ru4ls
2025-11-07 08:54:09 +07:00
parent 4363f7d1ad
commit f596b8ff23
4 changed files with 319 additions and 32 deletions
+69 -31
View File
@@ -39,6 +39,7 @@ This gives you the best of both worlds: the authenticity of a real photograph co
- **Clean Output:** No UI overlays, just the pure image.
- **Panorama Mode:** Stitch multiple images together to create ultra-wide cinematic landscapes.
- **(New in v1.01) Animation Mode:** Animate camera parameters over time to create smooth transitions and camera movements.
- **(New in v1.02) Cubemap Mode:** Generate 3D environment maps with six images representing all directions (front, back, left, right, up, down) for use in 3D applications and game engines.
---
@@ -70,23 +71,19 @@ This node requires a Google Cloud API key to function. Google provides a generou
### Step-by-Step Guide
**Part A: Create Project & Enable API**
1. Go to the [Google Cloud Console](https://console.cloud.google.com/).
1. **Create Project & Enable API**: Go to the [Google Cloud Console](https://console.cloud.google.com).
2. Create a **New Project**. Give it a name like `ComfyUI-API`.
3. In the new project, search for "API Library".
4. In the library, search for and **Enable** the **"Street View Static API"**.
**Part B: Set Up Billing**
5. You will be prompted to link a billing account. This is required, but you will **not be charged** unless you exceed the $200 free monthly credit.
5. **Set Up Billing**: You will be prompted to link a billing account. This is required, but you will **not be charged** unless you exceed the $200 free monthly credit.
**Part C: Create and Secure Your API Key**
6. In the Cloud Console search bar, navigate to **"Credentials"**.
6. **Create and Secure Your API Key**: In the Cloud Console search bar, navigate to **"Credentials"**.
7. Click **"+ Create Credentials"** and select **"API key"**.
8. **Copy this key immediately.**
9. **(IMPORTANT!)** Click **"Edit API key"**. Under "API restrictions," select **"Restrict key"** and add **"Street View Static API"** to the list. This protects your account. Click **"Save"**.
**Part D: Configure the Node**
10. In your `ComfyUI/custom_nodes/ComfyUI_StreetView-Loader/` folder, create a new file named `.env` (or rename the existing `.env.example` file).
10. **Configure the Node**: In your `ComfyUI/custom_nodes/ComfyUI_StreetView-Loader/` folder, create a new file named `.env` (or rename the existing `.env.example` file).
11. Open this `.env` file and add your copied API key in the following format:
```
GOOGLE_STREET_VIEW_API_KEY="your_actual_api_key_goes_here"
@@ -95,22 +92,24 @@ This node requires a Google Cloud API key to function. Google provides a generou
---
## 3. How to Use & Upscaling
## 3. How to Use & Workflow
The recommended workflow is to use the **URL Parser** node to feed information into the **Loader** node.
The recommended workflow is to use the **URL Parser** node to feed information into the other nodes.
1. **Find your view** in [Google Maps](https://maps.google.com) and enter Street View.
2. Frame the perfect shot, then **copy the entire URL** from your browser's address bar.
3. In ComfyUI, add the **`Street View URL Parser`** node and paste the URL into it.
4. Add the **`Street View Loader`** node.
5. Connect the outputs of the Parser to the inputs of the Loader (`location` to `location`, etc.).
4. Add one of the loader nodes.
5. Connect the outputs of the Parser to the inputs of the loader (`location` to `location`, etc.).
### Understanding the Nodes and Image Size Limit
---
#### `Street View URL Parser`
### Node Descriptions
**`Street View URL Parser`**
This node takes a full Google Maps URL as input and outputs the camera parameters (`location`, `heading`, `pitch`, `fov`).
#### `Street View Loader`
**`Street View Loader`**
This is the main node that fetches the image.
- **`aspect_ratio`**: Choose your desired output aspect ratio from the dropdown. This replaces manual width/height inputs.
- **API Limit & Upscaling:** The Google Street View API has a maximum output size of **640x640 pixels**. For high-resolution images (like 1080p or 4K), you **must** use an upscaling workflow.
@@ -119,10 +118,7 @@ This is the main node that fetches the image.
2. Connect its `IMAGE` output to an **`Upscale Image (using model)`** node.
3. Use a `Load Upscale Model` node (e.g., `4x-UltraSharp`) to get a final, high-quality **2560x1440** image.
---
## 4. Panoramic Loader Node
**`Street View Pano Loader`**
For users who need to create wide, cinematic landscapes, the project includes the **Street View Pano Loader** node.
This node overcomes the API's FOV limitations by using a sophisticated stitching algorithm. It fetches multiple overlapping image "tiles" and then uses the OpenCV library to analyze, warp, and seamlessly blend them into a single, perspective-corrected panoramic image.
@@ -158,14 +154,7 @@ Once you have your stitched result, you have two great options to create a final
- **Goal:** To use AI to intelligently fill in the missing areas, creating a larger, natural-looking scene.
- **How:** Feed the panoramic image into your main workflow (`VAE Encode`, `KSampler`, etc.) with a descriptive prompt of the scene and a **low denoise** (e.g., 0.3-0.5). The AI will use the existing pixels as a guide to generate new details in the black corners.
### Important Notes
- **API Usage:** This node makes multiple API calls. A panorama with **3 images** will count as **3 requests** against your free monthly Google Cloud credit.
- **Stitching Process:** If the OpenCV algorithm cannot find enough matching features, it will automatically **fall back to a simple side-by-side stitch** to ensure you always get an output. If this happens, the best solution is to increase the `overlap_percentage`.
- **Resolution:** The output image will be very wide but only 640px tall. It is **highly recommended** to chain the output of this node into an **Upscale Image** node to increase the final resolution for your projects.
## 5. Animation Node (New in v1.01)
**`Street View Animator`** (New in v1.01)
Version 1.01 introduces the **Street View Animator** node, which allows you to create animated sequences by smoothly transitioning camera parameters over time.
This node enables you to create dynamic camera movements like slow rotations, pitch changes, or field-of-view adjustments that can be used as input for video generation workflows or simply to create smooth transitions between different viewpoints of the same location.
@@ -207,13 +196,62 @@ https://github.com/user-attachments/assets/7edbbdf8-2dcd-4e0c-aae0-ccc5ad1be679
- **Tilt Effects:** Combine pitch changes with heading changes for dynamic camera movements
- **Frame Count:** Total frames = duration × fps (higher values = smoother but may increase API usage costs)
**`Street View Cubemap Loader`** (New in v1.02)
Version 1.02 introduces the **Street View Cubemap Loader** node, which enables the generation of 3D environment maps from Street View locations. This node fetches six images at specific orientations to create a complete cubemap suitable for 3D applications and environment mapping in game engines or rendering software.
![Street View Cubemap Loader Node in ComfyUI](https://github.com/ru4ls/ru4ls-public-media/blob/main/comfyui-streetview-loader/images/StreetView_cubemap.png)
### How to Use & Parameter Suggestions
1. Add the **"Street View Cubemap Loader"** node to your canvas.
2. Provide a `location` (same as other nodes) for the center point of your environment map.
3. Adjust the parameters as needed:
- **`face_resolution`**: Select the resolution for each of the 6 cubemap faces. Options include 256x256, 512x512, 640x640, and 1024x1024. Higher resolutions provide better quality but require more API requests and resources.
- **`output_mode`**: Choose how to output the cubemap faces:
- **individual_faces**: Outputs six separate image tensors (front, back, left, right, up, down)
- **merged_cross**: Combines all faces into a single cross-shaped layout texture
- **merged_hstrip**: Combines all faces into a horizontal strip layout
- **merged_vstrip**: Combines all faces into a vertical strip layout
4. The node will generate outputs based on the selected output mode:
- **Individual faces** (when selected): Six separate image outputs representing the six faces of the cube map:
- **Front**: The view facing forward from the location (0° heading, 0° pitch)
- **Back**: The view facing backward from the location (180° heading, 0° pitch)
- **Left**: The view facing left from the location (270° heading, 0° pitch)
- **Right**: The view facing right from the location (90° heading, 0° pitch)
- **Up**: The view looking upward from the location (0° heading, 85° pitch) - note slightly less than 90° to work with API limitations
- **Down**: The view looking downward from the location (0° heading, -85° pitch) - note slightly less than -90° to work with API limitations
- **Merged texture** (when selected): A single image tensor containing all six faces arranged according to the selected layout
![Street View Cubemap Merged Hstrip Loader Node in ComfyUI](https://github.com/ru4ls/ru4ls-public-media/blob/main/comfyui-streetview-loader/images/StreetView_cubemap-Hstrip.png)
![Street View Cubemap Merged Hstrip Loader Node in ComfyUI](https://github.com/ru4ls/ru4ls-public-media/blob/main/comfyui-streetview-loader/images/StreetView_cubemap-Vstrip.png)
### Use Cases for Cubemap Output
- **3D Environment Mapping**: Use the cubemap output as an environment map for 3D scenes in Blender, Unity, or Unreal Engine
- **Background Texturing**: Create realistic backgrounds for 3D scenes with accurate real-world lighting information
- **VR Applications**: Generate real-world environments for virtual reality experiences
- **Reflection Probes**: Use in rendering pipelines for accurate reflections and lighting calculations
- **Easy 3D Integration**: The merged output modes make it simple to use cubemaps directly in 3D engines without manual texture assembly
### Important Notes
- **API Usage:** This node makes multiple API calls equal to the total number of frames generated. Each frame is a separate API request.
- **Performance:** Animation rendering time increases with duration and fps. Start with low settings and increase as needed.
- **Memory:** Large frame sequences can consume significant memory. Consider using in smaller batches if needed.
- **API Usage**: This node makes six API calls (one for each face of the cube). Each cubemap generation will count as **6 requests** against your free monthly Google Cloud credit.
- **API Limitations**: The Street View API might not return valid images for extreme pitch angles. The node uses 85° and -85° for the up and down faces to avoid common API limitations at exactly 90° vertical pitch.
- **Resolution Constraints**: Each face will be limited by the Street View API's maximum output size of 640x640 pixels. For higher resolutions, you'll need to upscale the results using ComfyUI's upscaling nodes after generation.
- **Memory Considerations**: The merged output modes will create larger textures (e.g., cross layout is 4x width by 3x height of individual faces) so consider your system's memory limitations when choosing high resolutions.
## 6. Troubleshooting
### Important Notes for All Nodes
- **API Usage:** All nodes make API requests against your Google Cloud monthly credit.
- **API Limitations:** The Street View API may not have coverage for all locations or may return black images for extreme angles or unavailable locations.
- **Resolution Constraints:** The maximum resolution from the API is 640x640 pixels per image. For higher-resolution results, use ComfyUI's upscaling nodes.
## 7. Troubleshooting
- **`ValueError: API key not found`:** Your `.env` file is missing, in the wrong location, or the variable name is not `GOOGLE_STREET_VIEW_API_KEY`.
- **Black Image Output:** This usually means Google has no Street View imagery for that coordinate, or your API key is invalid/restricted. Check your key's restrictions on the Google Cloud Console.
+3
View File
@@ -4,12 +4,14 @@ from .nodes.streetview_loader import StreetViewLoader
from .nodes.streetview_url_parser import StreetViewURLParser
from .nodes.streetview_pano_loader import StreetViewPanoLoader
from .nodes.streetview_animator import StreetViewAnimator
from .nodes.streetview_cubemap_loader import StreetViewCubemapLoader
NODE_CLASS_MAPPINGS = {
"StreetViewLoader": StreetViewLoader,
"StreetViewURLParser": StreetViewURLParser,
"StreetViewPanoLoader": StreetViewPanoLoader,
"StreetViewAnimator": StreetViewAnimator,
"StreetViewCubemapLoader": StreetViewCubemapLoader,
}
NODE_DISPLAY_NAME_MAPPINGS = {
@@ -17,6 +19,7 @@ NODE_DISPLAY_NAME_MAPPINGS = {
"StreetViewURLParser": "Street View URL Parser",
"StreetViewPanoLoader": "Street View Pano Loader",
"StreetViewAnimator": "Street View Animator",
"StreetViewCubemapLoader": "Street View Cubemap Loader",
}
print("------------------------------------------")
+246
View File
@@ -0,0 +1,246 @@
# file: ComfyUI_StreetView-Loader/nodes/streetview_cubemap_loader.py
import torch
import numpy as np
import os
from dotenv import load_dotenv
from PIL import Image
from ..utils.connect_api_utils import fetch_streetview_image
# --- Load API Key ---
current_dir = os.path.dirname(os.path.abspath(__file__))
parent_dir = os.path.dirname(current_dir)
dotenv_path = os.path.join(parent_dir, '.env')
load_dotenv(dotenv_path=dotenv_path)
API_KEY_FROM_ENV = os.getenv("GOOGLE_STREET_VIEW_API_KEY")
class StreetViewCubemapLoader:
"""
A ComfyUI node that creates a cubemap by fetching 6 Street View
images at specific orientations (front, back, left, right, up, down)
to form the 6 faces of a cube map for 3D applications.
"""
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"location": ("STRING", {"multiline": False, "default": "46.6237597,8.0305018"}),
"face_resolution": ([
"256x256", "512x512", "640x640", "1024x1024"
], {"default": "512x512"}),
"output_mode": ([
"individual_faces",
"merged_cross",
"merged_hstrip",
"merged_vstrip"
], {"default": "merged_cross"}),
}
}
RETURN_TYPES = ("IMAGE", "IMAGE", "IMAGE", "IMAGE", "IMAGE", "IMAGE", "IMAGE", "STRING")
RETURN_NAMES = ("front", "back", "left", "right", "up", "down", "merged_cubemap", "metadata")
FUNCTION = "load_cubemap"
CATEGORY = "Ru4ls/StreetView/Environment"
def pil_to_tensor(self, image: Image.Image):
image_np = np.array(image).astype(np.float32) / 255.0
return torch.from_numpy(image_np)[None,]
def is_valid_image(self, image: Image.Image):
"""
Check if the image is valid (not completely black or filled with a single color).
This helps detect when the Street View API returns placeholder images.
"""
# Convert to numpy array
img_array = np.array(image)
# Check if it's mostly black (typical for failed API requests)
# If more than 90% of pixels are very dark, consider it invalid
dark_pixels = np.sum(img_array < 30) # pixels with values under 30/255
total_pixels = img_array.size
# If more than 90% of pixels are very dark, it's likely a failed request
if dark_pixels / total_pixels > 0.9:
return False
# Also check if the image is filled with a uniform color (another failure indicator)
# Calculate the standard deviation of pixel values; low std indicates uniform color
if np.std(img_array) < 5: # If standard deviation is very low, it's likely a uniform color
return False
return True
def create_merged_cubemap(self, face_images, output_mode):
"""
Create a merged cubemap texture based on the selected output mode.
"""
if not face_images:
return None
# Get the dimensions of the faces
width = face_images["front"].width
height = face_images["front"].height
if output_mode == "merged_cross":
# Create a cross layout: up in center top, left-front-right-back in middle row, down in center bottom
# Cross layout: 3 columns x 4 rows, with the side faces in the middle row
cross_width = 4 * width
cross_height = 3 * height
merged_image = Image.new('RGB', (cross_width, cross_height), (0, 0, 0))
# Place faces in cross layout
# Up face at top center (column 1, row 0)
merged_image.paste(face_images["up"], (width, 0))
# Middle row: left, front, right, back
merged_image.paste(face_images["left"], (0, height))
merged_image.paste(face_images["front"], (width, height))
merged_image.paste(face_images["right"], (width * 2, height))
merged_image.paste(face_images["back"], (width * 3, height))
# Down face at bottom center (column 1, row 2)
merged_image.paste(face_images["down"], (width, height * 2))
elif output_mode == "merged_hstrip":
# Horizontal strip layout: all 6 faces in a single row
strip_width = 6 * width
strip_height = height
merged_image = Image.new('RGB', (strip_width, strip_height), (0, 0, 0))
# Order: right, left, up, down, front, back (common for some 3D engines)
merged_image.paste(face_images["right"], (0, 0)) # Right
merged_image.paste(face_images["left"], (width, 0)) # Left
merged_image.paste(face_images["up"], (2 * width, 0)) # Up
merged_image.paste(face_images["down"], (3 * width, 0)) # Down
merged_image.paste(face_images["front"], (4 * width, 0)) # Front
merged_image.paste(face_images["back"], (5 * width, 0)) # Back
elif output_mode == "merged_vstrip":
# Vertical strip layout: all 6 faces in a single column
strip_width = width
strip_height = 6 * height
merged_image = Image.new('RGB', (strip_width, strip_height), (0, 0, 0))
# Order: right, left, up, down, front, back
merged_image.paste(face_images["right"], (0, 0)) # Right
merged_image.paste(face_images["left"], (0, height)) # Left
merged_image.paste(face_images["up"], (0, 2 * height)) # Up
merged_image.paste(face_images["down"], (0, 3 * height)) # Down
merged_image.paste(face_images["front"], (0, 4 * height)) # Front
merged_image.paste(face_images["back"], (0, 5 * height)) # Back
else:
# Default to individual faces if mode is invalid
return None
return merged_image
def load_cubemap(self, location, face_resolution, output_mode):
if not API_KEY_FROM_ENV:
raise ValueError("Google Street View API key not found in .env file.")
# Determine width and height for each face based on resolution selection
if face_resolution == "256x256":
width, height = 256, 256
elif face_resolution == "512x512":
width, height = 512, 512
elif face_resolution == "640x640":
width, height = 640, 640
elif face_resolution == "1024x1024":
width, height = 1024, 1024
else:
width, height = 512, 512
# Define the 6 face orientations for the cubemap
# For perfect cubemap geometry, we use 90° FOV for all faces
# Note: For up/down faces, extreme pitch values (+90/-90) may not work with Street View API,
# so we use near-vertical values that are more likely to succeed while maintaining 90° FOV
face_orientations = {
"front": (0, 0), # Facing forward
"back": (180, 0), # Facing backward
"left": (270, 0), # Facing left
"right": (90, 0), # Facing right
"up": (0, 90), # Looking up (with 90° FOV - may fail but geometrically correct)
"down": (0, -90) # Looking down (with 90° FOV - may fail but geometrically correct)
}
face_images = {}
face_metadata = []
successful_fetches = 0
print(f"StreetView Cubemap: Fetching 6 images for cubemap faces at resolution {width}x{height}.")
# Fetch each face of the cubemap using 90° FOV for all faces (optimal for cubemap geometry)
for face_name, (heading, pitch) in face_orientations.items():
print(f" - Fetching {face_name} face at heading {heading}°, pitch {pitch}°, fov 90°...")
try:
image_pil, metadata_url = fetch_streetview_image(
API_KEY_FROM_ENV, location, heading, pitch, 90, width, height # Always use 90° FOV for proper cubemap geometry
)
# Check if the image is valid (not all black, which often indicates API failure at extreme angles)
if image_pil and self.is_valid_image(image_pil):
face_images[face_name] = image_pil
face_metadata.append(f"{face_name}: {metadata_url}")
successful_fetches += 1
else:
# For up/down faces, try a fallback pitch if the extreme pitch failed
if face_name in ["up", "down"]:
fallback_pitch = 85 if face_name == "up" else -85
print(f" - {face_name} face failed with {pitch}° pitch, trying {fallback_pitch}°...")
fallback_image, fallback_metadata_url = fetch_streetview_image(
API_KEY_FROM_ENV, location, heading, fallback_pitch, 90, width, height
)
if fallback_image and self.is_valid_image(fallback_image):
face_images[face_name] = fallback_image
face_metadata.append(f"{face_name}: (fallback) {fallback_metadata_url}")
successful_fetches += 1
print(f" - Successfully fetched {face_name} face with fallback pitch {fallback_pitch}°")
else:
print(f" - Failed to fetch {face_name} face with fallback, using gray placeholder.")
face_images[face_name] = Image.new('RGB', (width, height), color=(64, 64, 64)) # Gray fallback
else:
# For non-up/down faces, use fallback immediately
print(f" - Failed to fetch {face_name} face, using fallback image.")
face_images[face_name] = Image.new('RGB', (width, height), color=(64, 64, 64)) # Gray fallback
except Exception as e:
print(f" - Error fetching {face_name} face: {str(e)}")
face_images[face_name] = Image.new('RGB', (width, height), color=(64, 64, 64)) # Gray fallback
if successful_fetches == 0:
empty_tensor = torch.zeros((1, height, width, 3), dtype=torch.float32)
return (empty_tensor,) * 6 + (empty_tensor, "Failed to fetch any cubemap faces.")
# Convert each face to tensor if we're only outputting individual faces
# Convert individual faces to tensors
front_tensor = self.pil_to_tensor(face_images["front"])
back_tensor = self.pil_to_tensor(face_images["back"])
left_tensor = self.pil_to_tensor(face_images["left"])
right_tensor = self.pil_to_tensor(face_images["right"])
up_tensor = self.pil_to_tensor(face_images["up"])
down_tensor = self.pil_to_tensor(face_images["down"])
# Create merged cubemap based on selected mode
if output_mode in ["merged_cross", "merged_hstrip", "merged_vstrip"]:
merged_image = self.create_merged_cubemap(face_images, output_mode)
merged_tensor = self.pil_to_tensor(merged_image) if merged_image else torch.zeros((1, height, width, 3), dtype=torch.float32)
else: # individual_faces
# For individual faces mode, create a blank merged tensor
merged_tensor = torch.zeros((1, height, width, 3), dtype=torch.float32)
# Prepare metadata string
if output_mode == "individual_faces":
metadata = f"Successfully created cubemap with {successful_fetches}/6 faces. Resolution: {width}x{height}, FOV: 90°, Output mode: {output_mode}\n"
else:
metadata = f"Successfully created cubemap with {successful_fetches}/6 faces. Resolution: {width}x{height}, FOV: 90°, Output mode: {output_mode}\n"
metadata += "\n".join(face_metadata)
return (front_tensor, back_tensor, left_tensor, right_tensor, up_tensor, down_tensor, merged_tensor, metadata)
+1 -1
View File
@@ -1,7 +1,7 @@
[project]
name = "comfyui-streetview-loader"
description = "A custom node for ComfyUI that allows you to load images directly from Google Street View using Google Street Static API to use as backgrounds, textures, or inputs in your workflows."
version = "1.0.1"
version = "1.0.2"
license = {file = "LICENSE"}
# classifiers = [
# # For OS-independent nodes (works on all operating systems)