Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
de679db043 | ||
|
|
1d0ae6dc61 | ||
|
|
3bedb49949 | ||
|
|
b67ae9a0a2 | ||
|
|
1d072713b4 | ||
|
|
8b331c6422 | ||
|
|
90bcfd720f | ||
|
|
9c02934b76 | ||
|
|
e8c4642104 | ||
|
|
f375b69bf0 | ||
|
|
892cfa474a | ||
|
|
d604ec8646 | ||
|
|
6f0fe189cd | ||
|
|
c6ab9e2cbb | ||
|
|
a32489234a | ||
|
|
38d18f19b2 | ||
|
|
c52390afa8 | ||
|
|
4771b9ee45 | ||
|
|
962027f75a | ||
|
|
65b7671bca | ||
|
|
296175de54 |
@@ -0,0 +1,19 @@
|
||||
.PHONY: install install_all install_modules download_flux download_vae
|
||||
|
||||
# install everything except WAN‑VACE downloads
|
||||
install:
|
||||
./install.sh install
|
||||
|
||||
# install everything + WAN‑VACE + HF login
|
||||
install_all:
|
||||
./install.sh all
|
||||
|
||||
# lower‑level helpers
|
||||
install_modules:
|
||||
./install.sh modules
|
||||
|
||||
download_flux:
|
||||
./install.sh flux
|
||||
|
||||
download_vae:
|
||||
./install.sh vae
|
||||
@@ -1,5 +1,5 @@
|
||||
# camera-comfyUI
|
||||
|
||||
[](https://deepwiki.com/Alexankharin/camera-comfyUI)
|
||||

|
||||
|
||||
> Custom ComfyUI nodes for advanced reprojections, point cloud processing, and camera-driven workflows.
|
||||
@@ -80,18 +80,25 @@ A collection of ComfyUI custom nodes to handle diverse camera projections (pinho
|
||||
* ### Reprojection Nodes
|
||||
|
||||
* `ReprojectImage`, `ReprojectDepth`, `OutpaintAnyProjection`
|
||||
|
||||
* ### Matrix Nodes
|
||||
|
||||
* `TransformToMatrix`, `TransformToMatrixManual`
|
||||
|
||||
* ### Depth Nodes
|
||||
|
||||
* `DepthEstimatorNode`, `DepthToImageNode`, `ZDepthToRayDepthNode`
|
||||
* `CombineDepthsNode`, `DepthRenormalizer`
|
||||
* `CombineDepthsNode`, `DepthRenormalizer`, `FisheyeDepthEstimator`
|
||||
|
||||
* ### Point Cloud Nodes
|
||||
|
||||
* `DepthToPointCloud`, `TransformPointCloud`, `ProjectPointCloud`
|
||||
* `PointCloudUnion`, `PointCloudCleaner`, `LoadPointCloud`, `SavePointCloud`
|
||||
* `DepthToPointCloud`, `TransformPointCloud`, `ProjectPointCloud`, `PointCloudUnion`
|
||||
* `PointCloudCleaner`, `LoadPointCloud`, `SavePointCloud`, `ProjectAndClean`
|
||||
|
||||
* ### Trajectory Nodes
|
||||
|
||||
* `CameraMotionNode`, `CameraInterpolationNode`, `CameraTrajectoryNode`
|
||||
* `SaveTrajectory`, `LoadTrajectory`, `PointcloudTrajectoryEnricher`
|
||||
|
||||
---
|
||||
|
||||
@@ -110,7 +117,7 @@ A collection of ComfyUI custom nodes to handle diverse camera projections (pinho
|
||||
| `ZDepthToRayDepthNode` | Converts Z-depth (output of metric-depth-anything) to ray depth to compensate lens curvature. |
|
||||
| `TransformPointCloud` | Applies 4×4 rotation matrix to point cloud |
|
||||
| `ProjectPointCloud` | Z-buffer–based projection of point cloud into image + mask. |
|
||||
| `CameraMotionNode` | Generates image sequences by moving camera along a trajectory. |
|
||||
| `CameraMotionNode` | Generates image and mask sequences along a camera trajectory with optional mask dilation/inversion. |
|
||||
| `CameraInterpolationNode` | Builds a trajectory tensor from two poses. |
|
||||
| `CameraTrajectoryNode` | Interactive Open3D GUI for recording camera waypoints. |
|
||||
| `PointCloudCleaner` | Removes isolated points via voxel filtering. |
|
||||
@@ -132,6 +139,7 @@ A set of JSON workflows illustrating typical use cases. Each workflow lives in `
|
||||
| **Pointcloud.json** | Metric‐depth‐anything v2 → point cloud → camera view synthesis |
|
||||
| **pointcloud\_inpaint.json** | Inpaint + backproject to 3D for dynamic camera motion videos |
|
||||
| **Pointcloud\_walker.json** | GUI‐based camera control via Open3D |
|
||||
| **sbs180\_workflow.json** | Generate stereo (side-by-side) wide-angle/fisheye/equirectangular stereo pairs from a high-res input |
|
||||
|
||||
---
|
||||
|
||||
@@ -190,10 +198,53 @@ Inpaint image with shifted camera and backproject for dynamic camera‐driven vi
|
||||
<img src="demo_images/Fisheye_camera_pointcloud_moved_outpainted.png" alt="PointCloud Inpaint" width="40%" />
|
||||
<img src="demo_images/Camera_interpolation_pointcloud.gif" alt="PointCloud Inpaint Video" width="40%" />
|
||||
|
||||
### 9. `sbs180_workflow.json`
|
||||
|
||||
Take a wide-angle (fisheye or equirectangular) high-resolution (e.g., 4096×4096) image and generate a stereo pair by moving the camera horizontally. The output is a wide-angle stereo pair (side-by-side), simulating a fisheye or equirectangular stereo camera.
|
||||
|
||||
<img src="demo_images/equirect_stereo.gif" alt="Equirectangular Stereo Demo" width="80%" />
|
||||
|
||||
### 10. `Pointcloud_walker.json`
|
||||
|
||||
Interactive Open3D-based GUI for walking and setting camera trajectory inside pointcloud.
|
||||
|
||||
### 11. `wan-vace_ref_to_video.json`
|
||||
|
||||
Integrate the [wan2.1-vace] video generation model to inpaint empty or newly revealed regions during camera movement or view synthesis. This workflow demonstrates how to use the camera-comfyUI nodes to generate camera trajectories and masks, then fill missing areas with the video inpainting model for smooth, high-quality results.
|
||||
|
||||
<img src="demo_images/wan-vace-camera.gif" alt="wan2.1-vace Camera Inpainting Demo" width="80%" />
|
||||
|
||||
---
|
||||
|
||||
## Trajectory Concept
|
||||
|
||||
A **trajectory** in camera-comfyUI is a sequence of camera poses, each represented as a 4×4 transformation matrix. This set of matrices defines the path and orientation of the camera through 3D space, enabling smooth and complex camera movements for view synthesis, point cloud rendering, and video generation.
|
||||
|
||||
### Creating Trajectories
|
||||
|
||||
There are two main ways to create a trajectory:
|
||||
|
||||
- **Camera Matrices Interpolation:**
|
||||
Define two or more camera poses (as matrices), and interpolate between them to generate a smooth path. The `CameraInterpolationNode` automates this process, producing a trajectory tensor for use in camera motion nodes.
|
||||
|
||||
- **Walking in Open3D Environment:**
|
||||
Use the interactive Open3D GUI (`CameraTrajectoryNode`) to "walk" through the point cloud. As you move the camera, waypoints (poses) are recorded, forming a trajectory that can be exported and reused.
|
||||
|
||||
### Using Trajectories
|
||||
|
||||
The `CameraMotionNode` takes a trajectory (set of matrices) and interpolates camera positions and orientations along it, producing smooth camera movements for rendering sequences or videos.
|
||||
|
||||
---
|
||||
|
||||
## Point Cloud Formats
|
||||
|
||||
Point clouds can be saved and loaded in two formats:
|
||||
|
||||
- **.npy**: Numpy array format (fast, preserves all tensor data, recommended for internal pipelines).
|
||||
- **.ply**: Polygon File Format (widely supported, viewable in external 3D tools).
|
||||
|
||||
Use the `SavePointCloud` and `LoadPointCloud` nodes to handle I/O operations in either format.
|
||||
|
||||
---
|
||||
|
||||
## Contributing
|
||||
@@ -202,11 +253,12 @@ Contributions welcome! Please open issues or PRs to add features, improve docs,
|
||||
|
||||
## TODO List
|
||||
|
||||
* [ ] Add processing to pointcloud or depthmap to remove outlier and lonely points at depth borders.
|
||||
* [x] Add processing to pointcloud or depthmap to remove outlier and lonely points at depth borders.
|
||||
* [x] Use built-in comfyUI mask type an image.
|
||||
* [x] Unite nodes into groups to simplify workflows.
|
||||
* [ ] Create a single workflow for view synthesis.
|
||||
* [x] Implement easier and more flexible camera control - more complex camera movements with more than 2 points.
|
||||
* [x] Add more examples and documentation for each node.
|
||||
* [x] Add pointcloud union
|
||||
* [ ] Fix imports for renamed folders (e.g., inpainting_flux)
|
||||
* [x] Fix imports for renamed folders (e.g., inpainting_flux)
|
||||
* [x] Integrate camera movement pipeline with video models (e.g., wan2.1) for smooth, high-quality inpainting along camera trajectories.
|
||||
|
||||
+2
-1
@@ -3,6 +3,7 @@ from .reprojection_nodes import NODE_CLASS_MAPPINGS as NCM2
|
||||
from .metric_depth_nodes import NODE_CLASS_MAPPINGS as NCM3
|
||||
from .flux_fisheye_filling_nodes import NODE_CLASS_MAPPINGS as NCM4
|
||||
from .complex_nodes import NODE_CLASS_MAPPINGS as NCM5
|
||||
NODE_CLASS_MAPPINGS = {**NCM1, **NCM2, **NCM3, **NCM4, **NCM5}
|
||||
from .video_nodes import NODE_CLASS_MAPPINGS as NCM6
|
||||
NODE_CLASS_MAPPINGS = {**NCM1, **NCM2, **NCM3, **NCM4, **NCM5, **NCM6}
|
||||
|
||||
__all__ = ["NODE_CLASS_MAPPINGS"]
|
||||
+2
-2
@@ -64,7 +64,7 @@ class FisheyeDepthEstimator:
|
||||
RETURN_TYPES = ("TENSOR","MASK")
|
||||
RETURN_NAMES = ("depthmap","mask")
|
||||
FUNCTION = "estimate_fisheye_depth"
|
||||
CATEGORY = "Camera/depth"
|
||||
CATEGORY = "Camera/Depth"
|
||||
|
||||
def estimate_fisheye_depth(
|
||||
self,
|
||||
@@ -241,7 +241,7 @@ class PointcloudTrajectoryEnricher:
|
||||
RETURN_TYPES = ("TENSOR","IMAGE","TENSOR")
|
||||
RETURN_NAMES = ("enriched_pointcloud","debug_image","debug_depth")
|
||||
FUNCTION = "enrich_trajectory"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/Trajectory"
|
||||
|
||||
def enrich_trajectory(
|
||||
self,
|
||||
|
||||
Binary file not shown.
Binary file not shown.
|
After Width: | Height: | Size: 224 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 13 MiB |
+125
@@ -0,0 +1,125 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
# ----------------------------------------
|
||||
# Functions
|
||||
# ----------------------------------------
|
||||
install_pytorch() {
|
||||
echo "Installing PyTorch, TorchVision, TorchAudio..."
|
||||
pip3 install -U torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
|
||||
pip3 install bitsandbytes
|
||||
pip3 install accelerate
|
||||
}
|
||||
|
||||
install_system_deps() {
|
||||
echo "Updating apt and installing system dependencies..."
|
||||
sudo apt-get update
|
||||
sudo apt-get install -y build-essential ffmpeg libsm6 libxext6
|
||||
}
|
||||
|
||||
clone_and_install_comfyui() {
|
||||
echo "Cloning ComfyUI..."
|
||||
git clone https://github.com/comfyanonymous/ComfyUI.git
|
||||
echo "Installing ComfyUI requirements..."
|
||||
pip3 install -r ComfyUI/requirements.txt
|
||||
}
|
||||
|
||||
install_camera_node() {
|
||||
echo "Cloning camera‑ComfyUI..."
|
||||
git clone https://github.com/Alexankharin/camera-comfyUI.git \
|
||||
ComfyUI/custom_nodes/camera-comfyUI
|
||||
echo "Installing camera‑ComfyUI requirements..."
|
||||
pip3 install -r ComfyUI/custom_nodes/camera-comfyUI/requirements.txt
|
||||
}
|
||||
|
||||
install_image_filters() {
|
||||
echo "Cloning Image‑Filters node..."
|
||||
git clone https://github.com/spacepxl/ComfyUI-Image-Filters.git \
|
||||
ComfyUI/custom_nodes/ComfyUI-Image-Filters
|
||||
echo "Installing Image‑Filters requirements..."
|
||||
pip3 install -r ComfyUI/custom_nodes/ComfyUI-Image-Filters/requirements.txt
|
||||
}
|
||||
|
||||
clone_flux_inpainting() {
|
||||
echo "Cloning ComfyUI‑Flux‑Inpainting..."
|
||||
git clone https://github.com/rubi-du/ComfyUI-Flux-Inpainting.git \
|
||||
ComfyUI/custom_nodes/Flux-Inpainting
|
||||
echo "Renaming Flux‑Inpainting folder..."
|
||||
mv ComfyUI/custom_nodes/Flux-Inpainting \
|
||||
ComfyUI/custom_nodes/inpainting_flux
|
||||
}
|
||||
|
||||
install_hf_hub() {
|
||||
echo "Installing huggingface_hub..."
|
||||
pip3 install huggingface_hub
|
||||
}
|
||||
|
||||
download_vae_models() {
|
||||
echo "Downloading WAN‑VACE models via wget..."
|
||||
wget -O ComfyUI/models/vae/wan_2.1_vae.safetensors \
|
||||
"https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/vae/wan_2.1_vae.safetensors?download=true"
|
||||
|
||||
wget -O ComfyUI/models/text_encoders/umt5_xxl_fp16.safetensors \
|
||||
"https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/text_encoders/umt5_xxl_fp16.safetensors?download=true"
|
||||
|
||||
wget -O ComfyUI/models/diffusion_models/wan2.1_vace_14B_fp16.safetensors \
|
||||
"https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/diffusion_models/wan2.1_vace_14B_fp16.safetensors"
|
||||
}
|
||||
|
||||
login_hf_hub() {
|
||||
echo "Logging in to Hugging Face Hub..."
|
||||
huggingface-cli login
|
||||
}
|
||||
|
||||
# ----------------------------------------
|
||||
# Main CLI
|
||||
# ----------------------------------------
|
||||
# default to "install" if no arg given
|
||||
MODE="${1:-install}"
|
||||
|
||||
case "$MODE" in
|
||||
modules)
|
||||
install_pytorch
|
||||
install_system_deps
|
||||
clone_and_install_comfyui
|
||||
install_camera_node
|
||||
install_image_filters
|
||||
install_hf_hub
|
||||
;;
|
||||
|
||||
flux)
|
||||
clone_flux_inpainting
|
||||
;;
|
||||
|
||||
vae)
|
||||
download_vae_models
|
||||
;;
|
||||
|
||||
install)
|
||||
install_pytorch
|
||||
install_system_deps
|
||||
clone_and_install_comfyui
|
||||
install_camera_node
|
||||
clone_flux_inpainting
|
||||
install_image_filters
|
||||
install_hf_hub
|
||||
;;
|
||||
|
||||
all)
|
||||
install_pytorch
|
||||
install_system_deps
|
||||
clone_and_install_comfyui
|
||||
install_camera_node
|
||||
clone_flux_inpainting
|
||||
install_image_filters
|
||||
install_hf_hub
|
||||
download_vae_models
|
||||
login_hf_hub
|
||||
echo "All done! 🎉"
|
||||
;;
|
||||
|
||||
*)
|
||||
echo "Usage: $0 {install|modules|flux|vae|all}"
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
@@ -41,7 +41,7 @@ class DepthEstimatorNode:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("depth tensor",)
|
||||
FUNCTION = "estimate_depth"
|
||||
CATEGORY = "Camera/depth"
|
||||
CATEGORY = "Camera/Depth"
|
||||
|
||||
def estimate_depth(
|
||||
self,
|
||||
@@ -110,7 +110,7 @@ class DepthToImageNode:
|
||||
RETURN_TYPES = ("IMAGE",)
|
||||
RETURN_NAMES = ("depth image",)
|
||||
FUNCTION = "depth_to_image"
|
||||
CATEGORY = "Camera/depth"
|
||||
CATEGORY = "Camera/Depth"
|
||||
|
||||
def depth_to_image(
|
||||
self,
|
||||
@@ -158,7 +158,7 @@ class ZDepthToRayDepthNode:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("ray depth",)
|
||||
FUNCTION = "depth_to_ray_depth"
|
||||
CATEGORY = "Camera/depth"
|
||||
CATEGORY = "Camera/Depth"
|
||||
|
||||
def depth_to_ray_depth(
|
||||
self,
|
||||
@@ -229,7 +229,7 @@ class CombineDepthsNode:
|
||||
RETURN_TYPES = ("TENSOR","MASK")
|
||||
RETURN_NAMES = ("combined_depth","combined_mask")
|
||||
FUNCTION = "combine_depths"
|
||||
CATEGORY = "Camera/depth"
|
||||
CATEGORY = "Camera/Depth"
|
||||
|
||||
def combine_depths(
|
||||
self,
|
||||
@@ -358,7 +358,7 @@ class DepthRenormalizer:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("depth tensor",)
|
||||
FUNCTION = "renormalize_depth"
|
||||
CATEGORY = "Camera/depth"
|
||||
CATEGORY = "Camera/Depth"
|
||||
|
||||
def renormalize_depth(
|
||||
self,
|
||||
|
||||
+126
-56
@@ -113,7 +113,7 @@ def XYZ_to_equirect(X: torch.Tensor, Y: torch.Tensor, Z: torch.Tensor, fov: floa
|
||||
Convert XYZ coordinates to normalized UV and depth using equirectangular projection.
|
||||
"""
|
||||
# full 360°×180°
|
||||
fov_rad = math.radians(fov)
|
||||
fov_rad = math.radians(fov) / 2
|
||||
depth = torch.sqrt(X**2 + Y**2 + Z**2)
|
||||
lon = torch.atan2(X, Z) # –π → +π
|
||||
lat = torch.asin(Y / depth) # –π/2 → +π/2
|
||||
@@ -167,7 +167,7 @@ class DepthToPointCloud:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("pointcloud",)
|
||||
FUNCTION = "depth_to_pointcloud"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/PointCloud"
|
||||
|
||||
def depth_to_pointcloud(
|
||||
self,
|
||||
@@ -273,7 +273,7 @@ class TransformPointCloud:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("transformed pointcloud",)
|
||||
FUNCTION = "transform_pointcloud"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/PointCloud"
|
||||
|
||||
def transform_pointcloud(
|
||||
self,
|
||||
@@ -321,7 +321,7 @@ class ProjectPointCloud:
|
||||
RETURN_TYPES = ("IMAGE", "MASK", "TENSOR")
|
||||
RETURN_NAMES = ("image", "mask", "depth")
|
||||
FUNCTION = "project_pointcloud"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/PointCloud"
|
||||
|
||||
def project_pointcloud(
|
||||
self,
|
||||
@@ -393,6 +393,7 @@ class ProjectPointCloud:
|
||||
img4 = flat.view(output_height, output_width, 4)
|
||||
rgb = img4[..., :3].clamp(0, 255)
|
||||
alpha = (img4[..., 3] > 0).float()
|
||||
mask_init= (img4[..., 3] > 0) # initial mask
|
||||
rgb *= alpha.unsqueeze(-1)
|
||||
depth_img = z_front.view(output_height, output_width)
|
||||
rgb_HR = rgb
|
||||
@@ -406,7 +407,6 @@ class ProjectPointCloud:
|
||||
idxbuf.fill_(-1)
|
||||
idxbuf.scatter_reduce_(0, pix, order_m, reduce='amax', include_self=True)
|
||||
win_back = idxbuf[pix] >= 0
|
||||
|
||||
flat.fill_(0)
|
||||
flat[pix[win_back]] = colors[win_back]
|
||||
back4 = flat.view(output_height, output_width, 4)
|
||||
@@ -430,8 +430,29 @@ class ProjectPointCloud:
|
||||
# merge only at hole locations
|
||||
rgb[hole] = rgb_med[hole]
|
||||
# alpha already set to 1.0 for holes
|
||||
# 8 apply median blur to mask if point_size > 1 and to initial image
|
||||
mask_t = alpha.unsqueeze(0).unsqueeze(0) # [1,1,H,W]
|
||||
pad = point_size // 2
|
||||
ksize = (point_size, point_size)
|
||||
|
||||
# 8) Pack and return with original script shapes
|
||||
# b) grow (dilate) mask by max‑pool
|
||||
mask_grow = F.max_pool2d(mask_t, kernel_size=ksize, stride=1, padding=pad)
|
||||
|
||||
# c) shrink (erode) by inverting, max‑pool, then inverting back
|
||||
mask_shrink = 1.0 - F.max_pool2d(1.0 - mask_grow, kernel_size=ksize, stride=1, padding=pad)
|
||||
|
||||
# d) back to [H,W] and use as our new alpha
|
||||
alpha = mask_shrink.squeeze(0).squeeze(0)
|
||||
print(1)
|
||||
# e) median‑filter the *whole* RGB image
|
||||
# prep for kornia: [B,C,H,W]
|
||||
rgb_t_full = rgb.permute(2,0,1).unsqueeze(0) # [1,3,H,W]
|
||||
rgb_med_full = median_blur(rgb_t_full, ksize) # [1,3,H,W]
|
||||
rgb_med_full = rgb_med_full.squeeze(0).permute(1,2,0) # [H,W,3]
|
||||
|
||||
# g) refill *only* the original holes with the median result
|
||||
rgb[~mask_init] = rgb_med_full[~mask_init]
|
||||
# 9) Pack and return with original script shapes
|
||||
img = rgb.unsqueeze(0) # [1,H,W,3]
|
||||
mask_out = alpha # [H,W]
|
||||
depth4 = depth_img.unsqueeze(0).unsqueeze(-1) # [1,H,W,1]
|
||||
@@ -456,7 +477,7 @@ class PointCloudUnion:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES =("merged pointcloud",)
|
||||
FUNCTION = "union_pointclouds"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/PointCloud"
|
||||
|
||||
def union_pointclouds(
|
||||
self,
|
||||
@@ -492,7 +513,7 @@ class LoadPointCloud:
|
||||
}
|
||||
}
|
||||
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/PointCloud"
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("loaded pointcloud",)
|
||||
FUNCTION = "load_pointcloud"
|
||||
@@ -504,24 +525,45 @@ class LoadPointCloud:
|
||||
arr = np.load(file_path)
|
||||
tensor_pc = torch.from_numpy(arr)
|
||||
return (tensor_pc,)
|
||||
coords = []
|
||||
colors = []
|
||||
with open(file_path, 'r') as f:
|
||||
line = f.readline().strip()
|
||||
while not line.startswith("end_header"):
|
||||
|
||||
if o3d is None:
|
||||
logging.warning("[camera-comfyUI] open3d is not installed. Falling back to manual PLY parser.")
|
||||
coords = []
|
||||
colors = []
|
||||
with open(file_path, 'r') as f:
|
||||
line = f.readline().strip()
|
||||
for line in f:
|
||||
parts = line.strip().split()
|
||||
if len(parts) < 7:
|
||||
continue
|
||||
x, y, z = map(float, parts[0:3])
|
||||
r, g, b, a = map(int, parts[3:7])
|
||||
coords.append((x, y, z))
|
||||
colors.append((r, g, b, a))
|
||||
np_coords = np.array(coords, dtype=np.float32)
|
||||
np_colors = np.array(colors, dtype=np.float32)/255.0
|
||||
combined = np.concatenate([np_coords, np_colors], axis=1)
|
||||
tensor_pc = torch.from_numpy(combined)
|
||||
while not line.startswith("end_header"):
|
||||
line = f.readline().strip()
|
||||
for line in f:
|
||||
parts = line.strip().split()
|
||||
if len(parts) < 7:
|
||||
continue
|
||||
x, y, z = map(float, parts[0:3])
|
||||
r, g, b, a = map(float, parts[3:7])
|
||||
coords.append((x, y, z))
|
||||
colors.append((r, g, b, a))
|
||||
np_coords = np.array(coords, dtype=np.float32)
|
||||
np_colors = np.array(colors, dtype=np.float32)
|
||||
# if colors are > 1, normalize them to [0,1]
|
||||
if np_colors.max() > 1.0:
|
||||
np_colors = np_colors / 255.0
|
||||
else:
|
||||
pc = o3d.t.io.read_point_cloud(file_path)
|
||||
np_coords = pc.point["positions"].numpy().astype(np.float32)
|
||||
if "colors" in pc.point:
|
||||
cols = pc.point["colors"].numpy().astype(np.float32)
|
||||
else:
|
||||
cols = np.ones((np_coords.shape[0], 3), dtype=np.float32)
|
||||
if "alpha" in pc.point:
|
||||
alpha = pc.point["alpha"].numpy().astype(np.float32)
|
||||
else:
|
||||
alpha = np.ones((np_coords.shape[0], 1), dtype=np.float32)
|
||||
np_colors = np.concatenate([cols, alpha], axis=1)
|
||||
if np_colors.max() > 1.0:
|
||||
np_colors = np_colors / 255.0
|
||||
# combine coords and colors into a single tensor
|
||||
combined = np.concatenate([np_coords, np_colors], axis=1)
|
||||
tensor_pc = torch.from_numpy(combined)
|
||||
return (tensor_pc,)
|
||||
|
||||
@classmethod
|
||||
@@ -569,7 +611,7 @@ class SavePointCloud:
|
||||
RETURN_TYPES = ()
|
||||
FUNCTION = "save_pointcloud"
|
||||
OUTPUT_NODE = True
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/PointCloud"
|
||||
DESCRIPTION = "Saves the input point cloud to your ComfyUI output directory as .ply or .npy."
|
||||
|
||||
def save_pointcloud(self, pointcloud: torch.Tensor, filename_prefix: str, save_as: str = "ply"):
|
||||
@@ -586,24 +628,36 @@ class SavePointCloud:
|
||||
os.makedirs(full_output_folder, exist_ok=True)
|
||||
base_name = filename.replace("%batch_num%", "0")
|
||||
if save_as == "ply":
|
||||
ply_name = f"{base_name}_{counter:05}.ply"
|
||||
ply_path = os.path.join(full_output_folder, ply_name)
|
||||
coords = pointcloud[:, :3].cpu().numpy()
|
||||
colors = pointcloud[:, 3:].cpu().numpy().clip(0,1)
|
||||
with open(ply_path, 'w') as f:
|
||||
f.write("ply\n")
|
||||
f.write("format ascii 1.0\n")
|
||||
f.write(f"element vertex {coords.shape[0]}\n")
|
||||
f.write("property float x\n")
|
||||
f.write("property float y\n")
|
||||
f.write("property float z\n")
|
||||
f.write("property uchar red\n")
|
||||
f.write("property uchar green\n")
|
||||
f.write("property uchar blue\n")
|
||||
f.write("property uchar alpha\n")
|
||||
f.write("end_header\n")
|
||||
for (x,y,z), (r,g,b,a) in zip(coords, colors):
|
||||
f.write(f"{x} {y} {z} {int(r*255)} {int(g*255)} {int(b*255)} {int(a*255)}\n")
|
||||
ply_name = f"{base_name}_{counter:05}.ply"
|
||||
ply_path = os.path.join(full_output_folder, ply_name)
|
||||
coords = pointcloud[:, :3].cpu().numpy().astype(np.float32)
|
||||
colors = pointcloud[:, 3:].cpu().numpy().clip(0, 1).astype(np.float32)
|
||||
|
||||
if o3d is None:
|
||||
logging.warning("[camera-comfyUI] open3d is not installed. Falling back to manual ASCII PLY writer.")
|
||||
with open(ply_path, 'w') as f:
|
||||
f.write("ply\n")
|
||||
f.write("format ascii 1.0\n")
|
||||
f.write(f"element vertex {coords.shape[0]}\n")
|
||||
f.write("property float x\n")
|
||||
f.write("property float y\n")
|
||||
f.write("property float z\n")
|
||||
f.write("property float red\n")
|
||||
f.write("property float green\n")
|
||||
f.write("property float blue\n")
|
||||
f.write("property float alpha\n")
|
||||
f.write("end_header\n")
|
||||
for (x, y, z), (r, g, b, a) in zip(coords, colors):
|
||||
f.write(f"{x} {y} {z} {r} {g} {b} {a}\n")
|
||||
else:
|
||||
pc = o3d.t.geometry.PointCloud()
|
||||
pc.point["positions"] = o3d.core.Tensor(coords, o3d.core.float32)
|
||||
pc.point["colors"] = o3d.core.Tensor(colors[:, :3], o3d.core.float32)
|
||||
if colors.shape[1] > 3:
|
||||
pc.point["alpha"] = o3d.core.Tensor(colors[:, 3:], o3d.core.float32)
|
||||
else:
|
||||
pc.point["alpha"] = o3d.core.Tensor(np.ones((coords.shape[0], 1), dtype=np.float32), o3d.core.float32)
|
||||
o3d.t.io.write_point_cloud(ply_path, pc)
|
||||
file_name = ply_name
|
||||
else:
|
||||
npy_name = f"{base_name}_{counter:05}.npy"
|
||||
@@ -640,12 +694,15 @@ class CameraMotionNode:
|
||||
"output_width": ("INT", {"default":512, "min":8, "max":16384}),
|
||||
"output_height": ("INT", {"default":512, "min":8, "max":16384}),
|
||||
"point_size": ("INT", {"default":1, "min":1}),
|
||||
"widen_mask": ("INT", {"default":0, "min":0, "max":64}),
|
||||
"invert_mask": ("BOOLEAN", {"default": False}),
|
||||
"points_to_mask": ("BOOLEAN", {"default": False, "tooltip": "Output mask frames of projected points"}),
|
||||
}}
|
||||
|
||||
RETURN_TYPES = ("IMAGE",)
|
||||
RETURN_NAMES = ("motion_frames",)
|
||||
RETURN_TYPES = ("IMAGE", "MASK")
|
||||
RETURN_NAMES = ("motion_frames", "mask_frames")
|
||||
FUNCTION = "generate_motion_frames"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/Trajectory"
|
||||
|
||||
def generate_motion_frames(
|
||||
self,
|
||||
@@ -656,7 +713,10 @@ class CameraMotionNode:
|
||||
output_horizontal_fov: float,
|
||||
output_width: int,
|
||||
output_height: int,
|
||||
point_size: int = 1
|
||||
point_size: int = 1,
|
||||
widen_mask: int = 0,
|
||||
invert_mask: bool = False,
|
||||
points_to_mask: bool = False
|
||||
) -> Tuple[torch.Tensor]:
|
||||
# validate trajectory shape
|
||||
if trajectory.dim() != 3 or trajectory.shape[1:] != (4,4):
|
||||
@@ -683,9 +743,10 @@ class CameraMotionNode:
|
||||
proj_node = ProjectPointCloud()
|
||||
transform_node = TransformPointCloud()
|
||||
frames = []
|
||||
masks = []
|
||||
for M in tqdm(full_traj):
|
||||
pc_t, = transform_node.transform_pointcloud(pointcloud, M)
|
||||
img, _, _ = proj_node.project_pointcloud(
|
||||
img, mask, _ = proj_node.project_pointcloud(
|
||||
pc_t,
|
||||
output_projection,
|
||||
output_horizontal_fov,
|
||||
@@ -693,10 +754,19 @@ class CameraMotionNode:
|
||||
output_height,
|
||||
point_size
|
||||
)
|
||||
if widen_mask > 0:
|
||||
k = 2 * widen_mask + 1
|
||||
pad = widen_mask
|
||||
mask = F.max_pool2d(mask.float().unsqueeze(0).unsqueeze(0), kernel_size=k, stride=1, padding=pad).squeeze(0).squeeze(0)
|
||||
if invert_mask:
|
||||
mask = 1.0 - mask
|
||||
masks.append(mask)
|
||||
if points_to_mask:
|
||||
img = mask.unsqueeze(-1).repeat(1,1,1,3)
|
||||
frames.append(img[0])
|
||||
|
||||
# output as (T,H,W,3)
|
||||
return (torch.stack(frames, dim=0),)
|
||||
return (torch.stack(frames, dim=0), torch.stack(masks, dim=0))
|
||||
|
||||
class CameraInterpolationNode:
|
||||
"""
|
||||
@@ -715,7 +785,7 @@ class CameraInterpolationNode:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("trajectory",)
|
||||
FUNCTION = "interpolate"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/Trajectory"
|
||||
|
||||
def interpolate(
|
||||
self,
|
||||
@@ -750,7 +820,7 @@ class CameraTrajectoryNode:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("trajectory",)
|
||||
FUNCTION = "build_trajectory"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/Trajectory"
|
||||
|
||||
def build_trajectory(
|
||||
self,
|
||||
@@ -890,7 +960,7 @@ class PointCloudCleaner:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("cleaned_pointcloud",)
|
||||
FUNCTION = "clean_pointcloud"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/PointCloud"
|
||||
|
||||
def clean_pointcloud(
|
||||
self,
|
||||
@@ -966,7 +1036,7 @@ class ProjectAndClean:
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("cleaned_pointcloud",)
|
||||
FUNCTION = "project_and_clean"
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/PointCloud"
|
||||
|
||||
def project_and_clean(
|
||||
self,
|
||||
@@ -1080,7 +1150,7 @@ class SaveTrajectory:
|
||||
RETURN_TYPES = ()
|
||||
FUNCTION = "save_trajectory"
|
||||
OUTPUT_NODE = True
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/Trajectory"
|
||||
DESCRIPTION = "Saves the input trajectory tensor (N,4,4) to your ComfyUI output directory as .npy."
|
||||
|
||||
def save_trajectory(self, trajectory: torch.Tensor, filename_prefix: str):
|
||||
@@ -1130,7 +1200,7 @@ class LoadTrajectory:
|
||||
}
|
||||
}
|
||||
|
||||
CATEGORY = "Camera/pointcloud"
|
||||
CATEGORY = "Camera/Trajectory"
|
||||
RETURN_TYPES = ("TENSOR",)
|
||||
RETURN_NAMES = ("loaded_trajectory",)
|
||||
FUNCTION = "load_trajectory"
|
||||
|
||||
@@ -148,7 +148,7 @@ class ReprojectImage:
|
||||
RETURN_TYPES: Tuple[str, str] = ("IMAGE", "MASK")
|
||||
RETURN_NAMES = ("reprojected image", "reprojected mask")
|
||||
FUNCTION: str = "reproject_image"
|
||||
CATEGORY: str = "Camera/reproject"
|
||||
CATEGORY: str = "Camera/Reprojection"
|
||||
|
||||
def reproject_image(
|
||||
self,
|
||||
@@ -303,7 +303,7 @@ class TransformToMatrix:
|
||||
RETURN_TYPES: Tuple[str] = ("MAT_4X4",)
|
||||
RETURN_NAMES = ("transformation matrix",)
|
||||
FUNCTION: str = "generate_matrix"
|
||||
CATEGORY: str = "Camera/reproject"
|
||||
CATEGORY: str = "Camera/Matrix"
|
||||
|
||||
def generate_matrix(
|
||||
self,
|
||||
@@ -392,7 +392,7 @@ class TransformToMatrixManual:
|
||||
RETURN_TYPES: Tuple[str] = ("MAT_4X4",)
|
||||
RETURN_NAMES = ("transformation matrix",)
|
||||
FUNCTION: str = "generate_matrix"
|
||||
CATEGORY: str = "Camera/reproject"
|
||||
CATEGORY: str = "Camera/Matrix"
|
||||
|
||||
def generate_matrix(
|
||||
self,
|
||||
@@ -444,7 +444,7 @@ class ReprojectDepth:
|
||||
RETURN_TYPES: Tuple[str, str] = ("TENSOR", "MASK")
|
||||
RETURN_NAMES = ("reprojected_depth", "reprojected_mask")
|
||||
FUNCTION: str = "reproject_depth"
|
||||
CATEGORY: str = "Camera/reproject"
|
||||
CATEGORY: str = "Camera/Reprojection"
|
||||
|
||||
def reproject_depth(
|
||||
self,
|
||||
|
||||
+265
@@ -0,0 +1,265 @@
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
import numpy as np
|
||||
import os
|
||||
import sys
|
||||
from typing import Dict, Any, Tuple
|
||||
from tqdm import tqdm # Added tqdm import
|
||||
|
||||
# Import existing pointcloud nodes and projection definitions
|
||||
from .pointcloud_nodes import DepthToPointCloud, TransformPointCloud, ProjectPointCloud, Projection, PointCloudCleaner
|
||||
import folder_paths
|
||||
|
||||
# Ensure video_depth_anything is on path
|
||||
video_depth_path = os.path.join("/root/Video-Depth-Anything", "metric_depth")
|
||||
if video_depth_path not in sys.path:
|
||||
sys.path.append(video_depth_path)
|
||||
try:
|
||||
from video_depth_anything.video_depth import VideoDepthAnything
|
||||
print("video_depth_anything module loaded successfully.")
|
||||
except ImportError:
|
||||
VideoDepthAnything = None
|
||||
print("Warning: video_depth_anything module not found. Ensure it is installed correctly.")
|
||||
print("error: ", sys.exc_info()[1])
|
||||
|
||||
|
||||
class VideoCameraMotionSequence:
|
||||
"""
|
||||
Takes a sequence of RGB frames and corresponding depth maps,
|
||||
converts each frame+depth to a pointcloud, interpolates a camera
|
||||
trajectory to match video length, cleans the pointcloud if needed,
|
||||
and outputs reprojected images, masks, and depth maps per frame.
|
||||
"""
|
||||
@classmethod
|
||||
def INPUT_TYPES(cls) -> Dict[str, Any]:
|
||||
return {
|
||||
"required": {
|
||||
# Sequence of frames: Tensor [T, H, W, 3]
|
||||
"frames": ("IMAGE", {"shape_hint": [None, None, None, 3]}),
|
||||
# Sequence of depth maps: Tensor [T, H, W] or [T, H, W, 1]
|
||||
"depth_seq": ("TENSOR", {"shape_hint": [None, None, None]}),
|
||||
# Camera trajectory waypoints: Tensor [K, 4, 4]
|
||||
"trajectory": ("TENSOR", {"shape_hint": [None, 4, 4]}),
|
||||
# Input projection parameters
|
||||
"input_projection": (Projection.PROJECTIONS, {}),
|
||||
"input_horizontal_fov": ("FLOAT", {"default": 90.0}),
|
||||
"depth_scale": ("FLOAT", {"default": 1.0}),
|
||||
"invert_depth": ("BOOLEAN", {"default": False}),
|
||||
# Output projection parameters
|
||||
"output_projection": (Projection.PROJECTIONS, {}),
|
||||
"output_horizontal_fov": ("FLOAT", {"default": 90.0}),
|
||||
"output_width": ("INT", {"default": 512, "min": 1}),
|
||||
"output_height": ("INT", {"default": 512, "min": 1}),
|
||||
"point_size": ("INT", {"default": 1, "min": 1}),
|
||||
# Cleaning parameters
|
||||
"voxel_size": ("FLOAT", {"default": 1.0, "min": 1e-3}),
|
||||
"min_points_per_voxel": ("INT", {"default": 3, "min": 1}),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("IMAGE", "MASK", "TENSOR")
|
||||
RETURN_NAMES = ("video_frames", "mask_frames", "depths")
|
||||
FUNCTION = "process_sequence"
|
||||
CATEGORY = "Camera/Video"
|
||||
|
||||
def process_sequence(
|
||||
self,
|
||||
frames: torch.Tensor,
|
||||
depth_seq: torch.Tensor,
|
||||
trajectory: torch.Tensor,
|
||||
input_projection: str,
|
||||
input_horizontal_fov: float,
|
||||
depth_scale: float,
|
||||
invert_depth: bool,
|
||||
output_projection: str,
|
||||
output_horizontal_fov: float,
|
||||
output_width: int,
|
||||
output_height: int,
|
||||
point_size: int,
|
||||
voxel_size: float,
|
||||
min_points_per_voxel: int,
|
||||
) -> Tuple[torch.Tensor, torch.Tensor, torch.Tensor]:
|
||||
# frames: [T, H, W, 3]
|
||||
# depth_seq: [T, H, W] or [T, H, W, 1]
|
||||
T, H, W, _ = frames.shape
|
||||
|
||||
# Interpolate trajectory to match T
|
||||
K = trajectory.shape[0]
|
||||
if K < 2:
|
||||
interp_traj = trajectory.expand(T, 4, 4).clone()
|
||||
else:
|
||||
idxs = torch.linspace(0, K - 1, T, device=trajectory.device)
|
||||
lower = idxs.floor().long().clamp(max=K - 2)
|
||||
upper = lower + 1
|
||||
alpha = (idxs - lower.float()).unsqueeze(-1).unsqueeze(-1)
|
||||
traj_lower = trajectory[lower]
|
||||
traj_upper = trajectory[upper]
|
||||
interp_traj = traj_lower * (1 - alpha) + traj_upper * alpha
|
||||
|
||||
out_frames = []
|
||||
out_masks = []
|
||||
out_depths = []
|
||||
|
||||
# Add tqdm progress bar for the sequence
|
||||
for frame, depth, pose in tqdm(zip(frames, depth_seq, interp_traj), total=T, desc="Processing video frames"):
|
||||
if depth.dim() == 3 and depth.shape[-1] == 1:
|
||||
depth = depth.squeeze(-1)
|
||||
# to pointcloud
|
||||
pc, = DepthToPointCloud().depth_to_pointcloud(
|
||||
image=frame.permute(2, 0, 1),
|
||||
input_projection=input_projection,
|
||||
input_horizontal_fov=input_horizontal_fov,
|
||||
depth_scale=depth_scale,
|
||||
invert_depth=invert_depth,
|
||||
depthmap=depth,
|
||||
mask=None,
|
||||
)
|
||||
# optional cleaning
|
||||
if min_points_per_voxel > 1:
|
||||
pc, = PointCloudCleaner().clean_pointcloud(
|
||||
pointcloud=pc,
|
||||
width=output_width,
|
||||
height=output_height,
|
||||
voxel_size=voxel_size,
|
||||
min_points_per_voxel=min_points_per_voxel,
|
||||
)
|
||||
# transform and project
|
||||
pc_t, = TransformPointCloud().transform_pointcloud(pc, pose)
|
||||
img_t, mask_t, depth_t = ProjectPointCloud().project_pointcloud(
|
||||
pointcloud=pc_t,
|
||||
output_projection=output_projection,
|
||||
output_horizontal_fov=output_horizontal_fov,
|
||||
output_width=output_width,
|
||||
output_height=output_height,
|
||||
point_size=point_size,
|
||||
)
|
||||
|
||||
out_frames.append(img_t[0])
|
||||
out_masks.append(mask_t)
|
||||
out_depths.append(depth_t)
|
||||
|
||||
return (
|
||||
torch.stack(out_frames, dim=0), # [T, 3, H, W]
|
||||
torch.stack(out_masks, dim=0), # [T, H, W]
|
||||
torch.stack(out_depths, dim=0), # [T, H, W]
|
||||
)
|
||||
|
||||
|
||||
class DepthFramesToVideo:
|
||||
"""
|
||||
Converts a sequence of depth maps into video frame tensors for saving.
|
||||
"""
|
||||
@classmethod
|
||||
def INPUT_TYPES(cls) -> Dict[str, Any]:
|
||||
return {
|
||||
"required": {
|
||||
"depth_seq": ("TENSOR", {"shape_hint": [None, None, None]}),
|
||||
"mask_seq": ("MASK", {"shape_hint": [None, None, None]}),
|
||||
"normalize": ("BOOLEAN", {"default": True}),
|
||||
"invert_depth": ("BOOLEAN", {"default": False}),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("TENSOR", "IMAGE")
|
||||
RETURN_NAMES = ("video_frames", "depth_video")
|
||||
FUNCTION = "depth_to_video_frames"
|
||||
CATEGORY = "Camera/Video"
|
||||
|
||||
def depth_to_video_frames(
|
||||
self,
|
||||
depth_seq: torch.Tensor,
|
||||
normalize: bool,
|
||||
invert_depth: bool,
|
||||
mask_seq: torch.Tensor,
|
||||
) -> Tuple[torch.Tensor, torch.Tensor]:
|
||||
ds = depth_seq.clone().squeeze()
|
||||
if ds.dim() == 2:
|
||||
ds = ds.unsqueeze(0) # [H, W] -> [1, H, W]
|
||||
if ds.dim() != 3:
|
||||
raise ValueError(f"Expected ds to be 3D [T, H, W], got shape {ds.shape}")
|
||||
if invert_depth:
|
||||
ds= 1.0 / (ds + 1e-8) # Avoid division by zero
|
||||
if normalize:
|
||||
# Mask: only normalize where depth > 0
|
||||
mask = mask_seq>0.5
|
||||
if mask.any():
|
||||
#percentile first 10 percent min
|
||||
# sample
|
||||
minv = ds[mask]
|
||||
# sample 10000 and find 10% quantile
|
||||
if minv.numel() > 10000:
|
||||
minv = minv[torch.randperm(minv.numel())[:10000]]
|
||||
|
||||
minv = minv.quantile(0.2)
|
||||
minv = minv if minv > 0.1 else 0.1 # Avoid division by zero
|
||||
#percentile last 10 percent max
|
||||
maxv = ds[mask]
|
||||
if maxv.numel() > 10000:
|
||||
maxv = maxv[torch.randperm(maxv.numel())[:10000]]
|
||||
maxv = maxv.quantile(0.98)
|
||||
maxv = maxv if maxv < 100 else 100
|
||||
print(f"Normalizing depth: min={minv}, max={maxv}")
|
||||
ds_norm = (ds - minv) / (maxv - minv + 1e-8)
|
||||
ds = ds_norm.clamp(0, 1) # torch.where(mask, ds_norm, ds) # Only normalize valid values
|
||||
else:
|
||||
print("Warning: No valid depth values for normalization.")
|
||||
# expand to 3 channels: [T, H, W] -> [T, 3, H, W]
|
||||
raw = depth_seq.clone().squeeze()
|
||||
ds_u8 = (ds * 255.0).round().to(torch.uint8)
|
||||
raw_u8 = (raw.clamp(0, 255)).to(torch.uint8) # if raw is already in a displayable range
|
||||
|
||||
# expand to 3 channels and permute to HWC
|
||||
ds_color = ds_u8.unsqueeze(1).repeat(1, 3, 1, 1).permute(0, 2, 3, 1)
|
||||
raw_color = raw_u8.unsqueeze(1).repeat(1, 3, 1, 1).permute(0, 2, 3, 1)
|
||||
return raw_color, ds_color # [T, 3, H, W] -> [T, H, W, 3]
|
||||
|
||||
class VideoMetricDepthEstimate:
|
||||
"""
|
||||
Estimates metric depth for a sequence of frames using VideoDepthAnything.
|
||||
"""
|
||||
@classmethod
|
||||
def INPUT_TYPES(cls) -> Dict[str, Any]:
|
||||
# model files (.pth) in input directory
|
||||
model_dir = os.path.join(os.getcwd(), "models", "checkpoints")
|
||||
os.makedirs(model_dir, exist_ok=True)
|
||||
files = [f for f in os.listdir(model_dir) if f.lower().endswith(('.pth', '.ckpt', '.safetensors'))]
|
||||
return {
|
||||
"required": {
|
||||
"frames": ("IMAGE", {"shape_hint": [None, None, None, 3]}),
|
||||
"model_checkpoint": (files, {"file_chooser": True}),
|
||||
"input_size": ("INT", {"default": 518, "min": 64, "max": 2048}),
|
||||
"max_fps": ("INT", {"default": 60, "min": 1}),
|
||||
}
|
||||
}
|
||||
RETURN_TYPES = ("TENSOR", "FLOAT")
|
||||
RETURN_NAMES = ("metric_depths", "fps")
|
||||
FUNCTION = "estimate_metric_depth"
|
||||
CATEGORY = "Camera/Video"
|
||||
|
||||
def estimate_metric_depth(
|
||||
self,
|
||||
frames: torch.Tensor,
|
||||
model_checkpoint: str,
|
||||
input_size: int,
|
||||
max_fps: int,
|
||||
) -> Tuple[torch.Tensor, float]:
|
||||
if VideoDepthAnything is None:
|
||||
raise ImportError("VideoDepthAnything library not found")
|
||||
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||
# if max input<1.5 normalize to 0-255
|
||||
if frames.max() < 1.5:
|
||||
frames = (frames * 255)
|
||||
model = VideoDepthAnything(**{"encoder": "vitl", "features": 256, "out_channels": [256,512,1024,1024]})
|
||||
state = torch.load("/root/ComfyUI/models/checkpoints/{}".format(model_checkpoint), map_location='cpu')
|
||||
model.load_state_dict(state, strict=True)
|
||||
model = model.to(device).eval()
|
||||
np_frames = frames.cpu().numpy().astype(np.uint8)
|
||||
metric_depths, fps = model.infer_video_depth(np_frames, max_fps, input_size=input_size, device=device.type, fp32=False)
|
||||
return (torch.from_numpy(metric_depths), float(fps))
|
||||
|
||||
# Register nodes
|
||||
NODE_CLASS_MAPPINGS = {
|
||||
"VideoCameraMotionSequence": VideoCameraMotionSequence,
|
||||
"VideoMetricDepthEstimate": VideoMetricDepthEstimate,
|
||||
"DepthFramesToVideo": DepthFramesToVideo,
|
||||
}
|
||||
@@ -1 +1,289 @@
|
||||
{"id":"dd56c0bf-7405-406e-924f-42b2feacb73f","revision":0,"last_node_id":6,"last_link_id":4,"nodes":[{"id":1,"type":"TransformToMatrix","pos":[-337.9580993652344,1500.5211181640625],"size":[315,154],"flags":{},"order":0,"mode":0,"inputs":[],"outputs":[{"localized_name":"MAT_4X4","name":"MAT_4X4","type":"MAT_4X4","links":[1]}],"properties":{"Node name for S&R":"TransformToMatrix"},"widgets_values":[0,0,0,0,0]},{"id":2,"type":"CameraMotion","pos":[130.0218505859375,1516.7567138671875],"size":[367.79998779296875,218],"flags":{},"order":3,"mode":0,"inputs":[{"localized_name":"pointcloud","name":"pointcloud","type":"TENSOR","link":4},{"localized_name":"initial_matrix","name":"initial_matrix","type":"MAT_4X4","link":1},{"localized_name":"final_matrix","name":"final_matrix","type":"MAT_4X4","link":2}],"outputs":[{"localized_name":"IMAGE","name":"IMAGE","type":"IMAGE","links":[3]}],"properties":{"Node name for S&R":"CameraMotion"},"widgets_values":[24,"PINHOLE",90,1024,1024,2]},{"id":3,"type":"SaveWEBM","pos":[606.5880737304688,1515.5966796875],"size":[315,437],"flags":{},"order":4,"mode":0,"inputs":[{"localized_name":"images","name":"images","type":"IMAGE","link":3}],"outputs":[],"properties":{},"widgets_values":["ComfyUI","vp9",10.000000000000002,32]},{"id":5,"type":"TransformToMatrix","pos":[-262.9134521484375,1740.8876953125],"size":[315,154],"flags":{},"order":1,"mode":0,"inputs":[],"outputs":[{"localized_name":"MAT_4X4","name":"MAT_4X4","type":"MAT_4X4","links":[2]}],"properties":{"Node name for S&R":"TransformToMatrix"},"widgets_values":[0.10000000000000002,0,0,0,0]},{"id":6,"type":"LoadPointCloud","pos":[-357.24761962890625,1294.427734375],"size":[315,58],"flags":{},"order":2,"mode":0,"inputs":[],"outputs":[{"localized_name":"TENSOR","name":"TENSOR","type":"TENSOR","links":[4]}],"properties":{"Node name for S&R":"LoadPointCloud"},"widgets_values":["ComfyUIPointCloud_00001.ply"]}],"links":[[1,1,0,2,1,"MAT_4X4"],[2,5,0,2,2,"MAT_4X4"],[3,2,0,3,0,"IMAGE"],[4,6,0,2,0,"TENSOR"]],"groups":[],"config":{},"extra":{"ds":{"scale":1.351305709310409,"offset":[-187.6257577580669,-1567.2143321744395]}},"version":0.4}
|
||||
{
|
||||
"id": "dd56c0bf-7405-406e-924f-42b2feacb73f",
|
||||
"revision": 0,
|
||||
"last_node_id": 8,
|
||||
"last_link_id": 9,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 3,
|
||||
"type": "SaveWEBM",
|
||||
"pos": [
|
||||
606.5880737304688,
|
||||
1515.5966796875
|
||||
],
|
||||
"size": [
|
||||
315,
|
||||
437
|
||||
],
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"ComfyUI",
|
||||
"vp9",
|
||||
10.000000000000002,
|
||||
32
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CameraMotionNode",
|
||||
"pos": [
|
||||
176.07933044433594,
|
||||
1504.22705078125
|
||||
],
|
||||
"size": [
|
||||
278.75,
|
||||
270
|
||||
],
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "pointcloud",
|
||||
"type": "TENSOR",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "trajectory",
|
||||
"type": "TENSOR",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "motion_frames",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
5
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "mask_frames",
|
||||
"type": "MASK",
|
||||
"links": null
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CameraMotionNode"
|
||||
},
|
||||
"widgets_values": [
|
||||
10,
|
||||
"PINHOLE",
|
||||
90,
|
||||
512,
|
||||
512,
|
||||
1,
|
||||
0,
|
||||
false,
|
||||
false
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "LoadPointCloud",
|
||||
"pos": [
|
||||
-357.24761962890625,
|
||||
1294.427734375
|
||||
],
|
||||
"size": [
|
||||
315,
|
||||
58
|
||||
],
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "loaded pointcloud",
|
||||
"type": "TENSOR",
|
||||
"links": [
|
||||
6
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadPointCloud"
|
||||
},
|
||||
"widgets_values": [
|
||||
"ComfyUIPointCloud_00001.ply"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 1,
|
||||
"type": "TransformToMatrix",
|
||||
"pos": [
|
||||
-537.2319946289062,
|
||||
1491.40380859375
|
||||
],
|
||||
"size": [
|
||||
315,
|
||||
154
|
||||
],
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "transformation matrix",
|
||||
"type": "MAT_4X4",
|
||||
"links": [
|
||||
7
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "TransformToMatrix"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
0,
|
||||
0,
|
||||
0,
|
||||
0
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "TransformToMatrix",
|
||||
"pos": [
|
||||
-531.216796875,
|
||||
1701.1632080078125
|
||||
],
|
||||
"size": [
|
||||
315,
|
||||
154
|
||||
],
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "transformation matrix",
|
||||
"type": "MAT_4X4",
|
||||
"links": [
|
||||
8
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "TransformToMatrix"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.10000000000000002,
|
||||
0,
|
||||
0,
|
||||
0,
|
||||
0
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "CameraInterpolationNode",
|
||||
"pos": [
|
||||
-121.91971588134766,
|
||||
1597.3514404296875
|
||||
],
|
||||
"size": [
|
||||
200.21640014648438,
|
||||
46
|
||||
],
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "initial_matrix",
|
||||
"type": "MAT_4X4",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "final_matrix",
|
||||
"type": "MAT_4X4",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "trajectory",
|
||||
"type": "TENSOR",
|
||||
"links": [
|
||||
9
|
||||
]
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CameraInterpolationNode"
|
||||
},
|
||||
"widgets_values": []
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
5,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
6,
|
||||
6,
|
||||
0,
|
||||
7,
|
||||
0,
|
||||
"TENSOR"
|
||||
],
|
||||
[
|
||||
7,
|
||||
1,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"MAT_4X4"
|
||||
],
|
||||
[
|
||||
8,
|
||||
5,
|
||||
0,
|
||||
8,
|
||||
1,
|
||||
"MAT_4X4"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
7,
|
||||
1,
|
||||
"TENSOR"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {
|
||||
"ds": {
|
||||
"scale": 1.015255979947716,
|
||||
"offset": [
|
||||
636.7531305750655,
|
||||
-1186.3099424359816
|
||||
]
|
||||
},
|
||||
"frontendVersion": "1.21.7"
|
||||
},
|
||||
"version": 0.4
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user