Compare commits

...
Author SHA1 Message Date
Alexander Kharin 74d71a91f2 update splats stitching 2025-12-26 19:27:09 +03:00
Alexander Kharin 55820b53e8 add image to splat submodule 2025-12-26 15:01:42 +03:00
Alexander Kharin f6b6707dd8 add GS nodes 2025-12-25 01:32:59 +03:00
Alexander Kharin 49f9700880 Merge pull request #19 from gitcapoom/claude/camera-comf-modified-011CUt8BMsMPugiCMpz1j9ah
Fix equirectangular projection aspect ratio to 2:1
2025-12-24 23:34:29 +03:00
Claude b961c2988d Revert input vertical FOV to use grid aspect ratio
- Reverts input_vertical_fov calculation to use output grid aspect ratio
- Keeps output_vertical_fov fix for equirectangular (2:1 aspect)
- Fixes vertical squeezing when converting FROM equirectangular
- Input FOV now adapts to output dimensions as originally intended
2025-11-08 19:54:09 +00:00
Claude 2b1b716495 Fix vertical squeezing in equirectangular to pinhole conversion
- Remove output grid aspect ratio from input vertical FOV calculation
- For non-equirectangular inputs, use square aspect (horizontal FOV)
- The normalized_grid already handles output projection aspect ratio
- Fixes 2x vertical compression when converting from equirectangular to pinhole
2025-11-08 18:48:00 +00:00
Claude 839fa5883a Fix width/height swap in ReprojectImage meshgrid
- Corrected meshgrid arguments to use height for y-axis and width for x-axis
- With indexing="ij", first arg should be height, second should be width
- ReprojectDepth was already correct and didn't need changes
2025-11-08 08:43:20 +00:00
Claude 74f98be5c2 Fix equirectangular projection aspect ratio to 2:1
- Set vertical FOV to horizontal FOV / 2 for equirectangular projections
- Applies to both input and output projections
- Eliminates wasted computation on top/bottom quarters
- Reduces resource usage by ~50% for equirectangular outputs
- Affects both ReprojectImage and ReprojectDepth nodes
2025-11-08 07:46:45 +00:00
Alexander Kharin c1e1b55464 replace video with gif for readme 2025-07-13 20:05:49 +02:00
Alexander Kharin 08b48b2fcb Update camera movement on video 2025-07-13 20:03:09 +02:00
Alexander Kharin 4812973e0f Fix background-foregraound. Update video_camera workflow 2025-07-13 19:39:30 +02:00
Alexander Kharin 57fc1e8593 do not import videonodes if video-depth anything is not installed, update projection logics to fill the holes in foreground 2025-07-13 12:27:30 +02:00
Alexander Kharin abdfdc188e fix mask format 2025-07-07 22:59:41 +02:00
Alexander Kharin ca6f904fc3 update projection node to accept masks for processing non square videos. Add video workflow for camera movement in videos 2025-07-07 22:22:42 +02:00
Alexander Kharin 696a4ad763 refactor pointcloud projector for better hole-filling mechanism 2025-07-07 20:15:10 +02:00
Alexander Kharin d67cd61c3c update installation script (add video depth anything). Add some ugly path management for video-depth-anything 2025-06-28 18:30:39 +02:00
Alexander Kharin 17a0edb3c1 fix bug with wrong optional type 2025-06-28 17:29:11 +02:00
Alexander Kharin 5592382b44 Merge pull request #15 from Alexankharin/5-loadpointcloud-node-errors
add video nodes
2025-06-28 17:23:51 +02:00
Alexander Kharin de679db043 add vidoe nodes 2025-06-21 23:59:09 +02:00
Alexander Kharin bc358f8263 Merge pull request #13 from Alexankharin/5-loadpointcloud-node-errors
update test workflow and readme, add trajectory example to load
2025-06-12 21:44:03 +03:00
Alexander Kharin 1d0ae6dc61 Merge branch 'main' into 5-loadpointcloud-node-errors 2025-06-12 21:43:47 +03:00
Alexander Kharin 3bedb49949 update test workflow and readme, add trajectory example to load 2025-06-10 22:51:36 +02:00
Alexander Kharin b67ae9a0a2 Merge pull request #11 from Alexankharin/codex/refactor-node-structuring-by-categories
Improve node category organization
2025-06-10 23:35:34 +03:00
14 changed files with 4521 additions and 130 deletions
+3
View File
@@ -0,0 +1,3 @@
[submodule "submodules/ml-sharpt"]
path = submodules/ml-sharpt
url = https://github.com/apple/ml-sharp
+1542
View File
File diff suppressed because it is too large Load Diff
+49 -6
View File
@@ -1,6 +1,7 @@
# camera-comfyUI
[![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/Alexankharin/camera-comfyUI)
![ComfyUI Custom Nodes](demo_images/Camera_interpolation_pointcloud.gif)
![Camera Movement Demo](demo_images/camera_movement.gif)
> Custom ComfyUI nodes for advanced reprojections, point cloud processing, and camera-driven workflows.
@@ -115,12 +116,20 @@ A collection of ComfyUI custom nodes to handle diverse camera projections (pinho
| `DepthToPointCloud` | Converts Depth and image to → 3D point cloud tensor (N×7). |
| `DepthToImageNode` | Converts depth to image (N×3) using a color map. |
| `ZDepthToRayDepthNode` | Converts Z-depth (output of metric-depth-anything) to ray depth to compensate lens curvature. |
| `TransformPointCloud` | Applies 4×4 rotation matrix to point cloud |
| `TransformPointCloud` | Applies 4×4 rotation matrix to point cloud. |
| `ProjectPointCloud` | Z-buffer–based projection of point cloud into image + mask. |
| `PointCloudCleaner` | Removes isolated points via voxel filtering. |
| `PointCloudUnion` | Combines multiple point clouds into one. |
| `LoadPointCloud` | Loads a point cloud from `.npy` or `.ply` format. |
| `SavePointCloud` | Saves a point cloud to `.npy` or `.ply` format. |
| `CameraMotionNode` | Generates image and mask sequences along a camera trajectory with optional mask dilation/inversion. |
| `CameraInterpolationNode` | Builds a trajectory tensor from two poses. |
| `CameraTrajectoryNode` | Interactive Open3D GUI for recording camera waypoints. |
| `PointCloudCleaner` | Removes isolated points via voxel filtering. |
| `SaveTrajectory` | Saves a trajectory tensor to a file. |
| `LoadTrajectory` | Loads a trajectory tensor from a file. |
| `VideoCameraMotionSequence` | Processes video frames and depth maps along a camera trajectory, generating reprojected outputs. |
| `DepthFramesToVideo` | Converts a sequence of depth maps into video frame tensors for saving. |
| `VideoMetricDepthEstimate` | Estimates metric depth for a sequence of frames using VideoDepthAnything. |
---
@@ -140,6 +149,7 @@ A set of JSON workflows illustrating typical use cases. Each workflow lives in `
| **pointcloud\_inpaint.json** | Inpaint + backproject to 3D for dynamic camera motion videos |
| **Pointcloud\_walker.json** | GUI‐based camera control via Open3D |
| **sbs180\_workflow.json** | Generate stereo (side-by-side) wide-angle/fisheye/equirectangular stereo pairs from a high-res input |
| **video_camera.json** | Camera trajectory movement workflow using `wan-vace` for video inpainting. |
---
@@ -208,11 +218,44 @@ Take a wide-angle (fisheye or equirectangular) high-resolution (e.g., 4096×4096
Interactive Open3D-based GUI for walking and setting camera trajectory inside pointcloud.
### 11. `wan-vace_ref_to_video.json`
### 11. `video_camera.json`
Integrate the [wan2.1-vace] video generation model to inpaint empty or newly revealed regions during camera movement or view synthesis. This workflow demonstrates how to use the camera-comfyUI nodes to generate camera trajectories and masks, then fill missing areas with the video inpainting model for smooth, high-quality results.
This workflow demonstrates camera trajectory movement using the `wan-vace` video inpainting model. It generates smooth camera movements along a trajectory while filling missing regions with high-quality inpainting.
<img src="demo_images/wan-vace-camera.gif" alt="wan2.1-vace Camera Inpainting Demo" width="80%" />
<div style="display:flex; gap:10px;">
<img src="demo_images/camera_movement.gif" alt="Camera Movement Demo" width="80%" />
</div>
---
## Trajectory Concept
A **trajectory** in camera-comfyUI is a sequence of camera poses, each represented as a 4×4 transformation matrix. This set of matrices defines the path and orientation of the camera through 3D space, enabling smooth and complex camera movements for view synthesis, point cloud rendering, and video generation.
### Creating Trajectories
There are two main ways to create a trajectory:
- **Camera Matrices Interpolation:**
Define two or more camera poses (as matrices), and interpolate between them to generate a smooth path. The `CameraInterpolationNode` automates this process, producing a trajectory tensor for use in camera motion nodes.
- **Walking in Open3D Environment:**
Use the interactive Open3D GUI (`CameraTrajectoryNode`) to "walk" through the point cloud. As you move the camera, waypoints (poses) are recorded, forming a trajectory that can be exported and reused.
### Using Trajectories
The `CameraMotionNode` takes a trajectory (set of matrices) and interpolates camera positions and orientations along it, producing smooth camera movements for rendering sequences or videos.
---
## Point Cloud Formats
Point clouds can be saved and loaded in two formats:
- **.npy**: Numpy array format (fast, preserves all tensor data, recommended for internal pipelines).
- **.ply**: Polygon File Format (widely supported, viewable in external 3D tools).
Use the `SavePointCloud` and `LoadPointCloud` nodes to handle I/O operations in either format.
---
@@ -229,5 +272,5 @@ Contributions welcome! Please open issues or PRs to add features, improve docs,
* [x] Implement easier and more flexible camera control - more complex camera movements with more than 2 points.
* [x] Add more examples and documentation for each node.
* [x] Add pointcloud union
* [ ] Fix imports for renamed folders (e.g., inpainting_flux)
* [x] Fix imports for renamed folders (e.g., inpainting_flux)
* [x] Integrate camera movement pipeline with video models (e.g., wan2.1) for smooth, high-quality inpainting along camera trajectories.
+4 -2
View File
@@ -3,6 +3,8 @@ from .reprojection_nodes import NODE_CLASS_MAPPINGS as NCM2
from .metric_depth_nodes import NODE_CLASS_MAPPINGS as NCM3
from .flux_fisheye_filling_nodes import NODE_CLASS_MAPPINGS as NCM4
from .complex_nodes import NODE_CLASS_MAPPINGS as NCM5
NODE_CLASS_MAPPINGS = {**NCM1, **NCM2, **NCM3, **NCM4, **NCM5}
from .video_nodes import NODE_CLASS_MAPPINGS as NCM6
from .GS_nodes import NODE_CLASS_MAPPINGS as NCM7
NODE_CLASS_MAPPINGS = {**NCM1, **NCM2, **NCM3, **NCM4, **NCM5, **NCM6, **NCM7}
__all__ = ["NODE_CLASS_MAPPINGS"]
__all__ = ["NODE_CLASS_MAPPINGS"]
Binary file not shown.
Binary file not shown.
Binary file not shown.

After

Width:  |  Height:  |  Size: 12 MiB

+64 -27
View File
@@ -5,76 +5,104 @@ set -euo pipefail
# Functions
# ----------------------------------------
install_pytorch() {
echo "Installing PyTorch, TorchVision, TorchAudio..."
echo "==> Installing PyTorch, TorchVision, TorchAudio, bitsandbytes, accelerate…"
pip3 install -U torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
pip3 install bitsandbytes
pip3 install accelerate
pip3 install -U bitsandbytes accelerate
}
install_system_deps() {
echo "Updating apt and installing system dependencies..."
echo "==> Updating apt and installing system packages…"
sudo apt-get update
sudo apt-get install -y build-essential ffmpeg libsm6 libxext6
sudo apt-get install -y build-essential ffmpeg libsm6 libxext6 python3.10-dev
}
clone_and_install_comfyui() {
echo "Cloning ComfyUI..."
echo "==> Cloning ComfyUI…"
git clone https://github.com/comfyanonymous/ComfyUI.git
echo "Installing ComfyUI requirements..."
echo "==> Installing ComfyUI Python requirements…"
pip3 install -r ComfyUI/requirements.txt
}
install_camera_node() {
echo "Cloning camera‑ComfyUI..."
echo "==> Installing camera‑ComfyUI node…"
mkdir -p ComfyUI/custom_nodes
git clone https://github.com/Alexankharin/camera-comfyUI.git \
ComfyUI/custom_nodes/camera-comfyUI
echo "Installing camera‑ComfyUI requirements..."
pip3 install -r ComfyUI/custom_nodes/camera-comfyUI/requirements.txt
}
install_image_filters() {
echo "Cloning Image‑Filters node..."
echo "==> Installing Image‑Filters node…"
mkdir -p ComfyUI/custom_nodes
git clone https://github.com/spacepxl/ComfyUI-Image-Filters.git \
ComfyUI/custom_nodes/ComfyUI-Image-Filters
echo "Installing Image‑Filters requirements..."
pip3 install -r ComfyUI/custom_nodes/ComfyUI-Image-Filters/requirements.txt
}
clone_flux_inpainting() {
echo "Cloning ComfyUI‑Flux‑Inpainting..."
echo "==> Installing Flux‑Inpainting node…"
mkdir -p ComfyUI/custom_nodes
git clone https://github.com/rubi-du/ComfyUI-Flux-Inpainting.git \
ComfyUI/custom_nodes/Flux-Inpainting
echo "Renaming Flux‑Inpainting folder..."
mv ComfyUI/custom_nodes/Flux-Inpainting \
ComfyUI/custom_nodes/inpainting_flux
ComfyUI/custom_nodes/inpainting_flux
}
install_metric_video_depth_anything() {
echo "==> Installing Metric Video Depth Anything…"
# Clone into ComfyUI root, not custom_nodes
git clone https://github.com/DepthAnything/Video-Depth-Anything.git \
ComfyUI/Video-Depth-Anything
echo " • Installing easydict…"
pip3 install -U easydict
echo " • Copying util.py…"
mkdir -p ComfyUI/utils
cp ComfyUI/Video-Depth-Anything/metric_depth/utils/util.py \
ComfyUI/utils/util.py
echo " • Downloading Metric Video Depth checkpoint…"
mkdir -p ComfyUI/models/checkpoints
wget -q -O ComfyUI/models/checkpoints/metric_video_depth_anything_vitl.pth \
"https://huggingface.co/depth-anything/Metric-Video-Depth-Anything-Large/resolve/main/metric_video_depth_anything_vitl.pth"
}
install_comfyui_manager() {
echo "==> Installing ComfyUI-Manager extension…"
mkdir -p ComfyUI/custom_nodes
git clone https://github.com/Comfy-Org/ComfyUI-Manager.git \
ComfyUI/custom_nodes/ComfyUI-Manager
pip3 install -r ComfyUI/custom_nodes/ComfyUI-Manager/requirements.txt
}
install_hf_hub() {
echo "Installing huggingface_hub..."
pip3 install huggingface_hub
echo "==> Installing huggingface_hub…"
pip3 install -U huggingface_hub
}
download_vae_models() {
echo "Downloading WAN‑VACE models via wget..."
wget -O ComfyUI/models/vae/wan_2.1_vae.safetensors \
echo "==> Downloading WAN‑VACE models…"
mkdir -p ComfyUI/models/vae
wget -q -O ComfyUI/models/vae/wan_2.1_vae.safetensors \
"https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/vae/wan_2.1_vae.safetensors?download=true"
wget -O ComfyUI/models/text_encoders/umt5_xxl_fp16.safetensors \
mkdir -p ComfyUI/models/text_encoders
wget -q -O ComfyUI/models/text_encoders/umt5_xxl_fp16.safetensors \
"https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/text_encoders/umt5_xxl_fp16.safetensors?download=true"
wget -O ComfyUI/models/diffusion_models/wan2.1_vace_14B_fp16.safetensors \
mkdir -p ComfyUI/models/diffusion_models
wget -q -O ComfyUI/models/diffusion_models/wan2.1_vace_14B_fp16.safetensors \
"https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/diffusion_models/wan2.1_vace_14B_fp16.safetensors"
}
login_hf_hub() {
echo "Logging in to Hugging Face Hub..."
echo "==> Hugging Face login…"
huggingface-cli login
}
# ----------------------------------------
# Main CLI
# Main
# ----------------------------------------
# default to "install" if no arg given
MODE="${1:-install}"
case "$MODE" in
@@ -84,6 +112,7 @@ case "$MODE" in
clone_and_install_comfyui
install_camera_node
install_image_filters
install_comfyui_manager
install_hf_hub
;;
@@ -95,6 +124,10 @@ case "$MODE" in
download_vae_models
;;
depth)
install_metric_video_depth_anything
;;
install)
install_pytorch
install_system_deps
@@ -102,7 +135,9 @@ case "$MODE" in
install_camera_node
clone_flux_inpainting
install_image_filters
install_comfyui_manager
install_hf_hub
install_metric_video_depth_anything
;;
all)
@@ -112,14 +147,16 @@ case "$MODE" in
install_camera_node
clone_flux_inpainting
install_image_filters
install_comfyui_manager
install_hf_hub
download_vae_models
install_metric_video_depth_anything
login_hf_hub
echo "All done! 🎉"
echo "✅ All done!"
;;
*)
echo "Usage: $0 {install|modules|flux|vae|all}"
echo "Usage: $0 {install|modules|flux|vae|depth|all}"
exit 1
;;
esac
+137 -92
View File
@@ -325,120 +325,165 @@ class ProjectPointCloud:
def project_pointcloud(
self,
pointcloud: torch.Tensor,
pointcloud: torch.Tensor,
output_projection: str,
output_horizontal_fov: float,
output_width: int,
output_width: int,
output_height: int,
point_size: int = 1,
point_size: int = 1,
return_inverse_depth: bool = False,
) -> Tuple[torch.Tensor, torch.Tensor, torch.Tensor]:
"""
Projects an (N×6) XYZRGB point cloud into an image,
fills occlusion holes robustly, and returns:
• img: [1,H,W,3] RGB image
• mask: [H,W] foreground mask
• depth: [1,H,W,1] depth (or inverse depth)
"""
device = pointcloud.device
coords = pointcloud[:, :3]
colors = pointcloud[:, 3:].float()
xyz, rgb_raw = pointcloud[:, :3], pointcloud[:, 3:6].float()
# 1) Filter points in front of the camera
mask_front = coords[:, 2] > 0
coords = coords[mask_front]
colors = colors[mask_front]
# 1) Keep only points in front of camera
in_front = xyz[:, 2] > 0
xyz, rgb_raw = xyz[in_front], rgb_raw[in_front]
# 2) Project to normalized UV + depth
X, Y, Z = coords.unbind(1)
X, Y, Z = xyz.unbind(1)
if output_projection == "PINHOLE":
u, v, depth = XYZ_to_pinhole(X, Y, Z, output_horizontal_fov)
u, v, d = XYZ_to_pinhole(X, Y, Z, output_horizontal_fov)
elif output_projection == "FISHEYE":
u, v, depth = XYZ_to_fisheye(X, Y, Z, output_horizontal_fov)
u, v, d = XYZ_to_fisheye(X, Y, Z, output_horizontal_fov)
else:
u, v, depth = XYZ_to_equirect(X, Y, Z, output_horizontal_fov)
u, v, d = XYZ_to_equirect(X, Y, Z, output_horizontal_fov)
# 3) Rasterize to pixel indices
px = (u * (output_width - 1) / 2) + (output_width - 1) / 2
py = (v * (output_height - 1) / 2) + (output_height - 1) / 2
ix = px.round().clamp(0, output_width - 1).long()
iy = py.round().clamp(0, output_height - 1).long()
pix = iy * output_width + ix
M = output_width * output_height
W, H = output_width, output_height
ix = ((u * 0.5 + 0.5) * (W - 1)).round().clamp(0, W - 1).long()
iy = ((v * 0.5 + 0.5) * (H - 1)).round().clamp(0, H - 1).long()
pix = iy * W + ix
valid = (pix >= 0) & (pix < W * H)
pix, d, rgb_raw = pix[valid], d[valid], rgb_raw[valid]
# —— NEW: drop any invalid / NaN→int_min projections ——
valid = (pix >= 0) & (pix < M)
depth = depth[valid]
colors = colors[valid]
pix = pix[valid]
order = torch.arange(depth.size(0), device=device)
# rebuild your "order" to match
M = W * H
# 3a) Front‐layer (nearest) depth
z1 = torch.full((M,), float('inf'), device=device)
z1.scatter_reduce_(0, pix, d, reduce='amin', include_self=True)
# 4) Allocate or reuse buffers
if not hasattr(self, '_z_front') or self._z_front.numel() != M:
self._z_front = torch.empty((M,), device=device)
self._z_back = torch.empty((M,), device=device)
self._idx = torch.full((M,), -1, dtype=torch.long, device=device)
self._flat = torch.zeros((M, 4), device=device)
z_front = self._z_front
z_back = self._z_back
idxbuf = self._idx
flat = self._flat
# 3b) Second‐layer (background) depth
farther = d > z1[pix]
pix2, d2 = pix[farther], d[farther]
z2 = torch.full((M,), float('inf'), device=device)
z2.scatter_reduce_(0, pix2, d2, reduce='amin', include_self=True)
# 5) Front z-buffer pass (nearest)
z_front.fill_(float('inf'))
z_front.scatter_reduce_(0, pix, depth, reduce='amin', include_self=True)
sel_front = depth == z_front[pix]
order = torch.arange(depth.size(0), device=device)
order_m = torch.where(sel_front, order, depth.size(0))
idxbuf.fill_(depth.size(0))
idxbuf.scatter_reduce_(0, pix, order_m, reduce='amin', include_self=True)
win_front = order == idxbuf[pix]
# 3c) Foreground colour (from z1)
keep = d == z1[pix]
rgb = torch.zeros((M, 3), device=device)
rgb[pix[keep]] = rgb_raw[keep].clamp(0, 255)
flat.fill_(0)
flat[pix[win_front]] = colors[win_front]
img4 = flat.view(output_height, output_width, 4)
rgb = img4[..., :3].clamp(0, 255)
alpha = (img4[..., 3] > 0).float()
rgb *= alpha.unsqueeze(-1)
depth_img = z_front.view(output_height, output_width)
rgb_HR = rgb
# reshape to image
rgb = rgb.view(H, W, 3)
z1 = z1.view(H, W)
z2 = z2.view(H, W)
fg_mask = z1 < float('inf') # has front hit
occl = (z2 < float('inf')) # has any back hit
rear_only = occl & ~fg_mask # true holes
# 6) Back z-buffer pass (farthest) for hole-filling
# ── Iterative ring‐based in‐painting of rear‐only pixels ─────────────────────
ker3 = torch.ones((1,1,3,3), device=device)
ker3c = ker3.repeat(3,1,1,1)
for _ in range(max(W, H)):
# find rear_only pixels adjacent to current FG
neigh = (
F.max_pool2d(fg_mask.float()[None,None], 3, 1, 1).bool()[0,0]
& ~fg_mask
)
to_fill = rear_only & neigh
if not to_fill.any():
break
# average depth + colour from current FG frontier
d_t = z1.masked_fill(~fg_mask, 0)[None,None]
c_t = rgb.permute(2,0,1)[None] # [1,3,H,W]
m_t = fg_mask.float()[None,None]
sum_d = F.conv2d(d_t * m_t, ker3, padding=1)
cnt_d = F.conv2d(m_t, ker3, padding=1).clamp(min=1)
sum_c = F.conv2d(c_t * m_t, ker3c, padding=1, groups=3)
cnt_c = cnt_d.repeat(1,3,1,1)
avg_d = (sum_d / cnt_d).squeeze()
avg_c = (sum_c / cnt_c).squeeze().permute(1,2,0)
z1[to_fill] = avg_d[to_fill]
rgb[to_fill] = avg_c[to_fill]
fg_mask[to_fill] = True
rear_only[to_fill] = False
# ── Depth‐aware generic hole closure ─────────────────────────────────────────
# close_rad: radius of hole to close; depth_eps: depth jump tolerance
close_rad = max(1, point_size // 2)
depth_eps = 0.015
pad = close_rad
k = 2 * close_rad + 1
ker = torch.ones((1,1,k,k), device=device)
kerc = ker.repeat(3,1,1,1)
front_t = fg_mask.float()[None,None]
# binary closing: dilate then erode
D = F.max_pool2d(front_t, k, 1, pad)
E = 1 - F.max_pool2d(1 - D, k, 1, pad)
small_hole = E[0,0].bool() & ~fg_mask
if small_hole.any():
# compute local mean depth of FG
z_t = z1.masked_fill(~fg_mask, 0)[None,None]
cnt = F.conv2d(front_t, ker, padding=pad).clamp(min=1)
z_avg = (F.conv2d(z_t, ker, padding=pad) / cnt)[0,0]
# depth‐range test
z_near = F.max_pool2d(z1[None,None], 3,1,1)[0,0]
z_far = -F.max_pool2d(-z1[None,None],3,1,1)[0,0]
flat = (z_far - z_near) / z_avg.clamp(min=1e-6) < depth_eps
final = small_hole & flat
if final.any():
sum_d = F.conv2d(z_t, ker, padding=pad)
sum_c = F.conv2d(rgb.permute(2,0,1)[None] * front_t, kerc,
padding=pad, groups=3)
avg_d = (sum_d / cnt)[0,0]
avg_c = (sum_c / cnt.repeat(1,3,1,1))[0].permute(1,2,0)
z1[final] = avg_d[final]
rgb[final] = avg_c[final]
fg_mask[final] = True
# ── Optional morphological blur for larger point_size ───────────────────────
if point_size > 1:
z_back.fill_(-float('inf'))
z_back.scatter_reduce_(0, pix, depth, reduce='amax', include_self=True)
sel_back = depth == z_back[pix]
order_m = torch.where(sel_back, order, -1)
idxbuf.fill_(-1)
idxbuf.scatter_reduce_(0, pix, order_m, reduce='amax', include_self=True)
win_back = idxbuf[pix] >= 0
r = point_size // 2
k = 2 * r + 1
pad = r
ker = torch.ones((1,1,k,k), device=device)
kerc = ker.repeat(3,1,1,1)
d_t = z1[None,None]
c_t = rgb.permute(2,0,1)[None]
m_t = fg_mask.float()[None,None]
flat.fill_(0)
flat[pix[win_back]] = colors[win_back]
back4 = flat.view(output_height, output_width, 4)
rgb_back = back4[..., :3].clamp(0,255)
alpha_back = (back4[..., 3] > 0).float()
z1 = (F.conv2d(d_t * m_t, ker, padding=pad) /
F.conv2d(m_t, ker, padding=pad).clamp(min=1)).squeeze()
rgb = (F.conv2d(c_t * m_t, kerc, padding=pad, groups=3) /
F.conv2d(m_t, ker, padding=pad).repeat(1,3,1,1).clamp(min=1)
).squeeze().permute(1,2,0)
# fill holes where front missed
hole = (alpha == 0) & (alpha_back > 0)
rgb[hole] = rgb_back[hole]
alpha[hole] = 1.0
depth_img[hole] = z_back.view(output_height, output_width)[hole]
# 7) Median-filter _only_ in hole regions
if hole.any():
# prepare for kornia median_blur: [B,C,H,W]
rgb_t = rgb.permute(2,0,1).unsqueeze(0) # [1,3,H,W]
# apply median filter
rgb_med = median_blur(rgb_t, (point_size, point_size))
# back to HWC
rgb_med = rgb_med.squeeze(0).permute(1,2,0)
# merge only at hole locations
rgb[hole] = rgb_med[hole]
# alpha already set to 1.0 for holes
# 8) Pack and return with original script shapes
img = rgb.unsqueeze(0) # [1,H,W,3]
mask_out = alpha # [H,W]
depth4 = depth_img.unsqueeze(0).unsqueeze(-1) # [1,H,W,1]
# 10) Pack outputs
img = rgb.unsqueeze(0) # [1,H,W,3]
mask = fg_mask.float() # [H,W]
depth = z1.unsqueeze(0).unsqueeze(-1) # [1,H,W,1]
if return_inverse_depth:
depth4 = 1.0 / depth4.clamp(min=1e-6)
depth4 = depth4 * mask_out.unsqueeze(0).unsqueeze(-1)
return img, mask_out, depth4
depth = 1.0 / depth.clamp(min=1e-6)
depth *= mask.unsqueeze(0).unsqueeze(-1)
return img, mask, depth
class PointCloudUnion:
"""
@@ -792,7 +837,7 @@ class CameraTrajectoryNode:
"pointcloud": ("TENSOR",),
},
"optional": {
"initial_matrix": ("MAT_4X4"),
"initial_matrix": ("MAT_4X4",),
}
}
+9 -2
View File
@@ -42,7 +42,14 @@ def map_grid(
output_horizontal_fov = torch.tensor(output_horizontal_fov, device=grid_torch.device).float()
# Calculate vertical field of view for input and output projections
output_vertical_fov = output_horizontal_fov # Assuming square aspect ratio
# For equirectangular, use 2:1 aspect ratio (vertical FOV = horizontal FOV / 2)
if output_projection == "EQUIRECTANGULAR":
output_vertical_fov = output_horizontal_fov / 2.0
else:
output_vertical_fov = output_horizontal_fov # Assuming square aspect ratio for other projections
# Calculate input vertical FOV based on output grid aspect ratio
# This allows the input's vertical range to adapt to the output dimensions
input_vertical_fov = input_horizontal_fov * (grid_torch.shape[0] / grid_torch.shape[1])
# Normalize the grid for vertical FOV adjustment
@@ -222,8 +229,8 @@ class ReprojectImage:
)
grid_y, grid_x = torch.meshgrid(
torch.linspace(-1, 1, output_width, device=image_tensor.device),
torch.linspace(-1, 1, output_height, device=image_tensor.device),
torch.linspace(-1, 1, output_width, device=image_tensor.device),
indexing="ij"
)
grid_init = torch.stack((grid_x, grid_y), dim=-1)
+292
View File
@@ -0,0 +1,292 @@
import torch
import torch.nn.functional as F
import numpy as np
import os
import sys
from typing import Dict, Any, Tuple
from tqdm import tqdm # Added tqdm import
# Import existing pointcloud nodes and projection definitions
from .pointcloud_nodes import DepthToPointCloud, TransformPointCloud, ProjectPointCloud, Projection, PointCloudCleaner
import folder_paths
# Ensure video_depth_anything is on path
_here = os.path.dirname(os.path.abspath(__file__))
# climb up 3 levels: camera-comfyUI → custom_nodes → ComfyUI
COMFYUI_ROOT = os.path.abspath(os.path.join(_here, os.pardir, os.pardir))
# point at metric_depth inside the Video-Depth-Anything clone at the ComfyUI root
video_depth_path = os.path.join(COMFYUI_ROOT, "Video-Depth-Anything", "metric_depth")
# insert at front so it always wins
if video_depth_path not in sys.path:
sys.path.insert(0, video_depth_path)
NO_VIDEO_DEPTH_ANYTHING= False
try:
from video_depth_anything.video_depth import VideoDepthAnything
print("✅ video_depth_anything module loaded successfully.")
except ImportError as e:
NO_VIDEO_DEPTH_ANYTHING = True
print(
f"❌ Could not load video_depth_anything from {video_depth_path!r}: {e}"
)
class VideoCameraMotionSequence:
"""
Takes a sequence of RGB frames and corresponding depth maps,
converts each frame+depth to a pointcloud, interpolates a camera
trajectory to match video length, cleans the pointcloud if needed,
and outputs reprojected images, masks, and depth maps per frame.
"""
@classmethod
def INPUT_TYPES(cls) -> Dict[str, Any]:
return {
"required": {
# Sequence of frames: Tensor [T, H, W, 3]
"frames": ("IMAGE", {"shape_hint": [None, None, None, 3]}),
# Sequence of depth maps: Tensor [T, H, W] or [T, H, W, 1]
"depth_seq": ("TENSOR", {"shape_hint": [None, None, None]}),
# Camera trajectory waypoints: Tensor [K, 4, 4]
"trajectory": ("TENSOR", {"shape_hint": [None, 4, 4]}),
# Optional mask sequence: Tensor [T, H, W] or [T, H, W, 1]
"mask_seq": ("MASK", {"shape_hint": [None, None, None], "optional": True}),
# Input projection parameters
"input_projection": (Projection.PROJECTIONS, {}),
"input_horizontal_fov": ("FLOAT", {"default": 90.0}),
"depth_scale": ("FLOAT", {"default": 1.0}),
"invert_depth": ("BOOLEAN", {"default": False}),
# Output projection parameters
"output_projection": (Projection.PROJECTIONS, {}),
"output_horizontal_fov": ("FLOAT", {"default": 90.0}),
"output_width": ("INT", {"default": 512, "min": 1}),
"output_height": ("INT", {"default": 512, "min": 1}),
"point_size": ("INT", {"default": 1, "min": 1}),
# Cleaning parameters
"voxel_size": ("FLOAT", {"default": 1.0, "min": 1e-3}),
"min_points_per_voxel": ("INT", {"default": 3, "min": 1}),
}
}
RETURN_TYPES = ("IMAGE", "MASK", "TENSOR")
RETURN_NAMES = ("video_frames", "mask_frames", "depths")
FUNCTION = "process_sequence"
CATEGORY = "Camera/Video"
def process_sequence(
self,
frames: torch.Tensor,
depth_seq: torch.Tensor,
trajectory: torch.Tensor,
input_projection: str,
input_horizontal_fov: float,
depth_scale: float,
invert_depth: bool,
output_projection: str,
output_horizontal_fov: float,
output_width: int,
output_height: int,
point_size: int,
voxel_size: float,
min_points_per_voxel: int,
mask_seq: torch.Tensor = None,
) -> Tuple[torch.Tensor, torch.Tensor, torch.Tensor]:
# frames: [T, H, W, 3]
# depth_seq: [T, H, W] or [T, H, W, 1]
T, H, W, _ = frames.shape
# Interpolate trajectory to match T
K = trajectory.shape[0]
if K < 2:
interp_traj = trajectory.expand(T, 4, 4).clone()
else:
idxs = torch.linspace(0, K - 1, T, device=trajectory.device)
lower = idxs.floor().long().clamp(max=K - 2)
upper = lower + 1
alpha = (idxs - lower.float()).unsqueeze(-1).unsqueeze(-1)
traj_lower = trajectory[lower]
traj_upper = trajectory[upper]
interp_traj = traj_lower * (1 - alpha) + traj_upper * alpha
out_frames = []
out_masks = []
out_depths = []
# Add tqdm progress bar for the sequence
# If mask_seq is a single mask [H, W] or [H, W, 1], repeat it for all frames
if mask_seq is not None:
if mask_seq.dim() == 2 or (mask_seq.dim() == 3 and mask_seq.shape[0] == 1):
mask_seq = mask_seq.unsqueeze(0) if mask_seq.dim() == 2 else mask_seq
mask_seq = mask_seq.repeat(T, 1, 1, 1) if mask_seq.dim() == 4 else mask_seq.repeat(T, 1, 1)
for i, (frame, depth, pose) in enumerate(tqdm(zip(frames, depth_seq, interp_traj), total=T, desc="Processing video frames")):
if depth.dim() == 3 and depth.shape[-1] == 1:
depth = depth.squeeze(-1)
# Use mask if provided
mask = None
if mask_seq is not None:
mask = mask_seq[i]
if mask.dim() == 3 and mask.shape[-1] == 1:
mask = mask.squeeze(-1)
# to pointcloud
pc, = DepthToPointCloud().depth_to_pointcloud(
image=frame.permute(2, 0, 1),
input_projection=input_projection,
input_horizontal_fov=input_horizontal_fov,
depth_scale=depth_scale,
invert_depth=invert_depth,
depthmap=depth,
mask=mask,
)
# optional cleaning
if min_points_per_voxel > 1:
pc, = PointCloudCleaner().clean_pointcloud(
pointcloud=pc,
width=output_width,
height=output_height,
voxel_size=voxel_size,
min_points_per_voxel=min_points_per_voxel,
)
# transform and project
pc_t, = TransformPointCloud().transform_pointcloud(pc, pose)
img_t, mask_t, depth_t = ProjectPointCloud().project_pointcloud(
pointcloud=pc_t,
output_projection=output_projection,
output_horizontal_fov=output_horizontal_fov,
output_width=output_width,
output_height=output_height,
point_size=point_size,
)
out_frames.append(img_t[0])
out_masks.append(mask_t)
out_depths.append(depth_t)
return (
torch.stack(out_frames, dim=0), # [T, 3, H, W]
torch.stack(out_masks, dim=0), # [T, H, W]
torch.stack(out_depths, dim=0), # [T, H, W]
)
class DepthFramesToVideo:
"""
Converts a sequence of depth maps into video frame tensors for saving.
"""
@classmethod
def INPUT_TYPES(cls) -> Dict[str, Any]:
return {
"required": {
"depth_seq": ("TENSOR", {"shape_hint": [None, None, None]}),
"mask_seq": ("MASK", {"shape_hint": [None, None, None]}),
"normalize": ("BOOLEAN", {"default": True}),
"invert_depth": ("BOOLEAN", {"default": False}),
}
}
RETURN_TYPES = ("TENSOR", "IMAGE")
RETURN_NAMES = ("video_frames", "depth_video")
FUNCTION = "depth_to_video_frames"
CATEGORY = "Camera/Video"
def depth_to_video_frames(
self,
depth_seq: torch.Tensor,
normalize: bool,
invert_depth: bool,
mask_seq: torch.Tensor,
) -> Tuple[torch.Tensor, torch.Tensor]:
ds = depth_seq.clone().squeeze()
if ds.dim() == 2:
ds = ds.unsqueeze(0) # [H, W] -> [1, H, W]
if ds.dim() != 3:
raise ValueError(f"Expected ds to be 3D [T, H, W], got shape {ds.shape}")
if invert_depth:
ds= 1.0 / (ds + 1e-8) # Avoid division by zero
if normalize:
# Mask: only normalize where depth > 0
mask = mask_seq>0.5
if mask.any():
#percentile first 10 percent min
# sample
minv = ds[mask]
# sample 10000 and find 10% quantile
if minv.numel() > 10000:
minv = minv[torch.randperm(minv.numel())[:10000]]
minv = minv.quantile(0.2)
minv = minv if minv > 0.1 else 0.1 # Avoid division by zero
#percentile last 10 percent max
maxv = ds[mask]
if maxv.numel() > 10000:
maxv = maxv[torch.randperm(maxv.numel())[:10000]]
maxv = maxv.quantile(0.98)
maxv = maxv if maxv < 100 else 100
print(f"Normalizing depth: min={minv}, max={maxv}")
ds_norm = (ds - minv) / (maxv - minv + 1e-8)
ds = ds_norm.clamp(0, 1) # torch.where(mask, ds_norm, ds) # Only normalize valid values
else:
print("Warning: No valid depth values for normalization.")
# expand to 3 channels: [T, H, W] -> [T, 3, H, W]
raw = depth_seq.clone().squeeze()
ds_u8 = (ds * 255.0).round().to(torch.uint8)
raw_u8 = (raw.clamp(0, 255)).to(torch.uint8) # if raw is already in a displayable range
# expand to 3 channels and permute to HWC
ds_color = ds_u8.unsqueeze(1).repeat(1, 3, 1, 1).permute(0, 2, 3, 1)
raw_color = raw_u8.unsqueeze(1).repeat(1, 3, 1, 1).permute(0, 2, 3, 1)
return raw_color, ds_color # [T, 3, H, W] -> [T, H, W, 3]
class VideoMetricDepthEstimate:
"""
Estimates metric depth for a sequence of frames using VideoDepthAnything.
"""
@classmethod
def INPUT_TYPES(cls) -> Dict[str, Any]:
# model files (.pth) in input directory
model_dir = os.path.join(os.getcwd(), "models", "checkpoints")
os.makedirs(model_dir, exist_ok=True)
files = [f for f in os.listdir(model_dir) if f.lower().endswith(('.pth', '.ckpt', '.safetensors'))]
return {
"required": {
"frames": ("IMAGE", {"shape_hint": [None, None, None, 3]}),
"model_checkpoint": (files, {"file_chooser": True}),
"input_size": ("INT", {"default": 518, "min": 64, "max": 2048}),
"max_fps": ("INT", {"default": 60, "min": 1}),
}
}
RETURN_TYPES = ("TENSOR", "FLOAT")
RETURN_NAMES = ("metric_depths", "fps")
FUNCTION = "estimate_metric_depth"
CATEGORY = "Camera/Video"
def estimate_metric_depth(
self,
frames: torch.Tensor,
model_checkpoint: str,
input_size: int,
max_fps: int,
) -> Tuple[torch.Tensor, float]:
if VideoDepthAnything is None:
raise ImportError("VideoDepthAnything library not found")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# if max input<1.5 normalize to 0-255
if frames.max() < 1.5:
frames = (frames * 255)
model = VideoDepthAnything(**{"encoder": "vitl", "features": 256, "out_channels": [256,512,1024,1024]})
state = torch.load("/root/ComfyUI/models/checkpoints/{}".format(model_checkpoint), map_location='cpu')
model.load_state_dict(state, strict=True)
model = model.to(device).eval()
np_frames = frames.cpu().numpy().astype(np.uint8)
metric_depths, fps = model.infer_video_depth(np_frames, max_fps, input_size=input_size, device=device.type, fp32=False)
return (torch.from_numpy(metric_depths), float(fps))
# Register nodes
if NO_VIDEO_DEPTH_ANYTHING:
NODE_CLASS_MAPPINGS = {}
else:
NODE_CLASS_MAPPINGS = {
"VideoCameraMotionSequence": VideoCameraMotionSequence,
"VideoMetricDepthEstimate": VideoMetricDepthEstimate,
"DepthFramesToVideo": DepthFramesToVideo,
}
+289 -1
View File
@@ -1 +1,289 @@
{"id":"dd56c0bf-7405-406e-924f-42b2feacb73f","revision":0,"last_node_id":6,"last_link_id":4,"nodes":[{"id":1,"type":"TransformToMatrix","pos":[-337.9580993652344,1500.5211181640625],"size":[315,154],"flags":{},"order":0,"mode":0,"inputs":[],"outputs":[{"localized_name":"MAT_4X4","name":"MAT_4X4","type":"MAT_4X4","links":[1]}],"properties":{"Node name for S&R":"TransformToMatrix"},"widgets_values":[0,0,0,0,0]},{"id":2,"type":"CameraMotion","pos":[130.0218505859375,1516.7567138671875],"size":[367.79998779296875,218],"flags":{},"order":3,"mode":0,"inputs":[{"localized_name":"pointcloud","name":"pointcloud","type":"TENSOR","link":4},{"localized_name":"initial_matrix","name":"initial_matrix","type":"MAT_4X4","link":1},{"localized_name":"final_matrix","name":"final_matrix","type":"MAT_4X4","link":2}],"outputs":[{"localized_name":"IMAGE","name":"IMAGE","type":"IMAGE","links":[3]}],"properties":{"Node name for S&R":"CameraMotion"},"widgets_values":[24,"PINHOLE",90,1024,1024,2]},{"id":3,"type":"SaveWEBM","pos":[606.5880737304688,1515.5966796875],"size":[315,437],"flags":{},"order":4,"mode":0,"inputs":[{"localized_name":"images","name":"images","type":"IMAGE","link":3}],"outputs":[],"properties":{},"widgets_values":["ComfyUI","vp9",10.000000000000002,32]},{"id":5,"type":"TransformToMatrix","pos":[-262.9134521484375,1740.8876953125],"size":[315,154],"flags":{},"order":1,"mode":0,"inputs":[],"outputs":[{"localized_name":"MAT_4X4","name":"MAT_4X4","type":"MAT_4X4","links":[2]}],"properties":{"Node name for S&R":"TransformToMatrix"},"widgets_values":[0.10000000000000002,0,0,0,0]},{"id":6,"type":"LoadPointCloud","pos":[-357.24761962890625,1294.427734375],"size":[315,58],"flags":{},"order":2,"mode":0,"inputs":[],"outputs":[{"localized_name":"TENSOR","name":"TENSOR","type":"TENSOR","links":[4]}],"properties":{"Node name for S&R":"LoadPointCloud"},"widgets_values":["ComfyUIPointCloud_00001.ply"]}],"links":[[1,1,0,2,1,"MAT_4X4"],[2,5,0,2,2,"MAT_4X4"],[3,2,0,3,0,"IMAGE"],[4,6,0,2,0,"TENSOR"]],"groups":[],"config":{},"extra":{"ds":{"scale":1.351305709310409,"offset":[-187.6257577580669,-1567.2143321744395]}},"version":0.4}
{
"id": "dd56c0bf-7405-406e-924f-42b2feacb73f",
"revision": 0,
"last_node_id": 8,
"last_link_id": 9,
"nodes": [
{
"id": 3,
"type": "SaveWEBM",
"pos": [
606.5880737304688,
1515.5966796875
],
"size": [
315,
437
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 5
}
],
"outputs": [],
"properties": {},
"widgets_values": [
"ComfyUI",
"vp9",
10.000000000000002,
32
]
},
{
"id": 7,
"type": "CameraMotionNode",
"pos": [
176.07933044433594,
1504.22705078125
],
"size": [
278.75,
270
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "pointcloud",
"type": "TENSOR",
"link": 6
},
{
"name": "trajectory",
"type": "TENSOR",
"link": 9
}
],
"outputs": [
{
"name": "motion_frames",
"type": "IMAGE",
"links": [
5
]
},
{
"name": "mask_frames",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "CameraMotionNode"
},
"widgets_values": [
10,
"PINHOLE",
90,
512,
512,
1,
0,
false,
false
]
},
{
"id": 6,
"type": "LoadPointCloud",
"pos": [
-357.24761962890625,
1294.427734375
],
"size": [
315,
58
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "loaded pointcloud",
"type": "TENSOR",
"links": [
6
]
}
],
"properties": {
"Node name for S&R": "LoadPointCloud"
},
"widgets_values": [
"ComfyUIPointCloud_00001.ply"
]
},
{
"id": 1,
"type": "TransformToMatrix",
"pos": [
-537.2319946289062,
1491.40380859375
],
"size": [
315,
154
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformation matrix",
"type": "MAT_4X4",
"links": [
7
]
}
],
"properties": {
"Node name for S&R": "TransformToMatrix"
},
"widgets_values": [
0,
0,
0,
0,
0
]
},
{
"id": 5,
"type": "TransformToMatrix",
"pos": [
-531.216796875,
1701.1632080078125
],
"size": [
315,
154
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformation matrix",
"type": "MAT_4X4",
"links": [
8
]
}
],
"properties": {
"Node name for S&R": "TransformToMatrix"
},
"widgets_values": [
0.10000000000000002,
0,
0,
0,
0
]
},
{
"id": 8,
"type": "CameraInterpolationNode",
"pos": [
-121.91971588134766,
1597.3514404296875
],
"size": [
200.21640014648438,
46
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "initial_matrix",
"type": "MAT_4X4",
"link": 7
},
{
"name": "final_matrix",
"type": "MAT_4X4",
"link": 8
}
],
"outputs": [
{
"name": "trajectory",
"type": "TENSOR",
"links": [
9
]
}
],
"properties": {
"Node name for S&R": "CameraInterpolationNode"
},
"widgets_values": []
}
],
"links": [
[
5,
7,
0,
3,
0,
"IMAGE"
],
[
6,
6,
0,
7,
0,
"TENSOR"
],
[
7,
1,
0,
8,
0,
"MAT_4X4"
],
[
8,
5,
0,
8,
1,
"MAT_4X4"
],
[
9,
8,
0,
7,
1,
"TENSOR"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 1.015255979947716,
"offset": [
636.7531305750655,
-1186.3099424359816
]
},
"frontendVersion": "1.21.7"
},
"version": 0.4
}
File diff suppressed because it is too large Load Diff