diff --git a/README.md b/README.md index bae59bb..aaa468b 100644 --- a/README.md +++ b/README.md @@ -12,10 +12,11 @@

Huggingface Model - Github + Github Huggingface Space arXiv YouTube + Demo Page

--- @@ -70,40 +71,17 @@ vae = AutoencoderKLWan.from_pretrained("./vae", torch_dtype=torch.float16) transformer = Transformer3DModel.from_pretrained("./transformer", torch_dtype=torch.float16) scheduler = UniPCMultistepScheduler.from_pretrained("./scheduler") -# Load video and mask -def load_video(path): - vr = VideoReader(path) - imgs = vr.get_batch(list(range(video_length))).asnumpy() - return torch.from_numpy(imgs) / 127.5 - 1.0 # [-1, 1] - -def load_mask(path): - vr = VideoReader(path) - masks = vr.get_batch(list(range(video_length))).asnumpy() - masks = torch.from_numpy(masks)[:, :, :, :1] - masks[masks > 20] = 255 - masks[masks < 255] = 0 - return masks / 255.0 # [0, 1] - -images = load_video("./video.mp4") # images in range [-1, 1] -masks = load_mask("./mask.mp4") # masks in range [0, 1] +images = # images in range [-1, 1] +masks = # masks in range [0, 1] # Initialize the pipeline (pass the loaded weights as objects) -pipe = Minimax_Remover_Pipeline( - vae=vae, - transformer=transformer, - scheduler=scheduler, - torch_dtype=torch.float16 +pipe = Minimax_Remover_Pipeline(vae=vae, transformer=transformer, \ + scheduler=scheduler, torch_dtype=torch.float16 ).to(device) -result = pipe( - images=images, - masks=masks, - num_frames=video_length, - height=480, - width=832, - num_inference_steps=12, - generator=torch.Generator(device=device).manual_seed(random_seed), - iterations=6 +result = pipe(images=images, masks=masks, \ + num_frames=video_length, height=480, width=832, \ + num_inference_steps=12, generator=torch.Generator(device=device).manual_seed(random_seed), iterations=6 \ ).frames[0] export_to_video(result, "./output.mp4") ``` @@ -141,8 +119,8 @@ These control output resolution and the speed/quality tradeoff. Changing them ma |-----------------------|----------------------------------------------------------------|---------| | `iterations` | Mask dilation hyperparameter (controls inpainting margin) | 6 | | `num_frames` | Number of frames to process | 81 | -| `height`, `width` | Output video resolution | 480,832 | -| `num_inference_steps` | Diffusion steps per frame (higher = better quality, slower) | 12 | +| `height`, `width` | Output video resolution | 480x832 | +| `num_inference_steps` | Diffusion steps | 12 | Other parameters can be adjusted in the pipeline call. @@ -150,20 +128,9 @@ Other parameters can be adjusted in the pipeline call. ## 📂 Model Weights -Place model weights in the following directories: - -- `./vae/` -- `./transformer/` -- `./scheduler/` - -You should load each component separately (as shown above) and pass the loaded objects to the pipeline. - ---- - -## 💡 Notes - -- `transformer_minimax_remover` and `pipeline_minimax_remover` are included as project modules and do **not** require separate installation. -- This repository is intended for **academic research only**. **Commercial use is strictly prohibited.** +```shell +huggingface-cli download zibojia/minimax-remover --include vae transformer scheduler --local-dir . +``` ---