Add files via upload
|
After Width: | Height: | Size: 364 KiB |
|
After Width: | Height: | Size: 344 KiB |
@@ -0,0 +1,22 @@
|
||||
## 2024/07/10
|
||||
|
||||
**First, thank you all for your attention, support, sharing, and contributions to LivePortrait!** ❤️
|
||||
The popularity of LivePortrait has exceeded our expectations. If you encounter any issues or other problems and we do not respond promptly, please accept our apologies. We are still actively updating and improving this repository.
|
||||
|
||||
### Updates
|
||||
|
||||
- <strong>Audio and video concatenating: </strong> If the driving video contains audio, it will automatically be included in the generated video. Additionally, the generated video will maintain the same FPS as the driving video. If you run LivePortrait on Windows, you need to install `ffprobe` and `ffmpeg` exe, see issue [#94](https://github.com/KwaiVGI/LivePortrait/issues/94).
|
||||
|
||||
- <strong>Driving video auto-cropping: </strong> Implemented automatic cropping for driving videos by tracking facial landmarks and calculating a global cropping box with a 1:1 aspect ratio. Alternatively, you can crop using video editing software or other tools to achieve a 1:1 ratio. Auto-cropping is not enbaled by default, you can specify it by `--flag_crop_driving_video`.
|
||||
|
||||
- <strong>Motion template making: </strong> Added the ability to create motion templates to protect privacy. The motion template is a `.pkl` file that only contains the motions of the driving video. Theoretically, it is impossible to reconstruct the original face from the template. These motion templates can be used to generate videos without needing the original driving video. By default, the motion template will be generated and saved as a `.pkl` file with the same name as the driving video, e.g., `d0.mp4` -> `d0.pkl`. Once generated, you can specify it using the `-d` or `--driving` option.
|
||||
|
||||
|
||||
### About driving video
|
||||
|
||||
- For a guide on using your own driving video, see the [driving video auto-cropping](https://github.com/KwaiVGI/LivePortrait/tree/main?tab=readme-ov-file#driving-video-auto-cropping) section.
|
||||
|
||||
|
||||
### Others
|
||||
|
||||
- If you encounter a black box problem, disable half-precision inference by using `--no_flag_use_half_precision`, reported by issue [#40](https://github.com/KwaiVGI/LivePortrait/issues/40), [#48](https://github.com/KwaiVGI/LivePortrait/issues/48), [#62](https://github.com/KwaiVGI/LivePortrait/issues/62).
|
||||
@@ -0,0 +1,24 @@
|
||||
## 2024/07/19
|
||||
|
||||
**Once again, we would like to express our heartfelt gratitude for your love, attention, and support for LivePortrait! 🎉**
|
||||
We are excited to announce the release of an implementation of Portrait Video Editing (aka v2v) today! Special thanks to the hard work of the LivePortrait team: [Dingyun Zhang](https://github.com/Mystery099), [Zhizhou Zhong](https://github.com/zzzweakman), and [Jianzhu Guo](https://github.com/cleardusk).
|
||||
|
||||
### Updates
|
||||
|
||||
- <strong>Portrait video editing (v2v):</strong> Implemented a version of Portrait Video Editing (aka v2v). Ensure you have `pykalman` package installed, which has been added in [`requirements_base.txt`](../../../requirements_base.txt). You can specify the source video using the `-s` or `--source` option, adjust the temporal smoothness of motion with `--driving_smooth_observation_variance`, enable head pose motion transfer with `--flag_video_editing_head_rotation`, and ensure the eye-open scalar of each source frame matches the first source frame before animation with `--flag_source_video_eye_retargeting`.
|
||||
|
||||
- <strong>More options in Gradio:</strong> We have upgraded the Gradio interface and added more options. These include `Cropping Options for Source Image or Video` and `Cropping Options for Driving Video`, providing greater flexibility and control.
|
||||
|
||||
<p align="center">
|
||||
<img src="../LivePortrait-Gradio-2024-07-19.jpg" alt="LivePortrait" width="800px">
|
||||
<br>
|
||||
The Gradio Interface for LivePortrait
|
||||
</p>
|
||||
|
||||
|
||||
### Community Contributions
|
||||
|
||||
- **ONNX/TensorRT Versions of LivePortrait:** Explore optimized versions of LivePortrait for faster performance:
|
||||
- [FasterLivePortrait](https://github.com/warmshao/FasterLivePortrait) by [warmshao](https://github.com/warmshao) ([#150](https://github.com/KwaiVGI/LivePortrait/issues/150))
|
||||
- [Efficient-Live-Portrait](https://github.com/aihacker111/Efficient-Live-Portrait) by [aihacker111](https://github.com/aihacker111/Efficient-Live-Portrait) ([#126](https://github.com/KwaiVGI/LivePortrait/issues/126), [#142](https://github.com/KwaiVGI/LivePortrait/issues/142))
|
||||
- **LivePortrait with [X-Pose](https://github.com/IDEA-Research/X-Pose) Detection:** Check out [LivePortrait](https://github.com/ShiJiaying/LivePortrait) by [ShiJiaying](https://github.com/ShiJiaying) for enhanced detection capabilities using X-pose, see [#119](https://github.com/KwaiVGI/LivePortrait/issues/119).
|
||||
@@ -0,0 +1,12 @@
|
||||
## 2024/07/24
|
||||
|
||||
### Updates
|
||||
|
||||
- **Portrait pose editing:** You can change the `relative pitch`, `relative yaw`, and `relative roll` in the Gradio interface to adjust the pose of the source portrait.
|
||||
- **Detection threshold:** We have added a `--det_thresh` argument with a default value of 0.15 to increase recall, meaning more types of faces (e.g., monkeys, human-like) will be detected. You can set it to other values, e.g., 0.5, by using `python app.py --det_thresh 0.5`.
|
||||
|
||||
<p align="center">
|
||||
<img src="../pose-edit-2024-07-24.jpg" alt="LivePortrait" width="960px">
|
||||
<br>
|
||||
Pose Editing in the Gradio Interface
|
||||
</p>
|
||||
@@ -0,0 +1,75 @@
|
||||
## 2024/08/02
|
||||
|
||||
<table class="center" style="width: 80%; margin-left: auto; margin-right: auto;">
|
||||
<tr>
|
||||
<td style="text-align: center"><b>Animals Singing Dance Monkey 🎤</b></td>
|
||||
</tr>
|
||||
|
||||
<tr>
|
||||
<td style="border: none; text-align: center;">
|
||||
<video controls loop src="https://github.com/user-attachments/assets/38d5b6e5-d29b-458d-9f2c-4dd52546cb41" muted="false" style="width: 60%;"></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
|
||||
🎉 We are excited to announce the release of a new version featuring animals mode, along with several other updates. Special thanks to the dedicated efforts of the LivePortrait team. 💪 We also provided an one-click installer for Windows users, checkout the details [here](./2024-08-05.md).
|
||||
|
||||
### Updates on Animals mode
|
||||
We are pleased to announce the release of the animals mode, which is fine-tuned on approximately 230K frames of various animals (mostly cats and dogs). The trained weights have been updated in the `liveportrait_animals` subdirectory, available on [HuggingFace](https://huggingface.co/KwaiVGI/LivePortrait/tree/main/) or [Google Drive](https://drive.google.com/drive/u/0/folders/1UtKgzKjFAOmZkhNK-OYT0caJ_w2XAnib). You should [download the weights](https://github.com/KwaiVGI/LivePortrait?tab=readme-ov-file#2-download-pretrained-weights) before running. There are two ways to run this mode.
|
||||
|
||||
> Please note that we have not trained the stitching and retargeting modules for the animals model due to several technical issues. _This may be addressed in future updates._ Therefore, we recommend **disabling stitching by setting the `--no_flag_stitching`** option when running the model. Additionally, `paste-back` is also not recommended.
|
||||
|
||||
#### Install X-Pose
|
||||
We have chosen [X-Pose](https://github.com/IDEA-Research/X-Pose) as the keypoints detector for animals. This relies on `transformers==4.22.0` and `pillow>=10.2.0` (which are already updated in `requirements.txt`) and requires building an OP named `MultiScaleDeformableAttention`.
|
||||
|
||||
Refer to the [PyTorch installation](https://github.com/KwaiVGI/LivePortrait?tab=readme-ov-file#for-linux-or-windows-users) for Linux and Windows users.
|
||||
|
||||
|
||||
Next, build the OP `MultiScaleDeformableAttention` by running:
|
||||
```bash
|
||||
cd src/utils/dependencies/XPose/models/UniPose/ops
|
||||
python setup.py build install
|
||||
cd - # this returns to the previous directory
|
||||
```
|
||||
|
||||
To run the model, use the `inference_animals.py` script:
|
||||
```bash
|
||||
python inference_animals.py -s assets/examples/source/s39.jpg -d assets/examples/driving/wink.pkl --no_flag_stitching --driving_multiplier 1.75
|
||||
```
|
||||
|
||||
Alternatively, you can use Gradio for a more user-friendly interface. Launch it with:
|
||||
```bash
|
||||
python app_animals.py # --server_port 8889 --server_name "0.0.0.0" --share
|
||||
```
|
||||
|
||||
> [!WARNING]
|
||||
> [X-Pose](https://github.com/IDEA-Research/X-Pose) is only for Non-commercial Scientific Research Purposes, you should remove and replace it with other detectors if you use it for commercial purposes.
|
||||
|
||||
### Updates on Humans mode
|
||||
|
||||
- **Driving Options**: We have introduced an `expression-friendly` driving option to **reduce head wobbling**, now set as the default. While it may be less effective with large head poses, you can also select the `pose-friendly` option, which is the same as the previous version. This can be set using `--driving_option` or selected in the Gradio interface. Additionally, we added a `--driving_multiplier` option to adjust driving intensity, with a default value of 1, which can also be set in the Gradio interface.
|
||||
|
||||
- **Retargeting Video in Gradio**: We have implemented a video retargeting feature. You can specify a `target lip-open ratio` to adjust the mouth movement in the source video. For instance, setting it to 0 will close the mouth in the source video 🤐.
|
||||
|
||||
### Others
|
||||
|
||||
- [**Poe supports LivePortrait**](https://poe.com/LivePortrait). Check out the news on [X](https://x.com/poe_platform/status/1816136105781256260).
|
||||
- [ComfyUI-LivePortraitKJ](https://github.com/kijai/ComfyUI-LivePortraitKJ) (1.1K 🌟) now includes MediaPipe as an alternative to InsightFace, ensuring the license remains under MIT and Apache 2.0.
|
||||
- [ComfyUI-AdvancedLivePortrait](https://github.com/PowerHouseMan/ComfyUI-AdvancedLivePortrait) features real-time portrait pose/expression editing and animation, and is registered with ComfyUI-Manager.
|
||||
|
||||
|
||||
|
||||
**Below are some screenshots of the new features and improvements:**
|
||||
|
||||
|  |
|
||||
|:---:|
|
||||
| **The Gradio Interface of Animals Mode** |
|
||||
|
||||
|  |
|
||||
|:---:|
|
||||
| **Driving Options and Multiplier** |
|
||||
|
||||
|  |
|
||||
|:---:|
|
||||
| **The Feature of Retargeting Video** |
|
||||
@@ -0,0 +1,18 @@
|
||||
## One-click Windows Installer
|
||||
|
||||
### Download the installer from HuggingFace
|
||||
```bash
|
||||
# !pip install -U "huggingface_hub[cli]"
|
||||
huggingface-cli download cleardusk/LivePortrait-Windows LivePortrait-Windows-v20240806.zip --local-dir ./
|
||||
```
|
||||
|
||||
If you cannot access to Huggingface, you can use [hf-mirror](https://hf-mirror.com/) to download:
|
||||
```bash
|
||||
# !pip install -U "huggingface_hub[cli]"
|
||||
export HF_ENDPOINT=https://hf-mirror.com
|
||||
huggingface-cli download cleardusk/LivePortrait-Windows LivePortrait-Windows-v20240806.zip --local-dir ./
|
||||
```
|
||||
|
||||
Alternatively, you can manually download it from the [HuggingFace](https://huggingface.co/cleardusk/LivePortrait-Windows/blob/main/LivePortrait-Windows-v20240806.zip) page.
|
||||
|
||||
Then, simply unzip the package `LivePortrait-Windows-v20240806.zip` and double-click `run_windows_human.bat` for the Humans mode, or `run_windows_animal.bat` for the **Animals mode**.
|
||||
@@ -0,0 +1,9 @@
|
||||
## Precise Portrait Editing
|
||||
|
||||
Inspired by [ComfyUI-AdvancedLivePortrait](https://github.com/PowerHouseMan/ComfyUI-AdvancedLivePortrait) ([@PowerHouseMan](https://github.com/PowerHouseMan)), we have implemented a version of Precise Portrait Editing in the Gradio interface. With each adjustment of the slider, the edited image updates in real-time. You can click the `🔄 Reset` button to reset all slider parameters. However, the performance may not be as fast as the ComfyUI plugin.
|
||||
|
||||
<p align="center">
|
||||
<img src="../editing-portrait-2024-08-06.jpg" alt="LivePortrait" width="960px">
|
||||
<br>
|
||||
Preciese Portrait Editing in the Gradio Interface
|
||||
</p>
|
||||
@@ -0,0 +1,65 @@
|
||||
## Image Driven and Regional Control
|
||||
|
||||
<p align="center">
|
||||
<img src="../image-driven-image-2024-08-19.jpg" alt="LivePortrait" width="512px">
|
||||
<br>
|
||||
<strong>Image Drives an Image</strong>
|
||||
</p>
|
||||
|
||||
You can now **use an image as a driving signal** to drive the source image or video! Additionally, we **have refined the driving options to support expressions, pose, lips, eyes, or all** (all is consistent with the previous default method), which we name it regional control. The control is becoming more and more precise! 🎯
|
||||
|
||||
> Please note that image-based driving or regional control may not perform well in certain cases. Feel free to try different options, and be patient. 😊
|
||||
|
||||
> [!Note]
|
||||
> We recognize that the project now offers more options, which have become increasingly complex, but due to our limited team capacity and resources, we haven’t fully documented them yet. We ask for your understanding and will work to improve the documentation over time. Contributions via PRs are welcome! If anyone is considering donating or sponsoring, feel free to leave a message in the GitHub Issues or Discussions. We will set up a payment account to reward the team members or support additional efforts in maintaining the project. 💖
|
||||
|
||||
|
||||
### CLI Usage
|
||||
It's very simple to use an image as a driving reference. Just set the `-d` argument to the driving image:
|
||||
|
||||
```bash
|
||||
python inference.py -s assets/examples/source/s5.jpg -d assets/examples/driving/d30.jpg
|
||||
```
|
||||
|
||||
To change the `animation_region` option, you can use the `--animation_region` argument to `exp`, `pose`, `lip`, `eyes`, or `all`. For example, to only drive the lip region, you can run by:
|
||||
|
||||
```bash
|
||||
# only driving the lip region
|
||||
python inference.py -s assets/examples/source/s5.jpg -d assets/examples/driving/d0.mp4 --animation_region lip
|
||||
```
|
||||
|
||||
### Gradio Interface
|
||||
|
||||
<p align="center">
|
||||
<img src="../image-driven-portrait-animation-2024-08-19.jpg" alt="LivePortrait" width="960px">
|
||||
<br>
|
||||
<strong>Image-driven Portrait Animation and Regional Control</strong>
|
||||
</p>
|
||||
|
||||
### More Detailed Explanation
|
||||
|
||||
**flag_relative_motion**:
|
||||
When using an image as the driving input, setting `--flag_relative_motion` to true will apply the motion deformation between the driving image and its canonical form. If set to false, the absolute motion of the driving image is used, which may amplify expression driving strength but could also cause identity leakage. This option corresponds to the `relative motion` toggle in the Gradio interface. Additionally, if both source and driving inputs are images, the output will be an image. If the source is a video and the driving input is an image, the output will be a video, with each frame driven by the image's motion. The Gradio interface automatically saves and displays the output in the appropriate format.
|
||||
|
||||
**animation_region**:
|
||||
This argument offers five options:
|
||||
|
||||
- `exp`: Only the expression of the driving input influences the source.
|
||||
- `pose`: Only the head pose drives the source.
|
||||
- `lip`: Only lip movement drives the source.
|
||||
- `eyes`: Only eye movement drives the source.
|
||||
- `all`: All motions from the driving input are applied.
|
||||
|
||||
You can also select these options directly in the Gradio interface.
|
||||
|
||||
**Editing the Lip Region of the Source Video to a Neutral Expression**:
|
||||
In response to requests for a more neutral lip region in the `Retargeting Video` of the Gradio interface, we've added a `keeping the lip silent` option. When selected, the animated video's lip region will adopt a neutral expression. However, this may cause inter-frame jitter or identity leakage, as it uses a mode similar to absolute driving. Note that the neutral expression may sometimes feature a slightly open mouth.
|
||||
|
||||
**Others**:
|
||||
When both source and driving inputs are videos, the output motion may be a blend of both, due to the default setting of `--flag_relative_motion`. This option uses relative driving, where the motion offset of the current driving frame relative to the first driving frame is added to the source frame's motion. In contrast, `--no_flag_relative_motion` applies the driving frame's motion directly as the final driving motion.
|
||||
|
||||
For CLI usage, to retain only the driving video's motion in the output, use:
|
||||
```bash
|
||||
python inference.py --no_flag_relative_motion
|
||||
```
|
||||
In the Gradio interface, simply uncheck the relative motion option. Note that absolute driving may cause jitter or identity leakage in the animated video.
|
||||
@@ -0,0 +1,28 @@
|
||||
## The directory structure of `pretrained_weights`
|
||||
|
||||
```text
|
||||
pretrained_weights
|
||||
├── insightface
|
||||
│ └── models
|
||||
│ └── buffalo_l
|
||||
│ ├── 2d106det.onnx
|
||||
│ └── det_10g.onnx
|
||||
├── liveportrait
|
||||
│ ├── base_models
|
||||
│ │ ├── appearance_feature_extractor.pth
|
||||
│ │ ├── motion_extractor.pth
|
||||
│ │ ├── spade_generator.pth
|
||||
│ │ └── warping_module.pth
|
||||
│ ├── landmark.onnx
|
||||
│ └── retargeting_models
|
||||
│ └── stitching_retargeting_module.pth
|
||||
└── liveportrait_animals
|
||||
├── base_models
|
||||
│ ├── appearance_feature_extractor.pth
|
||||
│ ├── motion_extractor.pth
|
||||
│ ├── spade_generator.pth
|
||||
│ └── warping_module.pth
|
||||
├── retargeting_models
|
||||
│ └── stitching_retargeting_module.pth
|
||||
└── xpose.pth
|
||||
```
|
||||
|
After Width: | Height: | Size: 82 KiB |
|
After Width: | Height: | Size: 301 KiB |
@@ -0,0 +1,29 @@
|
||||
## Install FFmpeg
|
||||
|
||||
Make sure you have `ffmpeg` and `ffprobe` installed on your system. If you don't have them installed, follow the instructions below.
|
||||
|
||||
> [!Note]
|
||||
> The installation is copied from [SoVITS](https://github.com/RVC-Boss/GPT-SoVITS) 🤗
|
||||
|
||||
### Conda Users
|
||||
|
||||
```bash
|
||||
conda install ffmpeg
|
||||
```
|
||||
|
||||
### Ubuntu/Debian Users
|
||||
|
||||
```bash
|
||||
sudo apt install ffmpeg
|
||||
sudo apt install libsox-dev
|
||||
conda install -c conda-forge 'ffmpeg<7'
|
||||
```
|
||||
|
||||
### Windows Users
|
||||
|
||||
Download and place [ffmpeg.exe](https://huggingface.co/lj1995/VoiceConversionWebUI/blob/main/ffmpeg.exe) and [ffprobe.exe](https://huggingface.co/lj1995/VoiceConversionWebUI/blob/main/ffprobe.exe) in the GPT-SoVITS root.
|
||||
|
||||
### MacOS Users
|
||||
```bash
|
||||
brew install ffmpeg
|
||||
```
|
||||
|
After Width: | Height: | Size: 325 KiB |
|
After Width: | Height: | Size: 544 KiB |
|
After Width: | Height: | Size: 491 KiB |
|
After Width: | Height: | Size: 801 KiB |
|
After Width: | Height: | Size: 217 KiB |
|
After Width: | Height: | Size: 115 KiB |
|
After Width: | Height: | Size: 6.3 MiB |
|
After Width: | Height: | Size: 2.7 MiB |
@@ -0,0 +1,13 @@
|
||||
### Speed
|
||||
|
||||
Below are the results of inferring one frame on an RTX 4090 GPU using the native PyTorch framework with `torch.compile`:
|
||||
|
||||
| Model | Parameters(M) | Model Size(MB) | Inference(ms) |
|
||||
|-----------------------------------|:-------------:|:--------------:|:-------------:|
|
||||
| Appearance Feature Extractor | 0.84 | 3.3 | 0.82 |
|
||||
| Motion Extractor | 28.12 | 108 | 0.84 |
|
||||
| Spade Generator | 55.37 | 212 | 7.59 |
|
||||
| Warping Module | 45.53 | 174 | 5.21 |
|
||||
| Stitching and Retargeting Modules | 0.23 | 2.3 | 0.31 |
|
||||
|
||||
*Note: The values for the Stitching and Retargeting Modules represent the combined parameter counts and total inference time of three sequential MLP networks.*
|
||||
|
After Width: | Height: | Size: 113 KiB |
|
After Width: | Height: | Size: 96 KiB |
|
After Width: | Height: | Size: 525 KiB |
|
After Width: | Height: | Size: 102 KiB |
|
After Width: | Height: | Size: 486 KiB |
|
After Width: | Height: | Size: 742 KiB |
|
After Width: | Height: | Size: 62 KiB |
|
After Width: | Height: | Size: 140 KiB |
|
After Width: | Height: | Size: 130 KiB |
|
After Width: | Height: | Size: 105 KiB |
|
After Width: | Height: | Size: 137 KiB |
|
After Width: | Height: | Size: 222 KiB |
|
After Width: | Height: | Size: 432 KiB |