AnimateDiff for ComfyUI
Improved AnimateDiff integration for ComfyUI, initially adapted from sd-webui-animatediff but changed greatly since then. Please read the AnimateDiff repo README for more information about how it works at its core.
Examples shown here will also often make use of two helpful set of nodes:
- ComfyUI-Advanced-ControlNet for loading files in batches and controlling which latents should be affected by the ControlNet inputs (work in progress, will include more advance workflows + features for AnimateDiff usage later).
- comfy_controlnet_preprocessors for ControlNet preprocessors not present in vanilla ComfyUI; this repo is archived, and future development by the dev will happen here: comfyui_controlnet_aux. While most preprocessors are common between the two, some give different results. Workflows linked here use the archived version, comfy_controlnet_preprocessors.
Installation
🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥
IMPORTANT: if you already have ArtVentureX's version of AnimateDiff installed, either remove the custom_nodes/comfyui-animatediff folder, uninstall or disable it using ComfyUI-Manager, or add .disabled to the end of that folder's name! Otherwise things will go wrong!
🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥🟥
If using ComfyUI-Manager:
- Look for
AnimateDiff, and be sure it is(Kosinkadink version). Install it.
If installing manually:
- Clone this repo into
custom_nodesfolder.
How to Use:
- Download motion modules. You will need at least 1. Different modules produce different results.
- Original research models available from Google Drive | HuggingFace | CivitAI | Baidu NetDisk.
- Stabilized finetunes of mm_sd_v14 by manshoety from HuggingFace
- Higher resolution finetune by CiaraRowles from HuggingFace
- Place models in
ComfyUI/custom_nodes/ComfyUI-AnimateDiff-Evolved/models. They can be renamed if you want. More motion modules are being trained by the community - if I am made aware of any good ones, I will link here as well. (TODO: create .safetensor versions of the motion modules and share them here.)
- Get creative! If it works for normal image generation, it (probably) will work for AnimateDiff generations. Latent upscales? Go for it. ControlNets, one or more stacked? You betcha. Masking the conditioning of ControlNets to only affect part of the animation? Sure. Try stuff and you will be surprised by what you can do. Samples with workflows are included below.
Current Features:
- txt2img support; if it works for generating ComfyUI images, it will likely work for AnimateDiff. If something does not work, start a discussion or open an issue and I'll see if I can make it work
- ControlNet support - both per-frame, or "interpolating" between frames; can kind of use this as img2video (see workflows below)
- Long animation lengths using sliding context windows {via AnimateDiff Loader (Advanced)}, allowing for longer coherent animations
Upcoming features (aka TODO):
- Prompt travel, and in general more control over per-frame conditioning
- Nodes for saving videos, saving generated files into a timestamped folder instead of all over ComfyUI output dir.
Known Issues (and Solutions, please read!)
Large resolutions may cause xformers to throw a CUDA error concerning a misconfigured value despite being within VRAM limitations.
It is an xformers bug accidentally triggered by the way the original AnimateDiff CrossAttention is passed in. Eventually either I will fix it, or xformers will. When encountered, the workaround is to boot ComfyUI with the "--disable-xformers" argument.
GIF has Watermark (especially when using mm_sd_v15)
Training data used by the authors of the AnimateDiff paper contained Shutterstock watermarks. Since mm_sd_v15 was finetuned on finer, less drastic movement, the motion module attempts to replicate the transparency of that watermark and does not get blurred away like mm_sd_v14. Using other motion modules, or combinations of them using Advanced KSamplers should alleviate watermark issues.
Samples (download or drag images of the workflows into ComfyUI to instantly load the corresponding workflows!)
txt2img
txt2img w/ latent upscale (partial denoise on upscale)
txt2img w/ latent upscale (full denoise on upscale)
txt2img w/ ControlNet-stabilized latent-upscale (partial denoise on upscale, Scaled Soft ControlNet Weights)
txt2img w/ ControlNet-stabilized latent-upscale (full denoise on upscale)
txt2img w/ Initial ControlNet input (using LineArt preprocessor on first txt2img as an example)
txt2img w/ Initial ControlNet input (using OpenPose images) + latent upscale w/ full denoise
(open_pose images provided courtesy of toyxyz)