diff --git a/README.md b/README.md index 667aae0..a476f89 100644 --- a/README.md +++ b/README.md @@ -93,6 +93,27 @@ Everything else can be the same the same as in the example above. +## HunYuan DiT + +WIP implementation of [HunYuan DiT by Tencent](https://github.com/Tencent/HunyuanDiT) + +> [!WARNING] +> Only a proof of concept, most things don't work yet and only 1024x1024 is supported. +> +> The text encoder device/dtype selection also doesn't work. + +The initial work on this was done by [chaojie](https://github.com/chaojie) in [this PR](https://github.com/city96/ComfyUI_ExtraModels/pull/37). + +Instructions: +- Download the [first text encoder from here](https://huggingface.co/Tencent-Hunyuan/HunyuanDiT/blob/main/t2i/clip_text_encoder/pytorch_model.bin) and place it in `ComfyUI/models/clip` - rename to "chinese-roberta-wwm-ext-large.bin" +- Download the [second text encoder from here](https://huggingface.co/Tencent-Hunyuan/HunyuanDiT/blob/main/t2i/mt5/pytorch_model.bin) and place it in `ComfyUI/models/T5` - rename it to "mT5.bin" +- Download the [model file from here](https://huggingface.co/Tencent-Hunyuan/HunyuanDiT/blob/main/t2i/model/pytorch_model_module.pt) and place it in `ComfyUI/checkpoints` - rename it to "HunYuanDiT.pt" +- Download/use any SDXL VAE, for example [this one](https://huggingface.co/madebyollin/sdxl-vae-fp16-fix) + +![image](https://github.com/city96/ComfyUI_ExtraModels/assets/125218114/7a9d6e34-d3f4-4f67-a17f-4f2d6795e54e) + + + ## DiT [Original Repo](https://github.com/facebookresearch/DiT)