add group offloading
This commit is contained in:
@@ -58,6 +58,7 @@ or you can click <a href="https://github.com/PKU-YuanGroup/Helios-Page/blob/main
|
||||
|
||||
## 📣 Latest News!!
|
||||
|
||||
* `[2026.03.08]` 👋 Helios now fully supports [Group Offloading](#-group-offloading-to-save-vram) and [Context Parallelism](#-context-parallelism-on-multiple-gpus)! These features significantly optimize VRAM (**only ~6GB**) usage and enable inference across multiple GPUs with *Ulysses Attention*, *Ring Attention*, *Unified Attention*, and *Ulysses Anything Attention*.
|
||||
* `[2026.03.06]` 🚀 [Cache-DiT](https://github.com/vipshop/cache-dit/pull/834) now supports Helios, it offers Fully Cache Acceleration and Parallelism support for Helios! Special thanks to the Cache-DiT Team for their amazing work.
|
||||
* `[2026.03.06]` 🚀 We fix the Parallel Inference logits for Helios, and provide an example [here](#-parallel-inference-on-multiple-gpus). Thanks [Cache-DiT Team](https://github.com/vipshop/cache-dit/pull/836).
|
||||
* `[2026.03.06]` 👋 We official release the [Gradio Demo](https://huggingface.co/spaces/BestWishYsh/Helios-14B-RealTime), welcome to try it.
|
||||
@@ -183,8 +184,36 @@ Before trying your own inputs, we highly recommend going through the sanity chec
|
||||
| **T2V** | <video src="https://github.com/user-attachments/assets/14e10753-0366-4790-ad8f-7b66d821ed11" controls width="240"></video> | <video src="https://github.com/user-attachments/assets/c1778691-a80b-428c-8094-88bb1dd1d52b" controls width="240"></video> | <video src="https://github.com/user-attachments/assets/4ca28c79-9dfa-49de-9c3a-f4c7b6c766cd" controls width="240"></video> |
|
||||
| **V2V** | <video src="https://github.com/user-attachments/assets/420cb572-85c2-42d8-98d7-37b0bc24c844" controls width="240"></video> | <video src="https://github.com/user-attachments/assets/7d703fa6-dc1a-4138-a897-e58cfd9236d6" controls width="240"></video> | <video src="https://github.com/user-attachments/assets/45329c55-1a25-459c-bbf0-4e584ec5b23d" controls width="240"></video> |
|
||||
|
||||
|
||||
### ✨ Group Offloading to Save VRAM
|
||||
|
||||
Helios supports group offloading to significantly reduce VRAM consumption, allowing you to run on GPU with limited memory footprint. For more details on the underlying mechanics, please refer to the [documentation](https://huggingface.co/docs/diffusers/main/en/optimization/memory#group-offloading).
|
||||
|
||||
The Helios model below requires `~6GB of VRAM`.
|
||||
|
||||
<details>
|
||||
<summary>Click to expand the code</summary>
|
||||
|
||||
```bash
|
||||
CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--base_model_path "BestWishYsh/Helios-Distilled" \
|
||||
--transformer_path "BestWishYsh/Helios-Distilled" \
|
||||
--sample_type "t2v" \
|
||||
--prompt "A vibrant tropical fish swimming gracefully among colorful coral reefs in a clear, turquoise ocean. The fish has bright blue and yellow scales with a small, distinctive orange spot on its side, its fins moving fluidly. The coral reefs are alive with a variety of marine life, including small schools of colorful fish and sea turtles gliding by. The water is crystal clear, allowing for a view of the sandy ocean floor below. The reef itself is adorned with a mix of hard and soft corals in shades of red, orange, and green. The photo captures the fish from a slightly elevated angle, emphasizing its lively movements and the vivid colors of its surroundings. A close-up shot with dynamic movement." \
|
||||
--num_frames 240 \
|
||||
--guidance_scale 1.0 \
|
||||
--is_enable_stage2 \
|
||||
--pyramid_num_inference_steps_list 2 2 2 \
|
||||
--is_amplify_first_chunk \
|
||||
--output_folder "./output_helios/helios-distilled" \
|
||||
--enable_low_vram_mode \
|
||||
--group_offloading_type "leaf_level"
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
### ✨ Context Parallelism on Multiple GPUs
|
||||
Helios supports various Context Parallelism mechanisms, including Ulysses Attention, Ring Attention, Unified Attention, and Ulysses Anything Attention. For more details, please refer to the [documentation](https://huggingface.co/docs/diffusers/v0.37.0/en/training/distributed_inference#context-parallelism).
|
||||
Helios supports various Context Parallelism mechanisms, including `Ulysses Attention`, `Ring Attention`, `Unified Attention`, and `Ulysses Anything Attention`. For more details, please refer to the [documentation](https://huggingface.co/docs/diffusers/main/en/training/distributed_inference#context-parallelism).
|
||||
|
||||
For example, let's take Helios-Base with 4 GPUs.
|
||||
|
||||
|
||||
+24
-9
@@ -56,8 +56,6 @@ def parse_args():
|
||||
)
|
||||
parser.add_argument("--output_folder", type=str, default="./output_helios")
|
||||
parser.add_argument("--enable_compile", action="store_true")
|
||||
parser.add_argument("--low_vram_mode", action="store_true")
|
||||
parser.add_argument("--enable_parallelism", action="store_true")
|
||||
|
||||
# === Generation parameters ===
|
||||
# environment
|
||||
@@ -141,6 +139,7 @@ def parse_args():
|
||||
|
||||
# === Context parallelism ===
|
||||
# Please refer to https://huggingface.co/docs/diffusers/v0.37.0/en/training/distributed_inference#context-parallelism
|
||||
parser.add_argument("--enable_parallelism", action="store_true")
|
||||
parser.add_argument(
|
||||
"--cp_backend",
|
||||
type=str,
|
||||
@@ -149,14 +148,31 @@ def parse_args():
|
||||
help="Context parallel backend to use.",
|
||||
)
|
||||
|
||||
# === Group-Offloading ===
|
||||
# Please refer to https://huggingface.co/docs/diffusers/v0.37.0/en/optimization/memory#group-offloading
|
||||
parser.add_argument("--enable_low_vram_mode", action="store_true")
|
||||
parser.add_argument(
|
||||
"--group_offloading_type",
|
||||
type=str,
|
||||
choices=["leaf_level", "block_level"],
|
||||
default="leaf_level",
|
||||
help="Specifies the granularity for group CPU offloading. Choose between 'leaf_level' (individual modules) or 'block_level' (entire blocks).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--num_blocks_per_group",
|
||||
type=str,
|
||||
default="4",
|
||||
help="The number of blocks to bundle together in each offloading group. Only relevant when using block-level offloading.",
|
||||
)
|
||||
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main():
|
||||
args = parse_args()
|
||||
|
||||
assert not (args.low_vram_mode and args.enable_compile), (
|
||||
"low_vram_mode and enable_compile cannot be used together."
|
||||
assert not (args.enable_low_vram_mode and args.enable_compile), (
|
||||
"enable_low_vram_mode and enable_compile cannot be used together."
|
||||
)
|
||||
|
||||
if args.weight_dtype == "fp32":
|
||||
@@ -177,7 +193,7 @@ def main():
|
||||
device = torch.device("cuda", rank % torch.cuda.device_count())
|
||||
world_size = dist.get_world_size()
|
||||
torch.cuda.set_device(device)
|
||||
assert world_size == 1 or not args.low_vram_mode, "low_vram_mode is only for single GPU."
|
||||
assert world_size == 1 or not args.enable_low_vram_mode, "enable_low_vram_mode is only for single GPU."
|
||||
else:
|
||||
rank = 0
|
||||
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||
@@ -270,13 +286,12 @@ def main():
|
||||
pipe.vae.compile(mode="max-autotune-no-cudagraphs", dynamic=False)
|
||||
pipe.transformer.compile(mode="max-autotune-no-cudagraphs", dynamic=False)
|
||||
|
||||
if args.low_vram_mode:
|
||||
if args.enable_low_vram_mode:
|
||||
pipe.enable_group_offload(
|
||||
onload_device=torch.device("cuda"),
|
||||
offload_device=torch.device("cpu"),
|
||||
# offload_type="leaf_level",
|
||||
offload_type="block_level",
|
||||
num_blocks_per_group=1,
|
||||
offload_type=args.group_offloading_type,
|
||||
num_blocks_per_group=args.num_blocks_per_group if args.group_offloading_type == "block_level" else None,
|
||||
use_stream=True,
|
||||
record_stream=True,
|
||||
)
|
||||
|
||||
@@ -16,6 +16,9 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-base"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --use_cfg_zero_star \
|
||||
# --use_zero_init \
|
||||
# --zero_steps 1 \
|
||||
@@ -15,6 +15,9 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-base"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --use_cfg_zero_star \
|
||||
# --use_zero_init \
|
||||
# --zero_steps 1 \
|
||||
@@ -16,6 +16,9 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-base"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --use_cfg_zero_star \
|
||||
# --use_zero_init \
|
||||
# --zero_steps 1 \
|
||||
@@ -18,4 +18,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-distilled"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --pyramid_num_inference_steps_list 1 1 1 \
|
||||
@@ -17,4 +17,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-distilled"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --pyramid_num_inference_steps_list 1 1 1 \
|
||||
@@ -18,4 +18,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-distilled"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --pyramid_num_inference_steps_list 1 1 1 \
|
||||
@@ -20,4 +20,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-mid"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --pyramid_num_inference_steps_list 17 17 17 \
|
||||
@@ -19,4 +19,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-mid"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --pyramid_num_inference_steps_list 17 17 17 \
|
||||
@@ -20,4 +20,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
|
||||
--output_folder "./output_helios/helios-mid"
|
||||
|
||||
|
||||
# --enable_low_vram_mode \
|
||||
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
|
||||
# --num_blocks_per_group
|
||||
# --pyramid_num_inference_steps_list 17 17 17 \
|
||||
Reference in New Issue
Block a user