add group offloading

This commit is contained in:
SHYuanBest
2026-03-08 07:14:58 +00:00
parent c4ecf8e0d0
commit 5e5537c20f
11 changed files with 81 additions and 10 deletions
+30 -1
View File
@@ -58,6 +58,7 @@ or you can click <a href="https://github.com/PKU-YuanGroup/Helios-Page/blob/main
## 📣 Latest News!!
* `[2026.03.08]` 👋 Helios now fully supports [Group Offloading](#-group-offloading-to-save-vram) and [Context Parallelism](#-context-parallelism-on-multiple-gpus)! These features significantly optimize VRAM (**only ~6GB**) usage and enable inference across multiple GPUs with *Ulysses Attention*, *Ring Attention*, *Unified Attention*, and *Ulysses Anything Attention*.
* `[2026.03.06]` 🚀 [Cache-DiT](https://github.com/vipshop/cache-dit/pull/834) now supports Helios, it offers Fully Cache Acceleration and Parallelism support for Helios! Special thanks to the Cache-DiT Team for their amazing work.
* `[2026.03.06]` 🚀 We fix the Parallel Inference logits for Helios, and provide an example [here](#-parallel-inference-on-multiple-gpus). Thanks [Cache-DiT Team](https://github.com/vipshop/cache-dit/pull/836).
* `[2026.03.06]` 👋 We official release the [Gradio Demo](https://huggingface.co/spaces/BestWishYsh/Helios-14B-RealTime), welcome to try it.
@@ -183,8 +184,36 @@ Before trying your own inputs, we highly recommend going through the sanity chec
| **T2V** | <video src="https://github.com/user-attachments/assets/14e10753-0366-4790-ad8f-7b66d821ed11" controls width="240"></video> | <video src="https://github.com/user-attachments/assets/c1778691-a80b-428c-8094-88bb1dd1d52b" controls width="240"></video> | <video src="https://github.com/user-attachments/assets/4ca28c79-9dfa-49de-9c3a-f4c7b6c766cd" controls width="240"></video> |
| **V2V** | <video src="https://github.com/user-attachments/assets/420cb572-85c2-42d8-98d7-37b0bc24c844" controls width="240"></video> | <video src="https://github.com/user-attachments/assets/7d703fa6-dc1a-4138-a897-e58cfd9236d6" controls width="240"></video> | <video src="https://github.com/user-attachments/assets/45329c55-1a25-459c-bbf0-4e584ec5b23d" controls width="240"></video> |
### ✨ Group Offloading to Save VRAM
Helios supports group offloading to significantly reduce VRAM consumption, allowing you to run on GPU with limited memory footprint. For more details on the underlying mechanics, please refer to the [documentation](https://huggingface.co/docs/diffusers/main/en/optimization/memory#group-offloading).
The Helios model below requires `~6GB of VRAM`.
<details>
<summary>Click to expand the code</summary>
```bash
CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--base_model_path "BestWishYsh/Helios-Distilled" \
--transformer_path "BestWishYsh/Helios-Distilled" \
--sample_type "t2v" \
--prompt "A vibrant tropical fish swimming gracefully among colorful coral reefs in a clear, turquoise ocean. The fish has bright blue and yellow scales with a small, distinctive orange spot on its side, its fins moving fluidly. The coral reefs are alive with a variety of marine life, including small schools of colorful fish and sea turtles gliding by. The water is crystal clear, allowing for a view of the sandy ocean floor below. The reef itself is adorned with a mix of hard and soft corals in shades of red, orange, and green. The photo captures the fish from a slightly elevated angle, emphasizing its lively movements and the vivid colors of its surroundings. A close-up shot with dynamic movement." \
--num_frames 240 \
--guidance_scale 1.0 \
--is_enable_stage2 \
--pyramid_num_inference_steps_list 2 2 2 \
--is_amplify_first_chunk \
--output_folder "./output_helios/helios-distilled" \
--enable_low_vram_mode \
--group_offloading_type "leaf_level"
```
</details>
### ✨ Context Parallelism on Multiple GPUs
Helios supports various Context Parallelism mechanisms, including Ulysses Attention, Ring Attention, Unified Attention, and Ulysses Anything Attention. For more details, please refer to the [documentation](https://huggingface.co/docs/diffusers/v0.37.0/en/training/distributed_inference#context-parallelism).
Helios supports various Context Parallelism mechanisms, including `Ulysses Attention`, `Ring Attention`, `Unified Attention`, and `Ulysses Anything Attention`. For more details, please refer to the [documentation](https://huggingface.co/docs/diffusers/main/en/training/distributed_inference#context-parallelism).
For example, let's take Helios-Base with 4 GPUs.
+24 -9
View File
@@ -56,8 +56,6 @@ def parse_args():
)
parser.add_argument("--output_folder", type=str, default="./output_helios")
parser.add_argument("--enable_compile", action="store_true")
parser.add_argument("--low_vram_mode", action="store_true")
parser.add_argument("--enable_parallelism", action="store_true")
# === Generation parameters ===
# environment
@@ -141,6 +139,7 @@ def parse_args():
# === Context parallelism ===
# Please refer to https://huggingface.co/docs/diffusers/v0.37.0/en/training/distributed_inference#context-parallelism
parser.add_argument("--enable_parallelism", action="store_true")
parser.add_argument(
"--cp_backend",
type=str,
@@ -149,14 +148,31 @@ def parse_args():
help="Context parallel backend to use.",
)
# === Group-Offloading ===
# Please refer to https://huggingface.co/docs/diffusers/v0.37.0/en/optimization/memory#group-offloading
parser.add_argument("--enable_low_vram_mode", action="store_true")
parser.add_argument(
"--group_offloading_type",
type=str,
choices=["leaf_level", "block_level"],
default="leaf_level",
help="Specifies the granularity for group CPU offloading. Choose between 'leaf_level' (individual modules) or 'block_level' (entire blocks).",
)
parser.add_argument(
"--num_blocks_per_group",
type=str,
default="4",
help="The number of blocks to bundle together in each offloading group. Only relevant when using block-level offloading.",
)
return parser.parse_args()
def main():
args = parse_args()
assert not (args.low_vram_mode and args.enable_compile), (
"low_vram_mode and enable_compile cannot be used together."
assert not (args.enable_low_vram_mode and args.enable_compile), (
"enable_low_vram_mode and enable_compile cannot be used together."
)
if args.weight_dtype == "fp32":
@@ -177,7 +193,7 @@ def main():
device = torch.device("cuda", rank % torch.cuda.device_count())
world_size = dist.get_world_size()
torch.cuda.set_device(device)
assert world_size == 1 or not args.low_vram_mode, "low_vram_mode is only for single GPU."
assert world_size == 1 or not args.enable_low_vram_mode, "enable_low_vram_mode is only for single GPU."
else:
rank = 0
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
@@ -270,13 +286,12 @@ def main():
pipe.vae.compile(mode="max-autotune-no-cudagraphs", dynamic=False)
pipe.transformer.compile(mode="max-autotune-no-cudagraphs", dynamic=False)
if args.low_vram_mode:
if args.enable_low_vram_mode:
pipe.enable_group_offload(
onload_device=torch.device("cuda"),
offload_device=torch.device("cpu"),
# offload_type="leaf_level",
offload_type="block_level",
num_blocks_per_group=1,
offload_type=args.group_offloading_type,
num_blocks_per_group=args.num_blocks_per_group if args.group_offloading_type == "block_level" else None,
use_stream=True,
record_stream=True,
)
+3
View File
@@ -16,6 +16,9 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-base"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --use_cfg_zero_star \
# --use_zero_init \
# --zero_steps 1 \
+3
View File
@@ -15,6 +15,9 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-base"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --use_cfg_zero_star \
# --use_zero_init \
# --zero_steps 1 \
+3
View File
@@ -16,6 +16,9 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-base"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --use_cfg_zero_star \
# --use_zero_init \
# --zero_steps 1 \
@@ -18,4 +18,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-distilled"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --pyramid_num_inference_steps_list 1 1 1 \
@@ -17,4 +17,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-distilled"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --pyramid_num_inference_steps_list 1 1 1 \
@@ -18,4 +18,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-distilled"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --pyramid_num_inference_steps_list 1 1 1 \
+3
View File
@@ -20,4 +20,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-mid"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --pyramid_num_inference_steps_list 17 17 17 \
+3
View File
@@ -19,4 +19,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-mid"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --pyramid_num_inference_steps_list 17 17 17 \
+3
View File
@@ -20,4 +20,7 @@ CUDA_VISIBLE_DEVICES=0 python infer_helios.py \
--output_folder "./output_helios/helios-mid"
# --enable_low_vram_mode \
# --group_offloading_type "leaf_level" \ # ["leaf_level", "block_level"]
# --num_blocks_per_group
# --pyramid_num_inference_steps_list 17 17 17 \