Files
hao-ai-lab-FastVideo/examples/inference/basic/basic_cosmos2_5_v2w.py
T
Davids048andClaude Fable 5.1 a5aa64ab6b [refactor]: decide every config value before the config freezes, and drop the from_pretrained keywords
VideoGenerator.from_pretrained(model_path, config) takes the same nested
settings as from_config; the 24 flat keywords, from_pretrained_kwargs_to_config,
FROM_PRETRAINED_KWARGS, and every flat_field in the schema are deleted.
nvfp4_fa4 is the typed field engine.attention.nvfp4_fa4, applied by a
resolution step. 65 call sites and the docs use the nested form; the three
kwargs golden cases are config cases with byte-identical results.

Every value is decided during resolution; no runtime code calls
with_override. The device offload policy (unified memory, MPS, layerwise
conflicts, lazy module load) runs as resolution steps in the main process
(fastvideo/api/device_policy.py), and the worker runs on the config it
receives. The LTX-2 refine defaults from model_index.json and the MiniMax-H3
schedule from fastvideo_inference.json are resolution steps
(fastvideo/api/checkpoint_defaults.py) that read local checkpoints or download
only those files; the H3 pipeline validates the schedule against the loaded
schedulers but no longer writes it. The preprocessing entry point passes the
downloaded local path as run state instead of a config override. Teacher and
critic models load with override_transformer_cls_name as a load argument.
The test isolation turns the device policy off, as it blocks downloads, so
that the goldens do not depend on the machine.

Against the 3f6893a0 goldens, every field still matches except boundary_ratio
and ltx2_vae_tiling (DESIGN section 9) and the fields that no longer exist.
The launcher and trainer YAML verifications report 0 unexplained differences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 07:14:31 +00:00

54 lines
2.4 KiB
Python

# SPDX-License-Identifier: Apache-2.0
from fastvideo import VideoGenerator
def main():
# Point this to your local diffusers model dir (or replace with a HF model ID).
model_path = "KyleShao/Cosmos-Predict2.5-2B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_path,
{
"engine": {
"num_gpus": 1,
"use_fsdp_inference": False, # set True if GPU is out of memory
"offload": {
"dit": False,
"vae": False,
"text_encoder": True,
"pin_cpu_memory": True,
},
},
},
)
# video2world example from official repo
video_path = "assets/videos/robot_pouring.mp4"
prompt = (
"A robotic arm, primarily white with black joints and cables, is shown in a clean, modern indoor setting with a white tabletop. "
"The arm, equipped with a gripper holding a small, light green pitcher, is positioned above a clear glass containing a reddish-brown liquid and a spoon. "
"The robotic arm is in the process of pouring a transparent liquid into the glass. "
"To the left of the pitcher, there is an opened jar with a similar reddish-brown substance visible through its transparent body. "
"In the background, a vase with white flowers and a brown couch are partially visible, adding to the contemporary ambiance. "
"The lighting is bright, casting soft shadows on the table. "
"The robotic arm's movements are smooth and controlled, demonstrating precision in its task. "
"As the video progresses, the robotic arm completes the pour, leaving the glass half-filled with the reddish-brown liquid. "
"The jar remains untouched throughout the sequence, and the spoon inside the glass remains stationary. "
"The other robotic arm on the right side also stays stationary throughout the video. "
"The final frame captures the robotic arm with the pitcher finishing the pour, with the glass now filled to a higher level, while the pitcher is slightly tilted but still held securely by the gripper."
)
generator.generate({
"prompt": prompt,
"inputs": {"video_path": str(video_path)},
"output": {"output_path": "outputs_video/cosmos2_5_v2w.mp4", "save_video": True},
"extensions": {"num_cond_frames": 1},
})
generator.shutdown()
if __name__ == "__main__":
main()