frankleeeee/CausalForcing-Long-Wan2.1-T2V-1.3B-Diffusers

🤗 On Hugging Facetext-to-videoapache-2.01.4B params2.8 GBsafetensorsChecksums witnessedupdated today
Magnet

CausalForcing-Long-Wan2.1-T2V-1.3B-Diffusers

Diffusers-layout conversion of the minute-level long-video Causal Forcing

checkpoint (zhuhz22/Causal-Forcing, chunkwise/longvideo.pt,

generator_ema weights) from

thu-ml/Causal-Forcing.

This checkpoint is a Rolling Forcing model (rolling-window joint denoising,

3-frame attention-sink block, 24-latent-frame KV buffer with a 21-frame

attention window) retrained from Causal Forcing's ODE initialization — upstream

adopts the TencentARC/RollingForcing

framework and only changes the ODE init. It therefore runs on the Rolling

Forcing pipeline (WanRollingForcingPipeline), not the 5-second Causal Forcing

pipeline. Non-transformer components are copied from

Wan-AI/Wan2.1-T2V-1.3B-Diffusers.

Converted with

sglang.multimodal_gen.tools.convert_forcing_to_diffusers --preset rolling-forcing

for use with the SGLang diffusion runtime:

sglang generate --model-path frankleeeee/CausalForcing-Long-Wan2.1-T2V-1.3B-Diffusers \
  --prompt "A stylish woman walks down a Tokyo street..." \
  --width 832 --height 480 --num-frames 321 --save-output