CausalForcing-Long-Wan2.1-T2V-1.3B-Diffusers
Diffusers-layout conversion of the minute-level long-video Causal Forcing
checkpoint (zhuhz22/Causal-Forcing, chunkwise/longvideo.pt,
generator_ema weights) from
This checkpoint is a Rolling Forcing model (rolling-window joint denoising,
3-frame attention-sink block, 24-latent-frame KV buffer with a 21-frame
attention window) retrained from Causal Forcing's ODE initialization — upstream
adopts the TencentARC/RollingForcing
framework and only changes the ODE init. It therefore runs on the Rolling
Forcing pipeline (WanRollingForcingPipeline), not the 5-second Causal Forcing
pipeline. Non-transformer components are copied from
Wan-AI/Wan2.1-T2V-1.3B-Diffusers.
Converted with
sglang.multimodal_gen.tools.convert_forcing_to_diffusers --preset rolling-forcing
for use with the SGLang diffusion runtime:
sglang generate --model-path frankleeeee/CausalForcing-Long-Wan2.1-T2V-1.3B-Diffusers \
--prompt "A stylish woman walks down a Tokyo street..." \
--width 832 --height 480 --num-frames 321 --save-output