kandinskylab/KVAE-3D-1.0

🤗 Hugging Face 来源mit618M 参数2.5 GBsafetensors✓ 3 个校验和今天更新
需要做种者 →

GitHub | Habr article | KVAE-Image 1.0

KVAE 1.0: Video tokenizer

KVAE-Video 1.0 is the video tokenizer from KVAE 1.0, a family of image and video tokenizers designed for latent diffusion models. Its causal, fully convolutional architecture encodes videos into compact continuous latent representations and reconstructs them with high fidelity.

Model zoo

Model Modality Compression Latent channels
KVAE-Image 1.0 Image 8 x 8 16
KVAE-Video 1.0 Video 4 x 8 x 8 16

Inference

Run from the KVAE source repository root. The reference environment uses Python 3.11, PyTorch 2.8.0, and CUDA 12.8.

pip install -r requirements.txt
import torch

from data import VideoReader
from kvae import KVAEVideo

device = torch.device("cuda:0")
dtype = torch.bfloat16

model = (
    KVAEVideo.from_pretrained("kandinskylab/KVAE-3D-1.0")
    .eval()
    .to(device=device, dtype=dtype)
)
reader = VideoReader(stream_pattern="*.png", input_norm="m11")
video = reader.read_video("path/to/video_frames")["frames"].unsqueeze(0)
video = video.to(device=device, dtype=dtype)

with torch.no_grad():
    latent = model.encode(video, seg_len=16).latent_dist.mode()
    reconstruction = model.decode(latent, seg_len=16).clip(-1, 1)

Temporal segments are processed through internal block caches. Do not interleave independent videos on the same model instance; use one KVAEVideo instance per concurrent stream.

Evaluation

Reconstruction was evaluated on MCL-JCV downsampled to 540p because of the model's limitations at high resolutions. All compared models use 4 x 8 x 8 compression with 16 latent channels.

Model PSNR↑ SSIM↑ LPIPS↓
Wan 2.1 33.75 0.90 0.089
HunyuanVideo 1.0 33.91 0.91 0.103
KVAE-Video 1.0 35.63 0.92 0.088
Show reconstruction figures

Qualitative comparison

Columns from left to right: original video, KVAE-Video 1.0 reconstruction, and HunyuanVideo 1.0 reconstruction.

Citation

@misc{kvae_1_2025,
  author       = {Kirill Chernyshev and Andrey Shutkin and Ilia Vasiliev and Denis Parkhomenko and Ivan Kirillov and Dmitrii Mikhailov and Denis Dimitrov},
  title        = {KVAE 1.0: Image and Video Tokenizers for Image and Video Generation Models},
  howpublished = {\url{https://github.com/kandinskylab/kvae}},
  year         = {2025}
}

License

MIT