LiheYoung/SigLIP-HD

🤗 Hugging Face 来源image-feature-extractionmit429M 参数1.7 GBsafetensors✓ 1 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo LiheYoung/SigLIP-HD ./model-folder
需要做种者 →

SigLIP-HD

SigLIP-HD is a vision encoder fine-tuned from SigLIP 2-So400m/16-512px with fine-to-coarse supervision.

SigLIP-HD exhibits better performance than SigLIP 2 in MLLMs, especially for OCR scenarios.

This repository contains only the vision encoder (no text tower). It is a drop-in replacement for the SigLIP 2 vision tower: identical architecture and I/O. To use it in an MLLM, keep your existing SigLIP 2 pipeline and only change the vision-tower path to this checkpoint.

Usage

import torch
from PIL import Image
from transformers import SiglipVisionModel, AutoImageProcessor

model = SiglipVisionModel.from_pretrained("LiheYoung/SigLIP-HD").eval()
processor = AutoImageProcessor.from_pretrained("LiheYoung/SigLIP-HD")

image = Image.open("example.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")

with torch.no_grad():
    features = model(**inputs, output_hidden_states=True).hidden_states[-1]  # (1, 1024, 1152)

Citation

@inproceedings{sigliphd,
  title={SigLIP-HD by Fine-to-Coarse Supervision},
  author={Yang, Lihe and Zhao, Zhen and Zhao, Hengshuang},
  booktitle={ICLR},
  year={2026}
}

Acknowledgement

This work is built upon SigLIP 2. We sincerely thank the authors for open-sourcing their excellent vision encoder.