LiheYoung/SigLIP-HD

🤗 Hugging Face sourceimage-feature-extractionmit429M params1.7 GBsafetensors✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo LiheYoung/SigLIP-HD ./model-folder
Needs a seeder →

SigLIP-HD

SigLIP-HD is a vision encoder fine-tuned from SigLIP 2-So400m/16-512px with fine-to-coarse supervision.

SigLIP-HD exhibits better performance than SigLIP 2 in MLLMs, especially for OCR scenarios.

This repository contains only the vision encoder (no text tower). It is a drop-in replacement for the SigLIP 2 vision tower: identical architecture and I/O. To use it in an MLLM, keep your existing SigLIP 2 pipeline and only change the vision-tower path to this checkpoint.

Usage

import torch
from PIL import Image
from transformers import SiglipVisionModel, AutoImageProcessor

model = SiglipVisionModel.from_pretrained("LiheYoung/SigLIP-HD").eval()
processor = AutoImageProcessor.from_pretrained("LiheYoung/SigLIP-HD")

image = Image.open("example.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")

with torch.no_grad():
    features = model(**inputs, output_hidden_states=True).hidden_states[-1]  # (1, 1024, 1152)

Citation

@inproceedings{sigliphd,
  title={SigLIP-HD by Fine-to-Coarse Supervision},
  author={Yang, Lihe and Zhao, Zhen and Zhao, Hengshuang},
  booktitle={ICLR},
  year={2026}
}

Acknowledgement

This work is built upon SigLIP 2. We sincerely thank the authors for open-sourcing their excellent vision encoder.