shi-labs/pretrain_dsg_OLA-VLM-CLIP-ViT-Phi3-4k-mini

🤗 Hugging Face sourceimage-text-to-textapache-2.04.3B params8.9 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo shi-labs/pretrain_dsg_OLA-VLM-CLIP-ViT-Phi3-4k-mini ./model-folder
Needs a seeder →

pretrain_dsg_OLA-VLM-CLIP-ViT-Phi3-4k-mini Model Card

Note: This is the pretrained model used for OLA-VLM-CLIP-ViT-Phi3-4k-mini.

OLA-VLM distills target visual information into the intermediate representations of the LLM from a set of target encoders. It adopts a predictive embedding optimization approach at selected LLM layers during training to minimize the embedding losses along with the next token prediction (NTP) objective, resulting in a vision-centric approach to training the Multimodal Large Language Model.

Citation

If you found our work useful in your research, please consider starring ⭐ us on GitHub and citing 📚 us in your research!

@article{jain2024ola_vlm,
    title={{OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation}},
    author={Jitesh Jain and Zhengyuan Yang and Humphrey Shi and Jianfeng Gao and Jianwei Yang},
    journal={arXiv},
    year={2024}
}