shi-labs/vcoder_llava-v1.5-7b

🤗 Hugging Face sourcetext-generationapache-2.07B activated29 GBother✓ 4 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo shi-labs/vcoder_llava-v1.5-7b ./model-folder
Needs a seeder →

VCoder LLaVA-1.5-7b

VCoder LLaVA-1.5-7b was trained on COST training dataset in December 2023. It uses the pretrained LLaVA-1.5-7b model weights. It was introduced by Jain et al. in this repository.

VCoder is an adapter for improving existing Multimodal LLMs at object-level perception tasks with the use of perception modalities as control inputs while retaining performance on other tasks.

Citation

@article{jain2023vcoder,
    title={{VCoder: Versatile Vision Encoders for Multimodal Large Language Models}},
    author={Jitesh Jain and Jianwei Yang and Humphrey Shi},
    journal={arXiv},
    year={2023}
}