Ruggero1912/Patch-ioner_talk2dino_decap_COCO_Captions

🤗 Hugging Face 来源image-to-textapache-2.02.1 GBother✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Ruggero1912/Patch-ioner_talk2dino_decap_COCO_Captions ./model-folder
需要做种者 →

Patch-ioner_talk2dino_decap_COCO_Captions - Patch-ioner Configuration

This repository contains a pre-trained DECAP model from the Patch-ioner framework for dense image captioning and controllable visual description.

📝 Paper Information

Title: "One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework"
Authors: Lorenzo Bianchi, Giacomo Pacini, Fabio Carrara, Nicola Messina, Giuseppe Amato, Fabrizio Falchi
ArXiv: https://arxiv.org/abs/2510.02898 Project Page: https://paciosoft.com/Patch-ioner/ Code: https://github.com/Ruggero1912/Patch-ioner

🎯 Model Overview

  • Model Type: DECAP
  • Configuration: mlp.karpathy.yaml
  • Vision Backbone: dinov2_vitb14_reg
  • Language Model: GPT-2
  • Input Resolution: 518x518
  • Prefix Size: 768

DeCap Configuration

  • Memory Bank Size: 591,753 entries
  • Projection Type: /raid/datasets/im2txtmemories/coco_train_karpathy.json
  • Linear Talk2DINO: False

📊 Performance

| Task | METEOR | CIDEr | SPICE | |------|--------|-------|-------|
| Image Captioning | 0.239 | 0.885 | 0.182 |
| Narratives | 10.700 | 27.900 | 12.600 |

📈 Detailed Results

Image Captioning Results

  • METEOR: 0.2393
  • CIDEr: 0.8846
  • SPICE: 0.1821
  • BLEU_4: 0.2364
  • ROUGE_L: 0.4854
  • CLIP-S: 0.7602

Narratives Results

  • METEOR: 10.7000
  • CIDEr: 27.9000
  • SPICE: 12.6000
  • BLEU_4: 2.5000
  • ROUGE_L: 23.2000
  • CLIP-S: 68.0000

🚀 Quick Start

from patchioner import Patchioner
from transformers import AutoModel

MODEL_ID = "Ruggero1912/Patch-ioner_talk2dino_decap_COCO_Captions"

# Option 1: Using Patchioner's native API
model_native = Patchioner.from_config(MODEL_ID)
print("Model loaded via native Patchioner API.")

# Option 2: Using Hugging Face Transformers AutoModel
model_hf = AutoModel.from_pretrained(MODEL_ID, trust_remote_code=True)
print("Model loaded via Hugging Face Transformers AutoModel.")

# For full usage examples, including inference, please refer to the model's
# GitHub repository: https://github.com/Ruggero1912/Patch-ioner

📁 Repository Contents

  • config.yaml: Model configuration file

  • coco_karpathy-009.pt: Pre-trained model weights

  • coco_captions_text_embeddings-B16-ViT-B.16-591753.h5: Memory bank for text projection- README.md: This file

🔧 Installation

pip install git+https://github.com/Ruggero1912/Patch-ioner

💡 Usage Examples

Refer to the Patch-ioner GitHub repository for updated usage examples.

🎛️ Model Configuration

  • Prefix Size: 768
  • Memory Bank Size: 591,753 entries
  • Normalization: True

📈 Training Details

  • Training Dataset: COCO Captions
  • Training Epochs: TBD
  • Batch Size: TBD
  • Learning Rate: TBD
  • Optimizer: AdamW

📚 Citation

If you use this model in your research, please cite our paper:

@misc{bianchi2025patchcaptionallunified,
      title={One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework},
      author={Lorenzo Bianchi and Giacomo Pacini and Fabio Carrara and Nicola Messina and Giuseppe Amato and Fabrizio Falchi},
      year={2025},
      eprint={2510.02898},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2510.02898},
}

🤝 Contributing

We welcome contributions to improve the Patch-ioner framework. Please see the main repository for contribution guidelines.

📄 License

See the main repository for detailed license information.

🐛 Issues and Support

For issues related to this model or the Patch-ioner framework, please:

  1. Check the main repository for existing issues
  2. Open a new issue with detailed information about your problem
  3. Contact the authors.

🔗 Related Models

Explore other Patch-ioner model configurations:

More models available in Ruggero1912's models


This model is part of the Patch-ioner framework for dense image captioning and controllable visual description.