Ruggero1912/Patch-ioner_siglip2_base_patch16_512_COCO_Captions

🤗 Hugging Face 来源image-to-textapache-2.01.9 GBother✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Ruggero1912/Patch-ioner_siglip2_base_patch16_512_COCO_Captions ./model-folder
需要做种者 →

Patch-ioner_siglip2_base_patch16_512_COCO_Captions - Patch-ioner Configuration

This repository contains a pre-trained DECAP model from the Patch-ioner framework for dense image captioning and controllable visual description.

📝 Paper Information

Title: "One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework"
Authors: Lorenzo Bianchi, Giacomo Pacini, Fabio Carrara, Nicola Messina, Giuseppe Amato, Fabrizio Falchi
ArXiv: https://arxiv.org/abs/2510.02898 Project Page: https://paciosoft.com/Patch-ioner/

🎯 Model Overview

  • Model Type: DECAP
  • Configuration: siglip2_base_patch16_512.k.yaml
  • Vision Backbone: siglip2_base_patch16_512
  • Language Model: GPT-2
  • Input Resolution: 512x512
  • Prefix Size: 768

DeCap Configuration

  • Memory Bank Size: 500,000 entries
  • Projection Type: coco
  • Linear Talk2DINO: False

📊 Performance

Task METEOR CIDEr SPICE
Narratives TBD TBD TBD
Image Captioning TBD TBD TBD
Dense Captioning TBD TBD TBD
Controllable Captioning TBD TBD TBD

📈 Detailed Results

Detailed evaluation results will be added here, including:

  • Per-task performance metrics
  • Comparison with baseline methods
  • Ablation study results
  • Qualitative examples

🚀 Quick Start

from patch_ioner import load_model, Patchioner

# Load the model
config_path = "config.yaml"
model = load_model(config_path)

# Run inference
image_path = "your_image.jpg"
results = model.forward(image_path)
print(results)

📁 Repository Contents

  • config.yaml: Model configuration file
  • coco_karpathy-009.pt: Pre-trained model weights
  • README.md: This file

🔧 Installation

pip install git+https://github.com/Ruggero1912/Patch-ioner

💡 Usage Examples

Refer to the Patch-ioner repository for updated usage examples.

🎛️ Model Configuration

  • Prefix Size: 768
  • Memory Bank Size: 500,000 entries
  • Normalization: False
  • Resize Dimension: 512
  • Crop Dimension: 512

📈 Training Details

  • Training Dataset: COCO Captions
  • Training Epochs: TBD
  • Batch Size: TBD
  • Learning Rate: TBD
  • Optimizer: AdamW

📚 Citation

If you use this model in your research, please cite our paper, refer to the Project Page for updated citation template.

🤝 Contributing

We welcome contributions to improve the Patch-ioner framework. Please see the main repository for contribution guidelines.

📄 License

See the main repository for detailed license information.

🐛 Issues and Support

For issues related to this model or the Patch-ioner framework, please:

  1. Check the main repository for existing issues
  2. Open a new issue with detailed information about your problem
  3. Contact the authors.

🔗 Related Models

Explore other Patch-ioner model configurations:

More models available in Ruggero1912's models


This model is part of the Patch-ioner framework for dense image captioning and controllable visual description.