fahmisyaifudin/indobert-tweet-spam-classifier

🤗 Hugging Face 来源text-classificationmit111M 参数442 MBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo fahmisyaifudin/indobert-tweet-spam-classifier ./model-folder
需要做种者 →

IndoBERT Tweet Spam Classifier

Fine-tuned indolem/indobert-base-uncased for binary twitter (X) post spam / not-spam classification of Indonesian social media posts.

Model Description

  • Base model: indolem/indobert-base-uncased
  • Task: Text classification (binary)
  • Language: Indonesian (id)
  • Labels: spam (1), not_spam (0)

Usage

from transformers import pipeline

clf = pipeline("text-classification", model="fahmisyaifudin/indobert-tweet-spam-classifier")
clf("GIVEAWAY ALERT! Dapatkan undian berhadiah 100 juta, klik link bio")
# [{'label': 'spam', 'score': 0.98}]

Training Data

Fine-tuned on a labeled dataset of 4000+ indonesian twitter social media posts, cleaned to remove links and emojis prior to training. Labels: 0 = not_spam, 1 = spam.

Training Procedure

  • Epochs: 4
  • Batch size: 16
  • Learning rate: 2e-5
  • Max sequence length: 256
  • Optimizer: AdamW

Evaluation Results

Metric Score
Accuracy 0.975
F1 0.968
Precision 0.968
Recall 0.968