IndoBERT Tweet Spam Classifier
Fine-tuned indolem/indobert-base-uncased
for binary twitter (X) post spam / not-spam classification of Indonesian social media posts.
Model Description
- Base model: indolem/indobert-base-uncased
- Task: Text classification (binary)
- Language: Indonesian (id)
- Labels:
spam(1),not_spam(0)
Usage
from transformers import pipeline
clf = pipeline("text-classification", model="fahmisyaifudin/indobert-tweet-spam-classifier")
clf("GIVEAWAY ALERT! Dapatkan undian berhadiah 100 juta, klik link bio")
# [{'label': 'spam', 'score': 0.98}]
Training Data
Fine-tuned on a labeled dataset of 4000+ indonesian twitter social media posts, cleaned to
remove links and emojis prior to training. Labels: 0 = not_spam, 1 = spam.
Training Procedure
- Epochs: 4
- Batch size: 16
- Learning rate: 2e-5
- Max sequence length: 256
- Optimizer: AdamW
Evaluation Results
| Metric | Score |
|---|---|
| Accuracy | 0.975 |
| F1 | 0.968 |
| Precision | 0.968 |
| Recall | 0.968 |