BERT model for spam detection in Indonesian with 95% accuracy. This v3 model has been fine-tuned from v2 model with email dataset for optimal performance on Indonesian content.
Quick Start
python
1from transformers import pipeline
23# The easiest way to use the model4classifier = pipeline("text-classification",5 model="newreyy/spam-detection-v4",6 tokenizer="newreyy/spam-detection-v4")78# Test with text9texts =[10"lacak hp hilang by no hp / imei lacak penipu/scammer/tabrak lari/terror/revengeporn sadap / hack / pulihkan akun",11"Senin, 21 Juli 2025, Samapta Polsek Ngaglik melaksanakan patroli stasioner balong jalan palagan donoharjo",12"Mari berkontribusi terhadap gerakan rakyat dengan membeli baju ini seharga Rp 160.000. Hubungi kami melalui WA 08977472296"13]1415results = classifier(texts)16for text, result inzip(texts, results):17print(f"Text: {text}")18print(f"Result: {result['label']} (confidence: {result['score']:.4f})")19print("---")
Model Details
Base Model: nahiar/spam-detection-bert-v2
Task: Binary Text Classification (Spam vs Ham)
Language: Indonesian (Bahasa Indonesia)
Model Size: ~110M parameters
Max Sequence Length: 512 tokens
Training Epochs: 3
Batch Size: 16
Learning Rate: 2e-5
Performance
Metric
HAM
SPAM
Overall
Precision
98%
77%
95%
Recall
96%
85%
95%
F1-Score
97%
81%
95%
Overall Accuracy
-
-
95%
Confusion Matrix
True HAM correctly predicted: 953/988 (96%)
True SPAM correctly predicted: 115/135 (85%)
False Positives (HAM predicted as SPAM): 35
False Negatives (SPAM predicted as HAM): 20
Key Features
Fine-tuned from v2 model with email dataset
Good accuracy (95%) on spam detection with Indonesian context
Better handling for spam email content
Enhanced performance on Indonesian email text
Optimized for Indonesian email and social media spam detection
Label Mapping
0: "HAM" (not spam)
1: "SPAM" (spam)
Training Process
This model was retrained using:
Optimizer: AdamW
Learning Rate: 2e-5
Epochs: 3
Batch Size: 16
Max Length: 128 tokens
Train/Validation Split: 80/20
Usage Example
python
1import torch
2from transformers import AutoTokenizer, AutoModelForSequenceClassification
34# Load model and tokenizer5tokenizer = AutoTokenizer.from_pretrained("nahiar/spam-detection-bert-v3")6model = AutoModelForSequenceClassification.from_pretrained("nahiar/spam-detection-bert-v3")78defpredict_spam(text):9 inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)10 outputs = model(**inputs)11 probs = torch.softmax(outputs.logits, dim=1)12 predicted_label = torch.argmax(probs, dim=1).item()13 confidence = probs[0][predicted_label].item()14 label_map ={0:"HAM",1:"SPAM"}15return label_map[predicted_label], confidence
1617# Test18text ="Dapatkan uang dengan mudah! Klik link ini sekarang!"19result, confidence = predict_spam(text)20print(f"Prediksi: {result} (Confidence: {confidence:.4f})")