A manually curated benchmark was created to simulate actual YouTube comments.
The benchmark includes:
Internet slang
Emoji-heavy comments
Creator terminology
Sarcasm
Mixed sentiment
Short comments
Viral internet phrases
Results
Metric
Score
Accuracy
88.00%
Macro F1
87.68%
Weighted F1
87.73%
📉 Normalized Confusion Matrix
The model performs consistently across all sentiment classes and shows balanced classification behavior. Most classification errors occur between Neutral and Positive comments.
Negative comments are generally identified more reliably.
Normalized Confusion Matrix
🏋️ Training Details
Base Model
cardiffnlp/twitter-roberta-base-sentiment-latest
Fine-Tuning Dataset
1M+ YouTube Comments
Hardware
NVIDIA RTX 3050 Laptop GPU (6GB)
Training Features
Mixed Precision Training (AMP)
Layer-wise Learning Rate Decay (LLRD)
Gradient Accumulation
Cosine Learning Rate Scheduler
Warmup Scheduling
Gradient Checkpointing
Class Weighted Loss
Early Stopping
Resume Training Support
Dynamic GPU Configuration
📋 Internal Test Set Results
For reference, the held-out test split produced:
Metric
Score
Accuracy
77.43%
Macro F1
77.38%
Weighted F1
77.41%
The external benchmark is considered a better estimate of real-world deployment performance.