Views
No views yet

| Variant | Description |
|---|---|
pre-filtered-corpus | Cleaned dataset with noise and ambiguous tweets removed |
raw-corpus | Full unfiltered dataset as collected from Twitter/X |
| Pipeline | Description |
|---|---|
standard | Baseline tokenization and normalization |
irony | Adds [IRONIA] token before sarcastic/ironic tweets |
obfuscated | Named entities (substances, slang) masked via NER |
| Model | Type | Representation |
|---|---|---|
naive_bayes | Multinomial Naive Bayes | TF-IDF (5,000 features, bigrams) |
logistic_regression | Logistic Regression | TF-IDF (5,000 features, bigrams) |
svm | Linear SVC | TF-IDF (5,000 features, bigrams) |
random_forest | Random Forest | TF-IDF (5,000 features, bigrams) |
ffn | Feed-Forward Network (100→64→32→1) | Word2Vec 100-dim |
cnn | TextCNN (filters 3/4/5) | Word2Vec 100-dim |
rnn | BiLSTM (hidden=64, bidirectional) | Word2Vec 100-dim |
bert_base | Fine-tuned BETO | dccuchile/bert-base-spanish-wwm-cased |
├── pre-filtered-corpus/
│ ├── naive_bayes/{standard,irony,obfuscated}/
│ │ ├── model.joblib
│ │ └── vectorizer.joblib
│ ├── logistic_regression/{standard,irony,obfuscated}/ (same structure)
│ ├── svm/{standard,irony,obfuscated}/ (same structure)
│ ├── random_forest/{standard,irony,obfuscated}/ (same structure)
│ ├── ffn/{standard,irony,obfuscated}/
│ │ └── model.pt
│ ├── cnn/{standard,irony,obfuscated}/
│ │ └── model.pt
│ ├── rnn/{standard,irony,obfuscated}/
│ │ └── model.pt
│ ├── word2vec/{standard,irony,obfuscated}/
│ │ └── word2vec.model
│ └── bert_base/{standard,irony,obfuscated}/
│ ├── model/ (HuggingFace saved model)
│ └── tokenizer/ (HuggingFace tokenizer)
└── raw-corpus/ (same structure)Note: Forraw-corpus, Naive Bayes, SVM, and Random Forest models were not persisted — re-train inline with identical hyperparameters (seed=42) using the provided training splits.
| Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| BERT (Base) | 86.00% | 86.01% | 86.00% | 86.00% |
| Logistic Regression | 80.22% | 80.29% | 80.22% | 80.21% |
| SVM | 79.78% | 79.81% | 79.78% | 79.77% |
| Naive Bayes | 77.33% | 77.39% | 77.33% | 77.32% |
| CNN | 77.33% | 77.47% | 77.33% | 77.30% |
| Random Forest | 76.67% | 76.67% | 76.67% | 76.67% |
| RNN (BiLSTM) | 76.67% | 76.71% | 76.67% | 76.66% |
| FFN | 72.89% | 72.98% | 72.89% | 72.86% |
| Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| BERT (Base) | 85.33% | 85.34% | 85.33% | 85.33% |
| Logistic Regression | 80.44% | 80.50% | 80.44% | 80.43% |
| CNN | 79.33% | 79.40% | 79.33% | 79.32% |
| SVM | 79.11% | 79.12% | 79.11% | 79.11% |
| Random Forest | 79.11% | 79.12% | 79.11% | 79.11% |
| RNN (BiLSTM) | 78.44% | 78.51% | 78.44% | 78.43% |
| Naive Bayes | 77.78% | 77.83% | 77.78% | 77.77% |
| FFN | 76.67% | 76.68% | 76.67% | 76.66% |
| Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| BERT (Base) | 85.78% | 85.80% | 85.78% | 85.78% |
| SVM | 79.78% | 79.81% | 79.78% | 79.77% |
| Logistic Regression | 78.89% | 79.02% | 78.89% | 78.87% |
| CNN | 78.22% | 78.28% | 78.22% | 78.21% |
| Naive Bayes | 76.89% | 76.92% | 76.89% | 76.88% |
| Random Forest | 76.44% | 76.52% | 76.44% | 76.43% |
| FFN | 75.78% | 75.78% | 75.78% | 75.78% |
| RNN (BiLSTM) | 75.11% | 75.14% | 75.11% | 75.10% |
| Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| BERT (Base) | 85.56% | 85.71% | 85.56% | 85.54% |
| Logistic Regression | 80.67% | 80.74% | 80.67% | 80.66% |
| SVM | 80.00% | 80.01% | 80.00% | 80.00% |
| Naive Bayes | 79.78% | 79.83% | 79.78% | 79.77% |
| CNN | 79.78% | 79.99% | 79.78% | 79.74% |
| RNN (BiLSTM) | 79.11% | 79.51% | 79.11% | 79.04% |
| FFN | 78.00% | 78.16% | 78.00% | 77.97% |
| Random Forest | 74.44% | 74.50% | 74.44% | 74.43% |
| Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| BERT (Base) | 87.56% | 87.60% | 87.56% | 87.55% |
| Logistic Regression | 80.67% | 80.77% | 80.67% | 80.65% |
| SVM | 80.00% | 80.01% | 80.00% | 80.00% |
| Naive Bayes | 79.78% | 79.83% | 79.78% | 79.77% |
| CNN | 79.33% | 79.50% | 79.33% | 79.30% |
| RNN (BiLSTM) | 77.56% | 77.96% | 77.56% | 77.47% |
| Random Forest | 76.00% | 76.02% | 76.00% | 76.00% |
| FFN | 74.67% | 76.30% | 74.67% | 74.27% |
| Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| BERT (Base) | 87.78% | 87.80% | 87.78% | 87.78% |
| Logistic Regression | 81.78% | 81.84% | 81.78% | 81.77% |
| SVM | 80.22% | 80.24% | 80.22% | 80.22% |
| Naive Bayes | 80.22% | 80.25% | 80.22% | 80.22% |
| RNN (BiLSTM) | 80.22% | 80.29% | 80.22% | 80.21% |
| CNN | 80.00% | 80.01% | 80.00% | 80.00% |
| FFN | 77.78% | 77.80% | 77.78% | 77.77% |
| Random Forest | 76.22% | 76.37% | 76.22% | 76.19% |
| Corpus | Best Configuration | F1 |
|---|---|---|
| Pre-filtered | BERT (Base) — Standard | 86.00% |
| Raw | BERT (Base) — Obfuscated | 87.78% (best overall) |
Note: Naive Bayes, SVM, and Random Forest are not the top model in any raw-corpus or pre-filtered-corpus variant as of this iteration — BERT (Base) leads across all six corpus/variant combinations. See the evaluation log for the full iteration history and how this ranking changed as split/evaluation bugs were fixed.
1import joblib
2from huggingface_hub import hf_hub_download
3
4# Download model and vectorizer
5model_path = hf_hub_download(
6 repo_id="lhbelfanti/drug-use-models",
7 filename="pre-filtered-corpus/svm/standard/model.joblib"
8)
9vectorizer_path = hf_hub_download(
10 repo_id="lhbelfanti/drug-use-models",
11 filename="pre-filtered-corpus/svm/standard/vectorizer.joblib"
12)
13
14model = joblib.load(model_path)
15vectorizer = joblib.load(vectorizer_path)
16
17texts = ["me tomé una línea antes de salir"]
18X = vectorizer.transform(texts)
19print(model.predict(X)) # [1] → POSITIVE1from transformers import AutoTokenizer, AutoModelForSequenceClassification
2from huggingface_hub import snapshot_download
3import torch
4
5# Download model and tokenizer
6model_dir = snapshot_download(
7 repo_id="lhbelfanti/drug-use-models",
8 allow_patterns="pre-filtered-corpus/bert_base/standard/**"
9)
10
11tokenizer = AutoTokenizer.from_pretrained(f"{model_dir}/pre-filtered-corpus/bert_base/standard/tokenizer")
12model = AutoModelForSequenceClassification.from_pretrained(f"{model_dir}/pre-filtered-corpus/bert_base/standard/model")
13model.eval()
14
15text = "me tomé una línea antes de salir"
16inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=128)
17with torch.no_grad():
18 logits = model(**inputs).logits
19pred = torch.argmax(logits, dim=1).item()
20print("POSITIVE" if pred == 1 else "NEGATIVE")max_features=5000, ngram_range=(1,2)dccuchile/bert-base-spanish-wwm-cased, fine-tuned 3 epochs, lr=2e-5, batch=16. Checkpoint selected on the validation split; test is evaluated once, after training.