A FinBERT model fine-tuned on r/wallstreetbets text for three-class sentiment
classification (negative / neutral / positive).
Off-the-shelf financial sentiment models are trained on formal financial English —
earnings calls, analyst reports, newswire. Retail social-media text is a different
register: irony, slang, emoji, and self-deprecation. On the held-out set below,
vanilla FinBERT scores below the majority-class floor, i.e. worse than ignoring
the text entirely. This model is a domain-adaptation check on that gap.
This is the frozen checkpoint behind the results reported in Beyond the Volume of Attention:
Domain-Adapted Sentiment and the Content of Retail Investor Discussion (ICAIF '26) — not a
retrained copy. The replication package for that paper is distributed separately; this
repository is the model only.
Labels
id
label
0
negative
1
neutral
2
positive
Usage
You do not need to download anything by hand. transformers fetches the weights on first use
and caches them under ~/.cache/huggingface/, so the first call takes a moment (~440 MB) and
every call after that is instant.
pip install transformers torch
python
1from transformers import AutoTokenizer, AutoModelForSequenceClassification
2import torch
34model_id ="AnonymousResearchICAIF/wsb-finbert"5tok = AutoTokenizer.from_pretrained(model_id)6model = AutoModelForSequenceClassification.from_pretrained(model_id)78text ="NVDA printing again, loaded calls for next week"9inputs = tok(text, return_tensors="pt", truncation=True, max_length=128)10with torch.no_grad():11 probs = model(**inputs).logits.softmax(-1)[0]1213print({model.config.id2label[i]:round(p.item(),3)for i, p inenumerate(probs)})
A continuous sentiment score in [-1, 1] is formed as P(positive) - P(negative). This is the
sentiment_wsb variable used throughout the paper — the sign gives the direction and the
magnitude gives the confidence. Use max_length=128: that is what the model was trained
with, and longer inputs are truncated to it.
For many texts at once, batch them rather than looping:
2,503 r/wallstreetbets posts and comments mentioning a fixed universe of AI-related
tickers, labelled for sentiment by an LLM teacher (Claude) and split 70/15/15
stratified by label with seed 42 — 1,752 train / 375 validation / 376 test.
The sample is stratified across tickers and balanced 50/50 between the first and
second halves of the collection period.
Training procedure
Base model
ProsusAI/finbert
Max sequence length
128
Learning rate
2e-5
Epochs
4
Train batch size
8
Weight decay
0.01
Warmup ratio
0.1
Seed
42
Evaluation
On the 376 held-out texts:
Model
Accuracy
Macro F1
WSB-FinBERT (this model)
0.524
0.495
Vanilla ProsusAI/finbert
0.410
0.400
Majority-class baseline ("always positive")
0.434
—
Cohen's κ = 0.252.
Per-class (this model):
label
precision
recall
f1
support
negative
0.448
0.312
0.368
96
neutral
0.513
0.513
0.513
117
positive
0.557
0.656
0.603
163
Limitations
Absolute accuracy is modest. 52.4% on a three-class problem is only ~9 points
above the majority-class floor. The gain over vanilla FinBERT (+11.4 points) is the
meaningful result; the model is not a strong standalone classifier.
Labels are LLM-generated, not human gold standard. They inherit the teacher's
biases. They do correlate with past five-day returns, as a text-only annotator
should, and show no positive correlation with forward returns (no evidence of
look-ahead leakage).
Narrow domain. Trained on r/wallstreetbets text about a small universe of
AI-related tickers over a specific period. Generalisation to other subreddits,
other sectors, other periods, or to formal financial text is untested and unlikely.
Positive skew. The forum's labels run ~1.7 to 1 positive and the model inherits
this (51.1% of predictions positive vs 43.4% in truth). Use within-entity demeaning
if the level matters for your application.
Not investment advice; not suitable for trading decisions on its own.