Views
No views yet
mdeberta-v3-base-subjectivity-sentiment-italian, is part of the AI Wizards' submission to the CLEF 2025 CheckThat! Lab Task 1: Subjectivity Detection in News Articles. Its primary goal is to classify sentences as subjective or objective. A key innovation in its development involved enhancing transformer-based classifiers, specifically mDeBERTaV3-base, by integrating sentiment scores derived from an auxiliary model with sentence representations. This sentiment-augmented architecture, combined with decision threshold calibration to address class imbalance, consistently boosted performance, especially the subjective F1 score.transformers library for text classification:1import torch
2import torch.nn as nn
3from transformers import DebertaV2Model, DebertaV2Config, AutoTokenizer, PreTrainedModel, pipeline, AutoModelForSequenceClassification
4from transformers.models.deberta.modeling_deberta import ContextPooler
5
6sent_pipe = pipeline(
7 "sentiment-analysis",
8 model="cardiffnlp/twitter-xlm-roberta-base-sentiment",
9 tokenizer="cardiffnlp/twitter-xlm-roberta-base-sentiment",
10 top_k=None, # return all 3 sentiment scores
11)
12
13class CustomModel(PreTrainedModel):
14 config_class = DebertaV2Config
15 def __init__(self, config, sentiment_dim=3, num_labels=2, *args, **kwargs):
16 super().__init__(config, *args, **kwargs)
17 self.deberta = DebertaV2Model(config)
18 self.pooler = ContextPooler(config)
19 output_dim = self.pooler.output_dim
20 self.dropout = nn.Dropout(0.1)
21 self.classifier = nn.Linear(output_dim + sentiment_dim, num_labels)
22
23 def forward(self, input_ids, positive, neutral, negative, token_type_ids=None, attention_mask=None, labels=None):
24 outputs = self.deberta(input_ids=input_ids, attention_mask=attention_mask)
25 encoder_layer = outputs[0]
26 pooled_output = self.pooler(encoder_layer)
27 sentiment_features = torch.stack((positive, neutral, negative), dim=1).to(pooled_output.dtype)
28 combined_features = torch.cat((pooled_output, sentiment_features), dim=1)
29 logits = self.classifier(self.dropout(combined_features))
30 return {'logits': logits}
31
32model_name = "MatteoFasulo/mdeberta-v3-base-subjectivity-sentiment-italian"
33tokenizer = AutoTokenizer.from_pretrained("microsoft/mdeberta-v3-base")
34config = DebertaV2Config.from_pretrained(
35 model_name,
36 num_labels=2,
37 id2label={0: 'OBJ', 1: 'SUBJ'},
38 label2id={'OBJ': 0, 'SUBJ': 1},
39 output_attentions=False,
40 output_hidden_states=False
41)
42model = CustomModel(config=config, sentiment_dim=3, num_labels=2).from_pretrained(model_name)
43
44def classify_subjectivity(text: str):
45 # get full sentiment distribution
46 dist = sent_pipe(text)[0]
47 pos = next(d["score"] for d in dist if d["label"] == "positive")
48 neu = next(d["score"] for d in dist if d["label"] == "neutral")
49 neg = next(d["score"] for d in dist if d["label"] == "negative")
50
51 # tokenize the text
52 inputs = tokenizer(text, padding=True, truncation=True, max_length=256, return_tensors='pt')
53
54 # feeding in the three sentiment scores
55 with torch.no_grad():
56 outputs = model(
57 input_ids=inputs["input_ids"],
58 attention_mask=inputs["attention_mask"],
59 positive=torch.tensor(pos).unsqueeze(0).float(),
60 neutral=torch.tensor(neu).unsqueeze(0).float(),
61 negative=torch.tensor(neg).unsqueeze(0).float()
62 )
63
64 # compute probabilities and pick the top label
65 probs = torch.softmax(outputs.get('logits')[0], dim=-1)
66 label = model.config.id2label[int(probs.argmax())]
67 score = probs.max().item()
68
69 return {"label": label, "score": score}
70
71examples = [
72 "Per quanto riguarda le motivazioni, è importante chiedersi se l’intervento è realmente mirato a risolvere un complesso, a modificare un aspetto del corpo con cui non si riesce a convivere serenamente, oppure se è frutto di una moda passeggera o dell’influenza tossica del web che spesso induce a volere cose di cui non si ha assolutamente bisogno.",
73 "Un roulette di tensioni in cui alla fine spunta un match point per Sinner.",
74]
75for text in examples:
76 result = classify_subjectivity(text)
77 print(f"Text: {text}")
78 print(f"→ Subjectivity: {result['label']} (score={result['score']:.2f})\n")| Training Loss | Epoch | Step | Validation Loss | Macro F1 | Macro P | Macro R | Subj F1 | Subj P | Subj R | Accuracy |
|---|---|---|---|---|---|---|---|---|---|---|
| No log | 1.0 | 101 | 0.6392 | 0.7244 | 0.7284 | 0.7208 | 0.5913 | 0.6071 | 0.5763 | 0.7886 |
| No log | 2.0 | 202 | 0.5375 | 0.6731 | 0.7018 | 0.7579 | 0.6064 | 0.4548 | 0.9096 | 0.6867 |
| No log | 3.0 | 303 | 0.5731 | 0.7453 | 0.7373 | 0.7563 | 0.6349 | 0.5970 | 0.6780 | 0.7931 |
| No log | 4.0 | 404 | 0.5788 | 0.7522 | 0.7405 | 0.7752 | 0.6534 | 0.5848 | 0.7401 | 0.7916 |
| 0.4395 | 5.0 | 505 | 0.6922 | 0.7491 | 0.7400 | 0.7628 | 0.6423 | 0.5971 | 0.6949 | 0.7946 |
| 0.4395 | 6.0 | 606 | 0.6602 | 0.7437 | 0.7322 | 0.7690 | 0.6437 | 0.5696 | 0.7401 | 0.7826 |
1@misc{fasulo2025aiwizardscheckthat2025,
2 title={AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles},
3 author={Matteo Fasulo and Luca Babboni and Luca Tedeschini},
4 year={2025},
5 eprint={2507.11764},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL},
8 url={https://arxiv.org/abs/2507.11764},
9}