Views
No views yet
xlm-roberta-base for Part-of-Speech (POS) tagging on Ancient Greek. It was trained on the Universal Dependencies Ancient Greek Perseus treebank.xlm-roberta-basegrc)1from transformers import pipeline
2
3# Load the POS tagging pipeline
4nlp = pipeline(
5 "token-classification",
6 model="your-username/your-model-name",
7 aggregation_strategy="first"
8)
9
10# Example: Opening line of the Iliad
11text = "Μῆνιν ἄειδε θεά Πηληϊάδεω Ἀχιλῆος"
12predictions = nlp(text)
13
14for word in predictions:
15 print(f"{word['word']:<15} -> {word['entity_group']}")1Μῆνιν -> NOUN
2ἄειδε -> VERB
3θεά -> NOUN
4Πηληϊάδεω -> NOUN
5Ἀχιλῆος -> NOUNUD_Ancient_Greek-Perseus dataset.| Tag | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| ADJ | 0.7742 | 0.7617 | 0.7679 | 1796 |
| ADP | 0.9812 | 0.9881 | 0.9846 | 1423 |
| ADV | 0.9341 | 0.7602 | 0.8382 | 2481 |
| AUX | 0.9247 | 0.9214 | 0.9231 | 280 |
| CCONJ | 0.7611 | 0.8728 | 0.8131 | 668 |
| DET | 0.9979 | 0.9882 | 0.9930 | 2367 |
| INTJ | 0.9688 | 0.8857 | 0.9254 | 35 |
| NOUN | 0.9247 | 0.9361 | 0.9304 | 4489 |
| NUM | 0.0789 | 0.7500 | 0.1429 | 4 |
| PART | 0.0000 | 0.0000 | 0.0000 | 0 |
| PRON | 0.8624 | 0.8517 | 0.8570 | 1119 |
| PUNCT | 0.9991 | 0.9991 | 0.9991 | 2306 |
| SCONJ | 0.9161 | 0.9016 | 0.9088 | 315 |
| VERB | 0.9725 | 0.9659 | 0.9692 | 3406 |
| X | 0.0000 | 0.0000 | 0.0000 | 1 |
PART and X had 0-1 support in the evaluation split, resulting in scores of 0.