ANTILLES : An Open French Linguistically Enriched Part-of-Speech Corpus
Dataset Summary
ANTILLES is a part-of-speech tagging corpora based on UD_French-GSD which was originally created in 2015 and is based on the universal dependency treebank v2.0.
Originally, the corpora consists of 400,399 words (16,341 sentences) and had 17 different classes. Now, after applying our tags augmentation script transform.py, we obtain 60 different classes which add semantic information… See the full description on the dataset page: https://huggingface.co/datasets/qanastek/ANTILLES.