This repository provides frequencies for French verbal paradigm cells in
the French Open Subtitles Corpus.
To obtain these frequencies we:
turained a tagger based on Flaubert-base-cased (see FlauBERT)
on the French_GSD part of the Universal Dependency treebank. In order to improve performances
on rarer cells, the corpus was augmented with synthetic sentences which… See the full description on the dataset page:
https://huggingface.co/datasets/datasets-CNRS/French_verbal_frequencies_Open_Subtitle_corpus.