This corpus is a collection of 57170 potentially idiomatic expressions (PIEs) based on the British National Corpus, prepaired for NER task.
Each of the objects is comes with a contextual set of tokens, BIO tags and boolean label.
The data sources are:
MAGPIE corpus
PIE corpus
Detailed data preparation pipeline can be found here
Supported Tasks and Leaderboards
Token classification (NER)
Languages… See the full description on the dataset page: https://huggingface.co/datasets/Gooogr/pie_idioms.