AV task used in "Tokenization is Sensitive to Language Variation" paper, Arxiv link.
Note that "Contrastive Learning" Train/Dev files were used with contrastive learning (SupConLoss) to fine-tune BERT models. Then a threshold was chosen based on the Thresholding Dev file and the performance was calculated on the Thresholding Test file.
@article{wegmann2025tokenization,
title={Tokenization is Sensitive to Language Variation},
author={Wegmann, Anna and Nguyen, Dong and Jurgens, David}… See the full description on the dataset page:
https://huggingface.co/datasets/AnnaWegmann/AV.