This is a corpus of code-switching sentences and tweets. We provide automatic tokenization, token-based language identification, translation to monolingual sentences in each of the contributing languages, alignments to these monolingual sentences and parses of them.
The data creation is described in the following publication:
@inproceedings{sterner-2025-acs,
author = {Igor Sterner and Simone Teufel},
title = {Minimal Pair-Based Evaluation of Code-Switching},
booktitle =… See the full description on the dataset page:
https://huggingface.co/datasets/igorsterner/acs-corpus.