This repo contains CroCoSum, a dataset of cross-lingual code-switched summarization sourced from a technology-oriented news forum.
For more description and statistics of the dataset, please refer to our paper here.
Thank you for using our dataset! If you have used our data in your work, please consider using the citation below.
@inproceedings{zhang-eickhoff-2024-crocosum,
title = "{C}ro{C}o{S}um: A Benchmark Dataset for Cross-Lingual Code-Switched… See the full description on the dataset page:
https://huggingface.co/datasets/ruochenz/CroCoSum.