MLCS code-mix research dataset
Dataset Details
Dataset Description
This dataset is the MLCS Data Card for CodeMix research. It includes multilingual code-mixed content, intended to support research in multilingual language modeling, classification, and other NLP tasks involving language switching.
Curated by: Chippo Sekkabanja
Funded by: KDD
Shared by [optional]: [More Information Needed]
Language(s) (NLP): English, Chinese
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page:
https://huggingface.co/datasets/chipotswift/MLCS-test-1.