ComMT is a comprehensive dataset suite designed to support the development and evaluation of universal translation models.
It includes diverse translation-related tasks, providing a well-curated data resource for training and testing LLM-based machine translation systems.
The dataset is meticulously curated from over 60+ publicly available data sources.
The… See the full description on the dataset page:
https://huggingface.co/datasets/NiuTrans/ComMT.