This dataset contains parallel sentences of English and Cebuano extracted from the Bible corpus available at
https://github.com/christos-c/bible-corpus. The dataset is formatted for use in training machine translation models, particularly with the Transformers library from Hugging Face.
This dataset can be used to evaluate the performance of Large Language Models for English-Cebuano machine translation using libraries… See the full description on the dataset page:
https://huggingface.co/datasets/eemberda/english-ceb-bible-prompt.