This is a collection of Latin strings that, hopefully, can be directly fed to an LLM during fine-tuning.
I have tried to remove all unnecessary characters and strings that are non-latin.
Each row in this dataset has a string with ~128 words.
@dataset{CleanLatin2025,
author = {Ashraff Hathibelagal},
title = {Clean Latin},
year = {2025},
url = {
https://huggingface.co/datasets/hathibelagal/clean_latin}
}