676K passages (~1.4B characters) machine-translated into Ancient Greek from the Latin holdings of
Corpus Corporum (University of Zurich), filtered for length and for Greekness.
Role in training. This is the bronze tier: it is held apart from genuine Greek throughout
pretraining and is downweighted first during the staged anneal, ending with gold data only. It is
augmentation for a finite corpus, not a scholarly text, and no result in the… See the full description on the dataset page:
https://huggingface.co/datasets/Ericu950/SyntheticAncientGreek-CorpusCorporum.