Data is St. Thomas Aquinas Latin-English parallel corpus (207,689 segments) from the Aquinas Translation ProjectGet a random Latin-English pair from it.
Latin source (src) and English target (trg) files are formatted for training with Eole, OpenNMT-py's successor, using Byte-Pair Encoding (BPE), which is far superior to word tokenization; cf. Lane & Dyshel, Natural Language Processing in Action, §2.2 "Beyond word tokens".