A benchmark dataset for evaluating embedding models on Rabbinic Hebrew and Aramaic texts, with parallel English translations sourced from Sefaria.
This dataset contains 3,708 parallel text pairs spanning diverse Rabbinic literature across multiple centuries and genres. It is designed for evaluating cross-lingual embedding models on their ability to align Hebrew/Aramaic source texts with English… See the full description on the dataset page:
https://huggingface.co/datasets/Sefaria/Rabbinic-Hebrew-English-Pairs.