Frozen lite English embedding-evaluation dataset — the downsampled
MTEB(eng, v1) + bitext subsets used by the
tips-evaluation framework
(english/ runner, --lite), published so evaluation boxes download this
(~0.6 GB) instead of the ~122 GB original corpora. The English counterpart of
TIPS-Korean.
Built from the 56 MTEB(eng, v1) tasks + BUCC/Tatoeba with seeded,
model-independent downsampling (english/downsample.py, seed 42):
Retrieval… See the full description on the dataset page:
https://huggingface.co/datasets/Cartinoe5930/TIPS-English.