JMTEB is a benchmark for evaluating Japanese text embedding models. It consists of 5 tasks, currently involving 28 datasets in total. You can find the update history here.
JMTEB_DATASET_NAMES = (
'livedoor_news',
'mewsc16_ja',
'sib200_japanese_clustering'… See the full description on the dataset page:
https://huggingface.co/datasets/sbintuitions/JMTEB.