This repository contains the streaming dataset used for product-domain masked-language-model adaptation.
The data was prepared from product titles and descriptions and tokenized with:
BAAI/bge-m3-retromae
The training split is stored as fixed-length pre-tokenized binary shards for efficient streaming training. The validation split is stored as JSONL product text.
metadata.json — dataset/preparation metadata.
train_*.bin —… See the full description on the dataset page:
https://huggingface.co/datasets/mjaliz/product-td-25M-rw.