Stage-00 snapshot (2026-05-02). Intermediate MERCI / MaxEmbed / DLRM
preprocessing artifacts: filtered CSVs, vocab tables, embedding-table
indexes, partition outputs.
These are the outputs of MERCI_page_aware/analysis/preprocess_*.py
applied to each raw dataset; they are the input to trace generation
(research_data/traces/) and to the Cylon/MQSim/MaxEmbed simulators.
criteo_terabyte/ — 45 GB
criteo_kaggle/ —… See the full description on the dataset page:
https://huggingface.co/datasets/shadowcollecter/cxlssd-processed.