This repository contains the domain-specific continued-pretraining (CPT) data,
the tokenized and preprocessed datasets, and the aligned KNN distributions used
by MemoryDecoder at Scale.
Project Page: Memory Decoder at Scale
GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale
Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
The preprocessed datasets and KNN distributions in this repository use… See the full description on the dataset page:
https://huggingface.co/datasets/Rubin-Wei/MemoryDecoder-at-Scale-domain-data.