Synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models.
Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have
different number of follow-up interactions, could use different strategy, and have train and validation
splits.
After first experiments with MRL, we decided to abandon single step and two steps stages. That's because with single
step… See the full description on the dataset page:
https://huggingface.co/datasets/ReactiveAI/TinyStories-MRL.