This data repository corresponds to our paper BERTtime Stories: Investigating the Role of Synthetic Story Data in
Language Pre-training as part of the 2024 BabyLM Challenge.
The code implementation is released on github
Our trained models are released on HuggingFace
This repository contains the results of the LLM-evaluation of generative performance, conduncted with Claude-3.5 Sonnet.
Specifically, it contains the following files:
assistant_responses.json: contains model generations and… See the full description on the dataset page:
https://huggingface.co/datasets/nikitastheo/BERTtime_Stories_LLM_Eval.