Example of a chunk record dataset, made for debugging purposes.
Made by taking the step-1216 validation chunk records of reason-only-k=32-lr=0.00015-replay=0.5 and splitting the doc uuids in half into train and test. Thus, all of a document's chunks lie within exactly one of the splits.
There are (4096/32) * 4864 =622_592$ chunks (including the first chunk) collectively.
test
311_296
38_659
0.1242… See the full description on the dataset page:
https://huggingface.co/datasets/JackHsieh/chunk-records-example.