This repository contains the midtraining corpus used for ODYSSIM behavioral foundation model experiments. The corpus is stored as parquet shards grouped by dataset/source, with train shards and held-out test files for the corresponding sources.
The release mirrors the previously staged dataset Xuhui/sft_processed_large_split_v3 into the CMU-LTI organization for the paper release. Summary statistics from the audit pass: 21.4M train rows, approximately… See the full description on the dataset page:
https://huggingface.co/datasets/cmu-lti/osim-mid-training.