Everything needed to rebuild any stage of the sbob project without the original
episode files. Private, personal research only — derived from © Nickelodeon material.
corpus/
lines.jsonl — 129k speaker-attributed transcript lines from 1043 episodes (the scraped source corpus); episodes.jsonl index
text_datasets/
per-character LoRA train/val jsonls (10 chars) + _ensemble/ (103 speakers, 77k records) —… See the full description on the dataset page:
https://huggingface.co/datasets/weystrom/sbob-datasets.