A 19,997-row supervised fine-tuning corpus ({"prompt","completion"} JSONL) for
adapting a language model to Yorùbá, built for the Adaption AutoScientist
challenge (language track). Every row is human-written, from an approved,
train-split-only source. No synthetic or LLM-generated text.
SHA-256: a711692096719c0d11f8e9c3784061211733029bf32a20e36b7e920488f8cc0c
Rows: 19,997 · Format: JSONL, prompt + completion… See the full description on the dataset page:
https://huggingface.co/datasets/Enochid/foundry-y-yoruba-corpus.