Translated proj-persona/PersonaHub using nayohan/llama3-instrucTrans-enko-8b.
For this dataset, we only used data that is 5000 characters or less in length and has language of English.
Thanks for @proj-persona and @nayohan.
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various… See the full description on the dataset page:
https://huggingface.co/datasets/youjunhyeok/PersonaHub-ko.