This dataset is a part of the zh-tw-llm project.
Tokenizer: zh-tw-pythia-tokenizer-a8000-v1
Built with: translations, sharegpt
Rows: train 305958, test 195
Max length: 1024
Full config:{"build_with": ["translations", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English: {lang_1}\nChinese: {lang_2}"… See the full description on the dataset page:
https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_sg-302-c1024.