This dataset is a part of the zh-tw-llm-dev project.
Tokenizer: zh-tw-llm-dev-tokenizer-a8k-d40d11
Built with: translations, wikipedia, sharegpt, alpaca
Rows: train 500, test 50
Max length: 2048
Full config:{"build_with": ["translations", "wikipedia", "sharegpt", "alpaca"], "preview_length": 256, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en"… See the full description on the dataset page:
https://huggingface.co/datasets/zh-tw-llm-dv-dv/zh-tw-llm-dev-sample-ta8k-d40d11-only_embeddings-tr_wiki_sg_alp-396867-c2048.