TWLLM-Data is the first large-scale open dataset containing real user-generated conversation logs from TWLLM and TWLLM Arena, where over 80% of users are based in Taiwan. The dataset is designed to facilitate the development and evaluation of Traditional Mandarin Large Language Models (LLMs).
We extend our gratitude to Professor Yun-Nung (Vivian) Chen for her guidance and advisement.
Special thanks to Tzu-Han Lin, Kang-Chieh… See the full description on the dataset page:
https://huggingface.co/datasets/yentinglin/twllm-data.