本仓库集合了github上的各种古籍资料,主要来自殆知阁的资料,经过整理,做了简单的数据清洗。是互联网上能找到的比较齐全的文言文数据集。可以作为知识库,也可以作为训练集使用。能优化模型的文言文能力,希望我们共同把华夏文明传承下去,用大模型和AI技术赋能我们的文明。
This repository aggregates a variety of ancient Chinese texts from GitHub, primarily sourced from 殆知阁 (Almost Know Pavilion), and has undergone organization and basic data cleaning.
It is one of the most comprehensive classical Chinese (文言文) datasets available on the internet. It can serve as a knowledge base or a training dataset, helping to… See the full description on the dataset page:
https://huggingface.co/datasets/astra77/huaxia-lib.