The dataset consists of three sources of dataset described in the paper, including:
ChessCLIP dataset: Annotated PGNs for training CLIP.
ChessGPT Base dataset: Game dataset, language dataset and mixed dataset for training ChessGPT-Base.
ChessGPT Chat dataset: Conversational dataset for training ChessGPT-Chat.
Because of the legal issue, for ChessGPT dataset, we do not open-source the chess-book, chess-forum, chess-blog, and Youtube… See the full description on the dataset page: https://huggingface.co/datasets/Waterhorse/chess_data.