This project provides a large-scale Chinese GPT model pre-trained on the dataset
LCCC.
We present a series of Chinese GPT model that are first pre-trained on a Chinese novel dataset and then post-trained on our LCCC dataset.
Similar to
TransferTransfo, we concatenate all dialogue histories into one context sentence, and use this sentence to predict the response. The input of our model consists of word embedding, speaker embedding, and positional embedding of each word.
1from transformers import OpenAIGPTLMHeadModel, GPT2LMHeadModel, BertTokenizer
2import torch
3
4
5tokenizer = BertTokenizer.from_pretrained("thu-coai/CDial-GPT2_LCCC-base")
6model = GPT2LMHeadModel.from_pretrained("thu-coai/CDial-GPT2_LCCC-base")
For more details, please refer to our
repo. on github.