This is a Japanese RoBERTa large model pretrained on Japanese Wikipedia and the Japanese portion of CC-100.
1from transformers import AutoTokenizer, AutoModelForMaskedLM
2tokenizer = AutoTokenizer.from_pretrained("nlp-waseda/roberta-large-japanese")
3model = AutoModelForMaskedLM.from_pretrained("nlp-waseda/roberta-large-japanese")
4
5sentence = '早稲田 大学 で 自然 言語 処理 を [MASK] する 。' # input should be segmented into words by Juman++ in advance
6encoding = tokenizer(sentence, return_tensors='pt')
7...
You can fine-tune this model on downstream tasks.
The input text should be segmented into words by
Juman++ in advance. Juman++ 2.0.0-rc3 was used for pretraining. Each word is tokenized into tokens by
sentencepiece.
BertJapaneseTokenizer now supports automatic
JumanppTokenizer and
SentencepieceTokenizer. You can use
this model without any data preprocessing.
The vocabulary consists of 32000 tokens including words (
JumanDIC) and subwords induced by the unigram language model of
sentencepiece.
This model was trained on Japanese Wikipedia (as of 20210920) and the Japanese portion of CC-100. It took two weeks using eight NVIDIA A100 GPUs.