Our pre-training starts from the bert-base-chinese checkpoint released by Hugging Face. We pre-train the model with the optimizer AdamW, learning rate 1e-4 and batch size 384 for 20 epochs on the whole corpus. A scheduler of linear warmup in the first 20% steps followed by linear decay is applied. The masking rate is 15% for the masked language modeling (MLM) task; the maximum input length is 512.In
addition, we apply the whole word masking and dynamic masking strategies.