Views
No views yet
<doc-s>, </doc-s> (the last two tokens in the vocabulary), and <s>, </s>, respectively.1from transformers import AutoTokenizer, AutoModel
2# load model and tokenizer
3tokenizer = AutoTokenizer.from_pretrained('biu-nlp/cdlm')
4model = AutoModel.from_pretrained('biu-nlp/cdlm')1@article{caciularu2021cross,
2 title={Cross-Document Language Modeling},
3 author={Caciularu, Avi and Cohan, Arman and Beltagy, Iz and Peters, Matthew E and Cattan, Arie and Dagan, Ido},
4 journal={Findings of the Association for Computational Linguistics: EMNLP 2021},
5 year={2021}
6}