Views
No views yet
1from transformers import AutoModelForMaskedLM, AutoTokenizer
2
3# See the `MDLM` collection page on the hub for list of available models.
4tokenizer = transformers.AutoTokenizer.from_pretrained('howey/HDT-E')
5model_name = 'howey/HDT-E'
6model = AutoModelForMaskedLM.from_pretrained(model_name)8192 and is similar in size to BERT with approximately 110M parameters,
was trained on standard masked language modeling task with a Transformer-based architecture using our proposed hierarchical attention.
The training regimen comprised 24 hours on the ArXiv+Wikipedia+HUPD corpus, involving the processing of a total of 1.3 billion tokens.@inproceedings{He2024COLM,
title={HDT: Hierarchical Document Transformer},
author={Haoyu He and Markus Flicke and Jan Buchmann and Iryna Gurevych and Andreas Geiger},
year={2024},
booktitle={Conference on Language Modeling}
}