MRA model for masked language modeling (MLM) for sequence length 512.
The MRA model was proposed in
Multi Resolution Analysis (MRA) for Approximate Self-Attention by Zhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn M Fung, and Vikas Singh.
This model was contributed by
novice03.
The original code can be found
here.