Views
No views yet
Conversational RuBERT) and other distilled models.| Model name | # params, M | # vocab, K | Mem., MB |
|---|---|---|---|
rubert-base-cased-conversational | 177.9 | 120 | 679 |
distilrubert-base-cased-conversational | 135.5 | 120 | 517 |
distilrubert-small-cased-conversational | 107.1 | 120 | 409 |
cointegrated/rubert-tiny | 11.8 | 30 | 46 |
| distilrubert-tiny-cased-conversational | 10.4 | 31 | 41 |
PyTorchBenchmark from transformers to evaluate model's performance and compare it with other pre-trained language models for Russian. All tests were performed on Intel(R) Xeon(R) CPU E5-2698 v4 @ 2.20GHz and nVIDIA Tesla P100-SXM2.0 16Gb.rubert-base-cased-conversational | 1 | 512 | 0.147 | 0.014 | 897 | 1531 |
| distilrubert-base-cased-conversational | 1 | 512 | 0.083 | 0.006 | 766 | 1423 |
| distilrubert-small-cased-conversational | 1 | 512 | 0.03 | 0.002 | 600 | 1243 |
| cointegrated/rubert-tiny | 1 | 512 | 0.041 | 0.003 | 272 | 919 |
| distilrubert-tiny-cased-conversational | 1 | 512 | 0.023 | 0.003 | 206 | 855 |
| rubert-base-cased-conversational | 16 | 512 | 2.839 | 0.182 | 1499 | 2071 |
| distilrubert-base-cased-conversational | 16 | 512 | 1.065 | 0.055 | 2541 | 2927 |
| distilrubert-small-cased-conversational | 16 | 512 | 0.373 | 0.003 | 1360 | 1943 |
| cointegrated/rubert-tiny | 16 | 512 | 0.628 | 0.004 | 1293 | 2221 |
| distilrubert-tiny-cased-conversational | 16 | 512 | 0.219 | 0.003 | 633 | 1291 |@misc{https://doi.org/10.48550/arxiv.2205.02340,
doi = {10.48550/ARXIV.2205.02340},
url = {https://arxiv.org/abs/2205.02340},
author = {Kolesnikova, Alina and Kuratov, Yuri and Konovalov, Vasily and Burtsev, Mikhail},
keywords = {Computation and Language (cs.CL), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
title = {Knowledge Distillation of Russian Language Models with Reduction of Vocabulary},
publisher = {arXiv},
year = {2022},
copyright = {arXiv.org perpetual, non-exclusive license}
}