Views
No views yet
Aniemore/rubert-tiny-emotion-russian-cedr-m7 — multi-label emotion recognition for Russian text over seven classes: anger, disgust, enthusiasm, fear, happiness, neutral, sadness.| subfolder | scheme | weights | ROC AUC (macro) | macro-F1 | WA | UA |
|---|---|---|---|---|---|---|
| (original repo) | fp32 | 111 MiB | 0.9008 | 0.6036 | 0.7965 | 0.6121 |
int8 | W8A16 | 107 MiB | 0.9008 | 0.6040 | 0.7965 | 0.6119 |
fp8 | W8A16-float | 107 MiB | 0.9004 | 0.6026 | 0.7991 | 0.6142 |
int4 | W4A16_ASYM | 106 MiB | 0.9035 | 0.6021 | 0.7944 | 0.6143 |
Linear layers are quantized. In a BERT classifier the embedding matrix is not one of them, and on the smaller models it is most of the checkpoint — so the saving here scales with the encoder rather than with the parameter count. The large model compresses well; rubert-tiny barely moves, and the table above says so rather than quoting a ratio from the layers that did shrink.1import torch
2from transformers import AutoModelForSequenceClassification, AutoTokenizer
3
4repo = "Aniemore/rubert-tiny-emotion-russian-cedr-m7-quantized"
5model = AutoModelForSequenceClassification.from_pretrained(
6 repo, subfolder="int8").eval() # or "fp8", "int4"
7tok = AutoTokenizer.from_pretrained(repo, subfolder="int8")
8
9x = tok("мне сегодня очень грустно", return_tensors="pt")
10with torch.no_grad():
11 # multi-label: sigmoid per class, not softmax over classes
12 probs = model(**x).logits.sigmoid()[0]
13print({model.config.id2label[i]: round(p.item(), 3) for i, p in enumerate(probs)})cointegrated/rubert-tiny; the licence follows the base model.