Views
No views yet
dsfsi/m2m100_418m-nso-eng.facebook/m2m100_418M. Licence: Apache-2.0, as declared
by the original repository. This repository only converts the weights to
ONNX. All credit for the model belongs to DSFSI.optimum-cli export onnx --model dsfsi/m2m100_418m-nso-eng \
--task text2text-generation-with-past --no-post-process <outdir>.onnx.data -> .onnx_data rename
and proto location rewrite (encoder, decoder and decoder-with-past, plus
one Constant node tensor attribute per file). Verified with a real load
from a cache cleared beforehand.| Path | Precision | Size |
|---|---|---|
*.onnx (root) | fp32 | 4.6 GB |
int8/*.onnx + int8/*.onnx_data | int8 dynamic | 3.0 GB |
AutoModelForSeq2SeqLM.generate().1parity:
2 sample_size: 10
3 metric: exact_match
4 fp32_greedy: 1.00 # 10/10
5 fp32_beam4: 1.00 # 10/10
6 int8_greedy: 0.80 # 8/10
7 int8_beam4: 0.70 # 7/10| Decoding | fp32 (n=10) | int8 (n=10) |
|---|---|---|
| greedy | 100% (10/10) | 80% (8/10) |
| beam=4 | 100% (10/10) | 70% (7/10) |
ns (not nso). The target is
not baked into forced_bos_token_id in generation_config.json for
this checkpoint - you must set it explicitly:1tokenizer.src_lang = "ns"
2forced_bos_token_id = tokenizer.lang_code_to_id["en"]
3model.generate(**inputs, forced_bos_token_id=forced_bos_token_id, ...)1from transformers import AutoTokenizer
2from optimum.onnxruntime import ORTModelForSeq2SeqLM
3
4repo = "TigreGotico/m2m100_418m-nso-eng-onnx"
5tok = AutoTokenizer.from_pretrained(repo)
6tok.src_lang = "ns"
7model = ORTModelForSeq2SeqLM.from_pretrained(repo, use_cache=True, use_merged=False)
8
9enc = tok("Boso bo bonolo kudu mo lebakeng le.", return_tensors="pt")
10tgt_id = tok.lang_code_to_id["en"]
11out = model.generate(**enc, forced_bos_token_id=tgt_id, num_beams=4, max_new_tokens=64)
12print(tok.batch_decode(out, skip_special_tokens=True)[0])subfolder="int8".