Views
No views yet
Helsinki-NLP/opus-mt-ja-en, packaged for offline use in Playto.Helsinki-NLP/opus-mt-ja-en — MarianMT, transformer-align architecture, OPUS dataset1pip install ctranslate2 transformers sentencepiece
2ct2-transformers-converter \
3 --model Helsinki-NLP/opus-mt-ja-en \
4 --output_dir opus-mt-ja-en-ct2 \
5 --quantization int8 \
6 --copy_files source.spm target.spm
7tar czf opus-mt-ja-en-ct2.tar.gz opus-mt-ja-en-ct2opus-mt-ja-en-ct2.tar.gz)| File | Size | Purpose |
|---|---|---|
model.bin | ~74 MB | CTranslate2 int8 quantized weights |
shared_vocabulary.json | ~1.2 MB | CTranslate2 vocab |
source.spm | ~780 KB | SentencePiece source tokenizer |
target.spm | ~800 KB | SentencePiece target tokenizer |
config.json | ~250 B | CTranslate2 config |
ctranslate2 (Python)1import ctranslate2
2import sentencepiece
3
4translator = ctranslate2.Translator("opus-mt-ja-en-ct2", device="cpu", compute_type="int8")
5sp_source = sentencepiece.SentencePieceProcessor("opus-mt-ja-en-ct2/source.spm")
6sp_target = sentencepiece.SentencePieceProcessor("opus-mt-ja-en-ct2/target.spm")
7
8source_tokens = sp_source.encode("せっかくまた、なるほどくんに会えたのに", out_type=str) + ["</s>"]
9results = translator.translate_batch([source_tokens])
10print(sp_target.decode(results[0].hypotheses[0]))
11# → "I can't believe I met you again."ct2rs (Rust)1use ct2rs::{Translator, Tokenizer};
2
3let tokenizer = Tokenizer::new("opus-mt-ja-en-ct2")?;
4let translator = Translator::with_tokenizer("opus-mt-ja-en-ct2", tokenizer, /* config */)?;
5let result = translator.translate_batch(&["せっかくまた、なるほどくんに会えたのに".to_string()], /* options */)?;</s> appended to source token sequences. The ct2rs::Tokenizer wrapper handles this automatically; raw SentencePiece calls must add it manually.Helsinki-NLP/opus-mt-ja-en. License is CC-BY 4.0 inherited from upstream.Helsinki-NLP. OPUS-MT — Open Machine Translation Models.
https://github.com/Helsinki-NLP/Opus-MT