Views
No views yet
trust_remote_code=True and use the causal backend.pip install torch transformers safetensors sentencepiece pypinyin jiebasentencepiece is required for AutoTokenizer. pypinyin is required
for raw Mandarin-to-pinyin tokenization. jieba is required when
use_jieba is true; this export was created with use_jieba=true.1from transformers import AutoConfig, AutoModel, AutoModelForCausalLM, AutoTokenizer
2
3model_path = "PATH_OR_REPO_ID"
4
5config = AutoConfig.from_pretrained(model_path, trust_remote_code=True)
6tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
7base_model = AutoModel.from_pretrained(model_path, trust_remote_code=True)
8model = AutoModelForCausalLM.from_pretrained(model_path, trust_remote_code=True)causaltokenizer(text), tokenizer(text, add_special_tokens=False), and
tokenizer(texts, padding=True, truncation=True, return_tensors="pt").
It also accepts return_offsets_mapping=True for compatibility with
completion-ranking evaluators that need suffix masks. The model supports
output_hidden_states=True for representation extraction tasks.patch_pathlib_utf8_open=true in config.json. When loaded
with trust_remote_code=True, the config installs a narrow Windows
compatibility shim so later text-mode Path.open("r") calls without an
explicit encoding default to UTF-8. Set
PINYIN_CODE_DISABLE_UTF8_OPEN_PATCH=1 before loading the model to disable
that shim.pinyin-codetrue