The model is the instruction-tuned version of
rinna/nekomata-14b. It adopts the Alpaca input format.
-
Model architecture
A 40-layer, 5120-hidden-size transformer-based language model. Please refer to the
Qwen paper for architecture details.
-
Fine-tuning
The fine-tuning data is the subset of the following datasets.
- Databricks Dolly data
- Japanese Databricks Dolly data
- FLAN Instruction Tuning data and its Japanese translation
- Izumi lab LLM Japanese dataset
- The following sections are used
- alt
- aozora-txt
- CourseraParallel
- ParaNatCom
- Tab-delimited_Bilingual_Sentence_Pairs
- tanaka-corpus
- wikinews
- wordnet
- yasashi-japanese
- The remaining sections contain commonly used evaluation corpora so they are skipped to prevent data leak.
-
Contributors
-
Release date
December 21, 2023
1import torch
2from transformers import AutoTokenizer, AutoModelForCausalLM
3
4tokenizer = AutoTokenizer.from_pretrained("rinna/nekomata-14b-instruction", trust_remote_code=True)
5
6# Use GPU with bf16
7# model = AutoModelForCausalLM.from_pretrained("rinna/nekomata-14b-instruction", device_map="auto", trust_remote_code=True, bf16=True)
8
9# Use GPU with fp16
10# model = AutoModelForCausalLM.from_pretrained("rinna/nekomata-14b-instruction", device_map="auto", trust_remote_code=True, fp16=True)
11
12# Use CPU
13# model = AutoModelForCausalLM.from_pretrained("rinna/nekomata-14b-instruction", device_map="cpu", trust_remote_code=True)
14
15# Automatically select device and precision
16model = AutoModelForCausalLM.from_pretrained("rinna/nekomata-14b-instruction", device_map="auto", trust_remote_code=True)
17
18instruction = "次の日本語を英語に翻訳してください。"
19input = "大規模言語モデル(だいきぼげんごモデル、英: large language model、LLM)は、多数のパラメータ(数千万から数十億)を持つ人工ニューラルネットワークで構成されるコンピュータ言語モデルで、膨大なラベルなしテキストを使用して自己教師あり学習または半教師あり学習によって訓練が行われる。"
20prompt = f"""
21以下は、タスクを説明する指示と、文脈のある入力の組み合わせです。要求を適切に満たす応答を書きなさい。
22
23### 指示:
24{instruction}
25
26### 入力:
27{input}
28
29### 応答:
30"""
31token_ids = tokenizer.encode(prompt, add_special_tokens=False, return_tensors="pt")
32
33with torch.no_grad():
34 output_ids = model.generate(
35 token_ids.to(model.device),
36 max_new_tokens=200,
37 do_sample=True,
38 temperature=0.5,
39 pad_token_id=tokenizer.pad_token_id,
40 bos_token_id=tokenizer.bos_token_id,
41 eos_token_id=tokenizer.eos_token_id
42 )
43
44output = tokenizer.decode(output_ids.tolist()[0])
45print(output)
46"""
47以下は、タスクを説明する指示と、文脈のある入力の組み合わせです。要求を適切に満たす応答を書きなさい。
48
49### 指示:
50次の日本語を英語に翻訳してください。
51
52### 入力:
53大規模言語モデル(だいきぼげんごモデル、英: large language model、LLM)は、多数のパラメータ(数千万から数十億)を持つ人工ニューラルネットワークで構成されるコンピュータ言語モデルで、膨大なラベルなしテキストを使 用して自己教師あり学習または半教師あり学習によって訓練が行われる。
54
55### 応答:
56 A large language model (LLM) is a computer language model composed of artificial neural networks with many parameters (from tens of millions to billions) trained by self-supervised learning or semi-supervised learning using a large amount of unlabeled text.<|endoftext|>
57"""
Please refer to
rinna/nekomata-14b for tokenization details.
1@misc{rinna-nekomata-14b-instruction,
2 title = {rinna/nekomata-14b-instruction},
3 author = {Zhao, Tianyu and Sawada, Kei},
4 url = {https://huggingface.co/rinna/nekomata-14b-instruction}
5}
6
7@inproceedings{sawada2024release,
8 title = {Release of Pre-Trained Models for the {J}apanese Language},
9 author = {Sawada, Kei and Zhao, Tianyu and Shing, Makoto and Mitsui, Kentaro and Kaga, Akio and Hono, Yukiya and Wakatsuki, Toshiaki and Mitsuda, Koh},
10 booktitle = {Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)},
11 month = {5},
12 year = {2024},
13 pages = {13898--13905},
14 url = {https://aclanthology.org/2024.lrec-main.1213},
15 note = {\url{https://arxiv.org/abs/2404.01657}}
16}