The name "Chikuma" is inspired by the Chikuma River, the longest in Japan, known for its continuous flow and meandering path.
This metaphorically represents the model's depth, fluidity, and adaptability in processing and understanding language.
Dataset used for Fine Tuning
Dataset: /argilla/distilabel-intel-orca-dpo-pairs
The dataset was roughly ~3000 samples but they were high quality (according to the chosen_score).
The following filters were applied to the original dataset:
The chat template for Chikuma_10.7B - V2 is a modified version of ChatML, optimized for improved interaction and engagement:
<|im_start|>GPT4 Correct system:
{system} Always use <|end_of_turn|> when you want to end the answer. <|im_end|>
<|im_start|>GPT4 Correct user:
{user}<|im_end|>
<|im_start|>GPT4 Correct Assistant:
{asistant}<|im_end|>
Nous Benchmark Evaluation
Model
AGIEval
GPT4All
TruthfulQA
Bigbench
Average
SynthIQ-7b
42.67
73.71
56.51
44.59
54.37
openchat/openchat-3.5-0106
44.17
73.72
52.53
44.4
53.71
Chikuma_10.7B
42.41
73.41
56.69
43.5
54.00
Chikuma_10.7B_v2
42.77
73.81
58.83
44.83
55.06
OpenLLM Leaderboard
Benchmark Name
Performance
ARC
66.38
HellaSwag
85
MMLU
65.27
TruthfulQA
58.83
Winogrande
78.77
GSM8K
63.68
Average
69.65
Training Environment
Hardware: Single A100 80GB GPU in a runpod, utilized for approximately 1.5 hours.