Views
No views yet
A100 40GB GPUx4.OpenOrca, Ultrafeedback, and OpenHermes.
However, this approach may violate these private models' terms of service (ToS).
For instance, OpenAI's license explicitly states: "⚠️Use Limitation: Creating services that compete with OpenAI.⚠️"
This implies that using data generated by private models to create unrestricted, open LLMs is challenging.Open-Source based dataset, we use microsoft/WizardLM-2-8x22B through DeepInfra.Evolving system, which is propsed by WizardLM.
In training, we used 1849 training dataset, and 200 validation dataset.Validation loss (epoch 15; Learning rate: 1e-5): 1.0040
Logickor-v2 eval model.(GPT-4o occasionally makes errors when grading. For example, it sometimes assigns a score of 0 for English responses to questions that were supposed to be answered in English.)
| Model | 추론 | 수학 | 글쓰기 | 코딩 | 이해 | 문법 | 싱글턴 | 멀티턴 | Overall |
|---|---|---|---|---|---|---|---|---|---|
| OpenAI/gpt-4o-2024-05-13 | 9.50 | 8.71 | 9.42 | 9.21 | 9.71 | 9.42 | 9.42 | 9.23 | 9.33 |
| Anthropic/clauide-3-5-sonnet-20240620 | 8.64 | 8.42 | 9.85 | 9.78 | 9.92 | 9.21 | 9.26 | 9.35 | 9.30 |
| google/gemini-1.5-pro-001 | 9.07 | 8.57 | 9.57 | 9.78 | 9.57 | 9.21 | 9.40 | 9.19 | 9.23 |
| ---- | ---- | ---- | ---- | ---- | ---- | ---- | ---- | ---- | ---- |
| Gukbap-Qwen2-7B🍚 | 5.71 | 6.43 | 8.07 | 9.14 | 7.29 | 3.57 | 7.02 | 6.38 | 6.70 |
| mirlab/AkaLlama-llama3-70b-v0.1 | 5.14 | 5.35 | 4.14 | 9.00 | 7.85 | 7.50 | 5.97 | 7.02 | 6.50 |
| Qwen/Qwen2-7B-Instruct | 6.07 | 4.71 | 7.21 | 7.00 | 8.00 | 4.85 | 6.61 | 6.00 | 6.30 |
| yanolja/EEVE-Korean-Instruct-10.8B-v1.0 | 6.00 | 3.64 | 6.64 | 5.64 | 8.42 | 5.85 | 6.61 | 5.45 | 6.01 |
| Model (type) | 추론 | 수학 | 글쓰기 | 코딩 | 이해 | 문법 | 싱글턴 | 멀티턴 | Overall |
|---|---|---|---|---|---|---|---|---|---|
| Gukbap-Qwen2-7B🍚 (cot-1-shot) | 7.07 | 5.71 | 8.86 | 9.00 | 8.07 | 3.86 | 7.79 | 6.40 | 7.10 |
| Gukbap-Qwen2-7B🍚 (1-shot) | 7.50 | 6.00 | 7.86 | 8.71 | 7.21 | 3.57 | 7.10 | 6.52 | 6.81 |
| Gukbap-Qwen2-7B🍚 (0-shot) | 5.71 | 6.43 | 8.07 | 9.14 | 7.29 | 3.57 | 7.02 | 6.38 | 6.70 |
judge_template, prompt, etc.1<|im_start|>user
2Hello! My favorite food is Gukbap🍚!<|im_end|>
3<|im_start|>assistant
4(model answer)@article{HumanF-MarkrAI,
title={Gukbap-Qwen2-7B},
author={MarkrAI},
year={2024},
url={https://huggingface.co/HumanF-MarkrAI}
}