#1 among Korean-developed models / #6 overall on the Open Ko-LLM Leaderboard Season 2 (NIA, Jan 2025) — at 8B parameters, ranking alongside 14B–27B models within ~2 points of average score.
A Llama 3.1 8B model supervised-fine-tuned for Korean language, culture, and enterprise domains (law, finance, tax, accounting).
오픈 Ko-LLM 리더보드 시즌2(NIA, 2025년 1월)에서 한국 개발 모델 중 1위, 전체 6위. 8B 파라미터로 14B~27B 모델들과 평균 점수 약 2점 차 이내에서 경쟁. 법률·재무·세무·회계 등 한국 엔터프라이즈 도메인에 특화된 Llama 3.1 8B SFT 모델입니다.
Results / 평가 결과
On the Open Ko-LLM Leaderboard Season 2 (hosted by NIA, the National Information Society Agency of Korea), this 8B model scored an average of 53.94, ranking #1 among Korean-developed models and #6 overall out of 1,228 evaluated models (as of January 2025). The top five were larger Gemma-2 / Qwen-2.5-based models from non-Korean developers — this 8B model competed within ~2 points of 14B–27B models.
오픈 Ko-LLM 리더보드 시즌2(NIA 한국지능정보사회진흥원 주최)에서 본 8B 모델은 평균 53.94점으로, 평가된 1,228개 모델 중 한국 개발 모델 1위, 전체 6위를 기록했습니다 (2025년 1월 기준). 상위 5개는 해외 개발자의 더 큰 Gemma-2/Qwen-2.5 기반 모델이었으며, 본 8B 모델은 14B~27B 모델들과 약 2점 차 이내에서 경쟁했습니다.
Note on efficiency / 효율성 참고: ranks 6th at 8B parameters, matching or exceeding larger 9B–27B models. (8B로 9B~27B 모델들과 대등하거나 앞섭니다.)
Model Description / 모델 설명
This model is fine-tuned from Llama 3.1 8B Instruct using supervised fine-tuning (SFT), specialized for the Korean language and Korean cultural and professional contexts. It is trained on a self-built dataset spanning 53 domains, with a focus on Korean social values and enterprise use cases such as law, finance, tax, and accounting.
본 모델은 Llama 3.1 8B Instruct를 기반으로 SFT(지도 미세 조정)한 모델로, 한국어와 한국의 문화적·전문적 맥락에 특화되어 있습니다. 자체 구축한 53개 영역의 데이터로 학습되었으며, 법률·재무·세무·회계 등 엔터프라이즈 활용에 중점을 두었습니다.
Attribution / 기여 표기: Model design, fine-tuning, and dataset curation led individually by the author, with partial use of company (ktds/AIDX) resources and infrastructure.
기여 표기: 모델 설계·파인튜닝·데이터 큐레이션은 저자가 개인적으로 주도하였으며, 일부 회사(ktds/AIDX) 자원 및 인프라를 활용했습니다.
Training Data / 학습 데이터
A self-developed dataset of ~3.6 GB / 2.33M examples (QnA, summarization, classification):
1.33M multiple-choice items across 53 domains (Korean history, society, finance, law, tax, math, biology, physics, chemistry, etc.), formatted with Chain-of-Thought.
Enterprise / 엔터프라이즈: legal, financial, tax, and accounting Q&A and document summarization (법률·재무·세무·회계 질의응답 및 문서 요약)
Education / 교육: Q&A and explanation generation across history, math, science (역사·수학·과학 질의응답 및 설명)
Research & culture / 연구·문화: Korean-context NLP, sentiment analysis, document generation (한국 맥락 NLP, 감정 분석, 문서 생성)
Customer service / 고객 서비스: conversational responses (대화형 응답 생성)
Limitations / 한계
This model is specialized for Korean language and culture; accuracy may drop for other languages, the latest international sources, or highly specialized fields. It may show limited reasoning on complex logical tasks, and biased training data can lead to biased outputs.
본 모델은 한국어·한국 문화에 특화되어 있어, 다른 언어·최신 국제 자료·고도 전문 분야에서는 정확도가 떨어질 수 있습니다. 복잡한 논리적 추론에서 한계를 보일 수 있으며, 학습 데이터의 편향이 출력에 반영될 수 있습니다.
Citation / 인용
bibtex
1@misc{seokdong_llama31_korean_v1_1,
2 title = {Llama-3.1-Korean-8B-SFT (v1.1)},
3 author = {[Your Name]},
4 year = {[year]},
5 url = {https://huggingface.co/SEOKDONG/llama3.1_korean_v1.1_sft_by_aidx}
6}