Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
EZO-Humanities-9B-gemma-2-it-gguf – AI Model by grapevine-AI | AlphaNeural AI
You can deploy this model and start earning money today!
grapevine-AI
/
EZO-Humanities-9B-gemma-2-it-gguf
like
0
conversational
endpoints_compatible
gguf
template
gemma
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
What is this?
Axcxept(アクセプト)社のHODACHI氏によるgemma-2の日本語ファインチューニングモデル
EZO-Humanities-9B-gemma-2-it
をGGUFフォーマットに変換したものです。
imatrix dataset
日本語能力を重視し、日本語が多量に含まれる
TFMC/imatrix-dataset-for-japanese-llm
データセットを使用しました。
なお、imatrixの算出においてはf32精度のモデルを使用しました。これは、本来の数値精度であるbf16でのimatrix計算に現行のCUDA版llama.cppが対応していないためです。
Chat template
<start_of_turn>user ここにpromptを書きます<end_of_turn> <start_of_turn>model
Quants
各クオンツと必要と想定されるVRAM容量をまとめておきます。
クオンツ
VRAM
IQ4_XS
10GB
Q4_K_M
11GB
Q5_K_M
11GB
Q6_K
12GB
Q8_0
14GB
bf16
22GB
Note
llama.cpp-b3389以降でご利用が可能です。
なお、
-fa
オプションによる推論高速化に対応したllama.cpp-b3621以降の使用を推奨します。
Environment
Windows版llama.cpp-b3389およびllama.cpp-b3472同時リリースのconvert-hf-to-gguf.pyを使用して量子化作業を実施しました。
License
gemma license
Developer
Google & Axcxept co., ltd