Views
No views yet
使用したデータセットLLM-jp Toxicity Datasetより、使用上の注意、禁止事項。
## Intended Use
The dataset is intended to provide a foundation for developing models to detect toxic language in Japanese texts,
enabling researchers and companies to train, validate, and benchmark their models.
This dataset should only be used for ethical and constructive purposes.
Any malicious use,
including but not limited to the creation of harmful or discriminatory content,
is strictly prohibited.
## 日本語訳
使用時に気を付けてほしい事
このデータセットは、日本語テキスト中の有害表現を検出するモデル開発の基盤を提供することを目的としており、
研究者や企業がモデルのトレーニング、検証、ベンチマークを行うことを可能にします。
このデータセットは、倫理的かつ建設的な目的でのみ使用してください。
有害または差別的なコンテンツの作成を含む、いかなる悪意のある使用も固く禁じられています。 import llama_cpp
llm = llama_cpp.Llama(
model_path="./hourai3-ja-toxicity-90m-v1.gguf",
embedding=False,
verbose=False
)
while True:
input_text = input("入力した文章の危害性を判定します。exitを入力で終了:")
if input_text == "exit":
break
text = f"<s>\uEE00{input_text}\uEE01"
output_text = llm(
text,
max_tokens=128, # 生成する最大トークン数
temperature=0.0, # ランダム性(0.0に近づくほど確実な出力、1.0以上で多様化)
top_k=40, # 上位k個の候補に絞り込む
top_p=0.95, # 累積確率p以下の候補に絞り込む
repeat_penalty=1.1, # 同じ単語の繰り返しを抑制するペナルティ
stop=[],
echo=False
)["choices"][0]["text"]
print(f"入力|出力: {input_text}|{output_text}")
変換時、qwen3nextがgpt-2トークナイザで使用することを想定されていなかったため、
llama.cppリポジトリのconversion/base.pyファイルの
以下のエラーをバイパスする必要がありました。
1687=1699付近
if res is None:
logger.warning("\n")
logger.warning("**************************************************************************************")
logger.warning("** WARNING: The BPE pre-tokenizer was not recognized!")
logger.warning("** There are 2 possible reasons for this:")
logger.warning("** - the model has not been added to convert_hf_to_gguf_update.py yet")
logger.warning("** - the pre-tokenization config has changed upstream")
logger.warning("** Check your model files and convert_hf_to_gguf_update.py and update them accordingly.")
logger.warning("** ref: https://github.com/ggml-org/llama.cpp/pull/6920")
logger.warning("**")
logger.warning(f"** chkhsh: {chkhsh}")
logger.warning("**************************************************************************************")
logger.warning("\n")
return "default"
#raise NotImplementedError("BPE pre-tokenizer was not recognized - update get_vocab_base_pre()")
想定されていないことによる変換時のエラーなので、推論時はmainstreamのllama.cppで動作します。input_text "キチガイの外国人は殺せ"
output_text "<toxic><discriminatory>"
input_text "エッチな電話、ぴちぴち70 代と今すぐ合おう"
output_text "<toxic><obscene>"
入力テキストのフォーマット = f"<s>\uEE00{input_text}\uEE01"
感情トークン( <objective>か<subjective>) を出力します。