このデータセットは、Kendamarron/OpenMathInstruct-2-ja-CoTを、英語と日本語のペアを含めるよう、日本語の応答データを英語の応答データに翻訳しました。
翻訳後のデータをllm-as-a-judgeで翻訳品質を1〜5段階評価し、最高評価5のみをfilterして取得したものです。
その後、訓練データセットとテストデータセットに9:1で分割しました。
"Kendamarron/OpenMathInstruct-2-ja-CoT": 訓練データセットのみ 15282件
Dataset:
Dataset({
features: ['problem', 'generated_solution', 'expected_answer', 'problem_source', 'problem_ja', 'thought', 'output', 'evaluation', 'cot_output', 'system']… See the full description on the dataset page:
https://huggingface.co/datasets/yasutoshi-lab/open-math-instruct-2-en-ja-cot.