This evaluation dataset was created as part of the sft_gs__bontune_n1 experiment using the SkillFactory experiment management system.
{"model": "TAUR-dev/M-sft_gs__baseline_bon_tuning_n1-sft", "tasks": ["countdown_2arg", "countdown_3arg", "countdown_4arg", "countdown_5arg", "countdown_6arg", "commonsenseQA"… See the full description on the dataset page:
https://huggingface.co/datasets/TAUR-dev/D-EVAL__standard_eval_v3__sft_gs__bontune_n1-eval_sft.