Benchmark Evaluation Set for Daoism-Qwen3.5-9B and other LLMs on Daoist knowledge tasks.
由鼎稔道學館(lius.cc)發布。本 eval set 是 Daoism-QA-5K v0.1 中經 stratified sampling 抽出的 120 題 hold-out 集,永久切出不再用於任何 SFT 訓練。
完整評測方法論見本 repo 的 methodology.md 與 evaluator_prompt_v1.md。
項目
值
樣本數
120
抽樣方式
Stratified(5 task_type × 24 題)
分層
每類依 groundedness_score 取 top/mid/bottom 1/3 各 8 題
來源
Daoism-QA-5K v0.1(249 條 pilot)
切出狀態
Hold-out,永久不再用於 SFT 訓練