Benchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-
/entry.json — that configuration's per-paper scores
results-/raw// —… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2-results.