Benchmark of openai/gpt-oss-20b against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split.
Accuracy: 27.1%.
Metric
Value
Correct
664
Incorrect
1785
Errors
1
Total samples
2450
Total completion tokens
3,163,003
Raw stats:
{
"accuracy": 0.271,
"correct": 664,
"incorrect": 1785,
"error": 1,
"total": 2450,
"completion_tokens": 3163003
}