Benchmark of openai/gpt-oss-20b against fingertap/GPQA-Diamond dataset. Results are averaged over 4 runs to reduce variance.
Accuracy: 64.4% with Python tool.
Raw stats:
{
"accuracy": 0.644,
"correct": 510,
"incorrect": 280,
"error": 2,
"total": 792,
"python_tool_calls": 306,
"python_tool_errors": 13… See the full description on the dataset page:
https://huggingface.co/datasets/kth8/gpt-oss-20b-GPQA-Diamond-benchmark.