Benchmark of openai/gpt-oss-20b against kth8/ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 83.7% with Python tool.
Raw stats:
{
"accuracy": 0.837,
"correct": 8371,
"incorrect": 1611,
"error": 18,
"total": 10000,
"python_tool_calls": 11354,
"completion_tokens": 5817252… See the full description on the dataset page:
https://huggingface.co/datasets/kth8/gpt-oss-20b-ValleyBench-benchmark.