This dataset contains evaluation results for the model hf-inference-providers/openai/gpt-oss-20b:cheapest using the eval eval.py.
To browse the results interactively, visit this Space.
evals: Evaluation runs metadata (one row per evaluation run)
samples: Sample-level data (one row per sample)
evals = load_dataset('dvilasuero/simpleqa_verified-sample-2'… See the full description on the dataset page:
https://huggingface.co/datasets/dvilasuero/simpleqa_verified-sample-2.