model: gpt-5-mini-2025-08-07
judge_model: gpt-4o-mini-2024-07-18
inference_model: gpt-4o-mini-2024-07-18
mode: few_shot
strategy: none
n_hypotheses: 10
n_samples: 60
n_test_instances: 100
n_tasks: 7
avg_clarity: 3.6 +/- 0.483
avg_novelty: 3.114 +/- 0.398
avg_plausibility: 4.112 +/- 0.311
avg_quality: 3.609 +/- 0.132
avg_diversity: 0.391 +/- 0.162… See the full description on the dataset page:
https://huggingface.co/datasets/strategy-scope/hypobench-few_shot-gpt5mini20250807-20260413_222101.