OR Reasoning Benchmark Results
Execution-based evaluation for the OR Reasoning model across classification, formulation, and code generation tasks.
Metric
Description
compile_rate
Generated Pyomo code passes syntax/structure check
feasibility_rate
Model has objective, variables, and constraints
objective_match_rate
Objective within 15% of ground truth
optimal_rate
Objective within 2% of ground truth
Reward Scheme… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/or-reasoning-benchmark-results.