Evaluation results for cross-capability merging of OLMo-3 and OLMo-3.1 RL-Zero models on 454 coding problems.
We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible.
Code: pmahdavi/modal-eval
pmahdavi/Olmo-3-7B-Think-Math-Code… See the full description on the dataset page:
https://huggingface.co/datasets/pmahdavi/livecodebench-merging-leaderboard.