Paper | GitHub
SHOR (Simple Harness Optimization Ranking) is a collection of human-verified optimization scenarios designed to enable the direct evaluation of harness optimizers. Instead of evaluating optimizers indirectly via the end-performance of target agents, SHOR uses Priority Ranking to quantify an optimizer's ability to identify which harness components (e.g., tools, prompts, memory, or workflows)… See the full description on the dataset page:
https://huggingface.co/datasets/LangAGI-Lab/SHOR.