open-scholar-gpt-oss-120b
Deep research agent evaluation on data/drtulu_open_scholar.jsonl (normal split).
Metric
Value
pass@1
0.0%
avg@1
0.0%
Trajectory accuracy
0.0% (0/11854)
Questions
11854
Trajectories
11854 (1 per question)
Avg tool calls
17.3
Full conversations
✅
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
None
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/open-scholar-gpt-oss-120b.