MUSES (Mining Unexplored Scientific Evidence to Spark novel hypothesis generation) is the first million-instance benchmark for prospective intellectual-roots prediction. Given an author's documented publication history at time t, the task is to rank a fixed pool of 2.33M scientific papers by how likely each one is to enter the author's next paper's bibliography.
The benchmark is hard along two orthogonal axes: