WildTrace is a source-internal long-context multi-hop reasoning benchmark
built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are
mined in situ from long source documents before questions are written. The
strict481 release contains 481 locked tasks over 214 public long-form sources,
with full-document, evidence-withheld evaluation. The model under test receives
only the source document and the public question; evidence spans, clue… See the full description on the dataset page:
https://huggingface.co/datasets/CinderD/wildtrace.