This dataset is a drop-in, official-format variant of the BrowseComp-Plus 100k corpus used in our RISE Agent experiments. It keeps the same row count, document ids, URLs, and column names as the original BrowseComp-Plus corpus, but replaces each document's text field with a structured version that adds a generated table of contents and section headings.
data.parquet: the corpus in the same three-column schema as the… See the full description on the dataset page:
https://huggingface.co/datasets/Tevatron/browsecomp-plus-md-toc-gpt5.4-nano.