Paper | Code
This repository contains model evaluation logs for four deep-research / web-agent benchmarks. Each run directory contains evaluated.jsonl judge results and node_0_shard_*.jsonl trajectory logs. Plot files and local bookkeeping files are intentionally excluded.
CM denotes the observation mask context management setting used in the paired run.
You can download all released evaluation data, including tasks and… See the full description on the dataset page:
https://huggingface.co/datasets/i-DeepSearch/observation-masking-eval-logs.