CordisBench tests whether language models can reason about the consequences of
component lifecycle changes in dynamic agent harnesses. Each record contains an
exact, programmatically generated oracle. Set-valued tasks use Jaccard
similarity, sequence prediction uses per-observable accuracy, and executable
reconfiguration is checked by running the proposed lifecycle operations.
This repository packages the frozen V2.0.1 release from
sileod/cordis-bench.… See the full description on the dataset page:
https://huggingface.co/datasets/sileod/cordis-bench.