This dataset packages the public tables from a corpus-boundary study of a
document-grounded RAG agent over NIST SP 800-160 Vol. 1 Rev. 1. The research asks
whether an enterprise team can determine that a document agent is operating
from its designated corpus—and turn observed behavior into reliable release
evidence.
The benchmark creates difficult corpus-boundary tests. The RAG produces the
behavior. EvalEngine turns those executions into… See the full description on the dataset page:
https://huggingface.co/datasets/Plumloom/plumloom-rag-eval-benchmark.