This package provides a 30-example benchmark across three governance-first task sets, a trace schema, a reproducible harness, and evaluation utilities. It is designed to be fine‑tuned or extended by partner teams. Final testing can be executed with their runners without modifying the task pack.
tasks.jsonl — 30 tasks (10 per set)
trace_schema.yaml — execution trace schema for logs and tools
eval_utils.py — metric… See the full description on the dataset page:
https://huggingface.co/datasets/Raiff1982/AGI_test.