A benchmark dataset for measuring cost-per-correct-answer of small language
models on verifiable agentic tasks.
data/sample/examples.jsonl — tiny synthetic sample.
dataset_card.md — full dataset card.
See the GitHub repository data_schema.md and schemas.py.
MIT. All content is synthetic.
The official implementation is available at:… See the full description on the dataset page:
https://huggingface.co/datasets/AntonioVFranco/slm-efficiency-frontier-benchmark.