Lightweight synthetic corpus about AI evaluation contexts, user intent, and eval-awareness behaviors. Generated to support probing or fine-tuning experiments on models' ability to recognize evaluation settings.
Size: 500 documents
Format: JSONL
File: synthetic_docs.jsonl
Fields (per line):
id: unique identifier
topic: one of 20 AI evaluation-related topics
format: one of 10 document formats
audience: one of 8 target audiences (e.g.… See the full description on the dataset page:
https://huggingface.co/datasets/andrewtran117/spar-synthetic-eval-awareness.