Anyone can generate synthetic data. The hard part is knowing whether it's any good, or whether your eval set has leaked into your training set without you noticing. This small dataset is the demo for SynthKit, a tool that generates data and then grades it before you train on it. Try the grader in your browser: 🤗 huggingface.co/spaces/LaelaZ/synthkit.
The point isn't the size. It's the setup. The benchmark split overlaps the… See the full description on the dataset page:
https://huggingface.co/datasets/LaelaZ/synthkit-demo.