Seed data collected from technical Q&A platforms and programming communities, used as source material for the SETA (Synthetic Environment Terminal Agent) data synthesis pipeline.
The dataset is organised by source, then by seed ID:
{source}/
└── {seed_id}/
├── main.json # primary Q&A pair with metadata
├── related_1.json # related question/post #1
├── related_2.json # related question/post… See the full description on the dataset page:
https://huggingface.co/datasets/camel-ai/seta-env-seed2synth-seed.