Supervised fine-tuning and patch trajectories exported from bun-server-bench,
a benchmark for evaluating coding agents on real-world Bun server engineering tasks.
Every record comes from an agent run that passed both the public and hidden tests
for its task — these are verified solutions, not raw attempts. The benchmark engineers
each task so that a plausible-but-wrong implementation passes the visible tests and
fails the hidden ones, so a passing… See the full description on the dataset page:
https://huggingface.co/datasets/tinycomputerai/bun-server-bench-trajectories.