This is a filtered copy of osunlp/TACO-Cobalt for code-generation training
experiments. We ran Qwen/Qwen3-4B-Instruct-2507 three times per problem and excluded problems
where at least two of the three completions passed the tests. In short, this
keeps problems with 0/3 or 1/3 correct probe completions and filters out
problems with 2/3 or 3/3 correct completions.
train.jsonl: filtered training split.
validation.jsonl: filtered… See the full description on the dataset page:
https://huggingface.co/datasets/agurung/taco-cobalt-qwen3-4b-filtered-v1.