This dataset is derived from .cache/glotcond/materialized/swe_fixer_diff_train_full.jsonl.
It adds prompt_input_token_count, computed with Qwen/Qwen3-4B-Base using
add_special_tokens=False.
The split is repo-disjoint by metadata.repo and targets train/dev/test ratios
of 0.7/0.1/0.2 while balancing the prompt token-count distribution.
Column order: id, prompt, answer, prompt_input_token_count, metadata.
split
rows
row… See the full description on the dataset page:
https://huggingface.co/datasets/d4nieldev/swe-fixer-diff.