This dataset contains 47498 high-quality samples from Sera-4.6-Lite-T1 and Sera-4.6-Lite-T2. Training leads to a open-source SoTA 50.67% +/- 1.86% performance on SWE-Bench Verified at 32K context length, outperforming Devstral-Small-2 and GLM-4.5-Air.
Method:We keep only model-submitted train samples and then filter by truncation ratio at 32K tokens until a threshold ratio of 0.88.
Schema:
messages: Generated trajectory
instance_id: ID of trajectory
rollout_patch: Created patch to the codebase… See the full description on the dataset page:
https://huggingface.co/datasets/alucent/mirror-SERA-4.6-Lite-Best-Subset.