SynthCoNL-neardedup corpus is a dataset of (comment, code, code) triplets generated starting from CodeSearchNet for the human data.
We then generated the code in a secondary language using Qwen 2.5 Coder-7B-Instruct.
SynthCoNL-neardedup has been used to finetune ModularStarEncoder-finetuned.
This dataset followed the near-deduplication process in ''MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings'', by processing the SynthCoNL raw dataset.… See the full description on the dataset page:
https://huggingface.co/datasets/modularStarEncoder/SynthCoNL-neardedup.