AI-generated code documentation pairs for training code embedding / retrieval models.
positive: A rich natural-language documentation of what the code does
queries: 4 natural-language search queries a developer might use to find this code
label: A short semantic label (3-8 words)
This dataset is designed for training bi-encoder embedding models (e.g.… See the full description on the dataset page:
https://huggingface.co/datasets/archit11/assesment_embeddings.