A Benchmark Dataset for AI-generated and Human-written Code Classification
Description
This dataset contains code samples generated by various Large Language Models (LLMs), including CodeStral (Mistral AI), Gemini (Google DeepMind), and CodeLLaMA (Meta), along with human-written codes from CodeNet. The dataset is designed to support research on distinguishing LLM-generated code from human-written code.
Dataset Structure
1.… See the full description on the dataset page: https://huggingface.co/datasets/basakdemirok/AIGCodeSet.