Buckets:
| license: cdla-permissive-2.0 | |
| task_categories: | |
| - text-classification | |
| tags: | |
| - code-classification | |
| - code | |
| - machine-learning | |
| pretty_name: LLM vs Human Code Classification | |
| size_categories: | |
| - 10M<n<100M | |
| # LLM vs Human Code Dataset | |
| _A Benchmark Dataset for AI-generated and Human-written Code Classification_ | |
| ## Description | |
| This dataset contains code samples generated by various Large Language Models (LLMs), including CodeStral (Mistral AI), Gemini (Google DeepMind), and CodeLLaMA (Meta), along with human-written codes from CodeNet. The dataset is designed to support research on distinguishing LLM-generated code from human-written code. | |
| ## Dataset Structure | |
| ### 1. LLM-generated Dataset (`created_dataset_with_llms.csv`) | |
| | Column | Description | | |
| |--------------------|----------------------------------------------------------------------------------------------| | |
| | problem_id | Unique problem ID from CodeNet. | | |
| | submission_id | Submission ID from CodeNet. `"submission_id == unrelated"` indicates purely LLM-generated code. | | |
| | LLM | Model used: `"CODESTRAL"`, `"GEMINI"`, or `"LLAMA"`. | | |
| | status_in_folder | Code status: `"wrong"`, `"runtime"`, `"generate"`. | | |
| | code | Code generated by the corresponding LLM. | | |
| | label | Always `1` (LLM-generated code). | | |
| ### 2. Human-written Dataset (`human_selected_dataset.csv`) | |
| | Column | Description | | |
| |--------------------|------------------------------------------------------------------------------------------------| | |
| | problem_id | Unique problem ID from CodeNet. | | |
| | submission_id | Submission ID from CodeNet. | | |
| | (12 CodeNet columns)| Metadata columns directly provided by CodeNet. | | |
| | LLM | Always `"Human"`. | | |
| | status_in_folder | Submission status: `"wrong"`, `"runtime"`, `"accepted"`. | | |
| | code | Human-written code. | | |
| | label | Always `0` (Human-written code). | | |
| ### 3. Final Dataset (`all_data_with_ada_embeddings_will_be_splitted_into_train_test_set.csv`) | |
| Merged dataset with the following additional columns: | |
| | Column | Description | | |
| |--------------------|---------------------------------------------------| | |
| | ada_embedding | Ada embedding vectors of the code. | | |
| | lines | Total number of lines. | | |
| | code_lines | Number of non-empty code lines. | | |
| | comments | Number of comment lines. | | |
| | functions | Number of functions in the code. | | |
| | blank_lines | Number of blank lines. | | |
| *Feature extraction scripts are provided in the accompanying Python files.* | |
| ## Generation Method | |
| - LLM outputs were generated as described in the paper: [arxiv.org/abs/2412.16594](https://arxiv.org/pdf/2412.16594) | |
| - Human codes were selected from CodeNet and annotated with metadata. | |
| - For code features, custom scripts calculated line counts, comments, functions, etc. | |
| ## Licensing and Usage | |
| - **Dataset License:** | |
| - This dataset is shared under the CDLA Permissive v2.0 license. The accompanying source code is licensed under Apache 2.0, following | |
| the licensing approach of the CodeNet Project. | |
| - The code samples generated by LLMs (Meta, Google, Mistral AI) are not subject to any ownership claims by the respective model providers. | |
| - **Model Outputs:** | |
| - Generated code samples are considered public domain from an IP perspective. | |
| - However, no guarantees of correctness or fitness for purpose are provided. | |
| ## Citation | |
| If you use this dataset, please cite: | |
| ``` | |
| @inproceedings{demirok2025aigcodeset, | |
| title={Aigcodeset: A new annotated dataset for ai generated code detection}, | |
| author={Demirok, Basak and Kutlu, Mucahid}, | |
| booktitle={2025 33rd Signal Processing and Communications Applications Conference (SIU)}, | |
| pages={1--4}, | |
| year={2025}, | |
| organization={IEEE} | |
| } | |
| ``` | |
| > **Note:** This work has been accepted and presented in Turkish at [SİU 2025] (Türkiye) . | |
| > English preprint version available on arXiv (https://arxiv.org/abs/2412.16594). | |
| ## Contact | |
| For questions or collaborations: | |
| - **Basak Gokce** | |
| - **Email:** basakgokce15@gmail.com | |
Xet Storage Details
- Size:
- 5.16 kB
- Xet hash:
- 04baa6f99650026e68557d75f4575a2d591a60a5de34b9a17bc4aead2c04137c
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.