Text Generation
PEFT
Safetensors
code
code-review
bug-fixing
qwen
qwen2.5-coder
qlora
trl
static-analysis
conversational
Eval Results (legacy)
Instructions to use devanshty/Code-Autopsy with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use devanshty/Code-Autopsy with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct") model = PeftModel.from_pretrained(base_model, "devanshty/Code-Autopsy") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from devanshty/Code-Autopsy: direct link, hf CLI and curl.
- Browser
- Download file 5.66 kB
-
https://huggingface.co/devanshty/Code-Autopsy/resolve/main/README.md
- Command line
-
hf download hf://devanshty/Code-Autopsy/README.md
-
curl -L -o README.md https://huggingface.co/devanshty/Code-Autopsy/resolve/main/README.md
5.66 kB
| license: apache-2.0 | |
| base_model: Qwen/Qwen2.5-Coder-7B-Instruct | |
| library_name: peft | |
| pipeline_tag: text-generation | |
| tags: | |
| - code | |
| - code-review | |
| - bug-fixing | |
| - qwen | |
| - qwen2.5-coder | |
| - qlora | |
| - peft | |
| - trl | |
| - static-analysis | |
| model-index: | |
| - name: Code-Autopsy | |
| results: | |
| - task: | |
| type: text-generation | |
| name: Code Bug Diagnosis & Refactoring | |
| metrics: | |
| - name: Validation Loss | |
| type: loss | |
| value: 0.2442 | |
| - name: Token Accuracy | |
| type: accuracy | |
| value: 93.20% | |
| <div align="center"> | |
| # π¬ Code-Autopsy (QLoRA) | |
| ### Deep Structural Bug Diagnosis & Remediation Model | |
| [](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct) | |
| [-purple.svg)](https://github.com/huggingface/peft) | |
| [](https://wandb.ai/devanshtyagi1903-innothoughts/code-autopsy/runs/gc70q2q2) | |
| [](LICENSE) | |
| </div> | |
| --- | |
| ## π Model Summary | |
| **Code-Autopsy** is a specialized code intelligence model fine-tuned on top of **Qwen2.5-Coder-7B-Instruct** using 4-bit QLoRA. It operates like an autonomous forensic compiler: given buggy, defective, or vulnerable code snippets across Python, JavaScript, and other languages, it outputs a clean, structured diagnostic report: | |
| 1. **Bug Identified:** Exact forensic analysis of the flaw (e.g. mutable default arguments, ZeroDivisionError, unawaited asynchronous promises, race conditions). | |
| 2. **Root Cause:** In-depth explanation of *why* the defect occurs at the runtime/memory level. | |
| 3. **Fixed Code:** Corrected, refactored, and production-ready implementation. | |
| --- | |
| ## π Training Metrics & Cloud Logs | |
| The model was trained for **3 full epochs (246 steps)** on a curated dataset of code bugs and algorithmic repairs. | |
| | Metric | Initial (Epoch 0.06) | Final (Epoch 3.0) | Delta | | |
| | :--- | :---: | :---: | :---: | | |
| | **Training Loss** | `2.162` | **`0.255`** | **-88.2%** π | | |
| | **Validation Loss (`eval_loss`)** | `1.397` | **`0.2442`** | **-82.5%** π | | |
| | **Token Accuracy** | `60.29%` | **`93.20%`** | **+32.91%** π | | |
| | **Gradient Norm** | `0.27` | `0.39` | Stable | | |
| > π **Interactive Training Logs & Loss Curves:** | |
| > View the live dashboard, loss charts, and hardware telemetry on [Weights & Biases](https://wandb.ai/devanshtyagi1903-innothoughts/code-autopsy/runs/gc70q2q2). | |
| --- | |
| ## βοΈ Hyperparameters & Hardware Configuration | |
| * **Base Model:** `Qwen/Qwen2.5-Coder-7B-Instruct` | |
| * **Quantization:** 4-bit NF4 (`bitsandbytes` double quant) | |
| * **Compute Dtype:** `bfloat16` | |
| * **LoRA Rank ($r$):** `16` | |
| * **LoRA Alpha ($lpha$):** `32` | |
| * **LoRA Target Modules:** `q_proj`, `v_proj` | |
| * **Optimizer:** `adamw_8bit` | |
| * **Peak Learning Rate:** `2e-4` (with Cosine Decay and 5% Warmup) | |
| * **Effective Batch Size:** `8` (Per-device `1`, Gradient Accumulation `8`) | |
| * **Hardware:** NVIDIA GeForce RTX 5060 (8GB VRAM) | |
| --- | |
| ## π Quickstart: Running Inference | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig | |
| from peft import PeftModel | |
| BASE_MODEL = "Qwen/Qwen2.5-Coder-7B-Instruct" | |
| ADAPTER_REPO = "devanshty/Code-Autopsy" | |
| # 1. Load Tokenizer & 4-bit Base Model | |
| tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL, trust_remote_code=True) | |
| bnb_config = BitsAndBytesConfig( | |
| load_in_4bit=True, | |
| bnb_4bit_quant_type="nf4", | |
| bnb_4bit_compute_dtype=torch.bfloat16, | |
| bnb_4bit_use_double_quant=True | |
| ) | |
| base_model = AutoModelForCausalLM.from_pretrained( | |
| BASE_MODEL, | |
| quantization_config=bnb_config, | |
| device_map="auto", | |
| torch_dtype=torch.bfloat16, | |
| trust_remote_code=True | |
| ) | |
| # 2. Load Fine-Tuned Code-Autopsy Adapter | |
| model = PeftModel.from_pretrained(base_model, ADAPTER_REPO) | |
| model.eval() | |
| # 3. Format Diagnostic Prompt | |
| code_snippet = '''def append_item(val, lst=[]): | |
| lst.append(val) | |
| return lst''' | |
| prompt = f"""<|im_start|>system | |
| You are a code review expert. Analyze the provided code, identify any bugs or issues, explain the root cause, and provide a corrected version.<|im_end|> | |
| <|im_start|>user | |
| Language: python | |
| ```python | |
| {code_snippet} | |
| ```<|im_end|> | |
| <|im_start|>assistant | |
| """ | |
| inputs = tokenizer(prompt, return_tensors="pt").to(model.device) | |
| with torch.no_grad(): | |
| outputs = model.generate( | |
| **inputs, | |
| max_new_tokens=512, | |
| do_sample=False, | |
| pad_token_id=tokenizer.eos_token_id | |
| ) | |
| print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)) | |
| ``` | |
| --- | |
| ## π Diagnostic Output Format | |
| The model generates responses structured in Markdown: | |
| ```markdown | |
| ## Bug Identified | |
| Mutable default argument `lst=[]` used in function definition. | |
| ## Root Cause | |
| In Python, default arguments are evaluated once when the function is defined, not each time it is called. Modifying `lst` mutates the single shared list object across subsequent calls. | |
| ## Fixed Code | |
| ```python | |
| def append_item(val, lst=None): | |
| if lst is None: | |
| lst = [] | |
| lst.append(val) | |
| return lst | |
| ``` | |
| ``` | |
| --- | |
| ## π Citation & Credits | |
| * **Author:** Devansh Tyagi ([devanshty](https://huggingface.co/devanshty)) | |
| * **Base Architecture:** Alibaba Cloud Qwen Team (`Qwen2.5-Coder-7B-Instruct`) | |
| * **Frameworks:** π€ Hugging Face `transformers`, `peft`, `trl`, and Weights & Biases `wandb`. | |