--- license: mit library_name: vllm base_model: Qwen/Qwen2.5-Coder-0.5B pipeline_tag: token-classification tags: - code - code-pruning - context-pruning - vllm - qwen2 datasets: - nick007x/github-code-2025 metrics: - f1 --- # Fast Code Pruner Task-aware context pruning for coding agents with a 17-layer Qwen2.5-Coder-0.5B backbone and a native vLLM serving path. [GitHub repository](https://github.com/rotem154154/fast-code-pruner) · [Original SWE-Pruner](https://github.com/Ayanami1314/swe-pruner) ## Validation quality Evaluation uses the fixed 6,119-example seed-43 validation split. | Model | Accuracy ↑ | Precision ↑ | Recall ↑ | F1 ↑ | |---|---:|---:|---:|---:| | **fast-code-pruner** | **85.94%** | **81.49%** | **83.49%** | **82.48%** | | code-pruner | 84.07% | 80.02% | 80.91% | 80.46% | ## Architecture ```text Original Code-Pruner Fast Code Pruner Qwen3 layers 1–28 Qwen2.5-Coder layers 1–17 frozen backbone frozen weights + rank-8 LoRA │ │ hidden layers 7 + 14 + 28 normalized hidden layer 17 3 × 1024 → concatenate 3072 896 → gated PolyNorm → 2432 │ │ 1× bidirectional full attention 1× bidirectional full attention 8 heads, width 3072 16 heads, width 2432 │ │ CRF emissions: 3072 → 256 → 2 CRF emissions: 2432 → 128 → 2 ``` Fast Code Pruner uses only the normalized final-layer representation and physically skips the three removed attention branches. Rank-8 LoRA updates are merged into dense weights during export. Prefix caching is disabled because the bidirectional fusion head requires the complete input. ## Serving Install the native vLLM package directly from this model repository: ```bash pip install \ "fast-code-pruner[vllm] @ https://huggingface.co/irotem98/fast-code-pruner/resolve/main/package/fast_code_pruner-0.1.0-py3-none-any.whl" fast-code-pruner serve ``` Inputs longer than 8,192 tokens are split with overlap and merged automatically. ### Serving performance | Model | Backend | Concurrency 1 ↑ | Concurrency 16 ↑ | |---|---|---:|---:| | **fast-code-pruner** | vLLM 0.13.0 | **79.0 req/s** | **143.4 req/s** | | fast-code-pruner | Hugging Face | 16.01 req/s | 16.03 req/s | | code-pruner | Hugging Face | 9.83 req/s | 10.03 req/s | The vLLM results are medians of three 1,000-request end-to-end HTTP runs over the same 100 validation examples on one NVIDIA RTX PRO 6000 Blackwell Server Edition. Prefix caching was disabled. Median model-input throughput was 56.1k tokens/s at concurrency 1 and 101.8k tokens/s at concurrency 16. `fast-code-pruner` Hugging Face results are medians of three 100-request runs with an 8,192-token model length, matching the serialized legacy endpoint. The original `code-pruner` row uses its Hugging Face backend benchmark. ## Training The public repository includes a lightweight, hard-label launcher matching the released architecture: ```bash bash train/train_fast_code_pruner.sh ``` It contains no teacher-target generation or knowledge-distillation stage. ## Attribution This model builds on [SWE-Pruner](https://github.com/Ayanami1314/swe-pruner) and [Qwen2.5-Coder-0.5B](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B). ```bibtex @misc{wang2026sweprunerselfadaptivecontextpruning, title={SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents}, author={Yuhang Wang and Yuling Shi and Mo Yang and Rongrui Zhang and Shilin He and Heng Lian and Yuting Chen and Siyu Ye and Kai Cai and Xiaodong Gu}, year={2026}, eprint={2601.16746}, archivePrefix={arXiv}, primaryClass={cs.SE} } ```