| --- |
| license: mit |
| library_name: vllm |
| base_model: Qwen/Qwen2.5-Coder-0.5B |
| pipeline_tag: token-classification |
| tags: |
| - code |
| - code-pruning |
| - context-pruning |
| - vllm |
| - qwen2 |
| datasets: |
| - nick007x/github-code-2025 |
| metrics: |
| - f1 |
| --- |
| |
| # Fast Code Pruner |
|
|
| Task-aware context pruning for coding agents with a 17-layer |
| Qwen2.5-Coder-0.5B backbone and a native vLLM serving path. |
|
|
| [GitHub repository](https://github.com/rotem154154/fast-code-pruner) Β· |
| [Original SWE-Pruner](https://github.com/Ayanami1314/swe-pruner) |
|
|
| ## Validation quality |
|
|
| Evaluation uses the fixed 6,119-example seed-43 validation split. |
|
|
| | Model | Accuracy β | Precision β | Recall β | F1 β | |
| |---|---:|---:|---:|---:| |
| | **fast-code-pruner** | **85.94%** | **81.49%** | **83.49%** | **82.48%** | |
| | code-pruner | 84.07% | 80.02% | 80.91% | 80.46% | |
|
|
| ## Architecture |
|
|
| ```text |
| Original Code-Pruner Fast Code Pruner |
| |
| Qwen3 layers 1β28 Qwen2.5-Coder layers 1β17 |
| frozen backbone frozen weights + rank-8 LoRA |
| β β |
| hidden layers 7 + 14 + 28 normalized hidden layer 17 |
| 3 Γ 1024 β concatenate 3072 896 β gated PolyNorm β 2432 |
| β β |
| 1Γ bidirectional full attention 1Γ bidirectional full attention |
| 8 heads, width 3072 16 heads, width 2432 |
| β β |
| CRF emissions: 3072 β 256 β 2 CRF emissions: 2432 β 128 β 2 |
| ``` |
|
|
| Fast Code Pruner uses only the normalized final-layer representation and |
| physically skips the three removed attention branches. Rank-8 LoRA updates are |
| merged into dense weights during export. Prefix caching is disabled because |
| the bidirectional fusion head requires the complete input. |
|
|
| ## Serving |
|
|
| Install the native vLLM package directly from this model repository: |
|
|
| ```bash |
| pip install \ |
| "fast-code-pruner[vllm] @ https://huggingface.co/irotem98/fast-code-pruner/resolve/main/package/fast_code_pruner-0.1.0-py3-none-any.whl" |
| fast-code-pruner serve |
| ``` |
|
|
| Inputs longer than 8,192 tokens are split with overlap and merged |
| automatically. |
|
|
| ### Serving performance |
|
|
| | Model | Backend | Concurrency 1 β | Concurrency 16 β | |
| |---|---|---:|---:| |
| | **fast-code-pruner** | vLLM 0.13.0 | **79.0 req/s** | **143.4 req/s** | |
| | fast-code-pruner | Hugging Face | 16.01 req/s | 16.03 req/s | |
| | code-pruner | Hugging Face | 9.83 req/s | 10.03 req/s | |
|
|
| The vLLM results are medians of three 1,000-request end-to-end HTTP runs over |
| the same 100 validation examples on one NVIDIA RTX PRO 6000 Blackwell Server |
| Edition. Prefix caching was disabled. Median model-input throughput was 56.1k |
| tokens/s at concurrency 1 and 101.8k tokens/s at concurrency 16. |
| `fast-code-pruner` Hugging Face results are medians of three 100-request runs |
| with an 8,192-token model length, matching the serialized legacy endpoint. The |
| original `code-pruner` row uses its Hugging Face backend benchmark. |
|
|
| ## Training |
|
|
| The public repository includes a lightweight, hard-label launcher matching the |
| released architecture: |
|
|
| ```bash |
| bash train/train_fast_code_pruner.sh |
| ``` |
|
|
| It contains no teacher-target generation or knowledge-distillation stage. |
|
|
| ## Attribution |
|
|
| This model builds on |
| [SWE-Pruner](https://github.com/Ayanami1314/swe-pruner) and |
| [Qwen2.5-Coder-0.5B](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B). |
|
|
| ```bibtex |
| @misc{wang2026sweprunerselfadaptivecontextpruning, |
| title={SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents}, |
| author={Yuhang Wang and Yuling Shi and Mo Yang and Rongrui Zhang and |
| Shilin He and Heng Lian and Yuting Chen and Siyu Ye and Kai Cai and |
| Xiaodong Gu}, |
| year={2026}, |
| eprint={2601.16746}, |
| archivePrefix={arXiv}, |
| primaryClass={cs.SE} |
| } |
| ``` |
|
|