Fast Code Pruner

Task-aware context pruning for coding agents with a 17-layer Qwen2.5-Coder-0.5B backbone and a native vLLM serving path.

GitHub repository · Original SWE-Pruner

Validation quality

Evaluation uses the fixed 6,119-example seed-43 validation split.

Model Accuracy ↑ Precision ↑ Recall ↑ F1 ↑
fast-code-pruner 85.94% 81.49% 83.49% 82.48%
code-pruner 84.07% 80.02% 80.91% 80.46%

Architecture

Original Code-Pruner                         Fast Code Pruner

Qwen3 layers 1–28                           Qwen2.5-Coder layers 1–17
frozen backbone                             frozen weights + rank-8 LoRA
          │                                           │
hidden layers 7 + 14 + 28                   normalized hidden layer 17
3 × 1024 → concatenate 3072                 896 → gated PolyNorm → 2432
          │                                           │
1× bidirectional full attention             1× bidirectional full attention
8 heads, width 3072                         16 heads, width 2432
          │                                           │
CRF emissions: 3072 → 256 → 2               CRF emissions: 2432 → 128 → 2

Fast Code Pruner uses only the normalized final-layer representation and physically skips the three removed attention branches. Rank-8 LoRA updates are merged into dense weights during export. Prefix caching is disabled because the bidirectional fusion head requires the complete input.

Serving

Install the native vLLM package directly from this model repository:

pip install \
  "fast-code-pruner[vllm] @ https://huggingface.co/irotem98/fast-code-pruner/resolve/main/package/fast_code_pruner-0.1.0-py3-none-any.whl"
fast-code-pruner serve

Inputs longer than 8,192 tokens are split with overlap and merged automatically.

Serving performance

Model Backend Concurrency 1 ↑ Concurrency 16 ↑
fast-code-pruner vLLM 0.13.0 79.0 req/s 143.4 req/s
fast-code-pruner Hugging Face 16.01 req/s 16.03 req/s
code-pruner Hugging Face 9.83 req/s 10.03 req/s

The vLLM results are medians of three 1,000-request end-to-end HTTP runs over the same 100 validation examples on one NVIDIA RTX PRO 6000 Blackwell Server Edition. Prefix caching was disabled. Median model-input throughput was 56.1k tokens/s at concurrency 1 and 101.8k tokens/s at concurrency 16. fast-code-pruner Hugging Face results are medians of three 100-request runs with an 8,192-token model length, matching the serialized legacy endpoint. The original code-pruner row uses its Hugging Face backend benchmark.

Training

The public repository includes a lightweight, hard-label launcher matching the released architecture:

bash train/train_fast_code_pruner.sh

It contains no teacher-target generation or knowledge-distillation stage.

Attribution

This model builds on SWE-Pruner and Qwen2.5-Coder-0.5B.

@misc{wang2026sweprunerselfadaptivecontextpruning,
  title={SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents},
  author={Yuhang Wang and Yuling Shi and Mo Yang and Rongrui Zhang and
          Shilin He and Heng Lian and Yuting Chen and Siyu Ye and Kai Cai and
          Xiaodong Gu},
  year={2026},
  eprint={2601.16746},
  archivePrefix={arXiv},
  primaryClass={cs.SE}
}
Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for irotem98/fast-code-pruner

Finetuned
(36)
this model

Dataset used to train irotem98/fast-code-pruner

Paper for irotem98/fast-code-pruner