RLPF-Qwen3-32B-PerfCodeBench

This repository contains the LoRA adapter for RLPF: Reinforcement Learning from Performance Feedback for Code Generation.

The adapter is trained from a Qwen3-32B base model with GRPO on PerfCodeBench-style executable performance-optimization tasks. The model is intended to generate optimized replacement implementations under a fixed benchmark harness.

Model Type

  • Base model: Qwen/Qwen3-32B
  • Adapter type: PEFT LoRA
  • LoRA rank: 32
  • LoRA alpha: 64
  • Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Training objective: GRPO with executable performance feedback

Reward

RLPF evaluates generated code by extracting the solution, compiling or loading it, running correctness checks, and measuring runtime. The default reward is:

R = max(stage_shape, CGRE + alpha_F * FBR + alpha_R * RBR)

where:

  • CGRE: correctness-gated relative efficiency
  • FBR: faster-than-baseline indicator
  • RBR: reference-or-better indicator
  • alpha_F = 0.16
  • alpha_R = 0.10
  • stage shaping values are -0.05 / 0.00 / 0.05 / 0.10

Evaluation Summary

On the PerfCodeBench test split with 306 family-disjoint tasks:

Model CRR FBR RBR CGRE
Qwen3-32B base 11.11 8.17 5.56 8.11
RLPF adapter, checkpoint-1500 54.58 46.41 25.82 38.58

All values are percentages. CRR is correct-and-runnable rate; FBR is faster-than-baseline rate; RBR is reference-or-better rate; CGRE is correctness-gated relative efficiency.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = "Qwen/Qwen3-32B"
adapter = "Egbertjing/RLPF-Qwen3-32B-PerfCodeBench"

tokenizer = AutoTokenizer.from_pretrained(base, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    base,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)
model = PeftModel.from_pretrained(model, adapter)

For best compatibility with the training distribution, use the chat template with enable_thinking=False when supported.

Intended Use

This adapter is intended for research on code generation, executable feedback, and performance optimization. It is not a general software-engineering assistant and should not be used without correctness tests, sandboxing, and runtime validation.

Limitations

Generated code may be incorrect, unsafe, non-compiling, or slower than the baseline. Performance measurements depend on compiler versions, hardware, system load, and benchmark harness settings. Always evaluate generated code under the target environment.

Downloads last month
3
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Egbertjing/RLPF-Qwen3-32B-PerfCodeBench

Base model

Qwen/Qwen3-32B
Adapter
(443)
this model