π§ CodeGuard Environment
π Overview
CodeGuard is a deterministic RL-style environment for evaluating code-fixing agents.
It simulates tasks like lint fixing, vulnerability patching, and structural refactoring with strict reward and grading rules.
π§© Environment Schema
State
{
"score": "float",
"history": [
{
"step": "int",
"action": "str",
"reward": "float"
}
]
}
Action
- Type:
str - Constraints:
- Non-empty
- Max length: 1000 chars
Reward
- Range:
[-2.5, 1.0] - Components:
- Task score (grader)
-0.02step penalty-2.0destructive penalty
π§ͺ Tasks & Difficulty
π’ Easy β Lint Fix
- Goal: Fix syntax errors
- Grader:
ast.parse() - Output:
0.0or1.0
π‘ Medium β Vulnerability Patch
- Goal: Remove unsafe calls (
eval,exec, etc.) - Checks:
- Unsafe calls removed
- Function structure preserved
- Output:
0.0 β 1.0
π΄ Hard β Refactor + Type Hints
- Goal: Improve structure + add typing
- Checks:
- Structural preservation
- Type annotations present
- Output:
0.0 β 1.0
βοΈ Setup & Usage
Install Dependencies
pip install -r requirements.txt
Run API Server
uvicorn src.env:app --host 0.0.0.0 --port 7860
Run Inference Loop
export HF_TOKEN=your_token
python inference.py
π Baseline Scores (Mock)
| Task | Score |
|---|---|
| Easy | 0.88 |
| Medium | 0.71 |
| Hard | 0.54 |
Model: gpt-4.1-mini
π§ͺ Testing
Run validation tests:
pytest tests/
π Notes
- All graders are deterministic & stateless
- Execution time per grading <50ms
- No randomness in reward or evaluation