agentBox / AgentBox /README.md
Jivan01's picture
Setup
a5c9fd4
|
Raw
History Blame Contribute Delete
1.7 kB

🧠 CodeGuard Environment

πŸ“Œ Overview

CodeGuard is a deterministic RL-style environment for evaluating code-fixing agents.
It simulates tasks like lint fixing, vulnerability patching, and structural refactoring with strict reward and grading rules.


🧩 Environment Schema

State

{
    "score": "float",
    "history": [
        {
            "step": "int",
            "action": "str",
            "reward": "float"
        }
    ]
}

Action

  • Type: str
  • Constraints:
    • Non-empty
    • Max length: 1000 chars

Reward

  • Range: [-2.5, 1.0]
  • Components:
    • Task score (grader)
    • -0.02 step penalty
    • -2.0 destructive penalty

πŸ§ͺ Tasks & Difficulty

🟒 Easy β€” Lint Fix

  • Goal: Fix syntax errors
  • Grader: ast.parse()
  • Output: 0.0 or 1.0

🟑 Medium β€” Vulnerability Patch

  • Goal: Remove unsafe calls (eval, exec, etc.)
  • Checks:
    • Unsafe calls removed
    • Function structure preserved
  • Output: 0.0 – 1.0

πŸ”΄ Hard β€” Refactor + Type Hints

  • Goal: Improve structure + add typing
  • Checks:
    • Structural preservation
    • Type annotations present
  • Output: 0.0 – 1.0

βš™οΈ Setup & Usage

Install Dependencies

pip install -r requirements.txt

Run API Server

uvicorn src.env:app --host 0.0.0.0 --port 7860

Run Inference Loop

export HF_TOKEN=your_token
python inference.py

πŸ“Š Baseline Scores (Mock)

Task Score
Easy 0.88
Medium 0.71
Hard 0.54

Model: gpt-4.1-mini

πŸ§ͺ Testing

Run validation tests:

pytest tests/

πŸ“Œ Notes

  • All graders are deterministic & stateless
  • Execution time per grading <50ms
  • No randomness in reward or evaluation