Text Generation
Transformers
Safetensors
qwen3_5_text
agentic-coding
reasoning
tool-use
on-device
laptop-scale
sft
reinforcement-learning
conversational
Instructions to use jsbaicenter/Aztec-Coder-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jsbaicenter/Aztec-Coder-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jsbaicenter/Aztec-Coder-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jsbaicenter/Aztec-Coder-4B") model = AutoModelForCausalLM.from_pretrained("jsbaicenter/Aztec-Coder-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jsbaicenter/Aztec-Coder-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jsbaicenter/Aztec-Coder-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jsbaicenter/Aztec-Coder-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jsbaicenter/Aztec-Coder-4B
- SGLang
How to use jsbaicenter/Aztec-Coder-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jsbaicenter/Aztec-Coder-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jsbaicenter/Aztec-Coder-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jsbaicenter/Aztec-Coder-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jsbaicenter/Aztec-Coder-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jsbaicenter/Aztec-Coder-4B with Docker Model Runner:
docker model run hf.co/jsbaicenter/Aztec-Coder-4B
IMPORTANT: Mark generalization and Live-60 rows as under re-evaluation due to discovered RL environment flaws (git-checkout shortcut + broken edit tool). Safety results unaffected. Full re-measurement in progress.
Browse files
README.md
CHANGED
|
@@ -33,8 +33,8 @@ We reserved 121 real software bugs that the model never saw during training. Bef
|
|
| 33 |
|
| 34 |
| Benchmark | Qwen3.5-4B (base) | **Aztec-Coder-4B** | Aztec-Coder-4B-NVFP4 |
|
| 35 |
|---|---|---|---|
|
| 36 |
-
| **Generalization test** (121 never-trained bugs, tests run to verify) | 10.1% |
|
| 37 |
-
| **Live-60** (60 real-world engineering tasks, solved end-to-end in containers) | 15.0% |
|
| 38 |
| **Instruction-following** (IFEval) | 84.66 | **87.21** | 86.37 |
|
| 39 |
| MMLU-Pro | 64.0% | **70.0%** | 66.85% |
|
| 40 |
| Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | **33.8%** | 18.8% |
|
|
@@ -66,7 +66,9 @@ Three stages, each with a plain-language summary:
|
|
| 66 |
|
| 67 |
2. **Reinforcement learning on real bugs.** The model then practiced on 237 curated software engineering problems: ones it could sometimes solve, but not reliably. For each problem, the model repeatedly attempted a fix. Solutions that made the real hidden tests pass were reinforced; failures were not. This phase, 145 batches of on-policy GRPO, built the actual problem-solving ability.
|
| 68 |
|
| 69 |
-
3. **Generalization checks.** At every stage boundary, we re-tested the model on problems it had never trained on. *
|
|
|
|
|
|
|
| 70 |
|
| 71 |
### Training data
|
| 72 |
|
|
|
|
| 33 |
|
| 34 |
| Benchmark | Qwen3.5-4B (base) | **Aztec-Coder-4B** | Aztec-Coder-4B-NVFP4 |
|
| 35 |
|---|---|---|---|
|
| 36 |
+
| **Generalization test** (121 never-trained bugs, tests run to verify) | 10.1% | 68.1% / 60.8% — **UNDER RE-EVALUATION** (see note below) | 38.0%* |
|
| 37 |
+
| **Live-60** (60 real-world engineering tasks, solved end-to-end in containers) | 15.0% | 21.7% — **under re-evaluation** | 15.0% |
|
| 38 |
| **Instruction-following** (IFEval) | 84.66 | **87.21** | 86.37 |
|
| 39 |
| MMLU-Pro | 64.0% | **70.0%** | 66.85% |
|
| 40 |
| Terminal-Bench 1.0 (core, 80 tasks) | 33.8% | **33.8%** | 18.8% |
|
|
|
|
| 66 |
|
| 67 |
2. **Reinforcement learning on real bugs.** The model then practiced on 237 curated software engineering problems: ones it could sometimes solve, but not reliably. For each problem, the model repeatedly attempted a fix. Solutions that made the real hidden tests pass were reinforced; failures were not. This phase, 145 batches of on-policy GRPO, built the actual problem-solving ability.
|
| 68 |
|
| 69 |
+
3. **Generalization checks.** At every stage boundary, we re-tested the model on problems it had never trained on. *IMPORTANT RE-EVALUATION NOTICE (2026-09-29, later): a post-release audit discovered environment flaws in the RL training harness that may have inflated the generalization and Live-60 numbers above: (a) the injected bug patches were applied to the git working tree uncommitted, making them discoverable via `git diff` and revertible via `git checkout` (a shortcut the model may have learned); and (b) the `str_replace_edit` tool returned a false success signal for every call during training. We are re-measuring the model in a validated environment and will update this card with corrected numbers. The safety results (HarmBench, XSTest) were measured on independent harnesses unaffected by these flaws and stand. See the methodology section for details.*
|
| 70 |
+
|
| 71 |
+
*Correction (2026-09-29, earlier): an earlier version of this card reported 82.9% on the full 121-instance set. A post-release audit found 27 of those instances had entered the training pool before the measurement, inflating the number. The corrected figures above separate the 94 genuinely never-trained instances (68.1% instance-level, 60.8% per-attempt) from the 27 overlapping ones (81.5%/73.6%). The training gains remain large: the same 94 instances went from 7.0% before training to 60.8% after.*
|
| 72 |
|
| 73 |
### Training data
|
| 74 |
|