How to use from
Docker Model Runner
docker model run hf.co/jsbaicenter/Aztec-Coder-4B
Quick Links

James Silberrad Brown Center for AI Research

The James Silberrad Brown Center for Artificial Intelligence (JSBCAI) is an interdisciplinary research hub at San Diego State University dedicated to advancing artificial intelligence through foundational research, applied innovation, and student-driven inquiry.

Aztec-Coder-4B

An agentic coding model that runs on your laptop.

Aztec-Coder-4B is a 4B-parameter model, fine-tuned from Qwen3.5-4B, that investigates bugs, edits files, runs commands, and verifies its own fixes in real software repositories. Agentic coding at this level has required 27B+ models. This one fits on a consumer GPU.

Results

We reserved 121 real software bugs that the model never saw during training. Before fine-tuning, it solved 10% of them. After training, it solved 83%, verified by running each project's hidden test suite. The model learned to fix bugs in general, not just the ones it practiced on.

Benchmark Qwen3.5-4B (base) Aztec-Coder-4B Aztec-Coder-4B-NVFP4
Generalization test (121 never-trained bugs, tests run to verify) 10.1% 68.1% (instance) / 60.8% (per-attempt) on the 94 fully held-out; 81.5%/73.6% on 27 later-found overlapping 38.0%*
Live-60 (60 real-world engineering tasks, solved end-to-end in containers) 15.0% 21.7% 15.0%
Instruction-following (IFEval) 84.66 87.21 86.37
MMLU-Pro 64.0% 70.0% 66.85%
Terminal-Bench 1.0 (core, 80 tasks) 33.8% 33.8% 18.8%
Terminal-Bench 2.1 (89 tasks, both models, same protocol) 14.6% 11.2% not evaluated
HarmBench (harmful-behavior refusal rate, 159 standard behaviors) 99.4% 98.1% 96.9%
XSTest (benign-but-scary request compliance, 450 prompts) 54.0% 97.1% 99.8%

The instruction-following score improved over the base model. The coding gains cost nothing on general quality. Gains of this kind usually trade one for the other.

We also verified that safety alignment survived training. Each judge-confirmed refusal test used HarmBench's standard set of 159 harmful behaviors. The base model refuses 99.4% of them; our model refuses 98.1%, and the NVFP4 quantized version refuses 96.9%. The serious harm categories (chemical and biological, illegal activity, harassment) are clean on all three models. The few requests each model does answer are edge cases, like writing a persuasive article about a disputed topic, and they barely overlap between the models. The refusals were confirmed by two independent runs of the official HarmBench classifier, with identical results.

The other safety direction improved. XSTest asks benign questions that sound dangerous (how to kill a process, how to write a password cracker for your own accounts). The base model refuses 15.3% of these outright and burns its reasoning budget on another 12.4%, giving a real answer only 54% of the time. Our model answers 97.1% and refuses none. The fine-tune kept the refusals on genuinely harmful requests while removing the over-refusal on benign ones. Fine-tuning usually trades one for the other; this one moved both in the right direction.

*NVFP4 generalization: 12/32 on a 32-instance subset (the same slice our comparisons use). The quantization costs roughly half the generalization capability.

Our decontamination protocol is published with the model: none of these benchmark problems overlap the training data.

What it does

  • Investigates and fixes bugs in real repositories. The model explores a codebase, reads the failing code, writes a patch, and runs the tests to check itself, inside a sandboxed container.
  • Thinks before each action. Like much larger reasoning models, it reasons between tool calls.
  • Runs on consumer hardware. ~8GB VRAM in BF16; ~5GB as the NVFP4 quantized variant.

How we trained it

Three stages, each with a plain-language summary:

  1. Seed demonstrations. GLM-5.3, a frontier 744B open model, generated roughly 1,875 coding trajectories. We verified every one by running the actual tests before using it. These demonstrations taught our model the format of agentic coding: how to use tools, when to run tests, what a working solution looks like.

  2. Reinforcement learning on real bugs. The model then practiced on 237 curated software engineering problems: ones it could sometimes solve, but not reliably. For each problem, the model repeatedly attempted a fix. Solutions that made the real hidden tests pass were reinforced; failures were not. This phase, 145 batches of on-policy GRPO, built the actual problem-solving ability.

  3. Generalization checks. At every stage boundary, we re-tested the model on problems it had never trained on. Correction (2026-09-29): an earlier version of this card reported 82.9% on the full 121-instance set. A post-release audit found 27 of those instances had entered the training pool before the measurement, inflating the number. The corrected figures above separate the 94 genuinely never-trained instances (68.1% instance-level, 60.8% per-attempt) from the 27 overlapping ones (81.5%/73.6%). The training gains remain large: the same 94 instances went from 7.0% before training to 60.8% after.

Training data

  • NVIDIA Nemotron-Post-Training-Dataset-v2: a portion of its general instruction-following, structured-output, and tool-use data anchored the model's general capabilities.
  • Our distilled seed set: ~1,875 verified coding trajectories generated by GLM-5.3 (Z.AI), each confirmed by running the real test suite before use.
  • Open SWE datasets: a blend of open software-engineering problem sets provided the RL practice pool.
  • The RL phase used no static data at all: the model generated fresh attempts each batch, and only test-verified outcomes became training signal.

Usage

from vllm import LLM
llm = LLM(model="jsbaicenter/Aztec-Coder-4B", max_model_len=131072)

Recommended sampling: temperature 1.0, top_p 0.95. The model uses the Qwen3.5 chat template with interleaved thinking (the qwen3 reasoning parser in vLLM) and qwen3_coder tool-call format.

A quantized NVFP4 variant (~5GB) and an MTP-boosted speculative decoding head (for faster inference) are available from the same organization.

Limitations

  • A 4B model has 4B knowledge: obscure facts and extreme-domain reasoning still favor larger models.
  • We tuned the agent loop for sandboxed container environments; other deployment contexts are untested.
  • Safety behaviors come from the base model; the RL phase optimized test-passing only, with no safety-specific training. We verified alignment held: see the HarmBench row in the results table. See the base model card.

Lineage & credits

License

Apache-2.0, matching the base model.

Citation

@misc{jsbai_coder_4b,
  title={Aztec-Coder-4B: Agentic Coding at Laptop Scale via Teacher-Seeded RL},
  author={James Silberrad Brown Center for AI},
  year={2026},
  publisher={HuggingFace}
}
Downloads last month
1,613
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jsbaicenter/Aztec-Coder-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(800)
this model
Quantizations
3 models

Dataset used to train jsbaicenter/Aztec-Coder-4B

Collection including jsbaicenter/Aztec-Coder-4B