--- datasets: - nvidia/Nemotron-Post-Training-Dataset-v2 base_model: - jsbaicenter/Aztec-Coder-4B library_name: transformers tags: - text-generation-inference - llama-cpp - agentic-coding - reasoning - tool-use - on-device - laptop-scale - sft - reinforcement-learning license: apache-2.0 language: - en pipeline_tag: text-generation --- # **Aztec-Coder-4B-GGUF** > **Aztec-Coder-4B** is a 4-billion-parameter agentic coding model from San Diego State University's James Silberrad Brown Center for AI Research (JSBCAI), fine-tuned from Qwen3.5-4B to investigate bugs, edit files, run commands, and verify its own fixes in real software repositories — a capability class the authors note has typically required 27B+ models — while fitting on consumer hardware (~8GB VRAM in BF16, ~5GB as an NVFP4 quantized variant, which retains roughly half the generalization capability at 38.0%/12-32 on the same evaluation slice). It was trained in three stages: seed demonstrations from ~1,875 test-verified coding trajectories generated by GLM-5.3, reinforcement learning via 145 batches of on-policy GRPO over 237 curated software-engineering problems using test-pass/fail as the only reward signal, and repeated generalization checks against unseen bugs, alongside general instruction-following data drawn from NVIDIA's Nemotron-Post-Training-Dataset-v2. On 121 held-out, never-trained-on real bugs verified by running each project's hidden test suite, it jumped from 10.1% (base Qwen3.5-4B) to 82.9% solve rate, while also improving general capabilities rather than trading them off — IFEval rose from 84.66 to 87.21 and MMLU-Pro from 64.0% to 70.0% — alongside gains on Live-60 real-world engineering tasks (15.0% to 21.7%), with Terminal-Bench 1.0 held flat at 33.8% and Terminal-Bench 2.1 results still pending. It's served via vLLM with the Qwen3.5 chat template, `qwen3` reasoning parser, and `qwen3_coder` tool-call format, is tuned specifically for sandboxed container agent loops rather than general deployment, inherits its safety behavior unmodified from the base model (the RL phase optimized purely for test-passing with no safety-specific training), and is released under Apache-2.0. ## Model Files | File Name | Quant Type | File Size | File Link | Description | |-----------|------------|-----------|-----------|-------------| | Aztec-Coder-4B.BF16.gguf | BF16 | 8.42 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.BF16.gguf) | Full BF16 weights. Highest quality, largest file size. | | Aztec-Coder-4B.Q3_K_L.gguf | Q3_K_L | 2.42 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q3_K_L.gguf) | Lower quality but usable, good for low RAM availability. | | Aztec-Coder-4B.Q3_K_M.gguf | Q3_K_M | 2.26 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q3_K_M.gguf) | Low quality. | | Aztec-Coder-4B.Q4_K_M.gguf | Q4_K_M | 2.71 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q4_K_M.gguf) | Good quality, default size for most use cases, *recommended*. | | Aztec-Coder-4B.Q4_K_S.gguf | Q4_K_S | 2.56 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q4_K_S.gguf) | Slightly lower quality with more space savings, *recommended*. | | Aztec-Coder-4B.Q5_K_M.gguf | Q5_K_M | 3.07 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q5_K_M.gguf) | High quality, *recommended*. | | Aztec-Coder-4B.Q5_K_S.gguf | Q5_K_S | 2.99 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q5_K_S.gguf) | High quality, *recommended*. | | Aztec-Coder-4B.Q6_K.gguf | Q6_K | 3.46 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q6_K.gguf) | Very high quality, near perfect, *recommended*. | ## llama.cpp LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp