Aztec-Coder-4B-GGUF / README.md
prithivMLmods's picture
Update README.md
547ac47 verified
|
Raw History Blame Contribute Delete
4.02 kB
---
datasets:
- nvidia/Nemotron-Post-Training-Dataset-v2
base_model:
- jsbaicenter/Aztec-Coder-4B
library_name: transformers
tags:
- text-generation-inference
- llama-cpp
- agentic-coding
- reasoning
- tool-use
- on-device
- laptop-scale
- sft
- reinforcement-learning
license: apache-2.0
language:
- en
pipeline_tag: text-generation
---
# **Aztec-Coder-4B-GGUF**
> **Aztec-Coder-4B** is a 4-billion-parameter agentic coding model from San Diego State University's James Silberrad Brown Center for AI Research (JSBCAI), fine-tuned from Qwen3.5-4B to investigate bugs, edit files, run commands, and verify its own fixes in real software repositories β€” a capability class the authors note has typically required 27B+ models β€” while fitting on consumer hardware (~8GB VRAM in BF16, ~5GB as an NVFP4 quantized variant, which retains roughly half the generalization capability at 38.0%/12-32 on the same evaluation slice). It was trained in three stages: seed demonstrations from ~1,875 test-verified coding trajectories generated by GLM-5.3, reinforcement learning via 145 batches of on-policy GRPO over 237 curated software-engineering problems using test-pass/fail as the only reward signal, and repeated generalization checks against unseen bugs, alongside general instruction-following data drawn from NVIDIA's Nemotron-Post-Training-Dataset-v2. On 121 held-out, never-trained-on real bugs verified by running each project's hidden test suite, it jumped from 10.1% (base Qwen3.5-4B) to 82.9% solve rate, while also improving general capabilities rather than trading them off β€” IFEval rose from 84.66 to 87.21 and MMLU-Pro from 64.0% to 70.0% β€” alongside gains on Live-60 real-world engineering tasks (15.0% to 21.7%), with Terminal-Bench 1.0 held flat at 33.8% and Terminal-Bench 2.1 results still pending. It's served via vLLM with the Qwen3.5 chat template, `qwen3` reasoning parser, and `qwen3_coder` tool-call format, is tuned specifically for sandboxed container agent loops rather than general deployment, inherits its safety behavior unmodified from the base model (the RL phase optimized purely for test-passing with no safety-specific training), and is released under Apache-2.0.
## Model Files
| File Name | Quant Type | File Size | File Link | Description |
|-----------|------------|-----------|-----------|-------------|
| Aztec-Coder-4B.BF16.gguf | BF16 | 8.42 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.BF16.gguf) | Full BF16 weights. Highest quality, largest file size. |
| Aztec-Coder-4B.Q3_K_L.gguf | Q3_K_L | 2.42 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q3_K_L.gguf) | Lower quality but usable, good for low RAM availability. |
| Aztec-Coder-4B.Q3_K_M.gguf | Q3_K_M | 2.26 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q3_K_M.gguf) | Low quality. |
| Aztec-Coder-4B.Q4_K_M.gguf | Q4_K_M | 2.71 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q4_K_M.gguf) | Good quality, default size for most use cases, *recommended*. |
| Aztec-Coder-4B.Q4_K_S.gguf | Q4_K_S | 2.56 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q4_K_S.gguf) | Slightly lower quality with more space savings, *recommended*. |
| Aztec-Coder-4B.Q5_K_M.gguf | Q5_K_M | 3.07 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q5_K_M.gguf) | High quality, *recommended*. |
| Aztec-Coder-4B.Q5_K_S.gguf | Q5_K_S | 2.99 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q5_K_S.gguf) | High quality, *recommended*. |
| Aztec-Coder-4B.Q6_K.gguf | Q6_K | 3.46 GB | [Link](https://huggingface.co/prithivMLmods/JSBAI-Coder-4B-GGUF/blob/main/Aztec-Coder-4B.Q6_K.gguf) | Very high quality, near perfect, *recommended*. |
## llama.cpp
LLM inference in C/C++ β€” https://github.com/ggml-org/llama.cpp